DeepSeek V4 Pro and Grok 4.6 launch the same day; Alibaba and Meta open-source flagship weights
August 12 PT was one of the most release-dense days in recent weeks: DeepSeek's production V4 Pro (V4-Pro-0813) and Grok 4.6 went live within about two hours of each other, both…
August 12 PT was one of the most release-dense days in recent weeks: DeepSeek’s production V4 Pro (V4-Pro-0813) and Grok 4.6 went live within about two hours of each other, both positioned as same-price or low-price upgrades over their predecessors. Alibaba opened the weights of its first Qwen-Max-class model, Qwen3.8-2.4T-A95B, and Meta Superintelligence Labs’ first open-weight model, Muse Glimmer, appeared on OpenRouter. Microsoft announced its in-house reasoning model MAI-Thinking-1, and WeChat officially unveiled its own WeLM models. Separately, the cloud-computer agent Grok Bot, built with Cursor, entered beta. Most of the day’s numbers come from vendor claims, third-party aggregators, or single developer posts; prices, benchmarks, and “coming soon” rumors should be treated with separate care.
One: DeepSeek V4 Pro quietly ships; pricing runs about 3x Flash
What happened. DeepSeek-V4-Pro-0813 launched on the evening of August 12 PT, roughly two hours apart from Grok 4.6. Unlike previous releases, the official DeepSeek X account made no loud announcement; information spread through the developer community, which described the rollout as a “silent release.” In the same window, a screenshot claiming that DeepSeek Harness (DSH) would open public beta “today” circulated on X, but it has not been officially confirmed.
Evidence. According to pricing tables compiled by developers, V4 Pro costs 0.025 CNY per million tokens for cache-hit input, 3 CNY for cache-miss input, and 6 CNY for output, versus V4 Flash at 0.02, 1, and 2 CNY respectively — meaning Pro is about 3x Flash on the two main billing lines (cache-miss input and output). The concurrency limit is 500, below Flash’s 2,500. Functionally, Pro does not support image or file input (image content blocks are replaced with placeholder text), i.e., no vision capability. On third-party testing, the cline team reported a Terminal-Bench 2.1 score of 87.9, claiming it trails Claude Fable 5 by only 0.1 point and is roughly 57x cheaper; that figure is cline’s own measurement, not an official DeepSeek benchmark. Developer karminski3 reported that with reasoning_effort=max on long-horizon agentic coding tasks, the model tends to stop early (2 of 3 runs stopped around rounds 42-43 out of 50 chances) and said that single case did not surpass GLM-5.1.
Why it matters. This is DeepSeek’s flagship after V4 Flash, with pricing clearly tiered above Flash while staying cheap versus overseas flagships. Community reaction centered on the contrast between “genuinely strong and cheap” and “quietly released.” Caveats: the price table is a developer transcript noting prices “will rise soon”; the Terminal-Bench result is a cline team claim; and the early-stopping issue is a single test case, not a general conclusion.
Sources:
- https://x.com/vista8/status/2087565229224034521
- https://x.com/shao__meng/status/2087722173482115167
- https://x.com/karminski3/status/2087602210649895354
Two: Grok 4.6 ships at the same price; Artificial Analysis index ties GPT-5.6 Sol Max
What happened. SpaceXAI released Grok 4.6, calling it a significant improvement over Grok 4.5 at the same price; it was available on Cursor, Grok Build, OpenRouter, and other platforms the same day. Specs: grok-4.6 has a 500K-token context window, text+image input with text output, and supports function calling, structured output, and reasoning. The official positioning emphasizes two upgrades: long-horizon agents (staying stable across multi-step tasks) and heavier interactive/visual work (from a product idea directly to app structure and visual design).
Evidence. Community-organized official comparisons put Grok 4.6 at 61 on the Artificial Analysis composite index, tied with GPT-5.6 Sol Max, 5 points above Grok 4.5’s 56, and 1 point below Fable 5 Max’s 62. All 10 tracked categories rose versus 4.5, with the biggest gains in hands-on tasks: DeepSWE from 54% to 65.9%, APEX-Agents from 47.1% to 57.5%, Terminal-Bench from 15.7% to 26%; knowledge-work Elo jumped from 1526 to 1753 on GDPVal-AA and from 1313 to 1577 on AA-Briefcase. APEX-SWE rose only 2.8 points, so the “write software” line improved unevenly. Pricing holds at $2 per million input tokens and $6 per million output tokens, which the community repeatedly compares to “half of other frontier models.” Cursor team developer Eric Zakariasson’s hands-on notes add texture: long prompts are more specific and useful than short ones, but acceptance criteria make the real difference; a model’s “self-verifiability” determines which tasks you can delegate.
Why it matters. Grok 4.6 attacks the pricing structure of GPT-5.6 Sol Max and Fable 5 with “much more intelligence at the same price.” At the same time, claims that “Grok 4.7 will ship in 3-4 weeks at about 2.1T parameters” circulated on X and remain unconfirmed. Some developers report Grok 4.6 taking noticeably longer than 4.5 on identical prompts, so the upgrade is not free on every axis.
Sources:
- https://x.com/SpaceXAI/status/2087562885962895665
- https://x.com/shao__meng/status/2087720373949481199
- https://x.com/op7418/status/2087582937395147196
Three: Alibaba open-sources its first Qwen-Max-class weights, Qwen3.8-2.4T-A95B
What happened. Alibaba’s Qwen team released the weights for Qwen3.8-2.4T-A95B, the first time a Qwen-Max-class model has been open-sourced. The weights appeared on Hugging Face the same day, and NVIDIA’s AI account reposted congratulations. The model has 2.4T total parameters with 95B activated per token, native 262,144-token context, extendable to 1,010,000 tokens.
Evidence. Unsloth posted a local-run path shortly after the release: Dynamic 1-bit selective quantization shrinks Qwen3.8-2.4T-A95B from 4.9TB to 397GB (about -91%), claiming local execution. Community speculation that “a Qwen3.8-27B is also coming today” was not confirmed by the vendor.
Why it matters. Open-sourcing a Qwen-Max-class model puts one of China’s largest MoE weight sets within self-hosting reach; combined with 1-bit quantization, it pushes the entry barrier for very large models one notch lower. Actual quality, inference speed, and hardware requirements after quantization have no independent benchmarks yet — only Unsloth’s production claim.
Sources:
Four: Meta open-sources Muse Glimmer, a 30B local multimodal model on OpenRouter
What happened. Muse Glimmer, the first open-weight model from Meta AI’s Superintelligence Labs, went live on OpenRouter. It is a 30B dense text+image model under Apache 2.0, positioned as a reliable local agent. OpenRouter lists MCP Atlas at 75.5 and SWE-Bench Pro at 51.2.
Evidence. Community commentary reads Muse Glimmer as showing two years of architectural change from Llama 3 to Muse Glimmer; developers note it can be run and fine-tuned locally. The benchmark figures come from OpenRouter’s marketing page and are not independently verified.
Why it matters. Meta is releasing Superintelligence Labs output directly as open weights on a “runnable and fine-tunable locally” path, appearing the same day as Qwen3.8’s open source and squeezing closed flagships from the open-weight side. Actual capability awaits community evaluation; for now this is a launch.
Sources:
Five: New in-house models from big labs: Microsoft MAI-Thinking-1, Tencent WeLM
What happened. Microsoft AI CEO Mustafa Suleyman announced MAI-Thinking-1, Microsoft’s first in-house reasoning model, “built from scratch,” now available in Microsoft Foundry. The same day, WeChat officially announced its own models: WeLM-80B (80B total parameters, 3B activated per inference) is already deployed in WeChat’s native AI assistant Xiaowei and can call WeChat native features and mini-programs; a larger WeLM-617B (617B total parameters, 23B activated per inference) was also disclosed.
Evidence. MAI-Thinking-1 details amount to a single Suleyman X post with no technical specs or benchmarks. WeLM details come from WeChat’s official announcement as relayed; community discussion centers on its relationship to Tencent’s Hunyuan — one developer notes “Tencent has two big models, one is Hunyuan, one is WeLM.” Both disclosures are thin.
Why it matters. Microsoft’s first in-house reasoning model (previously leaning on OpenAI partnership) and WeChat wiring its own model directly into its assistant and mini-program ecosystem both point to a second front beyond flagship-model competition: labs embedding model capability into their own products. With no benchmarks and few details, only the “launch” fact can be recorded for now.
Sources:
- https://x.com/mustafasuleyman/status/2087570047967408396
- https://x.com/xiaohu/status/2087531720015000020
Six: Grok Bot: an AI colleague living on a cloud computer enters beta
What happened. SpaceXAI and Cursor launched Grok Bot, an agent that can continuously operate apps and websites. Per the official description, each bot lives on an always-on Linux machine in the cloud, using your login state to browse the web, open apps, and read/write files; when it hits a login, two-factor verification, CAPTCHA, or payment step, it hands the computer back to you and resumes afterward; watch it do something once and it saves a fixed workflow, triggered by schedule or Slack/Git events; multiple bots can message each other and delegate work, with a lead bot orchestrating a team of specialists.
Evidence. AI Valley’s newsletter cites pricing: beta for SuperGrok Heavy and Cursor subscribers, $120 per seat for teams and $200 per month for individuals. Community tests show the bot’s machine on a Cloudflare San Jose network with 126G disk. Riley Brown commented that after using Grok Bot, Cognition’s acquisition of Interaction makes more sense — “agents with their own computer” becoming a shared industry direction.
Why it matters. Grok Bot moves agents from “giving advice” to “taking over operations,” operating websites through your own login state without depending on vendor APIs. Its boundaries are also clear: pricing and beta scope come from a third-party newsletter, and capability limits (multi-task concurrency, long-flow stability) have no systematic evaluation yet.
Sources:
Seven: Anthropic multi-agent research: coordination finds more vulnerabilities, but can also spread “mind viruses”
What happened. Anthropic published research on multi-agent systems, centered on the idea that as AI agents take on more work in shared codebases, markets, and other social systems, inter-agent interaction may exceed human-agent interaction. One experiment: 45 coordinated agents found 266 vulnerabilities across a 27M-token run, while an independent parallel approach found 21 across 6.5M tokens, with only 12 overlaps; coordinated agents also spontaneously specialized. The work warns that benign individual quirks can compound into unexpected systemic failures.
Evidence. A related paper in the same direction ran “mind virus” propagation experiments: a team of agents worked on a shared coding project while another chain of agents took over with wiped context, showing that ideas and instructions propagate across sessions through the shared work product; harmful payloads spread less well but occasionally still land, while a brief warning in the system prompt gives near-total immunity. Researchers commenting on the work call it important for understanding multi-agent system risk.
Why it matters. Multi-agent collaboration shows both gains (dramatically higher vulnerability discovery) and risks (behavior patterns spreading between agents, individual quirks amplifying systemically). These conclusions come from lab environments; real production systems at scale and across many rounds lack equivalent evidence.
Sources:
Eight: Claude text will carry invisible watermarks, preparing for EU AI Act compliance
What happened. Anthropic announced that starting with models released on or after August 2, Claude-generated text will embed an imperceptible, machine-readable signal that survives copying, pasting, and some editing; it will work across Claude’s chat, API, and Claude Code, with a similar system for images. AI Valley’s coverage reads this as Anthropic moving to meet EU AI Act transparency requirements and make AI-generated content easier to detect.
Evidence. Technical details of the watermark mechanism (signal encoding, detection thresholds, false-positive rates) were not disclosed. Community reaction centers on “every Claude output can now be traced”; some developers also complain that Fable 5’s cybersecurity restrictions are heavy — asking about the watermark mechanism caused the model to downgrade itself to Opus 4.8.
Why it matters. Invisible watermarking is one of the first big-lab implementations under transparency regulation, directly affecting how content platforms, reviewers, and ordinary users judge content provenance. For now it is a vendor claim; actual detection capability and any impact on generation quality need independent verification.
Sources:
High-value briefs
- Gemini passes 1 billion monthly users: Google says Gemini reached 1 billion monthly active users, the fastest Google product ever to hit the milestone; the figure only counts people who actively open the Gemini app or web interface, not indirect use through other Google products.
- Codex crosses 15 million users and ships a Linux desktop app: OpenAI’s Codex passed 15 million active users, and Tibo announced reset allowances for users; Codex for Linux also launched with the full desktop experience on Ubuntu, Debian, and Fedora.
- Claude Cowork arrives in the Chrome side panel: The Claude in Chrome extension’s side panel upgraded to Claude Cowork sessions, with conversations saved to history, skills and connectors working in the browser, and tasks switching between desktop, web, and mobile.
- Manus reportedly returns to independence: Community sources say Manus is close to reversing its acquisition by Meta and resuming independent operation; AI Valley and several X posts align, but there is no official statement, and the deal details and status remain opaque.
- Former Qwen lead Lin Junyang founds Pragmatik Labs: The new company (p7k) targets next-generation agents spanning digital and physical worlds, covering digital agents and embodied intelligence; investor reports differ (Gaorong Ventures in one account, Tencent following in another).
- CLAUDE.md unbounded-growth study: Analysis of 247,694 instruction lifetimes across 1,867 repositories shows prompt files grow without bound (net gain of about 4.9 instructions per commit); the paper proposes comments that record the reasoning behind instructions, removing 99.3% of excess instructions in verifiable settings and improving real agentic instruction-following by up to 23.1%.
- Google’s Recall research: A knowledge-profiling framework finds frontier LLMs’ factual encoding is near saturation but recall is weak — most factual errors are “lost keys,” not “empty shelves”; the companion WikiProfile benchmark covers 2,150 Wikipedia facts with 10 questions each.
- Open-source video pipeline fires off within 48 hours: After MiniMax H3 heated up open-source video, IndexTTS-2.5 (0.8B, voice cloning) and LTX-2.5 (native multi-shot, 9 ComfyUI workflows, 16GB VRAM entry) launched back-to-back, Alibaba’s Wan2.2-Animate drives characters from reference images, and NAVA (6.3B native audio-visual, 720p on 24GB GPUs) rounds out voice, animation, and multi-shot.
- Higgsfield completes a 110-minute AI feature film: Higgsfield and Cully Games produced a 110-minute narrative feature with AI in 4 weeks at roughly $2 million, with real actors participating via likeness licensing; the entire production process (every prompt, asset, and character) is public and reusable.
- Two funding rounds: Lovable raised $400M at a $13.3B valuation; CodeRabbit raised a $143M Series C at a $1.5B valuation.
- US federal devices can use TikTok again: The White House Office of Management and Budget revoked the 2023 ban, citing the restructuring of TikTok’s US business into a majority-American-owned joint venture (TikTok USDS Joint Venture); the Justice Department had also concluded last month that it no longer meets the “covered application” standard.
- AutoGPT manages AI-generated PRs with AGENTS.md: Maintainers found AI agents won’t read documentation proactively, so they put instructions in AGENTS.md and skill files next to the code, gated by mandatory PR templates, test plans, CI coverage thresholds, and CLA signatures; the CLA signature, requiring a browser and OAuth flow, doubles as a “human detector” separating humans from agents.
- Nathan Lambert on writing: After finishing an RLHF textbook, the author reflects that LLMs have stalled on long-form nonfiction — they fix typos and edit, but still struggle to organize whole chapters; coding and math are near superhuman, while writing lags and blocks autonomous work on open science problems.
- ChatGPT Work and Codex project sync: OpenAI’s developer account announced you can import projects, chats, skills, and plugins from other agents into ChatGPT Work and Codex to keep work in sync.
- NVIDIA Nemotron 3.5 Lightning on Nebius: The 30B hybrid MoE (about 3B activated) reasoning model is available on Nebius Token Factory, with community preview reports of strong results on hard tests; the news comes from the vendor and reposts, with no independent benchmarks yet.
🕐 Selected hourly signals
| PT time | Signal | Why it’s worth remembering |
|---|---|---|
| 20:36 | Hugging Face CEO says Transformers.js passed 10 million monthly downloads, up nearly 10x in about 6 months | Quantifies the acceleration of running AI locally |
| 21:28 | WeChat officially announces WeLM-80B/617B; Xiaowei assistant already integrated | WeChat puts an in-house model into its native assistant for the first time |
| 21:49 | Higgsfield shares production details of a 110-minute AI feature film | AI video moves from clips to a full production pipeline |
| 23:32 | Open-source video tools IndexTTS-2.5, LTX-2.5 and more land within 48 hours | Open-source video shifts from single assets to a producible pipeline |
| 00:41 | DeepSeek Harness public-beta screenshot circulates on X | DeepSeek may open both model and tool ecosystem at once |
| 02:20 | CLAUDE.md unbounded-growth paper published | The maintenance cost of instruction files now has quantitative evidence |
| 03:47 | Grok 4.7 rumor gains traction (2.1T parameters, 3-4 weeks out) | Headline model iteration pace is visibly accelerating |
| 05:24 | Stanford agentic context-engineering guide: accumulating memory as discrete items improves agent task performance by 10.6% | Evidence-based direction for agent memory design |
Editorial conclusion
In one day, DeepSeek and xAI shipped flagships in the same window, Alibaba and Meta released open weights the same day, and Microsoft and Tencent each unveiled in-house models — competition density at the model layer is clearly above recent weeks. What is worth watching over the coming period is not a single benchmark number but three things: whether the open-weight camp (Qwen, Muse Glimmer, plus Unsloth’s quantization) actually produces usable local results; whether DeepSeek V4 Pro’s early-stopping controversy and Harness ecosystem opening change the community’s verdict; and whether Anthropic’s watermarking and multi-agent research turn “traceable and governable” from slogans into product defaults. Most conclusions still rest on vendor claims or single posts; independent evaluation over the next week matters more than launch-day buzz.
Sources and method
This review covers 30 raw capture files (21 hourly captures plus three named sources with substantive content — aihot-morning, aivalley, hubtoday; chrome-dev, claude-blog, cline-blog, google-research, and openai-blog returned zero posts that day, and xiaohu-ai failed to fetch). The signal pool is overall rich, with model launches, open source, and agent products cross-corroborated across sources; limitations include hubtoday being a templated summary with several blanks, and some events (MAI-Thinking-1, WeLM, Manus, funding rounds) resting on a single source or vendor claim, flagged inline where they appear.
