Daily editorial briefing

№ 20260813

DeepSeek's triple play: V4 Pro official release hits deployment trouble, Harness open-sourced, API prices jump; Gemini 3.7 Flash arrives at half price

August 13 (PT) was one of the most crowded model-release days of 2026. In a single window, DeepSeek did three things at once: shipped the V4 Pro official release (only to have t…

August 13 (PT) was one of the most crowded model-release days of 2026. In a single window, DeepSeek did three things at once: shipped the V4 Pro official release (only to have the community flag deployment anomalies, a suspected rollback, and a re-upload), open-sourced its agent runtime DeepSeek Harness v0.1, and announced broad API price increases. Google DeepMind released Gemini 3.7 Flash at half the price of its predecessor; OpenAI rolled out GPT-5.6’s Ultrafast tier and a new chief revenue officer; xAI’s Grok 4.6 and Grok Bot spread across the same timeline. The through-line is that agent competition has moved from “model capability” to the full stack of runtime plus cost — and DeepSeek’s launch mishap is a reminder that even when the model is ready, shipping can still go wrong.

Theme 1: DeepSeek V4 Pro 0813 official release: launch, rollback, re-upload

DeepSeek shipped the V4 Pro official release (model name deepseek-v4-pro) simultaneously to its app, web, and API on the evening of August 13 Beijing time. Officially published agent results include HLE (wo/w tools) 42.7/60.0, Terminal Bench 2.1 at 87.9, Cybergym 83.3, DeepSWE up from 12.8 in preview to 62.7, and a full-stack development benchmark of 71.1 — with the company claiming performance “above Opus 4.8 and close to Fable 5.” DeepSeek also said the model natively supports the OpenAI Responses API; the community noted you could swap Codex’s backend to DeepSeek “without changing a line of code.”

But hands-on results diverged sharply from the marketing. Multiple users reported that complex tasks were thinking for “absurdly short” periods — often under 10 seconds — shortly after launch, with a suspected emergency rollback in the early morning hours. DeepSeek’s site briefly removed the release announcement, which the community dubbed “possibly DeepSeek’s first failed launch.” API model fingerprints then flipped from readable formats (e.g., fp_v4pro_20260812_prod0820_fp8_kvcache) to plain 32-char hashes, for both Pro and Flash at the same time, pointing to a clear change in the production inference stack. Most dramatic: the community found that the V4-Pro-0813 repo uploaded to Hugging Face carried a config matching V4 Flash (hidden_size 4096, 43 layers, 256 routed experts, 64 attention heads, versus 7168, 61 layers, 384 experts, 128 heads for a proper Pro). The repo was taken down, and after re-upload the config was corrected but some safetensors weights also changed SHA256 hash and file size.

Evidence boundary: official benchmarks are vendor-reported; the launch mishap, rollback, and weight changes are all community observations (MaxForAI, karminski-牙医 and others), with no public explanation from DeepSeek. If “the wrong model shipped” holds, this looks like a deployment/release problem rather than a sign the model itself is a non-starter.

On the same day, DeepSeek announced broad API price increases: starting midnight August 17, peak/off-peak pricing kicks in, with cache-hit input up to 12x higher and output up about 4.5x; off-peak output is around $1.98 per million tokens and peak around $3.96. Even after the increase, pricing remains far below Fable 5 ($10 input / $50 output), and the community consensus is “still the cheapest at its tier.” Community analysis of the official release also highlighted: Terminal Bench 87.9 is only 0.1 below the priciest Fable 5, Cybergym 83.3 ties Fable 5, DeepSWE jumped from 12.8 in preview to 62.7, with native OpenAI Responses API support and Codex-specific optimization; the model adds three reasoning-strength tiers (low for simple tasks, high for daily agent work, max for complex tasks), tying reasoning budget directly to price. DeepSeek also disclosed its internal DSBench-FullStack / DSBench-Hard benchmarks: FullStack ranks Fable 5 (77.2) > Kimi K3 (73.7) > Opus 4.8 (71.6) > V4 Pro 0813 (71.1), while Hard ranks Opus 4.8 (71.7) > Fable 5 (68.3) > V4 Pro (67.2) > Kimi K3 (63.0) — on DeepSeek’s own hard agent tasks, Opus 4.8 beats the newer Fable 5. The benchmark’s tasks, data distribution, and harness are not yet public, so conclusions should be cautious.

Community verification continued after the mishap: one reviewer said they had burned roughly 100 million tokens to isolate what went wrong with V4 Pro 0813, promising a full evaluation and deep-dive video; others compared Pi + V4 Pro against Claude Code / Codex / Gemini / Grok by feel. These are single-post hands-on takes, not verdicts on the model — but “official release vs. community experience mismatch” is itself a major signal for the day.

Sources:

Theme 2: DeepSeek Harness v0.1: open-sourcing the agent runtime itself

DeepSeek Harness (DSH) v0.1 shipped as a developer preview on the evening of August 13 Beijing time, MIT-licensed, built on the Cordis meta-framework, with “everything is a plugin” as the core design: models, tools, skills, sessions, sandboxes, file systems, loops, orchestration, even the UI can be freely swapped, replaced, and extended. Within hours of release, GitHub stars climbed from ~24k to over 30k, and the community dubbed it “the emacs / foobar2000 of the agent era.”

From the leaked system prompt and community code reading, DSH is not just another chat wrapper: it defines several execution modes — standard mode (write code, modify repos), PTC mode (Programmatic Tool Calling, composing multi-step tool calls in TypeScript so five round-trips collapse into one), minimal mode (only a persistent bash and an absolute-path str_replace_editor), and create mode (running JavaScript written by the model on a live runtime to build custom agent presets). At the runtime level it includes KV-cache-aware design (never mutating historical prefixes, appending changes at the end), long-running Goal state, re-arming of Goals after session resume/fork, and a “Ralph Loop” — each round spawns a fully amnesiac fresh agent that retains long-term memory only through a shared Workspace. Sessions use event sourcing: the append-only log is the source of truth, state is a projection. It defaults to being driven by deepseek-v4-flash.

Community reaction is split. Supporters say “open-sourcing the harness you use to train models and run production is real open source”; critics argue it is “needlessly heavy” for everyday coding and report freezes in the frontend under parallel complex tasks plus API 404s. Developers also compared DSH to “the emacs of the agent era” (Pi being vim), and 玉伯 said it evokes foobar2000 — a highly composable software shape that grows through a community plugin ecosystem. Within a day, a DSH desktop build, a search plugin (dsh-find-plugin), a curated plugin list (awesome-dsh-plugin), and Tokei’s Day-0 support all appeared; the plugin ecosystem got off the ground quickly.

Evidence boundary: the system prompt and architecture analysis come from community reading, and although the paper is dense with category-theory notation, engineering judgments diverge. The significance: open agent runtimes are becoming a new competitive focus, and the view that “model × harness is the complete agent capability” is gaining acceptance across labs.

Sources:

Theme 3: Gemini 3.7 Flash: three-week iteration, half the price

Google DeepMind released Gemini 3.7 Flash just three weeks after 3.6 Flash, focused on coding and agent tasks, priced at $0.75 input / $3.75 output per million tokens — half the cost of 3.6 Flash. Google says the introductory price runs through year-end; the OpenRouter channel offers 50% off through August 27. The model supports a 1M-token context, reasoning, tool calling, multi-step planning, and multimodality, and tools like Cline integrated it the same day. DeepMind CEO Hassabis and multiple researchers lined up to vouch for it, calling it “lightning fast.”

Community take: compared with Grok 4.6 and DeepSeek V4 Pro on the same day, Gemini’s launch was “quiet to the point of being sad,” but the GA release has clear value for production users — Preview’s concurrency and stability had been the main pain points. Cline officially says “benchmark results claim near-frontier results” (that claim should be read against official benchmarks), researchers noted markedly better PDF understanding, and the community spotted a PDF demo even on the model card. Some observers question whether Gemini 3.5 Pro has been skipped (“straight to training 4.0”), which has no official confirmation. The significance: Flash’s release cadence (one generation every three weeks) plus a half-price cut keeps pushing down the unit cost of coding and agent work — moving in the opposite direction from DeepSeek’s price increase.

Sources:

Theme 4: OpenAI: GPT-5.6 Ultrafast tier and commercial leadership shake-up

OpenAI published the GPT-5.6 builder’s guide and the Ultrafast service tier: the same Sol model runs up to 14x faster than Standard in Ultrafast, with output up to 750 tokens per second, powered by Cerebras, currently in limited preview with Jane Street, Podium, Basis, and Rogo testing. By contrast, the Fast Mode launched in late July was only 2.5x speed at 2x price. The stated use cases cluster around latency-sensitive agent scenarios: real-time voice, incident response, financial research, customer service. The builder’s guide also disclosed: with retained reasoning and compaction enabled, Sol jumped from 13.3% to 38.3% on ARC-AGI-3 while emitting ~6x fewer output tokens; Luna hit 84.04% on BrowseComp, matching GPT-5.5 (84.36%), with per-run cost dropping from $33.27 to $1.33; new API capabilities include reasoning persistence, native multi-agent orchestration, and programmatic tool calling.

On the commercial side, OpenAI lost two senior executives within a week: COO Brad Lightcap left on August 11, and CRO Denise Dresser left on August 13 (only eight months after joining; in May she had publicly said enterprise customers contribute 40% of revenue, heading to 50% by year-end). Her successor is Dali Rajic, most recently President/COO at Google-owned Wiz, with prior roles at Zscaler and AppDynamics; reports also say Greg Brockman is reasserting influence over day-to-day operations. Evidence boundary: personnel news comes from WSJ/Axios reports relayed by the community (single-source); OpenAI has only confirmed the appointment. With IPO and organizational restructuring running in parallel, consecutive departures at the top of the commercial organization are worth watching.

Sources:

Theme 5: Grok 4.6 and Grok Bot: xAI’s two-track agent push

SpaceXAI and Cursor launched Grok Bot, a general agent that continuously operates apps and websites, completes tasks, and returns with finished work: it can research leads, update CRMs, process invoices, reproduce bugs, and run learned workflows across tools, even while the user’s computer is closed. Beta access is for select SuperGrok Heavy and Cursor subscribers, from $120 per seat per month for teams and $200 per month for individuals. The same day, Grok 4.6 spread: Perplexity CEO Aravind Srinivas said that in their Wide-And-Deep-Research benchmark, evaluated with the Perplexity Computer harness, Grok 4.6 “sits neatly on the Pareto frontier of performance vs cost”; Warp integrated it the same day; multiple reviewers called it “as smart as GPT-5.6 Sol, close to Opus 5 and Fable 5, but much cheaper and faster.” Elon Musk also teased Grok 4.7 within about a month, saying he “would be very surprised” if any model surpassed it in real engineering ability (vendor/founder self-report — treat accordingly).

The community ran direct comparisons on “who is actually cheaper”: one side says Grok 4.6 leads on cost-per-successful-task versus V4 Pro; the other side reports V4 Pro is twice as fast as Grok 4.6 on the same prompt. These are single-run hands-on results, not general conclusions. The significance: xAI is simultaneously accelerating on models, subscriptions, and agent products, and the Cursor partnership puts Grok directly in front of the most active coding-agent users.

Sources:

Theme 6: X open-sources the For You recommendation algorithm and ships a shadowban self-check tool

X open-sourced the recommendation algorithm behind the For You feed and simultaneously shipped a self-check tool: users can see whether their account or individual posts were labeled with visibility-limiting labels (“Limiting Labels”) last month, and download the data for their own analysis. The feature is currently a pilot: a random sample of eligible accounts, with the requirement of an account at least one year old and at least 10 posts in the prior month. Community reading of the open weights: predicted quotes carry 5x weight and predicted follows 4x; sharing via DM or copy-link is heavily rewarded; repeated text gets flagged as COPYPASTA_SPAM; “not interested,” blocks, mutes, and reports all hurt a post’s score, with reports carrying the strongest negative weight; likes feed a SimClusters discovery score with an 8-hour half-life; and the account-reputation score UserCredV2 is closer to a trust/authority measure, where follower quality matters more than counts.

Significance: putting the “cards on the table” for content distribution is both a transparency move and the first time creators can systematically see moderation decisions. One caveat: the published weights describe predicted events, not literal per-interaction point changes.

Sources:

Theme 7: Three open-weight launches: Qwen3.8-2.4T-A95B, MiniMax Music 3.0, Xiaohongshu dots.tts

Alibaba open-sourced Qwen3.8-2.4T-A95B: 2.4T parameters with 95B active, aimed at autonomous coding, deep research, and end-to-end agent execution; SiliconFlow offers Day-0 support at $2.00 input / $6.00 output / $0.25 cache-hit input per million tokens.

MiniMax released Music 3.0: open weights, positioned as a “production-ready” versatile music model that turns a creative concept plus optional lyrics into a complete song — composition, arrangement, performance, and production — in one pass, up to five minutes long; it already runs locally in ComfyUI, with the company saying “we will keep open until AGI arrives.” On the same day MiniMax H3 claimed the top of Video Edit Arena (vendor self-report; note the evidence boundary).

Xiaohongshu’s dots team open-sourced dots.tts, a 2-billion-parameter fully continuous end-to-end autoregressive speech synthesis model, claiming best average content accuracy and speaker similarity across three Seed-TTS-Eval subsets (vendor self-report).

The common thread: open weights are now covering modalities — music, speech — that were mostly closed before, and “open weights + Day-0 inference-platform support” is becoming the standard. Separately, the community is hands-on testing MiniMax H3 local deployment (a dual-4090 write-up) and comparing Seedance 2.5’s facial micro-expressions in video; these are single-post experiences, not benchmark conclusions.

Sources:

Theme 8: Agent research signals: watermarks, the cost of skills, and memory benchmarks

Anthropic is adding an imperceptible machine-readable watermark to Claude text (effective for models released on or after August 2), spanning the chatbot, API, and Claude Code, with a similar system for images; the stated rationale is EU AI Act transparency requirements. Some X users say they are canceling Claude subscriptions over the watermarks — an individual-behavior observation, not a general trend.

Two agent papers were widely shared the same day. First, work from Microsoft and colleagues attributes 307 agent failures to loaded skills (125 functional failures, 182 efficiency regressions), finding that “seemingly relevant skills push the agent to incorrectly implement or omit what the task required,” and that cost regressions are not explained by prompt length — the largest source is excessive verification (67 cases) followed by heavy implementation pipelines (30 cases): skills quietly turn validation checklists into mandatory work.

Second, Harness-IF scores 256 AGENTS.md rules one at a time: across 12 frontier models, raw accuracy runs 72.1%–85.9%, dropping to 66.1%–78.6% once “the model would have done it anyway” is stripped out (Against-Prior) — every model gets worse by 3.6–7.4 points. Precedence does not follow prompt depth: system prompts, project files, and user instructions outrank tool and skill descriptions. Separately, a UC San Diego team built an agent-memory benchmark splitting memory into four competencies (accurate retrieval, learning at test time, long-range retention, selective forgetting); every system tested — plain context stuffing, RAG, external memory modules — fell short on at least one. “Most systems are strong on retrieval and quiet about forgetting.”

Together these point to one engineering lesson: an agent’s capability depends not only on the model but on which rules, skills, and memory mechanisms are loaded — more is not automatically stronger.

Sources:

High-value briefs

🕐 Selected hourly signals

PT time Signal Why it matters
00:00 DeepSeek V4 Pro 0813 underperforms in community tests after launch; abnormally short thinking time on complex tasks First on-the-ground evidence of the launch mishap
01:00 DeepSeek’s site removes the V4 Pro 0813 announcement “First failed launch” discussion heats up
03:00 Internal DSBench benchmark surfaced; Opus 4.8 beats Fable 5 on Hard Self-built benchmarks reveal what a lab optimizes for
04:00 DeepSeek price charts spread: cache-hit input up to 12x Pricing shifts from floor-price to peak/off-peak
05:00 API fingerprints flip from readable formats to 32-char hashes; Pro and Flash switch together Indirect evidence of an inference-stack change
07:00 DeepSeek Harness open-sourced, past 24k stars within hours Open agent runtime becomes the day’s second storyline
08:00 DSH passes 30k stars; S&P 500 touches 7800 Hype and market sentiment peak together
09:00 Gemini 3.7 Flash officially released; Cline integrates same day Half-price workhorse joins the pricing war
10:00 OpenAI CRO Denise Dresser departure spreads (WSJ report) Consecutive shake-ups at the top of commercial org
11:00 Hugging Face V4-Pro-0813 config matches Flash; repo pulled and re-uploaded Hard evidence for “wrong model shipped”
12:00 X open-sources For You algorithm and ships shadowban checker Content-distribution rules become verifiable for the first time
13:00 Claude desktop adds auto-continue checkbox Small UX win for long agent tasks

Editorial conclusion

The most substantive change of the day is not any single model’s score but the coordinate system of agent competition: DeepSeek’s “official release + Harness + price increase” triple play put model capability, runtime substitutability, and per-call cost on the same table for discussion. The launch had its mishap, but the open-source Harness and the new pricing structure both matter more over the long run than one day’s noise. Google, OpenAI, and xAI all played cards the same day, converging on the same direction — around coding and agent scenarios, competing on model, runtime, and cost-per-success. For developers, two things are worth remembering: agent performance is the product of model × harness × the rules you load, and cheap-and-fast is becoming a harder competitive dimension than leaderboard scores.

Sources and method

Reviewed 30 raw capture files (20 hourly snapshots plus 10 named sources), of which 5 named sources had substantive content and 4 hourly snapshots were missing (hours 06/20/21/22). Major events have cross-source or strong single-source support; vendor self-reported benchmarks, single-post hands-on results, and personnel reports carry evidence boundaries where noted. Signal pool: rich.

WeChat QR code for 智简 Smart&Concise

FOLLOW ON WECHAT

智简 Smart&Concise

Search in WeChat for independent development and AI updates.