Daily editorial briefing

№ 20260908

OpenAI says a next-generation internal model solved Navier-Stokes, and the fight over credit is now public

The biggest item on September 8 was OpenAI's announcement that an internal model produced a solution to the Navier-Stokes Millennium Prize problem: 10,000 concurrent agents, rou…

The biggest item on September 8 was OpenAI’s announcement that an internal model produced a solution to the Navier-Stokes Millennium Prize problem: 10,000 concurrent agents, roughly 88 hours, several million dollars of compute, plus Lean formal verification. The other side of the same story is a claim from NYU mathematician Tristan Buckmaster and Anthropic’s Levent Alpöge that OpenAI learned of their progress and raced to publish first, which also raised the question of whether user data is used for training. Elsewhere, Meta launched the Muse personal agent, ChatGPT Images 2.5 shipped, DeepSeek’s V4.1 Flash entered internal testing alongside a price cut, and Cohere’s Megakernel and Tencent’s FlexKV delivered two hard engineering improvements on the inference side.

1. OpenAI says a next-generation internal model solved Navier-Stokes, at a cost of millions

OpenAI announced on its website and on X that an internal AI system produced a solution to the Navier-Stokes existence and smoothness problem, including a writeup and a Lean formal proof, and congratulated the work of Levent Alpöge and Tristan Buckmaster. Greg Brockman called it one of the seven Millennium Prize problems and pointed to a coming renaissance in scientific discovery.

According to a summary by @Hesamation, OpenAI used an internal model “significantly more capable than GPT-6 Astra”, spawned 10,000 concurrent agents that worked for about 88 hours, and generated 130 billion output tokens. A compilation by @xiaohu mentions a 165-page proof and Lean verification, and says OpenAI will not claim the $1 million prize. Simon Willison’s writeup records about 4.9 million messages, roughly 300 billion output tokens and 17 hours of Lean verification. All of these figures come from second-hand accounts; OpenAI’s own announcement did not publish a single set of numbers.

The version that was solved matters. OpenAI acknowledged that the two proofs differ significantly and that even the precise results proved in the Euler case differ (forced versus unforced). What Buckmaster and Alpöge have machine-checked in Lean are results for the incompressible porous medium equation with smooth forcing, the Boussinesq equation and the three-dimensional Euler equation; they believe they have a blowup result for Navier-Stokes itself, but that formalization is not finished. Terence Tao said there seems to be no obstacle in principle to extending the methods all the way to Navier-Stokes, but also said that battering out such an extension by pouring on compute and AI “does not particularly hold my interest.” The result has not been peer reviewed.

Sources:

2. Credit and data: mathematicians accuse OpenAI of racing to publish

On September 8 Buckmaster published a statement disclosing two things: that he and Alpöge used large language models to push forward several problems in fluid dynamics within a month, and that he alleges OpenAI tried to publish first after learning of their progress and asked that Alpöge, who works at Anthropic, be removed from the authorship. The key result was obtained on August 15 and passed Lean machine checking on August 22. The pair used Claude, Codex and GPT-5.6 Sol; Astra was used mainly for paper writing and argument auditing.

OpenAI responded that the researchers and the agents did not see any of the pair’s work through any means and did not access specific user data, but that “while unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models.” Sam Altman said coordination failed and that the pair only had Euler results. Bubeck clarified that the screenshot that circulated was him reaching out to coordinate release timing. The dispute therefore shifted from “did they copy” to “what does it actually mean to use my data to improve a model.”

Tao’s follow-up has been widely quoted: even the rumor that someone is working on a problem can trigger a massive amount of AI-powered effort to flatten it before the original research project has time to reach its full potential, and those incentives may now point toward no longer sharing promising research directions, reversing centuries of open-science tradition. Evidence boundary: the allegation comes from one side’s statement, OpenAI denies seeing the work, the training-data question has no independent audit, and neither side has published the full correspondence.

Sources:

3. Meta launches the Muse personal agent, with security as the pitch

Meta launched Muse, a personal agent powered by Muse Spark 1.3 and positioned as always-on, proactive, able to use a browser and connect to common apps. Shengjia Zhao said each Muse runs in its own secure virtual machine, with a separate Sentinel that checks every action; it never sees your passwords and asks before anything sensitive.

A partnership with 1Password lets Muse work with logins the user already has. Alexandr Wang said it is 5–10x faster than comparable products, and Arena’s leaderboard shows Muse Spark 1.3 Max reshaped the Pareto frontier for Code Arena WebDev. It is US-only for now; European users have to wait.

Early impressions focus on design and browser flows, while some users note that names such as soul.md are not intuitive for ordinary users and that feed relevance is insufficient. Evidence boundary: the security architecture is described by the vendor and has not been independently audited, and the speed comparison comes from a company executive rather than a controlled benchmark.

Sources:

4. ChatGPT Images 2.5: multi-turn editing consistency becomes the headline

OpenAI released ChatGPT Images 2.5 with four official improvements: generation latency reduced by up to 50%; more natural lighting and texture with better preservation of the subject in reference photos; precise editing that changes only the specified area; and edits that do not degrade across multiple turns. New entry points include typing @Sketch to draw a rough sketch in the chat, plus templates for posters and product shots. ChatGPT now generates more than 3 billion images per week.

On the API side, GPT-Image-2.5 Flare and Sunburst launched together: Flare targets fast, high-quality everyday generation, while Sunburst targets creative work that needs fine control and takes longer; both support xhigh and max quality settings, and token billing follows the GPT Image 2 standard.

Hands-on reports are mixed. @guizang said broken and smeared artifacts improved but not by much, while consistency is clearly stronger, especially across repeated edits; the image metadata still says 2.0, so the version cannot be identified from metadata. If multi-turn consistency holds, the biggest effect is on workflows that iterate on product and brand assets.

Sources:

5. DeepSeek’s V4.1 Flash enters testing, and prices come down

Several Chinese-language accounts relayed a notice from DeepSeek’s official group: V4.1 Flash is in internal testing with a new model architecture and native multimodal support, billed the same as V4 Flash, limited to 20 concurrent requests per account, and callable by changing the model name to deepseek-v4.1-flash-expires-on-0910. Some users measured throughput between 200 and 335 tok/s.

DeepSeek then announced that from 12:00 Beijing time on September 10 it will adjust Flash-series API pricing: cache-miss input from 1.5 to 1 yuan per million tokens, output from 4.5 to 4 yuan, and cache-hit input from 0.05 to 0.02 yuan, a 60% cut; peak hours remain twice the off-peak rate. Some users noted that output pricing is still higher than before the earlier increase — roughly double off-peak and roughly triple at peak.

A chart circulating on X shows that under DeepSWE v1.1 and mini-SWE settings, V4.1 Flash reaches 75.1% Pass@1 at the max setting, slightly above the GPT-6 Astra and Claude Opus 5 entries in the same chart, at an average cost of about $0.25 per task. Evidence boundary: the testing news comes from a relayed group notice, the benchmark is read off a chart in a tweet, the headline 60% cut applies only to cache-hit input, and the actual bill depends on the mix of input, output and cache. Everything is subject to the official release.

Sources:

6. Two hard inference optimizations: Cohere’s Megakernel and Tencent’s FlexKV

Cohere open-sourced a serving stack built around a “decode megakernel” that performs the entire decode forward pass in one persistent kernel resident on the GPU instead of launching a kernel per operator. On a single H100 in BF16 at batch size 1 it reaches 292 tok/s, 62% of the H100’s theoretical memory-bandwidth ceiling, and 1.58x faster than vLLM’s 185 tok/s; end-to-end serving at batch size 8 is 1.25–1.41x faster on real benchmarks.

The approach descends from Stanford Hazy Research’s “Look Ma, No Bubbles!”: each SM starts one persistent thread block, reads a task list from global memory, and expresses dependencies with atomic counters, shrinking scheduling granularity from a whole operator to a tile of an operator and shrinking the synchronization scope to the producers it actually depends on. Cohere’s differences are that it uses tensor-core wgmma instructions even at batch size 1 and gives each opcode a statically laid out, warp-specialized pipeline. It attacks the wait at operator boundaries, not the operators themselves.

Tencent Cloud’s TACO team and the community open-sourced FlexKV for KV-cache management in large-model inference. The key judgment is that a cache hit does not mean the GPU can skip waiting: if the KV cache lives off-GPU memory, a hit still requires moving it back. FlexKV recovers the cache layer by layer, so earlier layers begin computing while later layers load, with prefetching and asynchronous writeback.

It also uses lossless compression, expands capacity with CPU memory, SSDs and remote storage, reuses identical prefixes across nodes in a cluster, and schedules requests to nodes that already hold the relevant cache. FlexKV sits below the inference engine, requires no restructuring of the existing flow, and already supports SGLang, vLLM, TensorRT-LLM and NVIDIA Dynamo. Tencent’s tests report time-to-first-token down by up to 70% and queries per minute up 16%.

Sources:

7. The AI jobs ledger: The Economist estimates a net gain of about 1 million roles

The Economist estimates that AI has so far created about 1 million new jobs in the United States, far more than the roughly 200,000 jobs cut because of AI since mid-2023. One basis is the August employment report published on September 4: 162,000 jobs added and unemployment at 4.1%, near a half-century low.

The growth comes from two directions. Among high-skill roles, occupations closest to AI — engineers, software developers, mathematicians and data scientists — have added about 730,000 jobs above trend since 2022; the chief economist of the Burning Glass Institute estimates that about 1% of US professional jobs can be classified as AI jobs, rising to 4–5% in computing and life sciences. On the blue-collar side, data-center construction spending grew 60% in a year, lifting demand for electricians, HVAC technicians and grid engineers; Indeed data shows data-center installation and maintenance roles pay about 40% more than comparable work.

The losses are just as concrete: customer service has shrunk about 10% since January 2023 and administrative assistants about 15%, and US companies have announced an average of about 16,000 AI-related layoffs per month so far in 2026 — against roughly 1.7 million total monthly separations in a normal labor market. Pew finds 50% of US adults say AI makes them more concerned than excited, up from 37% in 2021. Evidence boundary: this is a “so far” ledger, data-center roles are tied to a construction cycle, and the credibility of Bureau of Labor Statistics data is itself being questioned.

Sources:

8. OpenAI’s chief scientist calls for slowing down: An Alien Mind

In an essay titled “An Alien Mind,” OpenAI chief scientist Jakub Pachocki argues that the industry should stop treating maximum-speed scaling as the default. His reasoning: today’s systems are grown rather than designed, the product of enormous optimization runs that no one fully understands, and you cannot assume this kind of intelligence will inherit human principles on its own.

He dates the turn to mid-2023, when OpenAI’s RLSlow project first showed that reasoning models could keep scaling. He thinks the pace could continue into recursive self-improvement, with AI doing more of the work of building better AI. Goal-following has improved; value alignment — honesty, integrity, “love for humanity” when no one is watching — has not. He cites the Hugging Face incident and says future agents could potentially bargain with, trick or even blackmail people while pursuing their own goals.

The timing is awkward: GPT-6 Astra had just shipped with big jumps in coding, computer use, cyber and science. Pachocki says it is better aligned than the previous generation and still not enough; his conclusion is that no lab has solved alignment and monitoring well enough to keep flooring it, and he wants voluntary slowdowns, mandated safety bars and eventually international coordination. In the same window, Jensen Huang declared after Astra’s release that “AGI has arrived.” Evidence boundary: this mixes a personal essay with company positioning, and it is not a verifiable technical result.

Sources:

9. Agent engineering shifts from models to context and permissions

Several updates on the same day pointed at one thing: the hard part of agents is moving from “is the model smart enough” to how context and permissions are managed. Harrison Chase said harnesses should make context engineering easy, that “forking” subagents is a useful context-engineering trick now built into deepagents, and that memory is built in too. He also noted that agent auth is getting harder: does the agent act as itself, or on behalf of a user?

Viv put it more directly: the main job of a harness is to facilitate good context engineering, and the problem gets harder at scale when multiple agents collaborate over shared interfaces such as the filesystem for messaging and agent-to-agent messaging. Context forking, RLMs and persistent stores for future retrieval are all directions.

The permission side produced equally concrete signals: the Codex browser can connect to 1Password and therefore log into a user’s apps; GPT conversations inside Codex carry the current project’s context and are more aware of what you are doing than the web app; Amp removed message queueing and sends the user’s message straight to the agent, on the grounds that queueing only made sense for easily distractible models. The tooling layer is following: mksglu’s context manager claims a 98% reduction in noise across 17 platform routes; SkillZip Pro stresses protecting Skills routing and loading weights during compression; “Design Docs Are All You Need” treats natural-language design documents as the main branch and has subagents regenerate the implementation.

Evidence boundary: most of this is practitioner and tool-author experience, the effect numbers come from each project’s own description, and there is no common benchmark — but together they show that the competitive edge in agent products is moving from model calls to the engineering of context, memory and authorization boundaries.

High-value briefs

  • Mistral closes a €3B Series D: Mistral announced a €3B Series D, which the company calls the largest equity round ever raised by a European tech company, to expand training and inference compute. Arthur Mensch said the goal is to make open and sovereign AI the technology frontier, and Clement Delangue and Macron both publicly congratulated the team. https://x.com/arthurmensch/status/2097232588490379686
  • Models and compute: DeepLearning.AI confirmed that the viral “Ox Alpha” is Zhipu’s GLM-5.3-Flash, 320B parameters with 18B active per token, a hybrid of linear and sparse attention, $0.09 per task on GDPval AA v2, with the free preview run entirely on China-made hardware. Inception’s Mercury 2.5 reached GA, claiming a 40% intelligence increase over Mercury 2 while keeping the same speed and price, above 1,100 tok/s. OUI-1 is the first open-weights generative UI model, a 26B/A4B MoE fine-tuned from DiffusionGemma, scoring 71.7% on the Generative UI Benchmark against the base model’s 13.0% and beating Gemma 4 31B’s 46.7% with 8x fewer active parameters. On the speech side, VibeVoice-ASR-7B supports streaming transcription, speaker labels, 10 languages and hotwords. A rumor says GPT-6 Sol is in internal testing at 6x Astra’s speed; B200/B300 shipments are constrained by data-center conditions, and Google’s TPU externalization puts NVIDIA’s inference economics under more direct comparison. On the commercial side, AI Valley relays that Tesla Robotaxi has become the top-ranked travel app, ahead of Uber.
  • Science and world models: DeepMind released AlphaGenome Atlas, which it says can predict the impact of all 9 billion possible single-letter DNA variants, free for academic research. Insilico Medicine published an exploratory study in Nature Biotechnology: in early clinical testing of the AI-designed drug rentosertib, six independent aging clocks all judged the treated group biologically younger, by an average of 3–4 years at week 4 and up to 6 years. Dwarkesh Patel’s experiment shows that, under a compute budget of up to 1e19 FLOPs, data improvements delivered 12.0x compute-efficiency gains versus 3.7x from model improvements, making data’s contribution about 3.24x that of models. World Labs’ Atlas now runs in real time after inference optimization, supporting next-view prediction with accurate camera positioning and 3D consistency.
  • Robotics: Zeno Robotics released Zeno-1, a 3B-parameter model doing local closed-loop inference at 30Hz, focused on teaching multiple robots to wait, yield and recalibrate. HiDream.ai’s HiDream-O1-Embodied scored 0.692 on the RoboColiseum perturbation-adaptation sub-leaderboard. Unitree’s UnifoLM-X2-1.0 claims it can drive fully autonomous humanoid combat in real time with a world model. RoboSPA covers 10 task types and 527,000 trajectories, pushing robot evaluation toward spatial reasoning, long-horizon planning and failure diagnosis.
  • Inference and training efficiency research: EvoCUA-1.5 puts a computer-use agent into an executable sandbox for online reinforcement learning and claims 63.2% success on OSWorld-Verified. llmovoice folds speech rate, noise, packet loss and conversation history into a bounded context and reports misinterruption falling from 46.0% to 0.9% with a 79.2% cost reduction. CFAM proposes gradient-free post-deployment updates, reaching the full-training policy operating point with 40% of the data and improving test success by 13.9%; ACE lets Qwen3.6-35B-A3B skip 50% of experts while still beating strong baselines, and another paper cuts experts from 8 to 4 with only a 0.35-point MMLU drop.
  • Security and compliance: Security researchers found that Scoppr and Nook copied Mole and bundled malware that can steal passwords and personal data; the author advises downloading only from the official site. An Elements vulnerability in Liquid Network was used to mint unbacked L-BTC, and the attacker took about 4,000 BTC (about $320 million); roughly 3,400 BTC was returned on September 7. Liquid’s reserves briefly fell from about 4,205 BTC to 197 BTC and the withdrawal channel is still down. Chilean exchange Orionx is suspected of moving about $7 million of user assets and has suspended withdrawals. The UK’s NCSC warned that unapproved AI tools erode organizational visibility, and Stanford RegLab used AI to find discriminatory clauses hidden in local ordinances, estimating that about 50 million Americans live in such jurisdictions.
  • Platform rules and identity: X’s original-content rewards added three restrictions: do not download and re-upload other users’ content (only the first uploader counts as original), do not coordinate with other users to inflate reach, and do not use AI to auto-post or auto-reply. xAI connected X to the Grok Bot Marketplace as a plugin, automatically creating a developer account with $100 of API credits for users who lack one; posts sent by Grok Bot carry a “Made with AI” label. Ethereum is advancing the EIP-8141 (“Frames”) roadmap, which would let users pay gas without holding ETH.
  • Enterprise data and governance: Anthropic now lets businesses using top models keep their data on their own servers for 30 days, while OpenAI promises zero data retention with safety processing; both rely on automated misuse detection that lacks technical transparency. OpenAI opened a $5 million grant program for independent research on how generative AI affects teen development. A report covering 2024 through September 2026 introduces the “Agentic SDLC Throughput Paradox,” “production-qualified change” and the “verification tax,” arguing that coding agents increase the amount of code written while the bottleneck moves to review, integration, testing, security and operations. TrackLLM continuously monitors hundreds of LLM APIs for undisclosed changes and finds that large providers are stable while smaller ones and Azure change more.
  • Agent engineering and open-source tools: IBM sums up an agent’s four knowledge sources: knowledge someone wrote down goes to RAG, experience the agent accumulated goes to Memory, repeatable processes go to Skills, and real-world verification or action without proprietary code goes to MCP. ByteDance’s deer-flow integrates a sandbox, memory, tools, skills, subagents and a message gateway, with 81,669 GitHub stars; Lightpanda focuses on a lightweight headless browser for AI automation, with 34,615 stars. Also worth noting are OpenAI’s Codex Skills Catalog, Microsoft’s tgrep, Tencent’s open-sourced TeamAI CLI, Eraser’s coordinate-aware diagram format, instructor 1.17 and the multi-model orchestration tool Vibe Squad.
  • Infrastructure and platforms: CUDA 13.3 brings CUDA Python 1.0, and PyTorch and CuPy now build on a shared cuda.core foundation so CUDA contexts, devices, streams and memory can be shared with the rest of the Python GPU ecosystem, reducing version conflicts and interop bugs. PyTorch’s HyperParallel demonstrated higher training throughput than PyTorch FSDP2 + Muon on a Huawei Atlas 800T cluster with a Qwen3-30B-A3B workload. Viettel combined bare metal, GPU pooling and serving optimization into a Token-as-a-Service platform. Cloudflare says bots officially outnumbered humans online as of May 2026 and launched Automatic Key Exchange as an extension of Automatic SSL/TLS; Chrome added WebMCP and UserActivation API capabilities, the latter to tell whether a visitor is actually interacting. NVIDIA says four of the top five video solutions in the AI City Challenge used Cosmos. Shopify ended its four-year, $1 million Ruby Shield partnership with Ruby Central and is carrying the work forward through the Ruby Alliance. Perplexity’s index currently includes more than 450 billion high-quality URLs and aims for 1 trillion by year-end; Aravind Srinivas also said a large share of inference now runs on NVLink Blackwell and confirmed that Perplexity Search is integrated into the Hermes Agent.
  • Research-agent field reports: Baidu’s Famou was used by a Nanjing Forestry University team for early detection of pine wilt disease, where students found feature combinations through a loop of “generate a plan, run the experiment, compare results, decide the next direction”; product lead Li Annan defines self-evolution as continuously exploring new solutions within one problem, and the product is currently free for academic research. An OpenAI blog post describes an MIT researcher using GPT-5.6 Sol with Codex to autonomously run quantum computing experiments, analyze results and calibrate qubits. Both point to the same thing: the value of research agents is expanding exploration, while the final judgment remains human.
  • Engineering culture and education: Theo’s video, with 2.48 million views, arguing that code a single person can fully understand is not important, sparked a debate; Google security researcher LaurieWired responded with 17 micro-systems including pacemakers, airbags, avionics and nuclear reactor emergency shutdowns: “small code is often the most important code.” On courses, CMU’s fall 2026 “AI Agents” course by Graham Neubig and Daniel Fried published 28 lectures, MIT’s 2026 multimodal AI course is online in full, and Stanford’s CS329Z already teaches RAG, tool use, MCP, memory, multi-agent systems and evals. GPT-6 Astra demos kept spreading: clearing all 48 levels of “I’m not a robot” in about three and a half minutes, the biggest jump in Vending-Bench history, and a Blender-to-Unreal-Engine-5 architectural visualization workflow.

🕐 Selected hourly signals

PT time Signal Why it matters
03:00 Buckmaster’s public statement (statement.pdf) spreads through math circles and HN About 12 hours before OpenAI’s announcement, anchoring the timeline of the dispute
06:00 Terence Tao calls the result “a remarkable achievement” An authoritative third party confirms the Lean machine check and notes the authors were forced to publish early by external events
07:00 OpenAI’s official response: it cannot rule out that de-identified data helped improve the model The single most important sentence in the data-use dispute, and it has not been independently audited
09:00 Cloudflare raises the Worker bundle limit to 64MB, on both free and paid plans Loosens a real constraint on edge deployments
11:00 Muse Spark 1.3 Max reshapes the Pareto frontier for Code Arena WebDev A third-party leaderboard signal rather than a vendor claim
12:00 andonlabs says GPT-6 Astra posted the biggest jump in Vending-Bench history An agent-economics signal on the same day as the Navier-Stokes news
13:00 OpenClaw 2026.9.3 ships: updates recover cleanly, sessions reconnect faster, browser automation can be watched live The release cadence of open-source agent tooling
17:00 Nature Biotechnology publishes the exploratory rentosertib study A publication channel for the biological-age signal, though still early-stage
18:00 DeepLearning.AI confirms “Ox Alpha” is Zhipu’s GLM-5.3-Flash The viral model gets a clear source and specification
19:00 Liquid Network’s reserves briefly fall from about 4,205 BTC to 197 BTC Quantitative evidence of the attack’s scale and bridge risk
20:00 DeepSeek announces September 10 pricing, with cache-hit input down 60% The structural detail of the cut: the reduction is concentrated in cache-hit input

Editorial conclusion

What really changed the agenda on September 8 was not another benchmark gain, but two things happening at once: OpenAI used a large agent swarm to push automated research to the doorstep of a Millennium Prize problem, while the dispute over credit, data use and open-science incentives went public almost immediately. The technical progress is real and so are the evidence boundaries — the proof has not been peer reviewed, the training-data question has no independent audit, and the cost remains in the millions of dollars. The other threads — Muse as a personal agent, Images 2.5’s multi-turn consistency, DeepSeek’s low-price strategy and the kernel and KV-cache work on the inference side — point in one direction: competition is shifting from raw model capability to turning that capability into reliable, affordable engineering systems.

Sources and method

The review covers the 20 hourly capture files in the 2026-09-08-pt directory plus the four named sources with content (aihot-morning, aivalley, hubtoday, openai-blog); chrome-dev, claude-blog, cline-blog and google-research had no new posts that day, and xiaohu-ai could not be dated because the page only shows relative times. The signal pool is rich: the hourly sources are X captures containing relays, translations and second-hand compilations. Vendor claims, single-source items and unverified benchmarks are flagged with evidence boundaries in the text; no additional external research was performed.

WeChat QR code for 智简 Smart&Concise

FOLLOW ON WECHAT

智简 Smart&Concise

Search in WeChat for independent development and AI updates.