Overview
10 stories in this issue. The first 3 are today's priorities.
Hot model updates
- Top · GPT Image 2.5 ships, with Flare and Sunburst splitting the API
- Top · DeepSeek cuts Flash prices, effective tomorrow noon (Chinese-language source)
- Top · Ant open-sources Ling-3.0-flash-VL: 124B total, 5.5B active
- Mercury 2.5 lands on OpenRouter at a claimed 1,107 tokens per second
Global AI news
- OpenAI says 10,000 agents solved Navier-Stokes in 88 hours
- Cognition raises more than $2B at a $48B valuation
- Meta drops AI usage from engineer performance reviews
- Samsung and Mistral partner on semiconductor-specific models
Regional and early signals
- Tmall's AI Space Station puts model subscriptions on retail shelves
- Wang Yunhe's team ships NeoHorse-1 in 4B and 9B sizes (Chinese-language source)
Hot model updates
01/10
GPT Image 2.5 ships, with Flare and Sunburst splitting the API
OpenAI released ChatGPT Images 2.5 with up to 50% lower generation latency than Images 2.0, added two API models priced at the same token rates as GPT Image 2, and published a system card with unsafe-generation rates for all three.
Start with the consumer side. From September 8, ChatGPT Images 2.5 is rolling out to every ChatGPT, ChatGPT Work and Codex tier across desktop, mobile and web. OpenAI describes more natural lighting and richer textures, better preservation of subjects in reference photos — distinctive facial features and other identifying details now carry more reliably across repeated generations and edits inside one conversation — and up to 50% lower generation latency than Images 2.0. The most visible new feature is Sketch: type @Sketch in the composer to open a canvas, draw a rough reference with a finger, stylus or mouse, and have the model build from it. The release also adds templates for common formats such as flyers and product shots, comments left directly on a generated image for targeted edits, and prompts that can be shared alongside the image. OpenAI says more than 3 billion images are created each week across ChatGPT Images and the GPT-Image models in the API.
The API splits in two, and OpenAI's model pages are blunt about the difference. Flare is the "fastest model for high-quality, everyday image generation" — social assets, product imagery, rapid prototyping; SiliconANGLE reports it generates two to four times faster than GPT-Image-2, with better transparent-background output. Sunburst is the "most capable model for image generation and editing," trading generation time for editing precision.
| Flare | Sunburst | |
|---|---|---|
| Positioning | Everyday volume | Precision editing |
| Performance / speed rating | Higher / Very fast | Highest / Medium |
| Model ID | gpt-image-2.5-flare | gpt-image-2.5-sunburst |
| Dated snapshot | -2026-09-08 | -2026-09-08 |
| Quality settings | low…auto, six steps | low…auto, six steps |
Both can be selected directly in the Image API or as the model behind the Responses API image generation tool.
Prices did not move. The model pages list one table for both tiers and state outright that token rates match GPT Image 2:
| Per million tokens | Flare / Sunburst |
|---|---|
| Text input | $5.00 |
| Text cached input | $1.25 |
| Image input | $8.00 |
| Image cached input | $2.00 |
| Image output | $30.00 |
| Text output | Not billed |
The line that will actually trip people up is the warning next to it: the GPT Image 2 token calculator does not estimate GPT-Image-2.5 token consumption, so carrying old estimates into a migration will be off. The same pages set the capability boundaries — text is input-only, images are input and output, audio and video are unsupported, and streaming, function calling, structured outputs, fine-tuning and predicted outputs are all unsupported.
The system card published with the release gives comparable numbers, measured on a purpose-built adversarial prompt set:
| Model | Safe generation | Blocked | Unsafe presented |
|---|---|---|---|
| Sunburst | 77.0% | 21.9% | 1.09% |
| Flare | 79.4% | 19.2% | 1.41% |
| Images 2.0 (baseline) | 75.2% | 23.1% | 1.64% |
By category, Sunburst blocked 89.1% of sexual-content prompts, generated safely on 90.4% of political imagery, and presented unsafe output 0.00% of the time on extremism. The safety stack is layered: LLM-based policy checks refuse violating requests before generation, then a safety reasoning model moderates text and image inputs and inspects the finished image before it is shown. For provenance, C2PA metadata continues, now with Google DeepMind's SynthID invisible watermarking layered on top. The card also names the risk directly: heightened realism makes convincing deepfakes easier absent safeguards, including political, sexual or otherwise sensitive imagery.
Third-party platforms moved the same day:
- OpenRouter — model pages for both Flare and Sunburst, marked released September 9 with a 400K context
- fal —
openai/gpt-image-2.5/flareand.../sunburst, each with a text-to-image and an image-editing endpoint - Adobe Firefly — both tiers available, Flare as the default and Sunburst for demanding production work
All three carry the same token rates as OpenAI's own model pages.
Limitations: the system card notes that automated policy labels may contain errors and that these figures hold only for a fixed adversarial test set. The speed gap between tiers is published only as qualitative ratings — Very fast against Medium — with no per-image timings.

Jiufeng graphic based on the sources cited in this issue.

Image source: OpenAI Deployment Safety Hub; mirrored on Jiufeng R2.
Source: OpenAI · System card · Flare model page · Sunburst model page · The Verge · SiliconANGLE · OpenRouter · fal
02/10
DeepSeek cuts Flash prices, effective tomorrow noon (Chinese-language source)
DeepSeek's platform notice sets new Flash-series prices from 12:00 Beijing time on September 10, with off-peak output at 4 yuan per million tokens.
The new rates split by time of day, with peak-hour prices at double the off-peak figures:
| Per million tokens | Cache-hit input | Cache-miss input | Output |
|---|---|---|---|
| Off-peak | 0.02 yuan | 1 yuan | 4 yuan |
| Peak | 0.04 yuan | 2 yuan | 8 yuan |
Peak is defined as 9:00-12:00 and 14:00-18:00 Beijing time on weekdays; everything else counts as off-peak. Under that split, moving batch jobs to the night brings cache-hit input down to roughly a tenth of the current price.
The change covers the Flash series, currently deepseek-v4-flash and the experimental deepseek-v4-flash-vision-exp, which additionally accepts image input.
Evidence boundary: this is a Chinese-language report of a platform pricing notice. DeepSeek published no accompanying performance or architecture claims, and no English-language announcement is available.
Source: ITHome · DeepSeek pricing docs
03/10
Ant open-sources Ling-3.0-flash-VL: 124B total, 5.5B active
Ant's first natively multimodal Ling model is open weights: a 124B-parameter MoE activating 5.5B per token, under MIT.
The model card lays out the specs:
- Architecture — an MoE extension of Ling-3.0-flash: a ViT visual encoder with a two-layer MLP projector, VideoRoPE for spatial and temporal encoding, and a 42-layer hybrid backbone alternating KDA and Gated MLA layers at a 5:1 ratio
- Modalities and context — native image, text and video input, with a context window up to 1M tokens
- License and scores — MIT; 42 on the Artificial Analysis Intelligence Index v4.1.1, four points above the base model
Its headline mechanism is a visual feedback loop: instead of one-shot generation from an image, the task runs as a continuing observe-act-verify-correct cycle, with medical report interpretation, front-end code generation and GUI automation given as target scenarios.
Limitations: two caveats here. The numbers disagree across sources — Chinese coverage reports a 256K context while the model card says up to 1M tokens, and this site follows the model card. And deployment is not light: the card suggests four to eight high-end GPUs, with thinking mode on by default and disabled per request.
Source: Hugging Face model card · ITHome
04/10
Mercury 2.5 lands on OpenRouter at a claimed 1,107 tokens per second
Inception's diffusion language model Mercury 2.5 shipped and is now on OpenRouter, with the company claiming 1,107 tokens per second on widely available NVIDIA GPUs.
A diffusion LLM does not emit tokens strictly one after another; it produces several in parallel and refines them. Inception's figures: 1,107 tokens per second on widely available NVIDIA GPUs, a 40% increase in intelligence over Mercury 2, and a level it describes as comparable to cost-optimized models such as Gemini 3.5 Flash-Lite and Claude Haiku 4.5. Context is 260K, with tunable reasoning, parallel tool calls and schema-aligned JSON. Pricing is $0.20 per million input tokens and $0.75 per million output, discounted 80% at launch to $0.04 and $0.15. Beyond Inception's own API, it is callable through Baseten and OpenRouter.
Limitations: both the speed and the intelligence gain come from Inception's own launch materials; no third-party reproduction or independent evaluation is available yet.
Source: Inception · OpenRouter
Global AI news
05/10
OpenAI says 10,000 agents solved Navier-Stokes in 88 hours
OpenAI says an unreleased internal model coordinating roughly 10,000 agents produced a solution to the Navier-Stokes existence and smoothness problem in about 88 hours, published with a 165-page writeup and a Lean 4 formalization.
OpenAI released the materials: a 165-page analytical proof plus a Lean 4 formalization that outside researchers can download, build and inspect. The paper argues that a three-dimensional incompressible fluid can start at rest and develop unbounded velocity in finite time while its kinetic energy stays bounded and the external force applied to it stays smooth. The company describes the underlying model as significantly more capable than the newly released GPT-6 Astra, with training still under way.
Limitations: the boundaries matter here. The Clay Mathematics Institute still lists Navier-Stokes as unsolved, and OpenAI says it does not intend to claim the associated $1 million prize. The work also comes with an authorship dispute: NYU mathematician Tristan Buckmaster says a parallel OpenAI effort built on his and Anthropic mathematician Levent Alpöge's work before it was public, and that an OpenAI researcher twice pushed to drop his Anthropic-employed co-author from the paper. OpenAI has responded to the allegations.
Source: OpenAI · The Verge · TechCrunch
06/10
Cognition raises more than $2B at a $48B valuation
The AI coding company closed a Series E of more than $2 billion at a $48 billion valuation, with run-rate revenue approaching $900 million.
For comparison with the last round: in May, Cognition raised more than $1 billion at a $26 billion valuation with run-rate revenue of $492 million. Four months later the valuation has nearly doubled and the company puts run-rate revenue close to $900 million. Its flagship product Devin works from high-level requirements rather than line-by-line prompts, planning tasks, writing and testing code, debugging and pushing deployments inside a sandbox. The company also owns Windsurf, the AI coding editor it bought in July 2025. Devin is reported to be in use for chip design at NVIDIA.
Limitations: revenue here is a company-reported run rate, not confirmed annual revenue.
Source: SiliconANGLE · TechCrunch
07/10
Meta drops AI usage from engineer performance reviews
An internal memo says AI dashboards and token counters no longer factor into engineer performance reviews; quality, speed and complexity of the work do.
The memo, seen by The Information, is signed by executives Maher Saba and Santosh Janardhan. Meta had previously made AI use a review criterion, which produced what staff called "tokenmaxxing" — burning tokens in bulk to look good on internal leaderboards. The bill showed up directly: internal AI use alone is heading toward billions in costs in 2026, which is why Meta plans budgets and a central dashboard starting in 2027.
Limitations: this comes from a report on an internal memo rather than a public Meta announcement. The same report notes Meta is testing an agent tool called Hatch, with some employees reluctant to connect it to personal accounts over privacy concerns.
Source: The Decoder
08/10
Samsung and Mistral partner on semiconductor-specific models
The two announced a strategic partnership applying Mistral's models and services to Samsung's chip design, process optimization and engineering workflows, including custom on-premises models.
Per Samsung's announcement, the collaboration spans chip design, manufacturing process optimization and engineering workflows, bringing Mistral's services and flagship model into Samsung's semiconductor operations and building customized on-premises models on top. A repeatedly stated constraint is keeping proprietary semiconductor information from leaking. The announcement was made during the South Korea-France state summit in Paris.
Limitations: reports put Samsung at the head of a $3.5 billion investment in Mistral as part of the deal; that figure currently comes from media coverage rather than either company's announcement and awaits formal disclosure.
Source: Samsung Newsroom · The Korea Herald
Regional and early signals
09/10
Tmall's AI Space Station puts model subscriptions on retail shelves
Tmall launched an AI Space Station selling Chinese model vendors' token and coding subscriptions as retail-style top-ups, with Zhipu already running a Coding Plan storefront.
According to the reporting, the storefront turns model subscriptions and token packs into retail goods much like phone credit top-ups: Zhipu has opened a Coding Plan storefront, while Moonshot and MiniMax are in talks to follow. For Chinese model vendors it treats e-commerce traffic as a distribution channel alongside developer sign-ups.
Evidence boundary: neither published report gives prices, subscription tiers, a rollout timetable or any vendor statement on sales.
10/10
Wang Yunhe's team ships NeoHorse-1 in 4B and 9B sizes (Chinese-language source)
TokenRhythm, founded by former Huawei Noah's Ark Lab director Wang Yunhe, released its first Agent-Native model, post-trained on Qwen3.5 bases.
NeoHorse-1 comes in 4B and 9B versions, post-trained on Alibaba's open Qwen3.5-4B and Qwen3.5-9B. What sets it apart is the training corpus: execution traces produced by the company's own Routing Harness — the open-source OpenSquilla project — retaining capability prediction, routing choices, model responses, tool calls and environment feedback, filtered by completeness checks and quality scoring for goal completion, evidence consistency and error recovery. Training combines execution supervision with on-policy distillation. The technical report evaluates on 11 benchmarks covering agent execution, tool interaction, code and instruction following; among the 4B-class models it lists, NeoHorse 4B has the highest unweighted average and beats the Qwen3.5-9B base on five of them.
Evidence boundary: this is a Chinese-language report of a vendor technical report, with no independent evaluation. The report itself scopes the gains to tasks with clear procedures, observable feedback and verifiable results, and says larger models still lead on complex state tracking, long-horizon debugging and failure recovery.
Write a prompt in your browser and get an image — 1K, 2K or 4K output, up to 15 reference images. 50 free generations on sign-up, no credit card.
Generate free

