Overview
10 stories in this issue. The first 3 are today's priorities.
Hot model updates
- Top · GPT-6 Sol and Luna arrive at half the previous token price
- Top · Claude Opus 5.5 lands on Amazon Bedrock at a claimed 40% lower cost
- Top · Ten Claude agents produce a shortest-path algorithm with a Lean proof
- Intel adds Qwen-Image-2.1 to OpenVINO without publishing benchmarks
Global AI news
- A neuron-derived video model promises 5x speed and 80% lower cost
- AWS measures skill-equipped agents: a fluent answer proves nothing
- European neocloud Verda raises $189M, over $450M in total
Regional and early signals
- A 30B MoE runs on-device on Snapdragon at 330 tokens/s prefill
- Alibaba updates its multimodal lineup, next video model due in November
- Multi-model routing keeps 99.96% of quality at 88.9% lower cost

Jiufeng graphic based on the sources cited in this issue.
Hot model updates
01/10
GPT-6 Sol and Luna arrive at half the previous token price
OpenAI pushes its newest family into coding and agent workloads and halves the price, but independent analyses see little performance gain.
OpenAI launched GPT-6 Sol and GPT-6 Luna on September 22nd. Both are rolling out in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu users; Free and Go subscribers can reach Luna through the ChatGPT desktop app. Sol is positioned for complex coding and agentic workflows, while Luna targets focused tasks that must run cheaply at high volume. OpenAI attributes the cuts to caching and inference improvements and says it is passing the savings to users.
| Model | Input ($ / 1M tokens) | Output ($ / 1M tokens) |
|---|---|---|
| GPT-6 Sol | 2 | 10 |
| GPT-5.6 Sol | 4 | 20 |
| GPT-6 Luna | 0.10 | 0.50 |
| GPT-5.6 Luna | 0.20 | 1.20 |
Limitations: independent analyses cited by The Decoder find barely any performance movement over the previous generation, with price as the main change. Terra, previously the cheapest model in the lineup, is gone, and neither new model is available in the regular Chat interface at launch.

Image source: OpenAI Developers; mirrored on Jiufeng R2.
Source: RuntimeWire · The Decoder · OpenAI announcement · GPT-6 Sol docs
02/10
Claude Opus 5.5 lands on Amazon Bedrock at a claimed 40% lower cost
The first model of the Claude 5.5 family goes on sale the same day, pitched as doing the same work for fewer tokens.
Claude Opus 5.5 is available on Amazon Bedrock and Claude Platform on AWS. Anthropic says it matches Claude Fable 5.1 on most tasks while costing about 40 percent less to run than Opus 5: lower per-token prices and cheaper cache reads stack on top of the model doing more with fewer tokens. Adaptive thinking is always on, with the model deciding how much reasoning a task needs, so developers use an effort control instead of manual thinking budgets. Claude Sonnet 5.5 and Haiku 5.5 are expected in the coming weeks.
| Model | Terminal-Bench 4.0 (%) | FrontierCode v1.1 Main (%) |
|---|---|---|
| Claude Opus 5.5 | 66.4 | 54.4 |
| Claude Fable 5.1 | 55.8 | 50.3 |
| Claude Opus 5 | 52.3 | 48.0 |
| GPT-6 Astra | 57.9 | 53.3 |
| GPT-5.6 Sol | 37.3 | 47.5 |
Limitations: the table comes from Anthropic's own benchmarks, and the 40 percent figure is the vendor's number, which the AWS post also attributes to Anthropic. The Decoder notes that requests flagged by classifiers for biology or frontier LLM development are routed to Opus 5.
Source: AWS Machine Learning Blog · The Decoder
03/10
Ten Claude agents produce a shortest-path algorithm with a Lean proof
An evaluation company published a checkable artifact instead of another agent score.
Vals AI researcher Geby Jaff said on September 20th that ten Claude Opus 5.5 agents devised a shortest-path algorithm called C-HD and produced a Lean proof of its correctness and runtime bound in about 15 hours. The yardstick is Dijkstra's algorithm with a bound of O(m + n log n) using a suitable priority queue, where n is the number of vertices and m the number of edges. Vals was founded by Rayan Krishnan and Langston Nashold on the premise that model buyers need evaluations grounded in actual work rather than scores picked by model makers; the full Lean proof package has been released.
Limitations: the claimed improvement is mathematical, not a measured speedup. The proved bound holds only within a specified range whose formula involves ⌊log2 n⌋^(3/4); outside that range the formal program falls back to Bellman-Ford rather than claiming C-HD's improved bound everywhere. Practical speed and novelty still need separate scrutiny.
Source: RuntimeWire · Lean proof package · Reference paper
04/10
Intel adds Qwen-Image-2.1 to OpenVINO without publishing benchmarks
The "day-zero support" headline resolves, in the documentation, to one Optimum Intel route plus a recorded failure.
Intel added an OpenVINO deployment path for Qwen-Image-2.1, the Diffusers image-generation and editing model, and Qwen's official account amplified the announcement.
- Route: Optimum Intel plus Diffusers, with model export handled by Optimum's OpenVINO exporter
- Compatibility basis: the Optimum Intel compatibility documentation lists its supported Diffusers architectures, the clearest listing behind Intel's claim
- Vendor framing: Intel Devs described the integration on X as day-zero OpenVINO support
Limitations: a separate OpenVINO GenAI test rejected the converted model with an unsupported QwenImage21Pipeline error, the accompanying notebook is still experimental, and neither launch material provides OpenVINO-specific performance results. RuntimeWire reads the evidence as a narrower integration than "OpenVINO support" suggests, leaving validation work before production use.
Source: RuntimeWire · Optimum Intel docs · Qwen-Image-2.1 repository · Intel Devs
Global AI news
05/10
A neuron-derived video model promises 5x speed and 80% lower cost
A San Francisco startup and AWS are selling a text-to-video model whose only shipped artifact so far is a signup page.
The Biological Computing Co. (TBC) announced a partnership with Amazon Web Services to launch what the companies call the first "neuron-derived" AI video model. According to the press release, the text-to-video model builds on an open-source video model and, thanks to a thin software layer derived from measurements of real nerve cells, is supposed to generate videos five times faster and at 80 percent lower inference cost than the base model, with better quality. TBC plans to run it on AWS Trainium chips, serve it through Amazon SageMaker AI and sell it in the AWS Marketplace.
Limitations: TBC does not say which open-source model it starts from. The Decoder notes all of this is still on paper — the only thing available now is a signup for early access.
Source: The Decoder
06/10
AWS measures skill-equipped agents: a fluent answer proves nothing
Skills are portable files under an open standard, but only evaluation shows whether the agent picked the right one and followed it.
AWS documents how to evaluate skill-equipped agents with Strands Evals and Amazon Bedrock AgentCore. A skill is a reusable set of instructions, usually stored in a SKILL.md file, that teaches an agent a domain task such as redacting a contract, reconciling an invoice or following a team's pull-request conventions. Because skills follow the open Agent Skills standard, they are portable across compatible harnesses, and the agent loads only the skill it needs at runtime instead of carrying every procedure in its core instructions.
A skill packages these parts:
- Instructions: domain procedures and constraints injected into context
- Tool bindings: the APIs, MCP servers or local commands it depends on
- Knowledge: reference material the task needs
- Workflow: the order in which the steps proceed
- Guardrails: the boundaries the agent must not cross
Limitations: this is AWS's own blog, offering method and examples rather than third-party comparisons, and it publishes no baseline pass rate for correct skill selection.
Source: AWS Machine Learning Blog
07/10
European neocloud Verda raises $189M, over $450M in total
The Helsinki company writes its own compilers and serving software and sells managed enterprise inference.
Verda Cloud Oy, formerly DataCrunch, announced early-stage funding to expand its AI cloud infrastructure, from compute capacity to platform services.
- This round: $189 million Series B led by Emergence Capital
- Participating: MUFG Innovation Partners, Supermicro, Varma, Lifeline Ventures, 6 Degrees Capital, byFounders and Tesi
- Angels: Ola Tørudbakken, director of AI systems at Meta, and Core Automation co-founder Mark Saroufim
- Total raised: more than $450 million in equity and debt
- The business: large-scale managed AI inference for enterprises, with in-house compilers, serving software and a performance and reliability engineering team
Limitations: the report discloses no valuation, no deployed compute figure and no customer count, and gives no pricing or performance metrics for the inference service.
Source: SiliconANGLE
Regional and early signals
08/10
A 30B MoE runs on-device on Snapdragon at 330 tokens/s prefill
The agent assistant demoed at the Snapdragon Summit never leaves the phone, with runtime memory cut by more than half.
At the 2026 Snapdragon Summit on September 22nd (Hawaii time), Qualcomm said it worked with StepFun, Wuliang Huo and Longsys to adapt and optimize StepFun's StepEdge-Omni 30B-MoE for on-device inference on the sixth-generation Snapdragon 8 Elite platform, and demonstrated an agent assistant on stage.
- Model: 30 billion parameters, mixture-of-experts, activating only the experts a task needs to cut compute and memory bandwidth
- Throughput: over 330 tokens per second prefill, over 28 tokens per second decode
- Memory: runtime memory requirement reduced by more than 50 percent versus conventional industry approaches, via inference engine, heterogeneous scheduling and storage optimizations
- Demo tasks: email comprehension, trip planning, calendar sync, flight and hotel recommendations, itinerary sharing and email drafting, all handled locally without cloud calls
Limitations: Chinese-language source only. Qualcomm published no test handset, quantization scheme, power draw or battery data, and no comparison against other on-device stacks. IT Home quotes industry observers saying real-world scale still depends on device cost, power consumption, model reliability and user acceptance.
Source: IT Home (Chinese-language source)
09/10
Alibaba updates its multimodal lineup, next video model due in November
Image, music and long-video engines all get a refresh, with no parameters or benchmarks in the public material.
Alibaba updated its multimodal generation family with the Qwen-Image-3.1 image model, the Happy Shrimp 1.1 music model and Agentic PE, a creation engine aimed at long video and complex productions; a next-generation video model is due in November. InfoQ traces the line through Wan 2.6 in late 2025 (shot scheduling, reference-to-video, multi-character consistency) and Happy Horse 1.0 in April this year, which used a single-stream Transformer to natively generate audio and video together. Zheng Bo, Alibaba ATH's VP of technology and chief scientist of Taotian Group, said on stage that a natively omnimodal unified model will appear within three years.
Limitations: Chinese-language source only. The report gives no parameter counts, license or benchmark scores for Qwen-Image-3.1. The claims that director Lu Chuan finished in five days what used to take months, and that the model's context understanding is "postdoc level," are statements made at the event by the creator and the vendor, with no verifiable comparison behind them.
Source: InfoQ China (Chinese-language source)
10/10
Multi-model routing keeps 99.96% of quality at 88.9% lower cost
Sending each step to a suitable model is the company's cost claim, qualified as holding under a specific test configuration.
Han Kai, co-founder and CTO of Jiyuan Lüdong, told an enterprise agent summit at Alibaba's Yunqi Conference that heterogeneous models and diversified compute supply will be a long-term structure, and that one complex agent task often involves ten to several dozen model calls. The company explores model routing through its open-source project OpenSquilla: in its published end-to-end agent evaluation, under a specific test configuration it retained 99.96 percent of the task quality of a fixed flagship-model baseline while cutting cost by 88.9 percent. In a separate DRACO deep-research evaluation, a multi-model configuration scored above the strongest single-model baseline in that experiment at about one third of the cost. As validation of its feedback loop, NeoHorse-1, post-trained on the Qwen3.5 base, raised its macro-average across ten benchmarks from 58.94 to 64.87 (4B) and from 65.60 to 69.04 (9B).
Limitations: Chinese-language source only. Every figure comes from the company's own report and conference talk, the contents of the "specific test configuration" are not public, and no independent reproduction exists. The 140-trillion daily token figure for March 2026 cited in the talk is a national statistic, not the company's own volume.
Source: QbitAI (Chinese-language source)
Write a prompt in your browser and get an image — 1K, 2K or 4K output, up to 15 reference images. 50 free generations on sign-up, no credit card.
Generate free

