Overview
10 stories in this issue. The first 3 are today's priorities.
Hot model updates
- Top · Grok 4.7 ships at Grok 4.6's $2/$6 token prices
- Top · GPT-6 Astra reaches 53.3% on the Rails coding benchmark
- Top · Kyutai releases speech-native math models built on GLM-4-Voice
- Grok 4.6 arrives in Amazon Bedrock with a 500K context window
- Amazon blocks Meta's Muse agent from its marketplace
Global AI news
- AWS open-sources Strands harness, claiming 28% lower token cost
- ByteDance launches Dramagic for short-drama production
- Bristol proposes a drug-trial-style check for medical AI
Regional and early signals
- Inspur says one node runs the 2.8T-parameter Kimi K3
- New Mac mini pitched as a machine that runs agents around the clock

Jiufeng graphic based on the sources cited in this issue.
Hot model updates
01/10
Grok 4.7 ships at Grok 4.6's $2/$6 token prices
SpaceXAI prices a larger base model exactly where the last one sat, competing on the cost of finishing long jobs rather than on token rates.
SpaceXAI, the AI operation inside Elon Musk's SpaceX, launched Grok 4.7 on September 21st, 40 days after Grok 4.6. The pitch is a larger base model that stays on difficult coding and knowledge-work assignments longer without raising standard token prices.
- API pricing: $2 per million input tokens, $6 per million output tokens — unchanged from Grok 4.6
- Distribution: Cursor, Grok Build and the x.ai API from day one
- Specs: developer documentation lists a 500,000-token context window and a knowledge cutoff of May 2026
- Agents: SpaceXAI's launch post says the model natively understands the Grok Bot agent harness, aimed at conversations, research and document production outside software repositories
- Corporate backdrop: SpaceX acquired xAI on February 2nd; its June prospectus called the deal the foundation of a newly formed AI segment and said substantial capital would go to compute infrastructure
The report devotes a separate section to scores against the previous model:
| Metric | Grok 4.7 | Grok 4.6 |
|---|---|---|
| CursorBench 4.0 | 46.3% (xhigh) | 40.4% (high) |
| Cursor's own run (same effort) | 43.9% | 40.4% |
| Cost per task | $4.69 | $5.20 |
| DeepSWE v1.1 | 71% | — |
The report also says Grok 4.7 gains more over 4.6 on Terminal-Bench 4.0, electrical engineering and Harvey.
Limitations: the first row is not a like-for-like comparison — 4.7 runs at xhigh while 4.6 runs at high; only Cursor's own run holds effort constant. "Stays on long assignments" is SpaceXAI's own positioning, not a third-party evaluation result.
Source: RuntimeWire · SpaceX–xAI announcement · Grok 4.7 developer docs
02/10
GPT-6 Astra reaches 53.3% on the Rails coding benchmark
Rails pushed every model to maximum reasoning effort; only OpenAI's models turned the extra spend into a higher score.
The Rails Foundation said in a September 21st X post that it re-ran its feature-development benchmark with every eligible model set to the highest reasoning effort available. GPT-6 Astra kept first place, moving from 35% at medium effort to 53.3% at maximum and solving 32 of 60 runs. Claude Fable 5.1 solved 19 of 60 at both high and maximum effort, leaving accuracy at 31.7%. Gemini regressed. A previously scoreless OpenAI model completed 16 runs this time.
| Metric (max effort) | GPT-6 Astra | Claude Fable 5.1 |
|---|---|---|
| Accuracy | 53.3% (35% at medium) | 31.7% (same as high) |
| Runs solved / 60 | 32 (was 21) | 19 (was 19) |
| Campaign cost | $397.62 (was $150.47) | — |
| Avg cost per run | — | $19.10 (was $9.14) |
Limitations: Rails' own conclusion is that additional reasoning did not consistently produce better code — gains ranged from zero to 27 points while the aggregate bill nearly doubled. The harness itself is constrained: 90 minutes, 400 steps and $60 per run, with no internet access and no additional agent scaffolding; tasks and verification material are frozen.

Image source: GitHub; mirrored on Jiufeng R2.
Source: RuntimeWire · Rails Foundation on X · rails/ai-evals
03/10
Kyutai releases speech-native math models built on GLM-4-Voice
Spoken question in, spoken answer out — no transcription step, no separate text LLM, and up to 77.1% on GSM8K.
Paris-based Kyutai announced Voice of Reason in a five-post X thread on September 21st: two speech-native models that take spoken math problems and answer aloud without routing inference through a transcription system or a separate text language model. The accompanying paper, accepted at COLM 2026, is led by Kyutai PhD student Timothee Weisselberger with Edouard Grave and head of research Alexandre Defossez.
The release extends Z.ai's GLM-4-Voice, which alternates text and audio tokens as it generates. Kyutai changed the post-training rather than the architecture, producing one checkpoint that answers without hidden reasoning tokens and another that silently reasons in chunks while speaking.
Limitations: the models reach up to 77.1% on GSM8K, but larger cascaded systems still lead that benchmark, and the reporting notes weaker general-knowledge results as the cost of specializing this aggressively.
Source: RuntimeWire · Kyutai on X · Research paper · Reasoning-free checkpoint
04/10
Grok 4.6 arrives in Amazon Bedrock with a 500K context window
xAI's second Bedrock model is reachable from both endpoints and supports the Converse API.
AWS announced on September 21st that xAI's Grok 4.6 is available in Amazon Bedrock, positioned for long-running agents, coding and knowledge work.
- Context: 500K tokens
- Reasoning effort: four configurable levels — low, medium, high, xhigh
- Surface area: both the bedrock-mantle and bedrock-runtime endpoints, with Converse alongside Chat Completions and Responses
- Sequence: xAI's second model in Bedrock; when Grok 4.3 went GA, xAI joined as a model provider and the model was reachable only through Bedrock Mantle, the OpenAI-compatible inference engine
Limitations: AWS states that the capability and training details in its post come from xAI's own launch announcement, not from AWS evaluation. The September 21st post also records the Bedrock launch date as August 18, 2026.
Source: AWS Machine Learning Blog · Introducing Grok 4.6
05/10
Amazon blocks Meta's Muse agent from its marketplace
An agent that passed 730,000 downloads in five days and topped the App Store free chart can no longer shop on Amazon.
Amazon has blocked Meta's Muse AI agent from accessing its e-commerce marketplace and notified users late Sunday. Meta made Muse available in the U.S. earlier this month as a mobile app; it topped 730,000 downloads within five days, putting it ahead of both ChatGPT and Claude, and by last Friday it had passed ChatGPT as the most popular free app on the App Store. Meta offers a free version alongside two paid editions with higher rate limits.
Amazon gave GeekWire a written statement on the move, and SiliconANGLE's report also draws on GeekWire's reporting about undisclosed third parties operating in customer accounts.
Limitations: the download count and chart position are third-party figures relayed by the report rather than Meta disclosures, and the report does not describe how the block is enforced or how far it extends.
Source: SiliconANGLE
Global AI news
06/10
AWS open-sources Strands harness, claiming 28% lower token cost
A fully assembled, general-purpose agent harness under Apache 2.0, in Python and TypeScript, starting from one line of code.
The Strands Agents team at AWS released Strands harness, a packaged agent harness that runs locally or deploys to a cloud provider. A harness is the system around the model — the loop, tools, context handling, memory and recovery; Strands previously exposed those pieces only through its SDK, and this release ships them as working defaults.
- License and languages: Apache 2.0, Python and TypeScript
- Model backends: Amazon Bedrock, Anthropic, OpenAI, Google, Ollama or LiteLLM
- Built-in tools: shell and file operations among others
- Cost claim: running the same Claude or GPT models, the team reports 28% lower cost than other harnesses across 6 benchmarks at near-equal accuracy — the average of ALFWorld, ContextBench, GAIA, WebShop, τ²-bench and one more
- Named comparisons: the chart names Claude Code, Codex, oh-my-pi, OpenCode and DeepSeek Harness as the competing harnesses, with a per-harness cost and accuracy table on Terminal-Bench 2.1 using the same model
- Deployment: a bundled skills file helps a coding agent generate deployment config for AWS, GCP, Azure, Cloudflare and Modal
Limitations: the 28% figure is self-reported by the AWS Strands team, not independently reproduced.
Source: MarkTechPost · GAIA benchmark
07/10
ByteDance launches Dramagic for short-drama production
One pipeline from script analysis and character creation through storyboards to video preview, sold through BytePlus by request.
ByteDance launched Dramagic, an AI platform for producing short dramas and videos that covers the full pipeline — script analysis, character creation, storyboarding and video previews. Multiple users can collaborate simultaneously, with built-in consistency checks. It is sold through BytePlus, ByteDance's enterprise platform, and access must be requested.
The Decoder cites the market context: according to the China Netcasting Services Association, roughly 128,000 short dramas were published in China in Q1 2026, three times the total for all of the previous year, and 95 percent were AI-generated. Tsinghua University professor Shen Yang puts one minute of AI video at $90 to $120, about a tenth of traditional production cost. The industry directly employs 690,000 people.
Limitations: Dramagic is request-only through an enterprise channel, and the report gives no pricing or underlying model details. The report also relays accounts of performers being pressed to hand over voice and likeness to AI tools before being laid off.
Source: The Decoder · BytePlus announcement
08/10
Bristol proposes a drug-trial-style check for medical AI
Researchers argue medicine already knows how to license a black box, and medical AI should borrow the method.
Researchers at the University of Bristol propose a framework, called Learning Ensemble, that lets developers systematically test how reliable an AI system is in medical use, modeled on the standards medicine uses to bring new drugs to market. It defines three areas to check: the system's operating limits and training data, its reliability across all patient groups, and its actual fit for daily clinical use.
Their stated motivation is that medical AI systems often look good in early tests but fail in the clinic because they latch onto features in the training data that have nothing to do with the diagnosis — while medicine has already developed ways to handle drugs whose exact effect in the body isn't fully understood.
Limitations: this is a proposed checking framework from a paper; the source lists the three areas but no specific system that has been vetted with it and no pass rates.
Source: The Decoder · Learning Ensemble paper
Regional and early signals
09/10
Inspur says one node runs the 2.8T-parameter Kimi K3
128 tightly coupled domestic AI chips, with token generation latency claimed at 5.85 milliseconds.
At the 2026 AI Computing Conference on September 21st, Inspur launched the Yuannao SD200 Ultra supernode AI server and the Yuannao HC2000 compute cluster. Inspur says SD200 Ultra is built on domestic AI chips, runs the 2.8-trillion-parameter Kimi K3 on a single machine, and supports single-machine operation of frontier models at the 10-trillion-parameter scale.
- Scale: 3D Hyper Mesh architecture, 128 tightly coupled domestic AI chips, 8TB of unified-addressed VRAM and 64TB of memory
- Latency: token generation latency down to 5.85 milliseconds, which Inspur converts to 170 tokens/s for a single user
- Interconnect: native memory-semantic communication at 0.69 microseconds; symmetric memory lets a GPU directly access remote GPU memory, cutting AllReduce communication time by 3.5x
- Operators: a super-operator agent built with Kimi K3 fuses base operators in KDA, Gated MLA and MoE, reducing operator count 10x and raising inference performance more than 3x
Inspur frames this against model growth: from DeepSeek R1's 671 billion parameters to Kimi K3's 2.8 trillion, context from 128K to over 1M, and agents moving from single instances to hundreds working together, with latency accumulating across each reason–call–feedback–replan loop.
Limitations: Chinese-language source, and every figure above is Inspur's own conference claim — neither "5x the industry average" nor "10x token capacity at equal investment" comes with a published test methodology or third-party reproduction.
Source: IT Home (Chinese-language source)
10/10
New Mac mini pitched as a machine that runs agents around the clock
Apple describes the box, starting at 6,999 yuan, as a desktop that can keep agents working full time.
Chinese outlet ifanr tested a new Mac mini with an M5 Pro, 48GB of memory and 2TB of storage, priced at 25,749 yuan, against an entry configuration of an M6 chip with 16GB and 256GB starting at 6,999 yuan. Its argument is that the AI PC contest has moved from "can it run a model" to "can an agent keep working on this machine": LM Studio Bionic, on-device models, agents and exo clusters are the new Mac mini's headline features, and Apple itself describes it as a desktop that runs agents around the clock. The review unit held a 27-billion-parameter local model.
Limitations: Chinese-language source and a hands-on review. The claims that OpenAI buys tens of thousands of Mac minis and Mac Studios for training and that Anthropic rents large numbers of Mac servers on AWS are relayed by ifanr without attribution, and the 27-billion-parameter figure applies to the 48GB unit — the 6,999-yuan 16GB version is framed in the article as a place for agents to keep running rather than a local large-model host.
Source: ifanr (Chinese-language source)
Write a prompt in your browser and get an image — 1K, 2K or 4K output, up to 15 reference images. 50 free generations on sign-up, no credit card.
Generate free

