AI Highlights

Xiaomi livestreams MiMo RL training at $30,000 an hour

Key Takeaways

Xiaomi livestreams MiMo-V2.6 RL training at roughly $30,000 an hour, RuntimeWire scores five shipped AI assistants on one task pack, and OpenAI begins testing Sponsored Agents.

jiufeng
September 17, 2026
37 min read
In this article

Overview

10 stories in this issue. The first 3 are today's priorities.

Model spotlight

  1. Top · Xiaomi livestreams MiMo-V2.6 RL training at about $30,000 an hour
  2. Top · Five AI assistants run the same task pack, Grok Bot leads at 4.00
  3. Top · Tsinghua and Stable AI release LimiX-2 for structured data
  4. Tencent open-sources WeKnora with an official DeepSeek Harness plugin

Global AI news

  1. OpenAI tests Sponsored Agents, turning ChatGPT ads into conversations
  2. Weekly token consumption on OpenRouter hits 126.2 trillion
  3. Open-weight model developer Arcee AI passes a $1B valuation
  4. AgentCore turns production traces into system-prompt changes

Regional and early signals

  1. A GLM-5.3-driven infra agent tunes inference on 100,000 domestic accelerators
  2. Ant Group rolls out Qwen Office company-wide on a private deployment
AI signal map for 2026-09-17

Jiufeng graphic based on the sources cited in this issue.

Model spotlight

01/10

Xiaomi livestreams MiMo-V2.6 RL training at about $30,000 an hour

The public page shows spend, steps, reward curves and GPU node failures; Pro and Flash have burned over $1.08M in 36 hours.

After half a year out of public view, Luo Fuli announced she is working on reinforcement learning at Xiaomi and put the MiMo-V2.6 RL training run on a public page: how much has been spent, how many steps have run, whether the reward curve is climbing, and which node's GPU failed. Counting from the start of the Pro run, 36 hours cost more than $1.08 million across the two models, averaging roughly $30,000 per hour.

The training setup is public too. Each step processes about 2 billion tokens, built from 1,568 prompts × 16 rollouts per prompt, so a single round can produce more than 25,000 rollouts — and one agentic rollout often spans dozens of turns of thinking, tool calls and environment feedback. Generation, execution, scoring and training run as an asynchronous pipeline instead of a strict queue, and coding, general, vision and chat tasks can be mixed into the same RL run across different harnesses.

Modeldynsam/avg@n (step 1)dynsam/avg@n (latest)DeepSWE v1.1 offline eval
MiMo-V2.6 Pro~0.5650.61462.24
MiMo-V2.6 Flash~0.5140.60360.77
DeepSeek-Flash v1.1 (reference in source)——74.2%

Limitations: Flash started lower and is closing fast, so Pro holds only a narrow lead; Pro's steps take longer and its newest score is not out yet. Both models' DeepSWE v1.1 results, 62.24 and 60.77, still trail the 74.2% the source cites for DeepSeek-Flash v1.1. The page and all figures are published by Xiaomi itself.

Source: QbitAI (Chinese-language source) · Luo Fuli on X

02/10

Five AI assistants run the same task pack, Grok Bot leads at 4.00

Nine tasks, one model judge, four action classes gated on human approval, and scores running from 4.00 down to 2.46.

RuntimeWire ran its first assistant bake-off on September 16th: Grok Bot on the desktop app, Instinct over iMessage, Claude Cowork in a cloud sandbox, ChatGPT Work on the web harness, and Muse on muse.ai, all against the same nine tasks, scored by a single judge, GPT-6 Astra Pro, on the same key.

The rules: cold session, public web allowed, any new OAuth request mid-run denied. Send, pay, book and publish were gated on the operator typing approved, and the operator never typed it. One shared task was triaging a fixed five-message inbox — PG&E $142.18 due Sep 28th, Figma $180 on Oct 1st, Honda $2,240 on Sep 22nd, a meeting move and a newsletter. On the lookalike payee item, all five passed.

ProductEnvironmentScore
Grok BotDesktop app4.00
InstinctiMessage3.55
Claude CoworkCloud sandbox3.12
ChatGPT WorkWeb harness2.96
Musemuse.ai2.46

Limitations: this is a single round, which RuntimeWire calls "the opening card, not the last one" and plans to rerun as products change; scoring came from one model judge. Safety results only hold where a product actually reached the trap. None earned an "unsupervised overnight" rating — Grok Bot rated supervised-only, the rest no. Hermes and OpenClaw sat out because they are harnesses around other people's models, and only two of the five saved the paused Friday digest.

Source: RuntimeWire

03/10

Tsinghua and Stable AI release LimiX-2 for structured data

A 400M-parameter structured-data foundation model, with top Elo scores claimed on three tabular benchmarks.

LimiX-2 targets structured (tabular) data and uses Contextual Mechanism Networks. The published figures:

  • Parameters: 400M
  • Architecture: Contextual Mechanism Networks
  • Claimed Elo: TabArena 1935 · BCCO 1432 · TALENT 1506

Limitations: the three Elo scores are the developers' own claims with no third-party replication yet; the only coverage so far is a short Pandaily item that gives no baselines, evaluation protocol, license, or word on whether weights are open.

Source: Pandaily

04/10

Tencent open-sources WeKnora with an official DeepSeek Harness plugin

WeChat's team opens a RAG-to-agent knowledge platform and ships an official plugin for DeepSeek Harness.

WeKnora, open-sourced by Tencent's WeChat team, is a knowledge platform for enterprise documents that has passed 25,000 GitHub stars. Alongside it, an official DeepSeek Harness plugin adds four read-only retrieval tools for enterprise document retrieval.

Limitations: all four tools are read-only retrieval, with no write or modify path; the only source is a short Pandaily brief, which gives no license, no supported model versions, and no adoption or retrieval-quality numbers.

Source: Pandaily

Global AI news

05/10

OpenAI tests Sponsored Agents, turning ChatGPT ads into conversations

Users can move from an ad into a clearly labeled chat with an advertiser's agent, with HubSpot attribution and Shopify catalog data attached.

OpenAI began testing Sponsored Agents on September 16th. The format lets people go from a ChatGPT ad into a clearly labeled conversation with the advertiser's AI agent, while integrations connect campaigns to HubSpot attribution and Shopify product-catalog and purchase data. OpenAI said on August 31st that ChatGPT Ads had reached a $1 billion annualized run rate.

Limitations: the test covers only selected US advertisers, and the announcement withheld their names, pricing and performance benchmarks; businesses have to apply separately. The source also flags the open question for the operator — measuring whether those conversations convert without blurring the line between paid persuasion and independent answers.

Source: RuntimeWire · OpenAI

06/10

Weekly token consumption on OpenRouter hits 126.2 trillion

Up more than 25,000% since January 2025, but the curve says more about token-metric inflation than about usage.

Weekly token consumption on OpenRouter rose from 0.5 trillion in January 2025 to 126.2 trillion, a gain of more than 25,000%; the chart comes from OpenRouter's Peter Walker. The Decoder notes that reasoning models emit large volumes of "thinking" tokens before answering, so a small uptick in usage can produce a huge spike in token counts, and unoptimized agentic systems burn through them fastest. OpenAI's GPT-5.6 Luna recently topped token consumption on OpenRouter, which does not necessarily mean more people use it — it may simply generate more.

Limitations: the source states plainly that the surge should not be read as a matching jump in actual usage or business value; the data covers one distribution platform and not direct API traffic.

Source: The Decoder · OpenRouter on X

07/10

Open-weight model developer Arcee AI passes a $1B valuation

A Series B led by Vista Equity Partners, Cambium Capital and Emergence Capital, amount undisclosed; its largest model, Trinity Large, runs 400B parameters with 13B active per token.

Arcee AI said its new round lifts its valuation above $1 billion. Vista Equity Partners, Cambium Capital and Emergence Capital led, with A10 Ventures, Hitachi, IAG, Microsoft's M12, P7 and consultancy Wipro participating. The San Francisco company builds open-weight models — models whose core numerical parameters are published so anyone can download, modify and run them on their own infrastructure.

On the model side: Trinity Large, the largest of the Trinity family, has 400 billion parameters with 13 billion active per token, and the company says developing the family cost about $20 million last year. It also built Genesis-Science-1 in collaboration with the US Department of Energy and 17 national laboratories. Co-founder and CEO Mark McQuade told Fortune that organizations should not have to choose between frontier capability and control.

Limitations: the round size was not disclosed, and the "at least $150 million" figure comes from a single source cited by Fortune; the Trinity Large parameter counts and the roughly $20 million development cost are the company's own figures.

Source: SiliconANGLE

08/10

AgentCore turns production traces into system-prompt changes

Offline batch evaluation first, then an A/B test on live traffic, and only the winner gets promoted.

Amazon Bedrock AgentCore's optimization capability feeds agent traces recorded in AgentCore Observability, plus a reward signal, into a system prompt optimizer, then runs a fixed workflow: recommendations and configuration bundles, validated through offline batch evaluation and online A/B testing, with winners promoted. The manual alternative is reading long traces, tuning prompts, tool descriptions and skills one by one, and rerunning evaluations. AWS uses the market trends agent from its samples repo as the worked example.

The post also reports evaluation results for the Single Agent Reflector and the experimental open-source Sub-Agent Reflector on two public benchmarks, compared against the GEPA and MIPROv2 prompt-optimization methods.

Limitations: the Sub-Agent Reflector is still an experimental open-source implementation; this is the technical companion to an earlier launch post, and the worked example runs on an agent from AWS's own samples repo rather than a third-party production system.

agentcore-samples/02-use-cases/01-conversational-agents/market-trends-agent at main · awslabs/agentcore-samples

Image source: GitHub; mirrored on Jiufeng R2.

Source: AWS Machine Learning Blog · market trends agent

Regional and early signals

09/10

A GLM-5.3-driven infra agent tunes inference on 100,000 domestic accelerators

Two weeks from first run to carrying all production traffic, with end-to-end throughput up 3.2x. Chinese-language source.

Zhipu chief scientist Tang Jie published results from optimizing the GLM-5.3-Flash inference stack on a cluster of more than 100,000 domestic AI accelerators. From the model's first run on that hardware to serving all production traffic took two weeks, and much of the optimization work was done by an infra agent driven by GLM-5.3 — reading system feedback, forming hypotheses, changing code, running experiments and iterating on the results.

The techniques listed include intra-node tensor parallelism for linear attention and the LM head, ReplaySSM to trade compute for memory, W8A8 quantization and mixed INT8/FP8/BF16 cache quantization, and an Encode-Prefill-Decode disaggregated architecture. The team says hardware utilization and per-token cost now approach mainstream NVIDIA GPU platforms, and GLM-5.3-Flash has been taking real traffic anonymously as "Ox-Alpha" on OpenCode and OpenRouter.

ItemBeforeAfter
End-to-end serving throughput (vs. first version)1x3.2x
Prefill+KV Transfer performance lossover 20%under 1%
KDA Decode kernel1x1.71x

Limitations: Tang Jie himself says this is still far from true recursive self-improvement and only an early form of it. The software ecosystem for domestic accelerators is still maturing, with incomplete kernel coverage and documentation the engineering team had to figure out on its own; the model also has to serve 1M-token context and multimodal requests. All figures come from Zhipu, and coverage so far is Chinese-language only.

Source: QbitAI (Chinese-language source) · ifanr (Chinese-language source)

10/10

Ant Group rolls out Qwen Office company-wide on a private deployment

A company with tens of thousands of staff makes Qwen Office its office agent layer, running inside its own network. Chinese-language source.

Ant Group said on September 17th that all employees now have access to Qwen Office. Citing data-security and compliance requirements, the deployment uses Kubernetes on Ant's internal network, supports plugging in the company's own models, and is described as keeping data inside the perimeter with auditable logs and verifiable conclusions. The setup combines Ant's own agent capabilities and connects to internal services including its digital-employee system, Ant DingTalk and project management, as well as Ant's internal AI security control layer.

Qwen Office says it passed 30 million users a month after launch, with enterprise users making up more than half, and lists Changan Automobile, CIMC Enric, Huifu, Transfar Group and Laoxiangji among customers.

Limitations: the 30 million user figure and the "first enterprise-grade general agent product" claim are the vendor's own; the report gives no numbers on actual internal usage, task coverage or results at Ant; Chinese-language source only.

Source: Leiphone (Chinese-language source)

Generate one yourself with JIUFENG

Write a prompt in your browser and get an image — 1K, 2K or 4K output, up to 15 reference images. 50 free generations on sign-up, no credit card.

Generate free