AI Highlights

MiMo-V2.6-Pro tops open model rankings at 46 points

Key Takeaways
  • •Xiaomi's MiMo-V2.6-Pro tops open models at 46 points amid an Anthropic data accusation
  • •OpenAI claims 100+ math problems solved
  • •DeepSeek invited to the UN.
jiufeng
September 22, 2026
38 min read
In this article

Overview

10 stories in this issue. The first 3 are today's priorities.

Hot model watch

  1. Top · MiMo-V2.6-Pro tops open model rankings at 46 points
  2. Top · OpenAI says an internal model cleared 100+ open math problems in a month
  3. Top · DeepSeek invited to speak at a UN Security Council AI meeting

Global AI news

  1. SoL-Pi cuts coding-agent token traffic by nearly half
  2. NVIDIA PAIR spreads inference requests across machines on your LAN
  3. Meta open-sources Astryx: 150+ components with an MCP toolchain for agents
  4. User says Instinct surfaced a stranger's insurance document

Regional and early signals

  1. Alibaba says Qwen4 is in training, with later versions aimed at 5–10T parameters
  2. Ancient Greek model Apollo due Wednesday, co-developed with Mistral
  3. Banks warn that letting AI shop for you widens fraud and privacy risk
AI signal map for 2026-09-22

Jiufeng graphic based on the sources cited in this issue.

Hot model watch

01/10

MiMo-V2.6-Pro tops open model rankings at 46 points

The larger of Xiaomi's two new models scores 46 on the Artificial Analysis Intelligence Index, the strongest openly available model right now — and Anthropic says Claude data helped get it there.

Xiaomi released its MiMo-V2.6 lineup. According to Xiaomi, the larger of the two new models, MiMo-V2.6-Pro, scores 46 points on the Intelligence Index from analysis firm Artificial Analysis, placing it at the top of openly available models while undercutting rivals sharply on price:

  • Intelligence Index: 46 points (Artificial Analysis)
  • API pricing: $0.435 per million input tokens, $0.87 per million output tokens
  • Cost per task: roughly $0.13 to run one test task, by Artificial Analysis's methodology
  • What's public: a technical report, the training framework, a smaller model for continued training, and about 7,000 training tasks

The Decoder reports the gain comes from a massively expanded round of reinforcement learning.

Limitations: In the same report, Anthropic accuses Xiaomi of improperly siphoning data through Claude to train its own model lineup — an accusation with no third-party adjudication so far.

Source: The Decoder · Hugging Face

02/10

OpenAI says an internal model cleared 100+ open math problems in a month

OpenAI announced that a new internal model solved more than 100 long-standing math problems, while backing an independent advisory group at the Institute for Advanced Study.

OpenAI says a new internal model, after just a month of training, solved more than 100 long-standing problems across most areas of mathematics — on top of its earlier claim about the Navier-Stokes Millennium Problem. Alongside that, OpenAI is backing an independent advisory group at the Institute for Advanced Study. The group includes Fields Medalist Timothy Gowers and will advise on how OpenAI shares results with researchers and the public.

Limitations: Mathematicians warn that AI-generated solutions could undermine conceptual understanding. And OpenAI has excluded the pace of its own research from the group's advisory role — the group gets a say in how results are shared, not how fast the work moves.

Source: The Decoder · OpenAI

03/10

DeepSeek invited to speak at a UN Security Council AI meeting

Reuters says DeepSeek was invited to make a statement at the September 23rd meeting, but attendance, delegate and remarks all remain unconfirmed.

Per RuntimeWire citing Reuters, Hangzhou-based DeepSeek was invited to make a statement at a United Nations Security Council meeting on artificial intelligence and international security scheduled for September 23rd; other Chinese AI developers, including Moonshot AI, also received invitations. DeepSeek distributes model weights publicly — the DeepSeek-R1 repository on GitHub, for example — while selling managed API access. The Associated Press reported that founder Liang Wenfeng started DeepSeek in 2023 after building High-Flyer, an AI-focused quantitative investment manager.

Limitations: An invitation is not confirmed attendance. As of September 22nd, Reuters had not identified DeepSeek's delegate or confirmed that it would attend or deliver a statement, and plans remained fluid a day before the meeting. DeepSeek's position on safety or open-weight governance is also unstated.

GitHub - deepseek-ai/DeepSeek-R1 at runtimewire

Image source: GitHub; mirrored on Jiufeng R2.

Source: RuntimeWire · GitHub · AP

Global AI news

04/10

SoL-Pi cuts coding-agent token traffic by nearly half

NVIDIA, NTU and MIT added four efficiency mechanisms to the open-source Pi coding agent: token traffic on EdgeBench drops 44.7%–49.0% with scores roughly unchanged.

The four mechanisms were not hand-tuned — an AI found them by running auto-research loops at the harness layer, across 535 environments according to MarkTechPost. The numbers:

  • Token traffic: on the 51-task EdgeBench evaluation, recorded token traffic falls 44.7% to 49.0% versus stock Pi
  • API cost: down roughly 33%
  • Task scores: stay close to Pi on both GPT-5.6 Sol and Opus 5
  • Availability: shipped on GitHub under NVlabs as an MIT-licensed extension that runs on an unmodified Pi release; tested with Pi 0.85.1 and Node.js 22.19 or newer

Limitations: What gets saved is tokens, not points — scores merely stay close to Pi rather than beating it. The result holds only on EdgeBench's 51 tasks, and it depends on matching Pi and Node versions.

Source: MarkTechPost · arXiv

05/10

NVIDIA PAIR spreads inference requests across machines on your LAN

The PAIR beta pools inference capacity from several local machines, and NVIDIA is explicit that it neither merges GPUs nor pools VRAM into one bigger accelerator.

NVIDIA Personal AI Router (PAIR) entered beta for local multi-agent workloads: when a main agent farms subtasks out to several sub-agents, a single GPU is easily overwhelmed, so PAIR receives requests through its proxy, identifies the engine and model requirements, and picks an eligible node to run the request end to end — the agent still sees just one connection. It plugs into local inference services such as Ollama and LM Studio without changes to the underlying architecture or agent framework, and runs on Windows 11, Linux and macOS across x64 and arm64, including pairing nodes on different operating systems. In NVIDIA's demo, Hermes Desktop split a task into five separate analyses across an RTX Spark, a DGX Spark and an RTX 5090, cutting completion time to roughly half of running the same workload on a single RTX Spark laptop.

Limitations: NVIDIA says the demo is not a performance guarantee — results depend on workload parallelism, model, engine settings, hardware, network and node availability — and states plainly that PAIR does not merge GPUs or consolidate VRAM. The announcement still drew confusion on social media, with some users reading PAIR as a way to share compute with third parties or to combine weak machines to run big models.

Source: InfoQ · NVIDIA Developer Blog · InfoQ China

06/10

Meta open-sources Astryx: 150+ components with an MCP toolchain for agents

A React design system built internally over eight years ships as an open-source beta, with tooling aimed at engineers and AI agents alike.

Astryx is built on React 19 and Meta's StyleX, and offers more than 150 accessible UI components, customizable CSS design tokens, and a dedicated CLI and MCP toolchain. Component behavior and accessibility compliance are decoupled from visual presentation, which lives in a central design token layer covering color palettes, type scale and border radii. The core package @astryxdesign/core publishes typed React components alongside precompiled CSS, and components support className natively, so teams can mix in Tailwind, CSS Modules or plain stylesheets — using the precompiled CSS and className requires no extra compiler setup.

Limitations: This is a beta. Customization has a boundary: design tokens and standard React composition cover appearance and behavior, but if a change requires reaching internals that are not exposed — private state, DOM structure, inaccessible event listeners — the only route is the CLI's swizzle command, which exports a component's full source into your own repository.

Source: GitHub · InfoQ China

07/10

User says Instinct surfaced a stranger's insurance document

Screenshots show the assistant producing details from a claim document and confidently explaining a crossed chat — for a document that may never have existed.

On September 21st, user Pritak (@prit4k) posted screenshots on X of an exchange with Noah Shinn's AI assistant Instinct. Pritak asked whose name appeared on what the assistant described as a Gerber claim document; Instinct replied with Pritak's name, a middle name and an address. Pritak said the middle name was wrong, that he had not lived at the address during the relevant period, and that he had never sent Instinct a photo. Instinct then claimed it had checked its records and found that Pritak's message contained only text.

Limitations: Two materially different failure modes fit the same screenshots — Instinct may have delivered another person's document, or it may have hallucinated the document along with a confident account of the supposed mix-up. The screenshots do not establish which, and the assistant's own admission proves nothing. TechCrunch reported on August 24th that Instinct was already facing scrutiny over its data controls.

Source: RuntimeWire · X · TechCrunch

Regional and early signals

08/10

Alibaba says Qwen4 is in training, with later versions aimed at 5–10T parameters

At its Apsara Conference, Alibaba made recursive self-improvement the headline and laid out a parameter roadmap beyond Qwen4.

Per QbitAI, Alibaba said on September 22nd that recursive self-improvement (RSI) has entered its training, inference and chip-model co-design work:

  • Self-training: Qwen3.8-Max built its own training pipeline, constructed data, designed experiments and located defects, iterating for over a month with zero human involvement across 33 effective rounds; its Artificial Analysis score rose from 40 to 45
  • Inference: on an unseen new T-Head GPU, it autonomously adapted the inference framework for the next-architecture Qwen3.8-Flash, raising single-instance throughput 96%
  • Chip co-design: from one real bus-module specification, it ran the full front-end, verification and back-end loop for over 60 hours with more than 10,000 EDA tool calls, reaching a physical implementation with 42% less area
  • Roadmap: Qwen4 is training on a new architecture, while Qwen4.5 and Qwen5 are planned at 5 to 10 trillion total parameters

Alibaba also said it has open-sourced more than 460 Qwen models, and that Qwen3.8-27B is the most-liked open model in Hugging Face's history, ahead of DeepSeek-R1 and Meta-Llama3.1.

Limitations: All of the above is Alibaba's own framing at its own conference — the 33 rounds, the 96% throughput gain and the 42% area reduction have no third-party verification. Qwen4 is only described as "in training," and 5–10T is a plan, not a shipped model. Chinese-language source.

Source: QbitAI (Chinese-language source)

09/10

Ancient Greek model Apollo due Wednesday, co-developed with Mistral

The Austrian Academy of Sciences worked with Mistral and Sail Reply on a model for ancient Greek, meant to fill the gaps in damaged papyrus fragments.

Per IT Home citing WIRED, the Austrian Academy of Sciences will release Apollo on Wednesday local time. It was trained on manuscripts, papyri and stone inscriptions totaling roughly 600 million words of historical ancient Greek, and researchers can use it free through a chatbot interface. For damaged texts, Apollo proposes the statistically most likely wording to fill blanks. Anna Dolganov, a papyrologist at the Academy, says the model switches register by text type: Homeric text gets Homeric-era Greek, a Doric-dialect inscription gets Doric.

Limitations: Stephen Colvin, professor at University College London, says: "Ordinary readers may think we'll suddenly recover several lost plays of Sophocles, but that isn't going to happen." The report notes that most unrestored papyri record private letters, marriage contracts and administrative paperwork. Completions are statistical proposals that still need scholarly judgment. Chinese-language source.

Source: IT Home (Chinese-language source)

10/10

Banks warn that letting AI shop for you widens fraud and privacy risk

NatWest, Bank of America and others say agentic shopping is outrunning industry standards and consumer protection.

Per IT Home citing a September 22nd Reuters report, a study involving NatWest, Bank of America, ING, New Zealand's ASB and Capital One warns that letting AI agents shop on consumers' behalf raises the risk of scams, fraud and data privacy exposure: agents may ask consumers for card details and enter them directly on shopping sites, or steer users toward payment methods with weaker consumer protection. Demand is real — UK retailer John Lewis said in September that AI-agent-originated searches rose from 0.3% a year ago to 2.5%, and still accelerating. The banks plan to discuss proposals with policymakers: clear disclosure of whether an AI agent took part in a transaction, more transparency into how decisions are made, customer data protection mechanisms, plus free choice of AI commerce services for consumers and merchants and interoperability between systems.

Limitations: These are bank proposals, with no regulator's response and no timeline; the report puts no figure on losses already caused by agentic shopping. Chinese-language source.

Source: IT Home (Chinese-language source)

Generate one yourself with JIUFENG

Write a prompt in your browser and get an image — 1K, 2K or 4K output, up to 15 reference images. 50 free generations on sign-up, no credit card.

Generate free