Overview
10 stories in this issue. The first 3 are today's priorities.
Hot model updates
- Top · OpenAI pauses new $200 ChatGPT Pro subscriptions, citing GPT-6 Astra load
- Top · Anthropic says Moonshot routed nearly 300,000 Kimi requests to Claude
- Top · RunningHub open-sources multi-GPU acceleration for MiniMax H3, about 12x faster
- Cambricon ships day-0 DeepSeek-V4.1-Flash support on vLLM
Global AI news
- Sakana AI launches Fugu Max and Fugu Ultra v2 for learned multi-model orchestration
- SageMaker HyperPod inference adds model caching to cut cold starts
- Amazon Quick goes generally available on macOS and Windows desktops
Regional and early signals
- Researcher says 6TB of Chinese LLM relay logs exposed enterprise credentials
- Doubao Work adds local Office editing and browser record-and-replay
- Kimi K3 lifts Moonshot's annualized revenue past $1B, targeting $2B by year-end

Jiufeng graphic based on the sources cited in this issue.
Hot model updates
01/10
OpenAI pauses new $200 ChatGPT Pro subscriptions, citing GPT-6 Astra load
A week after GPT-6 Astra launched, OpenAI stopped taking new subscriptions to its $200-a-month ChatGPT Pro plan; existing subscribers and the API are unaffected.
Thibault Sottiaux, the OpenAI executive who leads Codex, announced on X on September 10th that OpenAI is pausing new ChatGPT Pro subscriptions because that tier puts "the most strain on our systems."
Existing Pro subscribers keep their accounts and access to Astra, and OpenAI's other subscriptions and API remain available. Sottiaux described suspending the $200 tier as the smallest intervention that would preserve broad access while OpenAI adds capacity. A day earlier, on September 9th, he had warned that OpenAI might stop accepting new Pro subscribers, calling demand for Astra unprecedented and saying OpenAI would prioritize existing customers.
- Timeline: GPT-6 Astra launched on September 3rd; new Pro sign-ups were paused seven days later
- Rollout: after launch it went to ChatGPT Plus, Pro, Business and Enterprise, plus the API, Microsoft Azure and AWS Bedrock
- API standard pricing: $10 per million input tokens and $50 per million output tokens; prompts over 272,000 input tokens carry higher rates
Limitations: only new sign-ups are affected, and new users who want Pro can for now only use other plans or the metered API; TechCrunch describes it as a temporary pause while OpenAI adds capacity. RuntimeWire argues the move exposes the tension between compute-heavy agents and fixed-price subscriptions.

Image source: OpenAI Developers; mirrored on Jiufeng R2.
Source: RuntimeWire · TechCrunch · Sottiaux on X · OpenAI: GPT-6 Astra · GPT-6 Astra API docs
02/10
Anthropic says Moonshot routed nearly 300,000 Kimi requests to Claude
According to Bloomberg, Anthropic alleges that Moonshot sent Kimi user requests to Claude through 5,380 fraudulent accounts without telling users.
Citing a Bloomberg report published September 10th, RuntimeWire says Anthropic alleges that Yang Zhilin's Moonshot AI routed Kimi user requests through Anthropic's Claude models without telling those users. One cluster Anthropic identified:
- Volume: nearly 300,000 requests
- Target model: sent primarily to an Opus model
- Accounts: 5,380 fraudulent accounts, most of which appeared to operate from Singapore and Japan
RuntimeWire notes that Moonshot presents Kimi as its own assistant and agent platform, powered by its own model family. If Kimi passed ordinary user prompts to Claude and showed the answers as Kimi's work, the dispute goes beyond how Moonshot trained its models to which model handled users' data and produced the product's output. Moonshot is asking US clouds, developers and investors to trust Kimi as an independent model platform; in July, Together AI announced a strategic partnership with Moonshot AI to natively serve Kimi models.
Limitations: this is Anthropic's allegation, surfaced by Bloomberg and relayed by RuntimeWire; the request records and account attribution still await independent verification.
Source: RuntimeWire · Together AI
03/10
RunningHub open-sources multi-GPU acceleration for MiniMax H3, about 12x faster
QbitAI reports that RunningHub's H3 Lightning cut a 5-second video render on four RTX 6000D cards from 348.8 to 28.7 seconds while keeping BF16 precision. (Chinese-language source)
After MiniMax open-sourced its H3 video model, the one-stop AIGC creation platform RunningHub built an inference acceleration package called H3 Lightning and open-sourced it on GitHub (repository RH-RunningHub/MiniMax-H3-MultiGPU-Lightning). QbitAI's test figures:
| Setup | Video | GPUs | Time (s) |
|---|---|---|---|
| Original BF16, 50 steps | 5 s, 1344×768 | 4×RTX 6000D | 348.8 |
| H3 Lightning | 5 s, 1344×768 | 4×RTX 6000D | 28.7 |
| Lightning, text only | 15 s, 768×1344 | 8×RTX 6000D | ~48 |
| Lightning, 2 ref images | 15 s, 768×1344 | 8×RTX 6000D | ~73 |
The first two rows amount to roughly a 12x speedup and about 92% less generation time. QbitAI's breakdown attributes the savings to three places: cutting excess computation, speeding up each step, and reducing time the GPUs spend waiting on one another; stock H3 inference runs 50 steps by default.
Limitations: the figures come from RunningHub's package as reported by QbitAI, and we have seen no third-party reproduction; QbitAI points out that the dog's legs occasionally look stiff during fast jumps. This item has a single Chinese-language source.
Source: QbitAI
04/10
Cambricon ships day-0 DeepSeek-V4.1-Flash support on vLLM
Cambricon finished adapting DeepSeek-V4.1-Flash on vLLM on the model's release day, so Cambricon chips were available alongside the new model.
Pandaily reports that Cambricon completed same-day (day-0) enablement of DeepSeek-V4.1-Flash on vLLM, combining Torch-MLU-Ops and BangC kernels with NeuWare. The report focuses on chip-side availability arriving together with the model rather than recapping the launch.
The model being adapted, V4.1-Flash, is a 552B-parameter mixture-of-experts model, close to double V4-Flash's 284B; SiliconANGLE reports that a new causal encoder-decoder design keeps just 8B parameters active while processing a prompt and 16B while generating, with image understanding built into the model.
Limitations: day-0 enablement means the model runs on this hardware and software stack; we have seen no public throughput or latency figures on Cambricon chips, and Pandaily is the only outlet we have seen reporting this adaptation.
Source: Pandaily · SiliconANGLE
Global AI news
05/10
Sakana AI launches Fugu Max and Fugu Ultra v2 for learned multi-model orchestration
Both models share one learned orchestration architecture; the cost-focused Fugu Max is priced at $2/$6 per 1M tokens.
MarkTechPost reports that Sakana AI released Fugu Max and Fugu Ultra v2, two models built on the same learned orchestration architecture. OpenRouter's model page says Fugu is not a single monolithic model but a learned multi-agent orchestration system.
- Fugu Max: routes tasks to lean open and specialized models, including NVIDIA Nemotron; $2/$6 per 1M tokens
- Fugu Ultra v2: aimed at peak capability; OpenRouter lists a 1,000,000-token context
Limitations: Fugu is not a single model, and tasks are handed to underlying models; the pricing and capability claims come from Sakana's release and the OpenRouter listing, and we have seen no independent evaluation.
Source: MarkTechPost · OpenRouter: Fugu Ultra v2
06/10
SageMaker HyperPod inference adds model caching to cut cold starts
AWS now lets HyperPod clusters pre-stage model weights and container images on nodes, so pods read them from local NVMe at about 7 GB/s instead of downloading them.
In its official blog, AWS says the wait between requesting a pod on SageMaker HyperPod and serving traffic is dominated by two sequential downloads: the inference server container image from Amazon ECR, and model weights from Amazon S3, Amazon FSx for Lustre or Hugging Face Hub. Smaller models may take a few minutes; a model like DeepSeek-R1 at 600+ GB can take 30 minutes or more before serving a single request, and every scale-out event repeats the download, so autoscaling response time is gated by network throughput to the storage backend.
The new model caching pre-loads model weights and container images onto cluster nodes before pods need them; at startup, pods read from local NVMe at approximately 7 GB/s rather than downloading over the network.
Limitations: this is AWS's own product documentation, the cold-start improvement is AWS's own description with no third-party testing seen, and the feature applies only to SageMaker Inference deployments on HyperPod.
Source: AWS Machine Learning Blog
07/10
Amazon Quick goes generally available on macOS and Windows desktops
Amazon's enterprise AI assistant Quick is now GA as a desktop app, and its mobile apps gain an activity feed that pulls email, calendar, CRM and messaging into one view.
AWS announced that the Amazon Quick desktop application is generally available on macOS and Windows. The iOS and Android apps gain an activity feed that consolidates email, calendar, CRM and messaging into one prioritized view, surfacing the decisions that need a team while agents handle routine items in the background.
AWS positions Quick as an enterprise-grade AI assistant running on AWS, with data staying in the customer's environment and conversations kept private, and pitches it as IT's answer to shadow-AI risk: rather than restricting AI access, give teams a work assistant on infrastructure IT already controls.
Limitations: this is AWS's own product announcement; claims about privacy and data staying in the customer's environment are vendor statements, and we have seen no third-party evaluation.
Source: AWS Machine Learning Blog
Regional and early signals
08/10
Researcher says 6TB of Chinese LLM relay logs exposed enterprise credentials
Security researcher Chaofan Shou says he obtained about 6TB of invocation logs from a Chinese large-model relay, containing enterprise-linked SSH keys, cloud credentials and tokens.
Pandaily reports that security researcher Chaofan Shou said he obtained roughly 6TB of invocation logs from a Chinese large-model relay service and found SSH keys, cloud credentials and tokens tied to enterprises. Pandaily presents it as a supply-chain risk claim pending confirmation.
Limitations: this is still the researcher's own claim; we have seen no independent verification of which relay was involved or how many companies are affected, so it should not yet be treated as a confirmed breach. This item has a single source, Pandaily.
Source: Pandaily
09/10
Doubao Work adds local Office editing and browser record-and-replay
Doubao Work can now edit local PPT and Excel files, turn a demonstrated web task into a reusable Skill, and switch between local and cloud computers mid-task. (Chinese-language source)
Leiphone reports three new features:
- Local Office editing: fully rolled out; pick files from the side workbench or right-click a local PPT or Excel file to have Doubao work on it; files the AI generates or edits can still be edited locally, and saved changes sync both online and locally
- Browser record and replay: after installing the latest Doubao desktop app, users demonstrate a task once in the Doubao browser; Doubao learns the clicks, form entries and uploads, turns them into a reusable Skill, and records when to invoke it, the inputs it needs and how to verify the result
- Execution environment switching: switch between "local computer" and "cloud computer" without starting a new conversation or interrupting the current one, avoiding lost context; a logged-in computer can also send instructions to other logged-in devices and use their files and compute
Limitations: local editing supports only PPT and Excel for now, with Word to follow, and record-and-replay covers only actions inside the Doubao browser. The feature details come from Leiphone's report, and we have seen no hands-on testing. This item has a single Chinese-language source.
Source: Leiphone
10/10
Kimi K3 lifts Moonshot's annualized revenue past $1B, targeting $2B by year-end
Moonshot's annual recurring revenue passed $1B in August, driven by Kimi K3, which launched in July, according to Bloomberg. (Chinese-language source)
Citing Bloomberg, IT Home reports that people familiar with the matter said Moonshot recently told investors its annual recurring revenue (ARR) passed $1B in August, and that it plans to raise it to $2B by the end of the year:
| Point in time | ARR ($B) |
|---|---|
| June 2026 | 0.3 |
| August 2026 | over 1 |
| Year-end 2026 target | 2 |
The report says the rapid ARR growth was driven by Kimi K3, released in July, which rose to the top of several benchmarks at a far lower cost than leading US rivals. People familiar with the matter said that, with consumer and enterprise interest in Kimi subscriptions and Moonshot's hosted models still rising, the company expects sales velocity to double again. Under the Kimi K3 license, companies that commercially host K3 and earn more than $20M cumulatively within 12 months must sign a separate agreement with Moonshot. Reuters reported last month that Moonshot is negotiating a 30% revenue share with Microsoft, Amazon and Google.
Limitations: the revenue figures come from people familiar with the matter, not a formal company disclosure. This item has a single Chinese-language source (IT Home relaying Bloomberg).
Source: IT Home
Write a prompt in your browser and get an image — 1K, 2K or 4K output, up to 15 reference images. 50 free generations on sign-up, no credit card.
Generate free

