Overview
11 stories in this issue. The first 3 are today's priorities.
Hot model updates
- Top · Nex-AGI ships N2.5 agent family, Pro weights still pending
- Top · Subscribers report Claude tokens draining with no work running
- Top · GPT-5.6 Sol is being used to run quantum computing experiments
Global AI news
- NVIDIA announces CUDA Rust for compile-time-safe GPU kernels
- Chrome moves to a two-week release cycle
- AlphaGenome Atlas predicts 9 billion single-letter DNA variants
- Pathway's BDH keeps reasoning in latent space, with no chain-of-thought
- MiniMax backs a Tokyo conference on paying Japanese rights holders
Regional and early signals
- Herdr raises $6M to keep coding agents running after you close the laptop
- Refuse the right subset of a topic, not the whole topic
- Roborock debuts a pool-cleaning robot at IFA (Chinese-language source)

Jiufeng graphic based on the sources cited in this issue.
Hot model updates
01/11
Nex-AGI ships N2.5 agent family, Pro weights still pending
Nex-AGI released a 35B Mini, a 397B Pro and a 1.6T Max in one go; Mini is downloadable while Pro's weights are still marked coming soon.
The three tiers split the work differently. Mini and Pro are multimodal models for operating computers and browsers, while Max drops visual input and concentrates the largest model on reasoning, coding, scientific work and longer-running agent tasks. Public materials describe Nex-AGI as a collaboration involving the Shanghai Innovation Institute, Shanghai Qiji Zhifeng, Mosi Intelligence and Kuafu Technology; the people behind the models identify collectively as the Nex-AGI Team, with no named founder or chief executive and no publicly established legal structure or headquarters. The earlier N1 technical report is dated December 4, 2025 and is likewise credited to a large research team, and the team's earlier open-source stack covered models, datasets, agent frameworks, reinforcement-learning tools and inference infrastructure.
Publishing weights and delivering a usable system are two different things: Mini runs in Nex's sample two-H100 deployment, Max's reference setup calls for 16 H200s across two nodes, and Pro's weights have yet to appear in the model repository.

Image source: huggingface; mirrored on Jiufeng R2.
Source: RuntimeWire · Hugging Face · GitHub
02/11
Subscribers report Claude tokens draining with no work running
A Claude Max 20x user watched usage climb from 45% to 55% during a controlled interval with no work at all; Anthropic has since warned users about hackers.
On August 4, Grant De Swardt, an independent AI consultant in East Sussex, U.K., noticed his Claude Max 20x account consuming tokens on a day he had not been working. The next day he disabled everything attached to Claude and did not use it, and consumption still increased. He described the clearest controlled interval to TechCrunch: usage went from 45% to 55% while he performed no work, scheduled Cowork tasks were paused or completed, Dispatch and cloud execution were disabled, and no local Claude Code task was running.
The evidence boundary matters here: the full controlled record made public comes from this one user, and the reporting is TechCrunch's alone. Anthropic's warning to users about hackers came afterwards.
Source: TechCrunch
03/11
GPT-5.6 Sol is being used to run quantum computing experiments
OpenAI published a case study saying an MIT researcher uses GPT-5.6 Sol with Codex to autonomously run quantum experiments, analyze results and calibrate qubits.
As described on OpenAI's page, the setup covers three parts of the workflow — running the experiments, analyzing the results and calibrating qubits — and is characterized as running autonomously.
This is OpenAI's own customer story, so it is a vendor account. We have seen no third-party reproduction or independent reporting; how transferable the setup is will depend on material published by the research side.
Source: OpenAI
Global AI news
04/11
NVIDIA announces CUDA Rust for compile-time-safe GPU kernels
Two NVlabs open-source projects cover the SIMT and Tile programming models, using Rust's ownership rules to reject aliasing bugs at compile time.
Rust code could already launch CUDA kernels, but the kernel body usually had to be written elsewhere; CUDA Rust closes that gap. cuda-oxide targets the classic SIMT model, where you describe what one thread does and launch thousands of them, while cutile-rs targets the newer Tile model, which is also available in C++ and Python. Both compile Rust kernels natively. NVIDIA's stated motivation is that the systems layer of AI is increasingly Rust: the Nova Linux driver is in Rust, NVIDIA Dynamo has a Rust core, and NVTX has Rust bindings — the GPU kernel was the exception.
Maturity differs sharply. cutile-rs is published on crates.io, runs on stable Rust 1.89+, and is already used in Hugging Face's Grout inference engine and in mistral.rs; cuda-oxide is labelled early alpha on GitHub.
Source: MarkTechPost · NVIDIA developer blog · GitHub
05/11
Chrome moves to a two-week release cycle
Starting with Chrome 153, desktop, iOS and Android releases ship every two weeks instead of every four.
The change Google promised earlier this year is now in effect: Tuesday's launch of Chrome 153 covered desktop, iOS and Android, and releases follow a two-week cadence from here. Google ties the shift to Chrome's evolving security strategy in the AI era — automated AI tools and community bug reports have pushed up the volume of patches and updates, and a shorter cycle makes security fixes easier to manage.
This is a cadence change rather than a fix announcement for any specific vulnerability, and "AI is pushing up patch volume" is Google's own framing.
Source: TechCrunch · Google security blog
06/11
AlphaGenome Atlas predicts 9 billion single-letter DNA variants
Google DeepMind's new platform holds predictions for every possible single-letter change in the human genome — 9 billion variants in all.
DeepMind calls it the most comprehensive catalogue of how genetic mutations affect molecular biology, available for academic research through a free website portal. It builds on the earlier AlphaGenome model and works around a hard constraint: with roughly 9 billion possible single-letter mutations, testing each one in the lab is practically impossible.
The limits are stated in the source too: the Atlas provides model predictions rather than experimental measurements, and access is scoped to academic research via the free portal.
Source: Google DeepMind
07/11
Pathway's BDH keeps reasoning in latent space, with no chain-of-thought
BDH reasons over a brain-inspired sparse neuron graph in latent space without emitting intermediate chain-of-thought tokens, and is trained and scaled on Amazon SageMaker HyperPod.
Pathway's Baby Dragon Hatchling (BDH) is described as a brain-inspired, post-transformer architecture, originally formulated as a graph of neurons that communicate through sparse, local interactions and keep state in synapse-like connections. Instead of externalizing reasoning as extra tokens generated sequentially and fed back into later steps, it learns from examples and refines a solution in latent space. Model states adapt in context without test-time weight updates, and the reasoning horizon is not bounded by a fixed-size context window.
The write-up is published on the AWS Machine Learning Blog and covers how BDH is developed and scaled on SageMaker HyperPod. The architectural claims come from Pathway and its paper, and we have seen no third-party evaluation.
Source: AWS Machine Learning Blog · arXiv paper
08/11
MiniMax backs a Tokyo conference on paying Japanese rights holders
The September 11 event covers how Japanese IP is used, protected and commercialized in generative AI, with a keynote slot for the H3 model.
MiniMax is sponsoring the conference, whose program puts compensation for rights holders and rights protection near the center of the discussion. The company reported that 73% of its 2025 revenue came from outside mainland China, and the Tokyo event tests whether it can turn interest in H3 into local relationships built around compensation, rights protection and the global commercialization of Japanese IP. Founder Yan Junjie is chairman, CEO and CTO; he previously spent more than six years at SenseTime, where he became a vice president and deputy head of its research institute.
The conference has not happened yet. What is public so far is the program and the sponsorship, not any licensing agreement or revenue-share terms, and the primary material for this item is MiniMax's post on X.
Source: RuntimeWire · MiniMax on X
Regional and early signals
09/11
Herdr raises $6M to keep coding agents running after you close the laptop
Bessemer led the seed round as founder Can Celik turns an open-source agent runtime into a developer-tools company.
Herdr addresses what long-running coding agents lack — persistence, status tracking and remote control — without tying those capabilities to any single model vendor's interface. Bessemer Venture Partners led the $6 million seed round, with Y Combinator, e2 and several angel investors participating. The financing was announced on September 8, one month after Celik joined YC's Fall 2026 batch.
Scale is the caveat: YC still lists Herdr as a one-person operation, Celik wrote in August that he was "the only person behind Herdr," and the funding announcement makes hiring the immediate use of the capital. The round is a test of whether open-source adoption can support a durable developer-tools business.
Source: RuntimeWire · Herdr on X
10/11
Refuse the right subset of a topic, not the whole topic
A Multiverse Computing team argues that treating harm as a property of a topic is why models refuse safe prompts containing a dangerous-looking word.
The article notes that most safety alignment work treats harm as a topic-level property — a prompt is unsafe because it falls into a general category such as weapons, fraud or self-harm — and that guard models like LlamaGuard-3 encode exactly this taxonomy. Benchmarks such as XSTest and OR-Bench probe the resulting failure mode, and refusal-calibration work tries to pull that number back down.
The authors' point is that real deployments do not fit that picture: the same base model may be adapted for a general assistant, an educational product, an enterprise system or a public-sector service, and each setting needs different boundaries within the same topic — a civics tutor and a public-sector assistant can share a model yet require opposite behaviour on politics. This is a vendor team's methodology post on Hugging Face; its evaluation numbers should be checked against the original, and we have seen no third-party reproduction.
Source: Hugging Face · XSTest paper
11/11
Roborock debuts a pool-cleaning robot at IFA (Chinese-language source)
The RockAqua P1 patrolled a pool built inside the Berlin hall, shown alongside the new RockNeo Q2 LiDAR lawn mower.
In Hall 9 at IFA Berlin on September 4, Roborock moved a pool into the exhibition space for the global debut of the RockAqua P1 pool-cleaning robot, with the new RockNeo Q2 LiDAR robotic lawn mower displayed opposite it. Rival booths from Ecovacs and Dreame drew crowds of their own. Leiphone cites IDC figures for 2025: 32.72 million home cleaning robots shipped worldwide, with Roborock at 5.8 million units and 17.7% share, followed by Ecovacs at 14.3%, Dreame at 10.5%, Xiaomi at 6.7% and Narwal at 5.3% — all five Chinese vendors, together above 54%. Former leader iRobot filed for bankruptcy protection in December 2025 and was taken over by its Chinese contract manufacturer.
Evidence boundary: this is a Chinese-language industry feature. On the product side it offers only trade-show debut information — no pricing, availability dates or autonomous-navigation performance figures.
Source: Leiphone
Write a prompt in your browser and get an image — 1K, 2K or 4K output, up to 15 reference images. 50 free generations on sign-up, no credit card.
Generate free

