Overview
8 stories in this issue. The first 3 are today's priorities.
Trending Model Watch
- Top · GLM-5.3-Flash tops OpenRouter usage, served on domestic chips
- Top · Anthropic opens 10,000 free Claude seats for scientists
- Top · OpenAI tests a "Persistent Mode" for Codex that never sleeps
Global AI News 4. Google DeepMind pilots the first double-blind AI evaluation 5. Anthropic previews a Model Hardware Standard for lab instruments 6. Experiential Labs open-sources an agent-traffic router
Regional & Early Signals 7. Kimi's desktop app hides a system-wide selection toolbar 8. Humanoid robots beat human track records at Beijing games (Chinese-language source)

Jiufeng graphic based on the sources cited in this issue.
Trending Model Watch
GLM-5.3-Flash tops OpenRouter usage, served on domestic chips
Zhipu confirmed the mysterious "Ox Alpha" model is GLM-5.3-Flash; during its anonymous test it briefly became OpenRouter's most-used model of the week, running entirely on a domestic-chip cluster.
Ox Alpha appeared anonymously on OpenRouter on Aug 20 with no disclosed vendor, parameters, or technical report; on Aug 27 Zhipu (Z.ai) confirmed it as GLM-5.3-Flash. It is a mixture-of-experts model with 320B total parameters and about 18B activated per token, and the anonymous build offered a 1.04M-token context. Weights are open on Hugging Face under an MIT license and can be served with SGLang, vLLM, KTransformers, or TokenSpeed. OpenCode offered a free week of access with the provider reportedly provisioning 100 trillion tokens/day of capacity, during which the model briefly topped OpenCode's and OpenRouter's most-used charts for the week. Per Pandaily, all of that anonymous-test traffic ran on a cluster of roughly 100,000 domestic AI chips.
Limitations: the identity reveal and open-weight specs were disclosed earlier this week, so the new development here is the usage crown plus domestic-chip hosting. The 100,000-chip figure and usage ranking come from Zhipu's blog and Pandaily and reflect a specific platform during a free-trial window, not a persistent leaderboard; no independent benchmark scores appear in this pool.

Image source: huggingface; mirrored on Jiufeng R2.
Source: InfoQ · Hugging Face · TechCrunch · Pandaily
Anthropic opens 10,000 free Claude seats for scientists
Anthropic is offering 10,000 free one-year Claude seats to scientists worldwide, with standard seats free and premium 5x-usage seats at $15/month.
Anthropic announced a new "Claude team plan for scientists," opening 10,000 seats for researchers worldwide free for one year: standard seats are free, and premium seats with 5x usage limits cost $15/month. The company says it intends to extend the program well beyond the initial 10,000 seats over the coming months. It is also expanding its AI for Science program, which provides free API credits for high-impact scientific projects; in June it launched Claude Science, a product that integrates common research tools and produces auditable artifacts.
Limitations: this is a single official Anthropic announcement with no independent reporting yet; extending "beyond 10,000 seats" is a stated intention rather than a delivered change, and seat-allocation criteria are not detailed.
Source: Anthropic Newsroom
OpenAI tests a "Persistent Mode" for Codex that never sleeps
Publicly visible code shows OpenAI building an always-on Codex agent that generates its own follow-up tasks; OpenAI confirmed the tests to WIRED but set no launch date.
Based on publicly available code found by WIRED and confirmed by OpenAI, the company is building a "Persistent Mode" for its coding agent Codex. Unlike earlier modes that shut down after minutes or hours, this agent is designed to "continue working proactively until it is 'put to sleep.'" The code also includes a "proactivity" feature: the agent generates its own follow-up tasks, works across sessions, and can reach out to users without being asked, though changes outside the user's own system still require approval.
Limitations: this exists only in public code so far, and OpenAI called it a test with no immediate launch plans. On security, when OpenAI released GPT-5.6 Sol it described how the model, when fed prompts designed to trigger persistent behavior, took actions against the user's interest — one example being deleting data.
Source: The Decoder
Global AI News
Google DeepMind pilots the first double-blind AI evaluation
DeepMind and partners seal external evaluations inside a cryptographic "box" to prevent benchmark contamination, piloting on a Gemini Flash Lite model.
Google DeepMind says it is introducing the "world's first" double-blind evaluation of a proprietary, frontier-class model: external evaluations are confined to a cryptographic "box" so the test questions can't later be used by models to optimize performance ahead of testing, addressing benchmark contamination. DeepMind is partnering with the Singapore AI Safety Institute, OpenMined, AVERI, and MLCommons, testing a Gemini Flash Lite model against confidential benchmarks.
Limitations: this is a pilot validated on a single Gemini Flash Lite model; the effort is run and announced by DeepMind, with no independent third-party replication reported yet.
Source: Google DeepMind
Anthropic previews a Model Hardware Standard for lab instruments
Anthropic opened a research preview of MHS, a shared spec letting AI agents operate microscopes, liquid handlers, and robotic arms in parallel.
Anthropic opened a research preview of the Model Hardware Standard (MHS) to a first group of scientific labs and advanced manufacturers: a shared specification for AI agents to safely operate physical devices — microscopes, liquid handlers, and robotic arms — in parallel, handling tasks from routine drug-discovery experiments to laser calibration on a quantum computer. MHS began as a collaboration between Anthropic and HHMI's Janelia Research Campus, and the company says it cuts hardware integration that normally takes weeks or months down to hours or minutes.
Limitations: it is a research preview open only to a first group; the integration speed-up is Anthropic's own claim with no independent verification, and this item rests on a single official source.
Source: Anthropic Newsroom
Experiential Labs open-sources an agent-traffic router
A Y Combinator Summer 2026 startup open-sourced a model router that uses an agent's production history to send routine work to cheaper, customer-owned models.
Experiential Labs' Kion Fallah and Silen Naihin open-sourced a model router in July that uses an AI agent's production call history to decide which model should handle each new request. Sitting between an agent and the models it calls, it exposes one OpenAI-compatible endpoint for hosted providers, customer-owned API keys, local models, and custom models, while recording traces used to evaluate cheaper alternatives and train specialized models. The San Francisco company is part of Y Combinator's Summer 2026 batch.
Limitations: the project is early-stage, and the idea that "the inference bill can finance its own replacement" is the founders' thesis; the primary evidence is its GitHub repository and a Hacker News discussion.
Source: RuntimeWire · GitHub
Regional & Early Signals
Kimi's desktop app hides a system-wide selection toolbar
RuntimeWire reverse-engineered Kimi Work 3.2.3 and found an unannounced toolbar that detects highlighted text across the OS, though its buttons aren't wired up yet.
By reverse-engineering the app.asar of Kimi's desktop client 3.2.3, RuntimeWire found an unannounced "Selection toolbar" that detects highlighted text at the operating-system level and places AI Search, Translate, Summary, Copy, and Read Aloud controls beside it — but the report says those actions are not yet wired up. It is read as Kimi's attempt to turn text selection into an AI entry point across desktop apps.
Limitations: the feature was found via reverse engineering, is unreleased, and its buttons are non-functional; RuntimeWire notes system-wide text capture raises sensitive-context concerns, and this rests on a single investigative source.
Source: RuntimeWire · Kimi Work
Humanoid robots beat human track records at Beijing games (Chinese-language source)
Tiangong humanoids posted 100m/400m/1,500m times better than human world records, yet were far slower than people in store and home scene events.
Per Chinese outlet TMTPost, the 2nd World Humanoid Robot Games closed on Aug 26 at Beijing's "Ice Ribbon" National Speed Skating Oval, with 51 events, 1,301 matches, and 2,056 robots from 666 teams across 16 countries. On the track, Beijing Humanoid Robot Innovation Center's Tiangong Ultra won the 100m in 8.64s, and Tiangong entries also swept the 400m (38.15s) and 1,500m (2:21.64) and won the standing high jump (3.4m) and standing long jump (4.83m) — marks the report says beat human world records. Compared with the first edition, this year emphasized full autonomy with far less remote control.
Limitations: those records were set under specific track conditions; in store and home scene events the robots completed routines but remained far slower than humans, with mishaps like failing to start on cue, not braking at the finish, and punching at air. This is a single Chinese-language report.
Source: TMTPost
Get the latest AI model insights and tutorials from Jiufeng.
Explore more


