In this article
AI Highlights

Z.ai gives GLM Coding Plan subscribers a quota refill

Key Takeaways
  • Z.ai refills GLM Coding Plan quotas
  • Claude Fable 5.1 hits AWS
  • OpenAI previews cyber-critical Astra
  • Gemini agentic video cuts tokens up to 88%.
jiufeng
September 2, 2026
29 min read
Z.ai gives GLM Coding Plan subscribers a quota refill

Overview

9 stories in this issue. The first 3 are today's priorities.

Hot model watch

  1. Top · Z.ai refills every GLM Coding Plan subscriber's quota at year one
  2. Top · Claude Fable 5.1 lands on AWS as a "Covered Model"
  3. Top · OpenAI previews Astra, its first "cyber-critical" model
  4. Gemini adds agentic video understanding, cutting tokens up to 88%

Global AI news 5. Visko launches Orbis, a "Live Model" for continuously running video 6. Nvidia and CrowdStrike unveil SafeMind, an agentic security system 7. Nvidia's DLSS 5 arrives September 3rd, RTX 50-series only 8. Ai2's BenchMIRT audits what LLM benchmarks actually measure

Regional & early signals 9. Dyson's AI toothbrush CameraJet uses a camera to find gaps and squirt rinse

AI signal map for 2026-09-02
AI signal map for 2026-09-02

Jiufeng graphic based on the sources cited in this issue.

Hot model watch

Z.ai refills every GLM Coding Plan subscriber's quota at year one

On the coding plan's first anniversary, Z.ai hands every current subscriber a card that resets both the five-hour and weekly limits.

Z.ai co-founder Zhang Peng said on September 1st that GLM Coding Plan, on its first anniversary, is giving every current subscriber a "Reset Card" that refills both the rolling five-hour allowance and the weekly quota, letting developers resume coding-agent sessions without waiting for the usual windows to refresh. The plan, reportedly launched September 1st, 2025, meters usage as a five-hour prompt pool plus a weekly cap, with supported tools sharing the same subscription.

Redemption is tied to a signed-in ZCode account. Per RuntimeWire, the perk doubles as distribution: it retains paying developers while pulling them into Z.ai's own coding agent, ZCode.

Source: RuntimeWire · Z.ai on X

Claude Fable 5.1 lands on AWS as a "Covered Model"

Anthropic's Fable 5.1 is now on Amazon Bedrock and Claude Platform on AWS, alongside new Enterprise Frontier Safeguards.

Claude Fable 5.1 became available on September 1st on Amazon Bedrock and Claude Platform on AWS, aimed at coding, scientific research and enterprise workflows. Anthropic says Fable 5.1 outperforms Fable 5 on its hardest reasoning tests but did not publish scores, and has designated it a "Covered Model" — a category carrying extra data-retention, safety-review and access policies. The same day Anthropic announced Enterprise Frontier Safeguards (EFS), which pairs zero data retention with misuse-detection safeguards and stores data in cloud infrastructure the customer controls, not Anthropic. EFS will span Claude Code, Claude Enterprise, the Claude Platform, Amazon Bedrock, Claude Platform on AWS, Google's Agent Platform and Microsoft Foundry.

Limitations: Anthropic offered only a qualitative "outperforms" claim with no benchmark numbers; EFS isn't live yet and will roll out in phases starting later this fall, with eligible customers getting ZDR on Fable 5 and 5.1 in the meantime.

Developing Enterprise Frontier Safeguards with our customers
Developing Enterprise Frontier Safeguards with our customers

Image source: anthropic; mirrored on Jiufeng R2.

Source: AWS Machine Learning Blog · Anthropic Newsroom

OpenAI previews Astra, its first "cyber-critical" model

OpenAI says Astra is the first LLM to hit its "Critical cybersecurity" threshold; it ships soon, but access to the top capabilities will be limited.

OpenAI shared more on its forthcoming Astra model, calling it the first large language model to meet the "Critical cybersecurity" capability threshold. TechCrunch reports Astra is "very good at breaking into computer systems." OpenAI wrote that it plans to make Astra available soon, but that "access to its most advanced cybersecurity capabilities will be more limited," and paired the release with stronger safeguards.

Limitations: OpenAI gave no firm release date (only "soon") and said it will restrict access to the advanced cyber capabilities, which cut both defensive and offensive ways.

Source: TechCrunch · OpenAI

Gemini adds agentic video understanding, cutting tokens up to 88%

Google DeepMind gives Gemini's Flash models a mode that scans video on demand, cutting tokens up to 88% and cost up to 66%.

On September 1st, Google DeepMind launched agentic video understanding for Gemini 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite. The model dynamically scans video segments, which Google says improves accuracy while cutting token usage by up to 88%, cutting costs by up to 66% and boosting quality by up to 7%. Developers enable it by setting the API configuration to "agentic" in Google AI Studio or the Gemini Enterprise Agent Platform.

Limitations: the 88% / 66% / 7% figures are Google's best-case "up to" numbers, not averages, and the feature is limited to the Flash-tier models.

Source: Google DeepMind · Google

Global AI news

Visko launches Orbis, a "Live Model" for continuously running video

Visko unveils Orbis, a model that keeps generating an interactive video world instead of stopping at a clip, and raised $10M.

Sunnyvale, California AI research firm Visko launched Orbis on September 1st and opened public access. Visko calls Orbis the first "Live Model" for continuously generated interactive video: rather than producing a finished clip and stopping, it keeps a generative model running to sustain a continuous environment, using persistent memory to stop characters, objects and scenery from drifting in identity over long sessions — a common failure mode in generative video. Visko also announced a $10M pre-seed round led by Llama Ventures. Founded in 2025, the 16-person company draws staff from Apple, Google DeepMind, Meta, Amazon and Tesla.

Limitations: whether the "always running" approach can be made "affordable and reliable" is unproven (RuntimeWire's framing); uses for robotics and simulation are so far a thesis, with no deployment data.

Source: RuntimeWire · Orbis 1.0 technical paper (arXiv)

Nvidia and CrowdStrike unveil SafeMind, an agentic security system

At Fal.Con 2026, the two companies announced SafeMind, which pits AI offense and defense in a continuous "co-evolution loop."

At CrowdStrike's Fal.Con 2026 in Las Vegas, Nvidia CEO Jensen Huang and CrowdStrike CEO George Kurtz announced CrowdStrike SafeMind — an agentic cybersecurity system from CrowdStrike's Cyber Superintelligence Lab — before a crowd of about 10,000 security professionals. SafeMind combines CrowdStrike's purpose-built "beyond frontier-capable" models and custom agentic harnesses with defensive models built on Nvidia Nemotron, in a continuous co-evolution loop where offense and defense repeatedly challenge and improve each other. CrowdStrike also announced Falcon IQ (operationalizing Project QuiltWorks via agentic workload automation) and expanded its Guardian AI safety solution. "The adversaries are going to be more armed than ever… all of you are going to be more armed than ever," Huang told the crowd.

Limitations: these are conference announcements; availability, pricing and independent benchmarks were not provided.

Source: NVIDIA AI Blog

Nvidia's DLSS 5 arrives September 3rd, RTX 50-series only

The divisive AI upscaling tech ships this week, but with only one supported game and only on 50-series GPUs.

Nvidia is officially launching DLSS 5 on September 3rd. It will officially run only on RTX 50-series desktop and laptop GPUs and via GeForce Now cloud gaming, and the only game supporting it so far is NBA 2K27, out September 3rd at 9PM PT on PC. Nvidia touts DLSS 5 as "the company's most significant breakthrough in computer graphics since the debut of real-time ray tracing."

Limitations: DLSS 5 has been divisive since its March reveal — The Verge likened it to a "real-time generative AI filter for video games" and "motion smoothing for video games, but worse" — and the launch is limited to a single game and to high-end 50-series cards.

Source: The Verge · Nvidia

Ai2's BenchMIRT audits what LLM benchmarks actually measure

Allen AI proposes auditing benchmarks prompt by prompt, showing that questions often test something other than the stated skill.

On September 1st, Allen AI (Ai2) introduced BenchMIRT, a method for auditing LLM benchmarks at the level of individual prompts — the specific questions a model is scored on. A benchmark is usually meant to measure one ability (safety, general reasoning, instruction following), but individual tasks often depend on more: in BBQ, meant to test reliance on social stereotypes, a question about a grandson and grandfather booking an Uber probes age bias but also requires tracking who's who and reasoning from the given evidence. Even within one benchmark, different groups of questions can measure different things — WildJailbreak, for instance, mixes harmful jailbreak prompts with benign ones. Ai2 released a tech report, dataset and code.

Limitations: BenchMIRT is an auditing method and tool, not a new capability leaderboard, and its conclusions still depend on interpreting individual tasks.

Source: Hugging Face Blog · BenchMIRT code (GitHub)

Regional & early signals

Dyson's AI toothbrush CameraJet uses a camera to find gaps and squirt rinse

Dyson packs an electric brush, water flosser, micro-camera and image recognition into one handle; presale September 7th at 3,899 yuan. (Chinese-language source)

According to ifanr, Dyson unveiled the CameraJet imaging toothbrush, with presale starting 10am on September 7th at 3,899 yuan in two colors. Below the brush head sits a roughly 1mm, 100,000-pixel macro camera plus a strobe light; the camera analyzes 28 frames of the mouth per second and feeds a machine-learning algorithm called Gap Optical Targeting to spot gaps between teeth, going from "sees the gap" to squirting rinse in as little as ~100 milliseconds, up to ~0.15ml per squirt. Dyson says the system was trained on 470,000 tooth images and that the full hardware-software stack runs to about 16 million lines of code; the 230g unit is IPX7-rated with a replaceable battery and heads and a 12.5ml rinse tank in the handle.

Limitations: all figures — including the recognition pipeline and training scale — are Dyson's own; there is no independent testing, and this is a presale announcement, so real-world cleaning and recognition performance remain unverified. (Chinese-language source)

Source: ifanr