AI Highlights

Claude now leads 26% of Anthropic's own AI R&D

Key Takeaways

Anthropic says Claude now leads 26% of its own AI R&D, plus an open NASA-IBM lunar model, OpenAI's $278B burn plan and Huawei's 4,096-card supernode.

jiufeng
September 19, 2026
31 min read
In this article

Overview

8 stories in this issue. The first 3 are today's priorities.

Hot model watch

  1. Top · Claude now leads 26% of Anthropic's AI R&D, with 30,000 agents running at once
  2. Top · NASA and IBM open a lunar foundation model: 11 modalities, 49,000 craters

Global AI news

  1. Top · OpenAI's five-year plan: $856B on compute, $278B of negative free cash flow
  2. Nscale files for a NYSE listing after a $44.6B Anthropic contract
  3. Bedrock AgentCore gets a new runtime that frees memory when a session ends
  4. DoorDash points multi-agent LLMs at 60,000 stale feature flags

Regional and early signals

  1. Huawei's Ascend 960 supernode scales to 4,096 cards; 960DT is ready three quarters early
  2. Tang Jie and the GLM team say GLM-5.3 has reached the RSI threshold
AI signal map for 2026-09-19

Jiufeng graphic based on the sources cited in this issue.

Hot model watch

01/08

Claude now leads 26% of Anthropic's AI R&D, with 30,000 agents running at once

Anthropic's first internal R&D automation index shows "leading" tasks going from under 1% to 26% in six months.

On September 17 Anthropic published internal data quantifying how much of its own model R&D Claude handles, using an in-house R&D Automation Index built on Epoch AI's six-level scale (AL0–AL5). AL3 is "collaboration"; AL4, "leading," means an engineer gives only a high-level goal while the model does most of the end-to-end work and humans supervise and make the final call; only AL5 is full autonomy. As of August 2026, Claude reached AL4 on about 26% of model R&D tasks and at least AL3 on more than 90% of them, and roughly 30,000 agents run research and engineering tasks concurrently on the company's most widely used internal agent platform.

Share of R&D tasks (weighted by human hours)Feb 2026Aug 2026
Claude at "leading" (AL4)under 1%26%
At least "collaboration" (AL3)not disclosedover 90%

The method: every week in July 2026, Anthropic randomly sampled 20% of staff in the departments inside the model R&D loop, had a Claude Research Agent read their Slack and internal documents, distilled roughly 15,000 fine-grained R&D tasks, and organized them into a 542-node task tree with 378 leaf nodes covering pretraining, reinforcement learning, eval-platform debugging, RL sandbox network policy and inference incident reviews, weighted by the human time each task consumed.

Limitations: the judge scoring Claude's automation level is also Claude. Cross-checked against the employees who own the work, model and human ratings matched exactly 59% of the time, while human raters agreed exactly with each other only 35% of the time; allowing a one-level gap, agreement reached 97%. The 26% figure does not mean a quarter of Anthropic's R&D is unmanned, or that a quarter of its researchers have been replaced — AL4 still needs humans to set goals, supervise and decide.

Measurements for understanding the pace of AI development inside frontier labs

Image source: anthropic; mirrored on Jiufeng R2.

Source: Anthropic · InfoQ China (Chinese-language source)

02/08

NASA and IBM open a lunar foundation model: 11 modalities, 49,000 craters

Pretrained from scratch on nearly two million geographically partitioned data bundles; downstream code is public, pretraining code is not.

NASA and IBM released the NASA-IBM Lunar Foundation Model on September 10 for crater mapping, volcanic-feature analysis and research into possible ice at the lunar poles. USRA detailed the contribution of planetary scientist Rachel Slank on September 18: working at NASA's Marshall Space Flight Center through USRA's Science and Technology Institute, she aligned observations collected by different instruments at different resolutions without discarding the physical context of each measurement, and labeled the crater data.

  • Training data: nearly two million geographically partitioned data bundles, pretrained from scratch
  • Modalities: 11, from camera imagery and topography to radar, mineralogy, gravity and illumination geometry
  • Resolution range: 1 meter to 20 kilometers per pixel
  • Human labels: 49,000 craters
  • Pretraining compute: about 1,100 H100 GPU-hours (per the model card)

Limitations: the GitHub release includes downstream code but omits the pretraining code, so full external reproduction is not possible.

Source: RuntimeWire · Hugging Face model card

Global AI news

03/08

OpenAI's five-year plan: $856B on compute, $278B of negative free cash flow

An investor deck shows revenue targets trailing the compute bill, with the gap left to capital markets.

The Financial Times reported on September 18, citing a recent presentation prepared for investors, that OpenAI expects negative free cash flow of $278 billion from 2026 through 2030, against roughly $856 billion in spending on computing power and infrastructure and a revenue target of $840 billion over the same period. Sam Altman is pairing that forecast with early talks for another private round: investors have discussed a valuation of at least $1.2 trillion and OpenAI is seeking more — $1.2 trillion would be 41% above the $852 billion post-money valuation set less than six months ago.

Five-year plan (2026–2030)Amount
Revenue target$840B
Compute and infrastructure spend$856B
Free cash flow−$278B

Limitations: these figures come from an internal investor deck seen by the FT, not from OpenAI disclosure, and the numbers are still moving — the revised plan already trims expected burn by $27 billion. The new round is only in discussion and the $1.2 trillion valuation is not done; the immediate runway comes from the financing the company announced on March 31.

Source: RuntimeWire · OpenAI's March announcement

04/08

Nscale files for a NYSE listing after a $44.6B Anthropic contract

The data center builder lost $1.02B in the first half while revenue grew more than 1,250%.

Nscale Global Holdings filed to go public on the New York Stock Exchange on September 18 without disclosing how much it hopes to raise; sources told Bloomberg last month that it is seeking up to $3 billion, and the company was valued at $14.6 billion in March. It is not yet profitable: net loss reached $1.02 billion in the first half, nearly three times the year-earlier figure, while revenue over the same period jumped more than 1,250% to $140.6 million. Last month Nscale signed a $44.6 billion infrastructure deal with Anthropic for about 460 megawatts at Monarch, its flagship West Virginia campus, which is powered entirely by an on-premises microgrid.

Limitations: the filing does not state the target raise, and the $3 billion figure comes from Bloomberg's sources; the "clear expansion path to 8GW+" is the company's own framing of a plan, not capacity under construction. This item rests on a single report.

Source: SiliconANGLE

05/08

Bedrock AgentCore gets a new runtime that frees memory when a session ends

AWS rebuilds the managed compute layer for long-running production agents and charges by actual usage.

AWS announced a new Amazon Bedrock AgentCore runtime on September 18. It is the managed compute layer for AgentCore, aimed at long-running production agents. The listed changes: memory is reclaimed as a session releases it instead of being held; resource allocation is finer-grained; starts are faster and more consistent; and pricing tracks actual usage. AWS says thousands of teams have used the runtime for production agents since launch, and the AgentCore samples repository carries a runnable example of hosting an agent over the HTTP protocol.

Limitations: this rests on AWS's own blog alone, with no published cold-start latency, throughput or cost figures and no third-party reproduction, so "faster, more consistent starts" remains a vendor claim.

Source: AWS runtime announcement · AgentCore samples repo

06/08

DoorDash points multi-agent LLMs at 60,000 stale feature flags

Flag cleanup across 623 repositories is handed to LLM agents, with experiment data pulled in over MCP.

InfoQ reports that DoorDash built a multi-agent LLM system to automate stale feature flag cleanup across more than 60,000 flags and 623 repositories. The workflow pulls live experimentation data in through MCP to judge which flags no longer serve a purpose, then combines that with engineer judgment before removal.

Limitations: this rests on a single InfoQ report, with no published accuracy, rollback rate or share of changes reviewed by hand, and no figure for how many flags were actually deleted.

Source: InfoQ

Regional and early signals

07/08

Huawei's Ascend 960 supernode scales to 4,096 cards; 960DT is ready three quarters early

Huawei moves the contest from a single chip to the whole system: 4,096-card supernodes, NPO optical interconnect, one generation a year.

Chinese outlet QbitAI reported on September 19 that Huawei's Wang Tao gave new Ascend numbers: the Ascend 960DT was ready three quarters ahead of schedule with double the per-chip compute, and the 960PR version is also expected early. Huawei put supernodes, NPO optical interconnect and the CANN software stack on the table together, arguing that "one good chip is nowhere near enough" — a large supernode involves chips, communications, IP networking, optical interconnect, power, cooling, software and fault management at once.

  • Ascend 960DT compute: 2 PFLOPS FP8, 4 PFLOPS FP4
  • Memory: up to 288GB HBM at 9.6TB/s
  • Supernode scale: 4,096 NPU cards, up to 8 EFLOPS FP8 and over 1PB HBM
  • Roadmap: Ascend 970 in 2028, Ascend 980 in 2029, one generation per year
SupernodeNPU cards
Ascend 910C384
Ascend 9501,024
Ascend 9604,096

Huawei also said that at the same 100,000-card cluster size, a cluster built from 4,096-card supernodes can raise MFU — the share of theoretical compute a model actually consumes — by 2.75x versus one built from traditional eight-card servers, and claims its NPO design, which pushes the optical engine close to the chip, is the first in volume commercial use, integrates 36 lanes of 200G and is the only one with a built-in light source.

Limitations: every figure comes from Huawei's own presentation, the 2.75x MFU gain is from Huawei's simulation rather than third-party measurement, and the 970 and 980 are roadmap items. Chinese-language source only.

Source: QbitAI (Chinese-language source)

08/08

Tang Jie and the GLM team say GLM-5.3 has reached the RSI threshold

Zhipu writes up its self-improvement progress and cites one inference-optimization case on a domestic-chip cluster.

InfoQ China published a long piece by Tang Jie and the GLM team on September 18 describing Zhipu's latest work on recursive self-improvement (RSI), saying GLM-5.3 has reached the threshold and using the phrase "step by step toward replacing us." The case the article cites took place on a 100,000-card domestic-chip cluster: in under two weeks, end-to-end throughput was raised to 3x the initial baseline, using an EPD-disaggregated architecture and W8A8 quantization.

Limitations: the progress and the numbers are the GLM team's own account, "reached the threshold" is their own judgment, and there is no third-party reproduction or independent benchmark behind either. Chinese-language source only.

Source: InfoQ China (Chinese-language source)

Generate one yourself with JIUFENG

Write a prompt in your browser and get an image — 1K, 2K or 4K output, up to 15 reference images. 50 free generations on sign-up, no credit card.

Generate free