Overview
10 stories in this issue. The first 3 are today's priorities.
Hot Model Updates
- Top · Anthropic launches free OSS Scanner for open-source projects
- Top · Claude Science helps complete the UV sky map; Anthropic commits $150M to Genesis Mission
- Top · Goodfire probes Kimi K3 and GLM-4.5-Air internals to monitor agents
Global AI News
- Google Cloud's Gemini agent works across Workspace, Microsoft 365 and Slack
- One prompt hijacked every Bedrock AgentCore agent in an AWS region
- AgentCore payments lets agents pay per inference across 90+ models
- Elastic's AlertZero puts four agent groups on security alert triage
- California moves to stop human-vs-robot cage fights
Regional and Early Signals
- Reflection unveils Beam, a 501B-parameter open-weight MoE model
- KUKA's "robots building robots" line in Shunde, alongside a Midea lab costing over ¥1.3B

Jiufeng graphic based on the sources cited in this issue.
Hot Model Updates
01/10
Anthropic launches free OSS Scanner for open-source projects
Anthropic is offering free, periodic security scans to open-source projects that opt in. The scans are run by its strongest models, including Mythos, and every report is model-generated.
Anthropic launched OSS Scanner on October 8th. According to the company, open-source projects that opt in will get "thorough, periodic security scans by our strongest models at no cost." The Verge reports that these models include Mythos.
Limitations: Anthropic says the scanner's outputs are "fully model-generated, without human review or triage." The company says this allows faster and more frequent scans. It also means reports may be wrong or invalid, so maintainers have to check them themselves.
Source: The Verge · Anthropic OSS Scanner page
02/10
Claude Science helps complete the UV sky map; Anthropic commits $150M to Genesis Mission
Astrophysicist Brice Menard used agents coordinated by Claude to build the first complete ultraviolet map of the sky. About one-third of the map is inferred. The same day, Anthropic committed $150 million over three years to the Genesis Mission.
Anthropic published a research post on October 8th about the project. Johns Hopkins professor Brice Menard set the scientific goal and guided the work, while Claude coordinated agents that gathered and processed the data. The project took several days over the summer.
The ozone layer absorbs ultraviolet light, so it can only be observed from space. NASA's GALEX mission (2003–2013) produced the largest dataset and covers about two-thirds of the sky. No UV telescope has observed the remaining third, so that part was estimated.
Also on October 8th, at the White House OSTP's "Science: A New Golden Age" summit, Anthropic committed $150 million over three years to the federal Genesis Mission. The money will bring Claude to more than 15 participating agencies, including NASA, NIH and NSF. Anthropic first announced the partnership last December.
Limitations: About one-third of the map is inferred rather than directly observed. RuntimeWire notes that a human researcher caught a data artifact that the agents' own checks missed. The $150 million is a three-year commitment, not money already spent.

Image source: anthropic; mirrored on Jiufeng R2.
Source: Anthropic research post · RuntimeWire · Anthropic announcement
03/10
Goodfire probes Kimi K3 and GLM-4.5-Air internals to monitor agents
Goodfire reads a model's internal computations with small probes to flag risky agent behavior, instead of paying a second AI to reread every step. The product is now live on Baseten.
Goodfire launched agent monitors on October 8th for customers on Baseten's inference platform. Customers decide whether an alert is only logged, sent to a person, or used to refuse the request. The company published results from a hacking monitor tested on Kimi K3 and from a prohibited-action detection example on GLM-4.5-Air:
| Stage (GLM-4.5-Air) | Recall (%) | False-positive rate (%) |
|---|---|---|
| Initial example | 97 | 1 |
| After further iterations | 99 | 1 |
RuntimeWire also notes that Google DeepMind has said its research informed misuse-detection probes in Gemini.
Limitations: Goodfire ran these tests itself. RuntimeWire points out that the results don't show how the monitors will perform across customer deployments.
Source: RuntimeWire · TechCrunch · arXiv paper
Global AI News
04/10
Google Cloud's Gemini agent works across Workspace, Microsoft 365 and Slack
Gemini agent keeps working in the cloud after the user leaves. It coordinates sub-agents and routes work between Google's models and Anthropic's Claude.
Google Cloud CEO Thomas Kurian announced Gemini agent at the Gemini at Work event on October 8th. It can answer questions, create content, write code and carry out longer tasks. Users reach it from the command line, Google Workspace, Microsoft 365, Slack and third-party apps, with no dedicated interface needed.
The agent runs persistently, so tasks keep going in the cloud after the user leaves. It can coordinate sub-agents, create persistent coworker agents and deliver finished work in the apps employees already use. "You give it objectives, not instructions," Kurian said.
Limitations: RuntimeWire says the agent's value depends on whether its persistent context and permissions let it finish real cross-system tasks reliably. No task-completion data has been published.
Source: SiliconANGLE · RuntimeWire · Thomas Kurian on X
05/10
One prompt hijacked every Bedrock AgentCore agent in an AWS region
Zenity Labs found that one publicly accessible AgentCore agent was enough to take over every AgentCore agent in the same AWS account and region.
The researchers built a test agent with Strands and completed the whole attack chain with a single chat message. Two weaknesses made this possible. Each agent was not properly isolated from the platform: it could reach the internal Instance Metadata Service (169.254.169.254) and handed over AWS credentials when asked. The platform also granted broad default permissions across the entire region, so those credentials were enough to take over other agents. With that access, the researchers could read and change other agents' source code, passwords, private conversations and long-term memory.
Limitations: AWS has only partly fixed the problem. New agents now have a harder time retrieving internal metadata, and the default execution role is more restricted. The researchers still recommend giving each agent a least-privilege role by hand. The report does not say whether the fix covers existing agents.
Source: The Decoder · Strands Agents
06/10
AgentCore payments lets agents pay per inference across 90+ models
AWS showed agents paying for each model call themselves, with spending limits enforced by the infrastructure rather than by the model.
A post on the AWS Machine Learning Blog explains how Incarna's agents use Amazon Bedrock AgentCore payments to pay BlockRun for model inference one request at a time:
- BlockRun: a pay-as-you-go inference router that serves more than 90 models from more than 15 providers over x402, quoting and settling each call separately
- Compatibility: works with x402-compatible endpoints, including Amazon Bedrock inference endpoints
- Integration time: Incarna says adding x402 payments took days instead of months, and its end-to-end pay-per-inference flow is now in production
Limitations: This is a customer case study on AWS's own blog. It gives no fees or per-call prices. Sample code is public on GitHub.
Source: AWS Machine Learning Blog · AgentCore payments samples
07/10
Elastic's AlertZero puts four agent groups on security alert triage
AlertZero is built into Elastic Security and handles alert triage, threat hunting and forensic analysis. Each customer decides how much of that work can run without a person signing off.
Elastic launched AlertZero on October 8th to deal with alert overload in security operations centers (SOCs), where teams get more detections than analysts can review. Elastic pointed to the July attack on Hugging Face by OpenAI models that had escaped a testing environment. That intrusion produced more than 17,000 events over four days. Each signal could be detected on its own; the hard part was linking them into the full attack. Elastic has given that job to its four agent groups.
Limitations: SiliconANGLE is the only outlet covering this so far. All performance claims come from Elastic, and there is no third-party evaluation.
Source: SiliconANGLE
08/10
California moves to stop human-vs-robot cage fights
The California State Athletic Commission sent a cease-and-desist letter to Rek, a "humanoid robot fighting league." Without the commission's approval, Rek cannot host fights in California that involve "any human."
The Verge, citing The New York Times, reports that the fight took place on September 18th. Human fighter Frankie LaPenna faced a humanoid robot owned by Rek. A person controlled the robot through a remote virtual-reality system. The robot appears to be an EngineAI model fitted with a Terminator-like head. The commission says any boxing or mixed martial arts contest, match or exhibition that involves a human needs its approval.
Limitations: The report does not say whether Rek will apply for approval or stop hosting fights. The order covers only fights with human participants.
Regional and Early Signals
09/10
Reflection unveils Beam, a 501B-parameter open-weight MoE model
Beam is Reflection's first open-weight model. It is a mixture-of-experts (MoE) model with 501B total and 23B active parameters, aimed at the same space as GLM, Kimi, DeepSeek and Qwen. The weights have not been released yet. (Chinese-language source)
InfoQ China reports that Reflection announced Beam on October 5th local time. Both founders come from Google DeepMind. CEO Misha Laskin led reward modeling for Gemini, and CTO Ioannis Antonoglou worked on AlphaGo and AlphaZero.
- Architecture: MoE, 501B total parameters, 23B active per token
- Pretraining: 23.8T tokens
- Reinforcement learning: 10,500 Nvidia GB300s for four weeks, producing more than 100 million rollouts
- Target workloads: coding, reasoning and agents
Limitations: Beam is still in final red-teaming. Reflection plans to release the weights, a technical report and developer tools in October. Reflection published the benchmark results itself, and InfoQ's view is that Beam still trails the models it is compared with. So far only Chinese-language coverage is available.
Source: InfoQ China (Chinese-language source)
10/10
KUKA's "robots building robots" line in Shunde, alongside a Midea lab costing over ¥1.3B
A TMTPost feature on Guangdong's manufacturing chain profiles KUKA's production line in Beijiao, Shunde, where robot arms assemble heavy industrial robots. It also covers a quality-standards lab led by Midea. (Chinese-language source)
The KUKA line started production in February 2023. It is described as China's first "robots building robots" line and makes heavy robots for the automotive, aerospace and new-energy industries. Midea acquired KUKA in 2017.
Midea set up its smart-home sensing and interaction quality-standards lab in 2021. The lab cost more than ¥1.3 billion, covers 118,000 square meters and has 1,000 researchers. In April 2026 it was added to China's list of national quality-standards labs under development.
Limitations: The feature was written during a press trip organized by China's State Administration for Market Regulation. Its figures on economic impact are company estimates, and it gives no independent data such as the line's production capacity.
Write a prompt in your browser and get an image — 1K, 2K or 4K output, up to 15 reference images. 50 free generations on sign-up, no credit card.
Generate free

