AI Highlights

Claude writes 80% of merged code, CI jobs up 25x

Key Takeaways

Anthropic says Claude writes ~80% of merged code as CI jobs jumped 25x, plus ChatGPT contractor review, Codex Handoff and Perplexity's local Windows agent.

jiufeng
September 15, 2026
39 min read
In this article

Overview

10 stories in this issue. The first 3 are today's priorities.

Model Watch

  1. Top · Claude writes 80% of Anthropic's merged code
  2. Top · OpenAI contractors review real ChatGPT chats and memory summaries
  3. Top · Codex demo moves a chat and working files between Macs

Global AI News

  1. Perplexity's local agent lands on Windows, accelerated by RTX
  2. Reward AI releases OM-1, trained only on human glove demonstrations
  3. Microsoft publishes a 37-page humanist AI code of conduct
  4. AgentCore adds a managed consent portal for OAuth session binding
  5. Abnormal AI uses the code interpreter as an ephemeral agent scratch pad

Regional & Early Signals

  1. Claude Code canary appears to route to Opus 5.2, with no Anthropic announcement
  2. Ant Group releases SingProbe, an in-model safety guardrail with sub-0.5% decode overhead

Model Watch

01/10

Claude writes 80% of Anthropic's merged code

Anthropic's own numbers: faster code output pushed the bottleneck downstream into testing and CI.

Anthropic says Claude now authors about 80% of the code merged into its repositories, and its engineers ship eight times as much code per quarter as they did during the 2021-to-2025 period. Addy Osmani (@addyosmani), the former Chrome developer-experience leader, surfaced the figures on September 14th, clarifying that the 80% measures lines merged into production that Anthropic can clearly attribute to Claude; the remaining 20% combines human-written code and other artifacts. The numbers come from a September 14th engineering post by engineer Sachin Malhotra.

MetricGrowthWindow
Code shipped per quarter8xvs. 2021-2025
Test count10xpast six months
CI job volume25xpast six months

Limitations: All figures are self-reported by Anthropic and not independently verified, and the 8x measures code volume — not a proportional increase in features, revenue or software quality. The volume also came at a downstream cost: the growth in tests and CI jobs forced a three-week rebuild of Anthropic's test-selection system.

Claude writes 80% of Anthropic's merged code: Growth

Jiufeng graphic based on the sources cited in this issue.

Source: RuntimeWire · Addy Osmani on X

02/10

OpenAI contractors review real ChatGPT chats and memory summaries

Hundreds of reviewers score live conversations under the codename Project Lily, behind a filter OpenAI admits can miss identifiers.

According to a 404 Media investigation published September 14th, OpenAI is using hundreds of outside contractors to read and evaluate real ChatGPT conversations, including chats containing sensitive personal information; internal documents identify the work by the codename "Project Lily." Reviewers get anonymized prompts and, in some cases, entire conversations. Usernames are removed, but the text itself can expose personal details. Some tasks also include a "user memories summary" describing previous interactions, location information and other context ChatGPT has retained about the user. OpenAI says conversations pass through its Privacy Filter first, which is built to detect and mask names, contact details, addresses, account numbers, passwords and other identifiers.

Limitations: OpenAI's own technical description of Privacy Filter says the model can miss uncommon identifiers or ambiguous references. The pipeline is on by default, the opt-out is buried in data controls, and per the report the pipeline can retain older chats even after a user opts out.

Source: RuntimeWire · OpenAI Privacy Filter · OpenAI consumer data FAQ

03/10

Codex demo moves a chat and working files between Macs

An untracked file arrived on the second machine with the same hash — on a modified build with feature flags on.

Developing Adventures (@DevAdventur3s) demonstrated Codex moving a coding chat and its project state from one Mac to another on September 13th, including an untracked file that arrived with the same hash. The 13-post thread showed two paths: sending work from a local machine to the cloud, and transferring it to another Mac. In the Mac-to-Mac test, the conversation continued on the destination machine and the untracked file moved without a Git commit or push — meaning a developer can leave a laptop, continue on a Mac mini or another workstation, and avoid packaging unfinished changes merely to change machines.

Limitations: The test used a modified Codex build with feature flags enabled, so it was a close look at an existing Handoff workflow rather than a new OpenAI release; OpenAI's documentation shows host-to-host Handoff has been available since June. The report notes developers still need a precise contract for files, secrets and failed transfers.

Source: RuntimeWire · OpenAI

Global AI News

04/10

Perplexity's local agent lands on Windows, accelerated by RTX

Work finished locally doesn't consume Perplexity Computer credits.

NVIDIA announced on September 14th that Perplexity is adding Portable Computer to its Windows app — a local version of the Perplexity Computer agent that plans and carries out multistep tasks using local models on compatible GeForce RTX PCs and RTX PRO Workstations: analyzing data, bringing together information across files and handling recurring work. The app ships with local models such as Qwen 3.8 27B, which was post-trained for Perplexity Computer and optimized for RTX GPUs. Sensitive information stays on device, locally completed work doesn't consume Perplexity Computer credits, and users can orchestrate work up to cloud models for more advanced research and reasoning. The release builds on existing support for NVIDIA DGX Spark systems and RTX PCs running Linux.

Limitations: The capability is tied to compatible NVIDIA RTX hardware — an RTX or RTX PRO GPU with 24GB or more of VRAM; heavier research and reasoning still has to be orchestrated up to cloud models.

Source: NVIDIA AI Blog

05/10

Reward AI releases OM-1, trained only on human glove demonstrations

No teleoperation data, no on-robot data — the policy still runs on industrial arms and humanoids.

Reward AI, a robotics startup whose team's prior work includes DexCap, HumanPlus and ALOHA, has released OM-1 (Omnibody Model 1), a general-purpose manipulation policy learned from humans wearing a 7-DoF sensorized glove, then run on industrial arms and humanoids at human speed. The company sums it up as "One Model, One Data Interface, Any Body." Most robot foundation policies train on teleoperated or self-collected robot data, which binds the dataset to one embodiment; Reward AI instead designs capture, learning and control as a single pipeline, so demonstrations recorded today can train robot bodies that do not exist yet.

Limitations: OM-1 is Reward AI's in-house policy. No weights, code, dataset or API have been released, so developers cannot run it on their own hardware yet.

Source: MarkTechPost

06/10

Microsoft publishes a 37-page humanist AI code of conduct

It states that "people matter more than AI" and rejects model personhood and model welfare.

Microsoft published a 37-page "humanist AI code of conduct" on September 14th. The document makes clear that "people matter more than AI," that AI models are not conscious and "should not be designed to imitate consciousness," and it rejects "the pursuit of legal personhood, or the idea that models might deserve welfare, or be entitled to rights." It lands amid growing safety concerns over AI model progress: Anthropic CEO Dario Amodei called for a coordinated slowdown of AI development over the weekend, after researchers warned that model progress could outpace our ability to safely deploy increasingly complex systems and to verify and control the actions of AI agents.

Limitations: This is a company-authored code, not a regulatory instrument; the report does not describe external auditing or consequences for violations.

Source: The Verge

07/10

Developers no longer have to host their own callback and session-binding infrastructure.

Amazon Bedrock AgentCore Identity now offers a Consent portal, a managed web experience and session binding endpoint for AgentCore Gateway. Before an agent can act on services such as GitHub and Slack, the user must authenticate with the provider and explicitly approve the requested access, and the application must securely associate the resulting OAuth grant with the user who authorized it — session binding. Customers using the three-legged OAuth (3LO, the OAuth 2.0 authorization code flow) previously had to build and host that infrastructure themselves: presenting the authorization URL, hosting a public HTTPS callback, authenticating the returning user, managing browser sessions and calling CompleteResourceTokenAuth. Now you create a portal for a gateway, share its URL, and users authenticate with your organization's identity provider and review the services available. AWS publishes end-to-end samples for Microsoft Entra ID and Okta.

Limitations: The portal only serves agents attached to an AgentCore Gateway, and the prerequisites stay on the customer: a registered GitHub OAuth App and Slack app, plus administrative access to the corporate IdP.

agentcore-samples/01-features/05-authenticate-and-authorize/07-consent-portal-auth-code-flow-targets at main · awslabs/agentcore-samples

Image source: GitHub; mirrored on Jiufeng R2.

Source: AWS Machine Learning Blog · AgentCore samples · GitHub OAuth Apps docs

08/10

Abnormal AI uses the code interpreter as an ephemeral agent scratch pad

Behind real-time threat detection on billions of daily messages, agents lean on throwaway compute for aggregation and verification.

Abnormal AI has deployed Amazon Bedrock AgentCore Code Interpreter as a compute scratch pad for the agents behind its real-time inline email threat detection — not only for coding tasks, but for data aggregation, analysis and verification, where semantic reasoning alone isn't enough. The company, which says it protects more than 25 percent of the Fortune 500, runs these systems in production today, processing billions of messages and executing agent-driven code at that same scale to block threats before they reach the inbox. It also disclosed that 80 percent of its code changes are built using an agent in some way, with 40 percent completed end-to-end by a background agent (fully AI built, not AI assisted).

Limitations: This is a customer story on AWS's own blog, with the scale and percentage figures self-reported by Abnormal AI and not independently verified.

Source: AWS Machine Learning Blog

Regional & Early Signals

09/10

Claude Code canary appears to route to Opus 5.2, with no Anthropic announcement

Developers say the underlying model slug already reads Opus 5.2 while the front-end name is unchanged.

IT Home reported on September 14th (Chinese-language source) that developers calling Opus 5 inside Claude Code saw behavior clearly different from web Opus 5. By inspecting requests and the /status output, they found the front-end name unchanged while the model slug behind it pointed to Opus 5.2 — suggesting Anthropic skipped 5.1. Testers ran an obscure-trivia probe prompt with web search disabled and reported that the two sides returned entirely different answers, which they read as evidence of different weights. Hands-on reports centre on faster generation and longer persistence on complex tasks, with fewer prompts asking the user to continue.

Limitations: The evidence is community self-testing relayed by a Chinese-language outlet; Anthropic has published no announcement or model card, and there is no official word on rollout scope or timing. The same report's claims about an internal model codenamed "Model 2" writing most of the company's code and replacing 85% of its research staff come from an unverified internal risk report and we could not confirm them.

Source: IT Home (Chinese-language source)

10/10

Ant Group releases SingProbe, an in-model safety guardrail with sub-0.5% decode overhead

It reuses internal signals already produced during inference to score risk while the answer is still being generated.

Leiphone reported on September 14th (Chinese-language source) that Ant Group's AI Security Lab released SingProbe during China's Cybersecurity Week. Today's applications typically screen inputs or outputs with a separate external guardrail model, which adds compute and deployment cost; checking only after a full answer is generated means risky content may already have reached the user, while checking more often adds more load. SingProbe instead runs the safety judgment alongside generation, emitting a continuous risk score through a lightweight detection module so the system can warn, halt generation or retry.

  • Overhead: measured at under 0.5% additional decode-stage cost in a Ling-3.0-flash production environment
  • Results: the team's evaluation says the flash version beats the selected public baselines on answer-safety classification and streaming safety detection, and roughly matches the reference baseline on hallucination detection
  • Medical: SingProbe-Med, with full intervention, corrected 25.03% of answers the baseline model originally got wrong on AntAngelMed-100B
  • Coverage: adapted to 29 mainstream open-source models including Ling-3.0, GLM-5.2/5.3, Qwen and DeepSeek V4, integrated with SGLang and vLLM, with code, models and benchmark open-sourced

Ant also released SingStreamBench, a benchmark for streaming generation that targets the exact moment a model turns from a normal answer to risky content, testing not just whether a guardrail catches the risk but whether it fires too early and how quickly it reacts.

Limitations: The evaluation was run by Ant's own team against "selected public baselines" with no third-party reproduction; the report gives neither the repository address and license nor the full list of the 29 supported models.

Source: Leiphone (Chinese-language source)

Generate one yourself with JIUFENG

Write a prompt in your browser and get an image — 1K, 2K or 4K output, up to 15 reference images. 50 free generations on sign-up, no credit card.

Generate free