NEWGPT-IMAGE-2 is now available in Jiufeng Hub Gallery!Try it out
In this article
AI Highlights

Anthropic's most capable model, Model 2, is internal-only

Key Takeaways
  • Anthropic reveals an internal-only model stronger than any public Claude, Alibaba ships Qwen-UI-Agent, and open-weight Kimi K3 nears Opus 5 through a harness.
jiufeng
August 20, 2026
30 min read
Anthropic's most capable model, Model 2, is internal-only

Overview

10 stories in this issue. The first 3 are today's priorities.

Hot model watch

  1. Top · Anthropic runs a stronger internal-only model, Model 2
  2. Top · Alibaba releases Qwen-UI-Agent, a foundation model for GUI agents
  3. Top · Open-weight Kimi K3 closes in on Opus 5 with a harness (Chinese-language source)
  4. OpenAI cuts GPT-5.6 Luna 80%, Replit bundles it into paid plans

Global AI news 5. Slack launches Code, putting coding agents in shared channels 6. Binance opens Agent OS to let AI agents trade 7. TrueFoundry open-sources TrueForge to sit beneath more agents 8. Terence Tao warns AI could trigger math's biggest crisis since Gödel

Regional & early signals 9. China's CAICT and Taobao Instant Retail issue a trust spec for retail agents (Chinese-language source) 10. Tencent's WorkBuddy and Baidu's Dazi lead China's desktop AI-office race (Chinese-language source)

AI signal map for 2026-08-20
AI signal map for 2026-08-20

Jiufeng graphic based on the sources cited in this issue.

Hot model watch

Anthropic runs a stronger internal-only model, Model 2

Anthropic's August risk report admits it runs an internal model more capable than any public Claude, codenamed Model 2, used only inside the company.

Per Anthropic's August 2026 Risk Report, the model sits in the Mythos class and is slightly stronger overall than Claude Mythos 5, scoring about 1.5 points higher on the internal capability index (AECI). The company leans on it heavily for coding, data generation, and research and engineering, sometimes through continuously running agents; the report says Claude now writes most of the code in Anthropic's production systems. Model 2 went through an internal review before deployment.

The report also notes Model 2 is only "slightly stronger" overall and is actually weaker than Mythos 5 in some areas; the 1.5-point gain is smaller than the earlier jump from Mythos Preview to Mythos 5, and far short of the Opus 4.6-to-Mythos leap. It is internal-only, with no public access or independent verification.

Source: The Decoder · Anthropic Risk Report

Alibaba releases Qwen-UI-Agent, a foundation model for GUI agents

Alibaba unveiled Qwen-UI-Agent, a real-world-centric foundation GUI agent spanning mobile, computer, web and DeepSearch, topping several flagships on a mobile benchmark.

According to its arXiv technical report (2607.28227), Qwen-UI-Agent operates on real devices, runs workflows across platforms, and combines GUI interaction with CLI execution for long-horizon tasks. On the reported benchmarks it scores 82.1% on MobileWorld, ahead of GPT-5.6 Sol, Claude and other flagship models.

The scores come from self-reported benchmarks in the technical report, without independent reproduction; release details such as weights and availability await official follow-up.

Source: arXiv

Open-weight Kimi K3 closes in on Opus 5 with a harness (Chinese-language source)

In a Prime Intellect nanoGPT speedrun of 18 frontier models, open-weight Kimi K3 paired with the Prime Agent Harness landed just 10 steps behind Claude Opus 5.

As reported by QbitAI, the experiment had each model autonomously tweak the optimizer, run training and compare loss, aiming to push a 124M-parameter GPT's validation loss below 3.28 in as few training steps as possible, with a 3290-step baseline. Kimi K3 with the Prime Agent Harness reached 2930 steps on the raw results page, beating GPT-5.6 Sol's 3042 and trailing Opus 5's 2920 by only 10. The 18 models also included Fable 5, Grok 4.5/4.6, GLM 5.2, Muse Spark 1.1/1.2, DeepSeek V4 Pro and Qwen 3.8.

The whole run was sandboxed to stop models from looking up answers. This is a single speedrun result relayed by QbitAI from Prime Intellect's experiment; the raw results page is not independently verified and does not equal a broad capability ranking.

Source: QbitAI — Chinese-language source

OpenAI cuts GPT-5.6 Luna 80%, Replit bundles it into paid plans

After OpenAI slashed its cheapest GPT-5.6 Luna by 80%, Replit built Free Mode on it so paid users can run routine agent tasks without spending credits.

Per RuntimeWire, Replit's Free Mode is powered by OpenAI's GPT-5.6 Luna and open to paid Core ($25/mo, $20/mo annual) and Pro ($100/mo, $95/mo annual) subscribers, letting routine prompts and edits skip the monthly credit meter; Core includes five-hour limits and up to 30 hours of monthly chat. OpenAI cut GPT-5.6 Luna's price by 80% on July 30th, and its docs now list $0.20 per million input tokens and $1.20 per million output.

Despite the name, Free Mode is available only on the paid Core/Pro plans and carries a five-hour window. RuntimeWire discloses it is built with Replit, a relationship with the company it covers.

Source: RuntimeWire · OpenAI

Global AI news

Slack launches Code, putting coding agents in shared channels

Salesforce's Marc Benioff launched Slack Code, letting teams assign work to coding agents from Anthropic, GitHub, Cognition and Vercel inside a Slack channel.

According to RuntimeWire, Benioff launched Slack Code on August 20th, so software teams can hand tasks to those four vendors' coding agents without leaving a Slack channel; he called it "multiplayer coding" and demoed it at Dreamforce 2026, with all four integrations live that day. Salesforce positions itself as the coordinator rather than the model provider.

So far this is a launch-day demo and integration; real collaboration quality, availability and stability remain unproven, and model capability still depends on each vendor's agent.

Source: RuntimeWire · Marc Benioff on X

Binance opens Agent OS to let AI agents trade

Binance launched Agent OS, letting AI agents built with ChatGPT, Claude Code or Cursor analyze markets and trade on users' behalf, while leaving oversight largely to users.

Per TechCrunch, Binance, with over 300 million registered users, launched Agent OS on Thursday, letting developers connect AI apps and agents to its financial infrastructure, including Binance APIs, the Wallet Agentic Hub, x402 transaction verification and payment APIs, and Skill Hub, plus new Model Context Protocol (MCP) support. It works with OpenAI's ChatGPT and Codex, Anthropic's Claude Code and Cursor; users can authorize agents to read market data, view accounts and execute trades.

TechCrunch notes that as AI shifts from chatbots to action-taking agents, Binance puts much of the responsibility for keeping agents in check on users, who decide what each agent can access.

Source: TechCrunch

TrueFoundry open-sources TrueForge to sit beneath more agents

TrueFoundry open-sourced TrueForge, an MIT-licensed agent runtime it claims completes tasks 30-75% cheaper than Claude Managed Agents — though its biggest saving comes from switching models.

Per RuntimeWire (citing VentureBeat), TrueFoundry founders Nikunj Bajaj, Abhishek Choudhary and Anuraag Gutgutia released TrueForge, which lets developers build and run agents with their own choice of models, tools and infrastructure. The strategy is to give away the runtime while placing TrueFoundry's paid gateway beneath deployed agents; Gutgutia said "all that traffic should still be flowing through our gateway."

RuntimeWire notes the largest chunk of the claimed 30-75% saving actually comes from switching to a cheaper model, not the runtime itself; the open-source runtime is also a funnel for the paid gateway.

GitHub - truefoundry/trueforge: The open-source agent harness - the runtime layer that turns an LLM into a working agent.
GitHub - truefoundry/trueforge: The open-source agent harness - the runtime layer that turns an LLM into a working agent.

Image source: GitHub; mirrored on Jiufeng R2.

Source: RuntimeWire · TrueForge (GitHub)

Terence Tao warns AI could trigger math's biggest crisis since Gödel

Fields medalist Terence Tao argues in a new essay that AI could push mathematics into a crisis on par with 1900, testing not mathematical truth but math's values and practices.

Per The Decoder, Tao writes in an essay for the 2026 International Congress of Mathematicians that the field should stop debating what AI can do and instead confront a long-ducked question: what are the goals of mathematical research? He draws a parallel to the 1900-1930 foundational crisis, when Russell's paradox and Gödel's incompleteness theorems forced mathematicians to make implicit assumptions explicit, producing a rigorous framework that held for a century. Today, he argues, the stress test has shifted to "the largely implicit framework of mathematical values and practices" — what counts as a contribution.

This is Tao's own argument and outlook, not an empirical result; the essay poses questions and analogies, and whether AI brings a "golden age or a deep crisis" is left open.

Source: The Decoder · arXiv

Regional & early signals

China's CAICT and Taobao Instant Retail issue a trust spec for retail agents (Chinese-language source)

China's CAICT and Taobao Instant Retail released the first trustworthiness spec for instant-retail AI agents, setting accuracy and interface-success thresholds.

Per Leiphone, the spec was released on August 18th at the 18th plenary of the AI Industry Alliance of China, initiated by its safety-governance committee and drafted by CAICT and Taobao Instant Retail, with Ant, Baidu, StepFun, MiniMax and Tongdun contributing. It spans three capability dimensions (completeness, execution accuracy, execution efficiency) and six domains (data, access and runtime security, vulnerability management, service management, tech ethics), and requires simple-task accuracy above 95%, complex-task above 90%, and interface success above 99%, with at least 100 test items each. Taobao Instant Retail opened MCP access on August 5th.

This is an industry self-governance spec, not a binding standard; adoption and testing rigor depend on each vendor. Only a Chinese-language source is available.

Source: Leiphone — Chinese-language source

Tencent's WorkBuddy and Baidu's Dazi lead China's desktop AI-office race (Chinese-language source)

Third-party MAU rankings put Tencent's WorkBuddy first at 11.15M and Baidu's Dazi second at 6.74M, with Dazi growing over tenfold month-on-month.

Per TMTPost, citing the latest AI Product Rankings, WorkBuddy leads desktop AI-office agents with 11.15M monthly active users, while Baidu's Dazi is second at 6.74M with 1063.79% month-on-month growth; in under six months, China has seen 20-plus desktop agent products launch. The article says WorkBuddy began as a weekend build by a few Tencent PMs on CodeBuddy's infrastructure and Agent SDK after Claude Cowork shipped, then opened up after the Spring Festival.

The MAU figures come from the third-party AI Product Rankings and are not independently verified; with Alibaba, ByteDance and more office-software vendors entering, the ranking is still early. Only a Chinese-language source is available.

Source: TMTPost — Chinese-language source