Overview
8 stories in this issue. The first 3 are today's priorities.
Hot model watch
- Top · PerceptionBench: no frontier model tops 60% at visual perception
- Top · Kimi Work quietly attaches raw records from five recent sessions to feedback
- Top · DeepSeek Harness sparks a plugin boom as V4-Pro prices rise Aug 16
- Study disputes claims that autonomous AI research is within reach
Global AI news 5. ChatGPT's desktop app hides a "finish setup, get credits" onboarding test 6. Cloudflare adds tracing for agents
Regional & early signals 7. Sota Wujie claims a 4D physical world model deployed in European supermarkets (Chinese-language source) 8. Lumos MOS2: a wheeled-arm robot with 50 kg dual-arm payload (Chinese-language source)

Jiufeng graphic based on the sources cited in this issue.
Hot model watch
PerceptionBench: no frontier model tops 60% at visual perception
Moonshot's new benchmark isolates "seeing" from reasoning — and every frontier model scores below 60%.
Moonshot AI (the team behind Kimi) has released PerceptionBench, a benchmark that isolates visual perception from logical reasoning and outside knowledge. Every question can be answered just by looking at the image, with no reasoning or external knowledge required; the categories are built from real-world perception errors, and each question is split into perception-only sub-questions to pinpoint exactly which visual ability fails. The dataset and evaluation code are open-sourced on GitHub (MoonshotAI/PerceptionBench).
Across the leading models — including GPT-5.6 Sol, Kimi K3 and Claude Fable 5 — none reaches 60% accuracy, and GPT-5.6 Sol leads only by a narrow margin. The authors conclude that many failures treated as "reasoning errors" actually occur at the image-reading stage, echoing earlier studies that even top models stumble on basic visual tasks.

Image source: moonshot; mirrored on Jiufeng R2.
Source: The Decoder · PerceptionBench blog · GitHub
Kimi Work quietly attaches raw records from five recent sessions to feedback
A RuntimeWire teardown found that submitting feedback in Kimi Work uploads raw records from your five most recently updated agent sessions — unlisted.
RuntimeWire reverse-engineered Kimi Desktop 3.1.5 and its bundled Daimon 0.5.49 service and found that when a user submits feedback in Kimi Work, the app also sorts all Work conversations by updatedAt, packages bounded raw-record archives from the five most recently updated agent sessions, and uploads the ZIPs to /file/upload_simple alongside a diagnostic log. The feedback form neither lists those attachments nor lets the user choose which sessions are included. In a test with the uploads blocked by Windows Firewall, Kimi selected five distinct conversation IDs that each reached the HTTP upload stage before failing with ERR_NETWORK_ACCESS.
RuntimeWire notes that Kimi Work handles private files, source code and shell output, while the company's public help page says in-product reports attach only "device and account context" and does not mention raw records from recent sessions. This is RuntimeWire's own finding from reverse engineering and testing, with no vendor response yet.
Source: RuntimeWire · Kimi help page
DeepSeek Harness sparks a plugin boom as V4-Pro prices rise Aug 16
Days after DeepSeek open-sourced its Harness runtime, GitHub already carries 700+ dsh-plugin repos — while V4-Pro API prices rise on August 16.
DeepSeek released its open-source coding-agent runtime, Harness v0.1, on August 13 under the MIT license (developer preview), letting developers swap models, tools and runtimes and pushing into the agent runtime layer held by Anthropic's Claude Code and OpenAI's Codex. Chinese outlet QbitAI reports that within days GitHub already hosts 700+ repos tagged dsh-plugin: browsers, long-term memory, databases and vision models, plus electronic pets, a pixel whale and 18 mini-games. Plugins like DSH Better Sidebar fold a file manager, code editor, real terminal and Git panel into one sidebar, while dsh-plugin-claude-bridge and dsh-claude-move migrate Claude Code's memory, skills, CLAUDE.md and old sessions into Harness.
The same launch brought the general-availability DeepSeek-V4-Pro (the official changelog confirms it hit the app, web and API on August 13), whose API prices rise on August 16. RuntimeWire notes the increase raises the cost of exactly the long, cache-heavy workflows Harness is built for. Harness remains a v0.1 developer preview, and the plugin counts are community-tracked on GitHub.
Source: RuntimeWire · QbitAI (Chinese-language source) · GitHub
Study disputes claims that autonomous AI research is within reach
Given six days, $3,000 in credits and GPUs, agents on Claude Opus 4.8 and GPT-5.6 Sol still failed the parts of research that matter.
A new paper from Princeton and the UK AI Security Institute (AISI) tests the claim that models can now do AI research on their own. Using a method they call "Shadow Evaluation," agents built on Claude Opus 4.8 and GPT-5.6 Sol each received the core research question of an unpublished paper, six days, $3,000 in API credits and GPU access to independently run the research and write a paper — then the original authors, who had spent months on the same question, graded them. The finding: today's frontier models can handle research engineering but fail at the parts that actually decide a study's outcome, often running undersized experiments.
The authors argue prior evidence for automated AI research has been weak — existing evaluations either test narrow, verifiable tasks or push AI-written papers through peer review, which they call "overstretched, stochastic," with "poor review quality." The result runs counter to lab claims, including Anthropic's June post "When AI Builds Itself." It is a single study with a specific methodology.
Source: The Decoder · Anthropic, "When AI Builds Itself"
Global AI news
ChatGPT's desktop app hides a "finish setup, get credits" onboarding test
RuntimeWire found an experiment-gated ChatGPT desktop flow that offers a remotely configured credit bonus for finishing the agent-permission setup.
RuntimeWire reverse-engineered ChatGPT desktop version 26.803.81509 and found a complete, experiment-gated "conversational onboarding" system: it asks the user's role, presents role-specific starter tasks, handles permissions and app connections, executes the chosen task and records the outcome. A separate remote configuration supplies a credit_amount; when it is greater than zero for a ChatGPT-authenticated account, the interface promises "Finish set up and get [X] credits," labels the reward an "Onboarding bonus" and says the credits will be "added to your balance," while skipping warns the offer is forfeited.
RuntimeWire reads this as OpenAI preparing to use credits as an onboarding incentive — paying users to complete the permission-heavy step that turns ChatGPT from a chatbot into a desktop agent. It is an experiment-gated, remotely configured feature discovered in the binary (the credit amount can be zero), not an announced product.
Source: RuntimeWire · OpenAI release notes
Cloudflare adds tracing for agents
Cloudflare now adds agent-level spans to Workers traces and replays sessions turn by turn — but with truncation limits and uneven payload defaults.
Cloudflare has launched agent tracing, adding spans for agent invocations, model calls, tool runs and approvals on top of its existing Workers traces, so a multi-step agent session can be replayed turn by turn for debugging.
Per InfoQ, the documentation cautions that these traces carry truncation limits and that payload-capture defaults are uneven across steps, so traces may be incomplete. The report currently rests on a single outlet.
Source: InfoQ
Regional & early signals
Sota Wujie claims a 4D physical world model deployed in European supermarkets (Chinese-language source)
A four-month-old startup says it will put 1,000+ embodied robots on real supermarket shelves over three years with Europe's largest grocery group.
Per Leiphone, embodied-AI startup Sota Wujie (索塔无界) announced a strategic partnership with Europe's largest supermarket group to deploy more than 1,000 embodied robots in real retail settings over the next three years. The company says it builds not hardware but a "native physical brain" — a World-Action model unifying 4D spatial representation, physical constraints, spatiotemporal reasoning and action output — and argues a physical foundation model should first master one high-complexity scenario (supermarket shelves) before generalizing. Chief scientist Liu Zhe's model team is drawn from HKU, Tsinghua and HKUST.
These figures come from the vendor and a single Chinese-language outlet; there is no third-party verification of deployment scale, success rate or quantified performance, and framing such as "world's first" is the company's own claim.
Source: Leiphone (Chinese-language source)
Lumos MOS2: a wheeled-arm robot with 50 kg dual-arm payload (Chinese-language source)
Lumos Robotics unveiled MOS2 on Aug 14, a heavy-duty wheeled-arm robot aimed at industrial lifting and mobile work.
Per Leiphone, Lumos Robotics on August 14 unveiled the Lumos MOS2, a wheeled-arm embodied robot pitched as a "heavy-load AI Worker" for industry: a 50 kg dual-arm payload, roughly 1.65 m tall with a ~0.5 m² footprint and a ~0.8 m aisle-width requirement, and 22 degrees of freedom (adding a two-DOF head and an omnidirectional base). Its sensing stack includes dual multi-line lidar on the base, dual wrist cameras, dual six-axis force sensors and a binocular head camera. The company positions it for flexible tasks with frequent line changeovers and complex assembly.
These are vendor launch figures reported by a single Chinese-language outlet; payload, DOF and other specs come from the company, with no third-party testing or quantified on-line throughput data yet.
Get the latest AI model insights and tutorials from Jiufeng.
Explore more


