In this article
AI Highlights

Z.AI Confirms Ox Alpha, Weight Release Planned

Key Takeaways
  • Z.AI confirms Ox Alpha as GLM
  • Kimi K3 seeks cloud distribution as inference chips and robotics face deployment tests.
jiufeng
August 26, 2026
22 min read
Z.AI Confirms Ox Alpha, Weight Release Planned

Overview

8 stories in this issue. The first 3 are today's priorities.

Popular Model Updates

  1. Top · Ox Alpha Confirmed as GLM, Weights Planned for Release
  2. Top · Moonshot Seeks 30% of Kimi K3 Cloud Sales
  3. Top · AI Models Still Stumble on Altered and Visual Puzzles

Global AI News 4. WhatsApp Tests On-Device Scam Alerts 5. OpenAI Reports Lower Latency for Jalapeno 6. Cerebras Details CS-4 and the Nexus Platform

Regional and Early Signals 7. Xiaomi Unveils Three On-Device AI Chips 8. WRC Shows Robots Performing Practical Tasks

AI signal map for 2026-08-26
AI signal map for 2026-08-26

Jiufeng graphic based on the sources cited in this issue.

Ox Alpha Confirmed as GLM, Weights Planned for Release

Z.AI confirmed that the anonymous Ox Alpha model is a new GLM iteration and planned to release its weights on August 26.

Ox Alpha previously appeared on OpenRouter as a free preview. A public investigation used tokenizer, multimodal-encoder and API-behavior evidence to identify it as part of the GLM-5 generation, most likely the original GLM-5 or a variant; Z.AI subsequently confirmed the model’s origin to Bloomberg.

The weights, license and hardware requirements were not available when the reports were published, leaving commercial modification and self-hosting terms unresolved. The public investigation was a pre-release third-party identification and does not replace independent verification after the weights become available.

ox-alpha-identification-public/investigations/independent-verification/REPORT.md at main · LuD1161/ox-alpha-identification-public
ox-alpha-identification-public/investigations/independent-verification/REPORT.md at main · LuD1161/ox-alpha-identification-public

Image source: GitHub; mirrored on Jiufeng R2.

Source: RuntimeWire · Public investigation · TechCrunch

Moonshot Seeks 30% of Kimi K3 Cloud Sales

Moonshot is discussing Kimi K3 distribution with Microsoft, Amazon and Google, seeking up to 30% of related service revenue.

The talks cover Microsoft Azure, Amazon Web Services and Google Cloud, with data access, revenue allocation and token-usage auditing still unresolved. Kimi K3 has 2.8 trillion total parameters, and its weights were released on July 27; cloud distribution would address enterprises unable to self-host a model of that scale.

The negotiations are preliminary and may end without agreements, so 30% is a requested ceiling rather than a signed rate. The technical report documents the model architecture but does not establish enterprise demand or finalized commercial terms.

Source: RuntimeWire · Kimi K3 technical report · Moonshot AI

AI Models Still Stumble on Altered and Visual Puzzles

Model performance on puzzles has advanced quickly, but altered riddles and visual-spatial tasks still expose failures.

A Columbia University team found in late 2024 that leading models solved only about 18% of New York Times Connections puzzles; by early 2025, some models were solving them nearly perfectly. The tests collected by MIT Technology Review show that subtle wording changes can still derail familiar reasoning patterns, while visual puzzles remain difficult.

The figures come from different dates and test settings and should not be treated as one current leaderboard. The linked SRBench repository uses the MIT License, and success on individual puzzles is not a general-intelligence measurement.

Source: MIT Technology Review · SRBench license

Global AI News

WhatsApp Tests On-Device Scam Alerts

WhatsApp is testing an opt-in Scam Alert that classifies suspicious messages from unknown contacts on the device.

The feature downloads a small machine-learning model and uses conversational structure and language signals to assess messages; warnings remain hidden from the sender, and recipients can block, report, continue or trust the chat. Aggregate metrics travel through anonymous credentials and Oblivious HTTP to a confidential virtual machine, where minimum cohort thresholds and differential privacy are applied; clients verify model signatures, freshness and SHA-256 hashes against a transparency ledger.

This remains a limited test, and users may optionally share the five most recent received messages from trusted chats to improve it. Meta plans to publish confidential-VM binaries and privacy-related source components for review, but those materials were not all available in the reporting.

Source: InfoQ · InfoQ Chinese

OpenAI Reports Lower Latency for Jalapeno

OpenAI reports that its Jalapeno inference chip delivered 1.7x to 3.6x lower end-to-end latency than selected Nvidia systems.

The company reports 2.1x to 4.1x higher interactivity and 1.5x to 1.9x more AI work per watt at peak throughput across three large open models. OpenAI plans to deploy Jalapeno in production by the end of 2026 to add inference capacity beyond outside chip suppliers.

These are OpenAI-run measurements on engineering hardware, without independent replication or production-scale evidence. The live report shows an August 25 date, while its supplied metadata says August 3, and no public revision history explains the discrepancy.

Source: RuntimeWire · OpenAI engineering report · OpenAI announcement

Cerebras Details CS-4 and the Nexus Platform

Cerebras detailed its CS-4 inference system at Hot Chips 2026 and previewed the CS-5 and CS-6 roadmap.

CS-4 is the first system built on the Nexus rack-scale platform, which supports three Wafer-Scale Engines in modular compute backpacks with dedicated power, cooling and I/O. Cerebras says the design allows components to evolve independently and provides a path to double token-generation speed annually for several years.

The disclosed performance and efficiency claims come from Cerebras and lack independent testing in the supplied material. CS-5 and CS-6 remain roadmap previews without verifiable delivery dates or production deployment results.

Source: Cerebras

Regional and Early Signals

Xiaomi Unveils Three On-Device AI Chips

Xiaomi introduced the O3, O100 and D100 on-device AI chips but did not define what qualifies a handset as an “AI phone.”

Pandaily reports that Xiaomi announced the O3, O100 and D100 at its Xuanjie briefing, with all three described as on-device AI chips. The report also says Xiaomi did not specify the capabilities or thresholds required for its AI-phone label.

The available material provides no process node, compute, power, model-compatibility or device-launch figures, preventing a practical comparison of the three chips.

Source: Pandaily

WRC Shows Robots Performing Practical Tasks

Robots at the 2026 World Robot Conference demonstrated cooking, parcel sorting and other practical tasks that one report views as a change in industry narrative.

Chinese-language reporting describes roughly 60,000 square meters of exhibition space and more than 1,000 robots performing tasks including cooking, ball games, coffee service, parcel sorting and pet-litter cleanup. The article’s author interprets the demonstrations as a narrative shift in embodied AI from staged feats toward practical work.

The evidence comes from one Chinese-language conference report focused on现场 demonstrations and the author’s industry assessment, without cross-environment success rates, continuous-operation duration or deployment-scale data.

Source: Leiphone, Chinese-language source