In this article
AI Highlights

Kimi K3 Legal Model Tenet Enters Research Preview

Key Takeaways
  • Harvey previews Kimi K3-based Tenet
  • Codex tests Luna Reserve, while Ox Alpha, agent governance, and regional products advance.
jiufeng
August 24, 2026
27 min read
Kimi K3 Legal Model Tenet Enters Research Preview

Overview

10 stories in this issue. The first 3 are today's priorities.

Popular Model Updates

  1. Top · Harvey Tenet Post-Trains Kimi K3 for Legal Agents
  2. Top · Codex Tests a Luna Reserve After Premium Limits
  3. Top · Fable 5 Reaches 11% of Anthropic Spend in Ramp Data

Global AI News 4. Anonymous Ox Alpha Offers a One-Million-Token Window 5. Netflix Open-Sources an Agentic Causal Workflow 6. Replit Packages Seven Agent Governance Features 7. An agent.md Template Enforces Test-First Bug Fixes

Regional and Early Signals 8. Wan3.0 Generates 30-Second Video from Documents 9. Shanghai AI Film Platform Adds 50 DGX Spark Systems 10. LEAPTIC Cube Records 8K Video in a 55.9-Gram Body

AI signal map for 2026-08-24
AI signal map for 2026-08-24

Jiufeng graphic based on the sources cited in this issue.

Harvey Tenet Post-Trains Kimi K3 for Legal Agents

Harvey has post-trained Kimi K3 into Tenet for long-horizon legal-agent work, with access currently limited to a research preview.

Harvey and Fireworks used asynchronous reinforcement learning with synthetic data, public legal data, and human-expert data; Harvey says no customer data was used. On Harvey’s Legal Agent Benchmark, Tenet completed nearly twice as many held-out tasks as the base Kimi K3 and 20% more LAB: Contracts tasks, raising all-pass rates by 9 and 2 percentage points respectively.

Tenet is not deployable yet, and no production service or weights have been released. Most results are self-reported, while the public Redline Bench leaderboard does not contain a corresponding result for independent review.

crosbylegal/RedlineBench · Datasets at Hugging Face
crosbylegal/RedlineBench · Datasets at Hugging Face

Image source: huggingface; mirrored on Jiufeng R2.

Source: MarkTechPost · GSPO paper · Redline Bench

Codex Tests a Luna Reserve After Premium Limits

Codex client code indicates that some Plus and Pro users may receive a separate Luna allowance after advanced-model access is constrained.

RuntimeWire’s reverse engineering of the production Codex app.asar found the gpt-reserve identifier, a reserve_enabled feature gate, Plus/Pro eligibility checks, and dedicated metering logic. The client preserves the user’s original model selection and restores it after Reserve ends; OpenAI’s documentation prices GPT-5.6 Luna at $0.20 per million input tokens and $1.20 per million output tokens.

OpenAI has not publicly announced Luna Reserve. The evidence establishes that feature-gated test code exists, but it does not confirm rollout scope or a release date.

Source: RuntimeWire · OpenAI model documentation

Fable 5 Reaches 11% of Anthropic Spend in Ramp Data

A Ramp spending sample suggests Anthropic customers are favoring cheaper models, with Fable 5’s share flattening after launch.

IT Home’s Chinese-language summary of Financial Times reporting says Ramp data from about 70,000 companies put Fable 5 at roughly 11% of Anthropic-related spending more than two months after release. The report attributes the pattern to Fable 5’s price and older models being sufficient for many business tasks; U.S. access resumed on July 1 after government approval.

The sample covers corporate payments visible to Ramp, not all Anthropic API, subscription, or cloud-channel usage. Hacker News comments document user concerns about pricing and allowance rules, but community discussion is not a substitute for complete adoption data.

Source: IT Home — Chinese-language source · Hacker News

Global AI News

Anonymous Ox Alpha Offers a One-Million-Token Window

Ox Alpha provides free access to a frontier-class coding model, but its developer, data destination, and long-term availability remain undisclosed.

Available through OpenRouter and the open-source OpenCode terminal agent since August 20, Ox Alpha has a 1,048,576-token context window and a 131,072-token output cap. It accepts text, images, and video but rejects audio; OpenCode advertised capacity of 100 trillion tokens per day.

OpenRouter indicated that free access would run through August 24, while OpenCode’s notice pointed to about August 27. No company has claimed the model, and the reporting could not establish where submitted code is sent, making it unsuitable for confidential code.

Source: SiliconANGLE

Netflix Open-Sources an Agentic Causal Workflow

Netflix’s oci-agent uses executor and reviewer agents to iterate on observational causal analyses, with a lightweight version available on GitHub.

The workflow turns an analysis plan into a specification, fills and runs a Jupyter Notebook, then assigns one of three ratings: not_satisfactory, satisfactory_with_caveats, or fully_satisfactory. In a case study connecting exposure to new entertainment formats with two-month retention, a Claude-run linear regression served as the baseline; oci-agent estimated only 25% of that effect and flagged early-adopter bias and a failed placebo test.

Netflix evaluated the workflow on ACIC competition data but did not publish one aggregate score covering every task. Human supervision remains part of the process, with analysts responsible for higher-level problem formulation and final judgment.

Source: InfoQ Chinese — Chinese-language source · InfoQ · oci-agent

Replit Packages Seven Agent Governance Features

Replit has grouped lower-cost chat, scheduled routines, portable instructions, live steering, model policies, and security checks into an Agent update.

The August 21 bundle covers Free Mode, repeatable routines, reusable instructions, in-run guidance, enterprise model controls, and a Level 3 security scan comprising three checks across seven releases. The package is designed to turn one-off prompt sessions into workflows that can be reused, scheduled, and constrained by enterprise policies.

Some components had already received separate rollouts, so the weekly release primarily shows how they fit together. The available material does not specify plan availability, usage limits, or security-scan coverage for every feature.

Source: RuntimeWire · Raouf Chebri

An agent.md Template Enforces Test-First Bug Fixes

A programming-agent template explicitly requires a failing test before a bug fix and a passing test afterward.

The agent.md instructs an assistant to write a test, observe it fail, implement the fix, and confirm the test passes; the linked TDD skill follows the same sequence. The article and discussion cover self-review, iterative implementation, and moving general coding standards into separate files to avoid loading irrelevant context.

No controlled results or before-and-after defect rates are provided. Hacker News commenters questioned whether the rules measurably improve model iteration, and one participant reported that the process can occasionally produce low-value tests.

Source: Original article · Hacker News · TDD skill

Regional and Early Signals

Wan3.0 Generates 30-Second Video from Documents

Alibaba Cloud has launched Wan3.0 with 30-second video generation and input support for five common document formats.

Wan3.0 accepts doc, xls, ppt, pdf, and md files and is available through Model Studio, Wanxiang, Qwen AI, and other Alibaba services. API pricing is RMB 0.30, RMB 0.60, and RMB 1.20 per second for 480P, 720P, and 1080P output, with a 30% discount from August 24 through September 23 on two platforms.

This item is based on a Chinese-language regional report relaying Alibaba Cloud’s announcement. No standardized video scores, generation speeds, or document-reconstruction accuracy figures were provided, while reported quality improvements remain qualitative company claims.

Source: IT Home — Chinese-language source

Shanghai AI Film Platform Adds 50 DGX Spark Systems

Shanghai Film’s Haopu community and Kuaizi.ai have expanded the Haozhen production platform with local compute and an annual trillion-token subsidy program.

The project adds 50 Nvidia DGX Spark systems to an existing hundred-card-scale film-production compute pool. It targets collaborative work across production, writing, art, and post-production, emphasizing private environments, predictable compute access, and centralized asset management.

This signal currently has only Chinese-language regional reporting, and “trillion-level” describes the subsidy allowance rather than measured consumption. The claim that it is China’s first industrial AI film-production system comes from the launch positioning and lacks independent industry-wide verification.

Source: APPSO/ifanr — Chinese-language source

LEAPTIC Cube Records 8K Video in a 55.9-Gram Body

Photon Leap demonstrated an 8K thumb-size action camera with voice control, subject tracking, and multi-device coordination at BIRTV 2026.

LEAPTIC Cube weighs 55.9 grams, uses a 1/1.3-inch sensor, records up to 8K at 30 frames per second, and includes a dual-screen design. Its demonstration featured the Moko voice assistant, real-time subject tracking, and a low-light algorithm, including voice-triggered recording controls.

The product currently has only Chinese-language regional coverage, with performance details derived from the exhibition and vendor material. Battery life, bitrate, storage, mass-production timing, and independent image-quality results were not supplied; “world’s first” is a vendor claim.

Source: Leiphone — Chinese-language source