AI Highlights

OpenAI opens Agents API beta with a managed Codex harness

Key Takeaways
  • •OpenAI opens a public beta of its Agents API, running the Codex harness as a managed service
  • •plus Fable 5.1 Build Days and Google's Dreambeans rollout.
jiufeng
September 11, 2026
36 min read
In this article

Overview

9 stories in this issue. The first 3 are today's priorities.

Hot model updates

  1. Top · OpenAI opens Agents API beta with a managed Codex harness
  2. Top · Claude community kicks off Fable 5.1 Build Days in cities worldwide

Global AI news

  1. Top · Google opens Dreambeans to all US adults, dropping the subscription gate
  2. Universal Music and ElevenLabs build an AI remix platform on licensed music
  3. SageMaker adds prefix-aware routing, cutting P50 TTFT by up to 77%
  4. Bedrock Knowledge Bases adds Marengo Embed 3.0 for semantic video search
  5. Redis LangCache semantic cache claims up to 90% lower LLM API costs
  6. NVIDIA's BioIR lifts Boltz-2 folding throughput 2.90x on 8xH100

Regional and early signals

  1. Ant Digital Technologies launches Agentar for finance with ten preset agent expert teams
AI signal map for 2026-09-11

Jiufeng graphic based on the sources cited in this issue.

Hot model updates

01/09

OpenAI opens Agents API beta with a managed Codex harness

OpenAI has turned the harness and cloud infrastructure behind Codex into a managed service: developers specify the task, model, tools and execution environment, and OpenAI runs the agent loop on its own infrastructure.

OpenAI launched the Agents API on September 10th for developers building long-running agents, and it is live for all developers in public beta. The API coordinates model calls, tool use and context while an agent works. The OpenAI Developers account summed it up this way: OpenAI handles orchestration, long-running sessions and context management. OpenAI says scaling Codex and ChatGPT for Work showed that long-running agents need a harness that manages context, uses tools efficiently and coordinates subagents, plus infrastructure that keeps them running reliably for days.

  • Core concepts: the official docs organize the API around four concepts, including Agent (the model, instructions, tools and MCP servers available to it), Environment (an optional sandbox where the agent accesses files, loads skills and runs commands) and Session (a durable agent instance)
  • Where code runs: an OpenAI-hosted sandbox, the developer's own infrastructure, or a partner sandbox
  • Billing: per the official docs, model usage is billed at the selected model's API rates, OpenAI tools at standard rates, and the hosted sandbox at standard container rates
  • Open source: the core Codex harness remains open source on GitHub, so developers can inspect the logic that coordinates models, tools and context

Limitations: this is still a public beta. According to MarkTechPost, data stays US-only and Zero Data Retention is not supported.

Agents API | OpenAI API

Image source: OpenAI Developers; mirrored on Jiufeng R2.

Source: OpenAI · Agents API docs · OpenAI-hosted sandbox docs · Codex (GitHub) · RuntimeWire · MarkTechPost

02/09

Claude community kicks off Fable 5.1 Build Days in cities worldwide

From September 11th to 25th, Anthropic is promoting Fable 5.1 Build Days run by local community organizers, inviting people to bring problems, ideas or unfinished projects.

According to RuntimeWire, Claude Community organizers are holding the buildathons from Austin to Melbourne. The image attached to the official announcement names Nairobi, Stockholm, Cape Town, Oslo and Mexico City, and the Claude Community calendar lists events hosted by local community members around the world. The invitation is deliberately loose: attendees can bring a problem or an idea, or simply come to see what others build. RuntimeWire argues this lowers the barrier for a model Anthropic otherwise pitches at the demanding end of coding, research and enterprise work.

Limitations: these are community-hosted events, not a change to Fable 5.1's capabilities, pricing or availability. RuntimeWire notes the payoff depends on whether short experiments turn into recurring API usage, products and workplace adoption.

Source: RuntimeWire · Claude on X

Global AI news

03/09

Google opens Dreambeans to all US adults, dropping the subscription gate

Google Labs' AI story-feed app has dropped the subscription gate from its June debut and builds a daily set of illustrated stories from the Google services a user chooses to connect.

Google opened Dreambeans to all users aged 18 and older in the US on Thursday, September 10th, and The Verge reported the expanded rollout the same day. Google Labs product manager Gozde Oznur introduced the app on June 3rd, initially limiting access to eligible Google AI Ultra subscribers in the US. Each day it assembles a finite collection of illustrated stories from the services a user connects, which can include Gemini, Gmail, Calendar, Google Photos, YouTube and Search history; results range from shopping and travel recommendations to event reminders and hobby-related suggestions.

The app is available on Google Play and Apple's App Store, requires a personal Google account, and only works once at least one supported service is connected.

Limitations: access is limited to US users aged 18 and over with personal accounts. RuntimeWire notes that Dreambeans acts without a prompt by combining personal data across services, so broad access will test whether that convenience outweighs the permissions it requires.

Source: The Verge · RuntimeWire · Google's June launch post

04/09

Universal Music and ElevenLabs build an AI remix platform on licensed music

Universal Music Group is developing, through a multiyear licensing agreement with ElevenLabs, an AI platform that lets users remix, mash up and reinterpret tracks from its licensed catalog.

According to The Verge, Universal Music Group (UMG) announced the platform on Thursday. Users will be able to draw on its licensed catalog to create remixes, mashups and new takes on tracks. The platform is being developed through a multiyear licensing agreement with ElevenLabs, which specializes in AI voice and music generation, and artists can choose whether to participate. It is another AI deal for the label: UMG is developing a separate AI music platform with Udio and has struck AI licensing deals with Spotify, Nvidia and Klay.

Limitations: The Verge describes it as an upcoming platform, so it is not yet usable, and participation is up to each artist. This item has a single source, The Verge.

Source: The Verge

05/09

SageMaker adds prefix-aware routing, cutting P50 TTFT by up to 77%

Requests that share a prompt prefix are sent to the same instance so its KV cache stays warm.

AWS announced that Amazon SageMaker Inference now offers prefix-aware routing, which routes requests sharing the same prompt prefix to the same instance. The starting point is that prompts usually combine a fixed part (instructions, reference documents, conversation history) with variable user input; in AWS's customer-service example, the instructions might be 3,000 tokens and the customer's question 50. Serving frameworks such as vLLM and TensorRT-LLM cache the key-value pairs computed for prefixes they have seen, so a new request only processes the new tokens at the end, which is known as prefix caching. In AWS benchmarks on Llama 3.1 70B, prefix-aware routing cut P50 time-to-first-token by up to 77% and raised the KV cache hit rate.

Limitations: the routing only helps requests that share the same prompt prefix, and 77% is an "up to" figure from AWS's own testing, with no third-party reproduction yet.

Source: AWS Machine Learning Blog

06/09

TwelveLabs' multimodal embedding model is now generally available in Amazon Bedrock Knowledge Bases, making video, image and audio searchable by meaning.

AWS announced the general availability of TwelveLabs Marengo Embed 3.0 as an embedding model in Amazon Bedrock Knowledge Bases. AWS pitches it at media, sports analytics, education, security and retail teams that need to find specific moments in hours of footage using natural language, with the example query "show me the penalty kick in the second half". AWS says building semantic video search yourself usually means stitching together transcription services, frame extraction pipelines, embedding models, vector databases and synchronization logic.

  • Model: Marengo Embed 3.0 jointly encodes video, audio, images and text into a compact 512-dimensional vector
  • Formats: video (MP4, MOV), images (JPEG, PNG) and audio tracks
  • Managed scope: Bedrock Knowledge Bases handles storage, ingestion, embedding, re-ranking and retrieval, with native connectors for Amazon S3, SharePoint, Confluence and more

Limitations: this comes from AWS's own blog and is a vendor account, with no independent evaluation yet.

Source: AWS Machine Learning Blog

07/09

Redis LangCache semantic cache claims up to 90% lower LLM API costs

A managed semantic cache between the application and the model returns a stored answer when the same question arrives in different words, instead of paying for another full model call.

According to MarkTechPost, Redis LangCache is a fully managed semantic caching service that matches incoming prompts against previously answered ones by meaning rather than exact text, and returns the stored response when a close enough match exists. The pitch starts from how support assistants and RAG pipelines field the same intents thousands of times a day in different phrasing, while most stacks treat each phrasing as a fresh, fully billed request. Redis reports API cost savings of up to 90% and cache-hit responses up to 15x faster than re-querying the model.

  • Availability: currently offered as a public preview on Redis Cloud
  • Access: REST API with Python and JavaScript SDKs

Limitations: these are all Redis's own "up to" figures, and they come from different places: per MarkTechPost, Redis's earlier public preview announcement cited 15x and 70%, while its current product page says 90%. The speed advantage applies only to cache hits; the product is still in public preview and Redis notes that features and behavior may change. This item has a single source, MarkTechPost.

Source: MarkTechPost

08/09

NVIDIA's BioIR lifts Boltz-2 folding throughput 2.90x on 8xH100

NVIDIA has detailed BioNeMo Inference Runtime (BioIR), a Python library that accelerates biomolecular structure-prediction models on NVIDIA GPUs while staying in plain PyTorch.

According to MarkTechPost, BioIR speeds up biomolecular structure-prediction models on NVIDIA GPUs without leaving plain PyTorch. In a matched benchmark on 1,000 human dimer targets across 8xH100 GPUs, BioIR-accelerated Boltz-2 reached 2.90x higher folding throughput and processed 58.5K residues per GPU-hour.

Limitations: both figures come from a single Boltz-2 benchmark on 8xH100, published by NVIDIA and relayed by MarkTechPost, with no third-party reproduction yet. This item has a single source.

Source: MarkTechPost

Regional and early signals

09/09

Ant Digital Technologies launches Agentar for finance with ten preset agent expert teams

A financial-agent platform unveiled at the Inclusion·Conference on the Bund packages agent expert teams, skills, MCP, evaluation and governance for financial institutions. (Chinese-language source)

On September 10th, Yu Bin, president of Ant Digital Technologies' AI business, launched the finance edition of Agentar at an insights forum of the 2026 Inclusion·Conference on the Bund. It offers ready-made financial agent expert teams, industry skills, MCP, agent evaluation tools and governance, covering agent development, deployment, collaboration, evaluation and management. The first release presets ten financial agent expert teams for scenarios including robo-advisory, wealth management and customer operations; each digital expert maps to a full job role and can break down tasks and dispatch multiple specialist agents.

  • Evaluation: more than 2,000 finance-specific evaluation tools covering expertise, compliance, accuracy and execution quality
  • Scale: Ant Digital says it has built more than 300 specialist agents for finance with partners, and serves all state-owned and joint-stock banks and more than 60% of local commercial banks
  • Pilot: its "digital account manager" raised pilot customers' total assets by 10% and customer activity by 15%, and increased the number of customers served more than fivefold

Limitations: the pilot and customer-coverage figures are Ant Digital's own, with no independent verification. Chinese-language source, single outlet (Leiphone).

Source: Leiphone

Generate one yourself with JIUFENG

Write a prompt in your browser and get an image — 1K, 2K or 4K output, up to 15 reference images. 50 free generations on sign-up, no credit card.

Generate free