AI Highlights

Claude Haiku 5.5 jumps to 72.4% on OSWorld

Key Takeaways
  • •Haiku 5.5 posts 72.4% on OSWorld and tops a small-model index
  • •OpenAI cites 40M Codex/Work users
  • •Scale AI's HSS stumps top models.
jiufeng
October 8, 2026
31 min read
In this article

Overview

10 stories in this issue. The first 3 are today's priorities.

Hot model updates

  1. Top · Claude Haiku 5.5 jumps to 72.4% on OSWorld and tops Artificial Analysis' small-model tier
  2. Top · GPT-6 rolls out in ChatGPT; OpenAI says Codex and Work have 40M active users
  3. Top · Scale AI's HSS visual-reasoning benchmark: top model GPT-6 Astra scores only 53.6%

Global AI news

  1. Architect launches Liquid Inference, a live auction for every LLM request
  2. Microsoft gives Copilot local file access and OS-wide actions under "Hybrid Intelligence"
  3. NVIDIA-led PivotOPD trains multi-turn agents to avoid and recover from pivotal mistakes
  4. AWS publishes an auto-remediation pattern for after DevOps Agent investigations
  5. Meta rolls out AI tools to find ads that secretly lead to child abuse material

Regional and early signals

  1. Leiphone on Apsara: Qwen3.8-Flash activates only 6B parameters per inference
  2. STEPX Neo agent phone from 阶跃终端 launches October 13

Hot model updates

01/10

Claude Haiku 5.5 jumps to 72.4% on OSWorld and tops Artificial Analysis' small-model tier

The benchmark data missing at launch is now out: Haiku 5.5 is far stronger at computer use than its predecessor and costs up to 90% less per token, but it also produces many more output tokens.

According to The Decoder, Haiku 5.5 scores 72.4% on the OSWorld computer-use test, up from 15.7% for Haiku 4.5. Token prices are up to 90% lower, and the model beats OpenAI's budget model GPT-6 Luna in areas such as agentic coding. Haiku 5.5 is available on AWS, Google Cloud and Azure. Anthropic also halved cache-read prices for Sonnet 5.5 and is starting monthly API credits for subscribers.

On October 8, Artificial Analysis published its small-model ranking on the Intelligence Index:

ModelIntelligence Index (points)
Claude Haiku 5.543
GLM-5.3 Flash42
Gemini 3.8 Flash41
GPT-6 Luna38

Limitations: Artificial Analysis notes that Haiku 5.5 uses far more tokens: about 162,000 output tokens per task at the highest effort level. The Decoder also points to a new tokenizer, so "up to 90% cheaper" may not translate into a 90% smaller bill. Test it on your own workload. Haiku 5.5 is still 13 points behind Claude Sonnet 5.5 on the index.

Claude Haiku 5.5 jumps to 72.4% on OSWorld and tops Artificial Analysis' small-model tier: Intelligence Index (points)

Jiufeng graphic based on the sources cited in this issue.

Introducing Claude Haiku 5.5

Image source: anthropic; mirrored on Jiufeng R2.

Source: The Decoder · Anthropic

02/10

GPT-6 rolls out in ChatGPT; OpenAI says Codex and Work have 40M active users

On the day GPT-6 started reaching ChatGPT's Chat tab, OpenAI shared a combined user count for its agent products.

GPT-6 began rolling out in ChatGPT's Chat tab on October 7. The same day, OpenAI product lead Thibault Sottiaux posted on X that Codex and ChatGPT Work together have 40 million active users. He also said paid accounts will get a banked usage reset.

  • Earlier figure 1: in June, OpenAI said Codex had passed 5 million weekly users
  • Earlier figure 2: in an August 25 interview, Sottiaux said ChatGPT Work had 20 million users

Limitations: The 40M figure comes from OpenAI, combines two products, and has no stated measurement window. RuntimeWire notes that the earlier numbers use different metrics. The new total doesn't show that Work doubled or that Codex grew eightfold, and it says nothing about repeat use.

Source: RuntimeWire · Thibault Sottiaux on X · OpenAI

03/10

Scale AI's HSS visual-reasoning benchmark: top model GPT-6 Astra scores only 53.6%

On free-form visual reasoning, the best model trails humans by almost 40 points.

Scale AI and Elorian released Humanity's Sixth Sense (HSS) on October 7. It tests inference from 288 images and 234 videos, with free-form answers. The dataset is on Hugging Face. Elorian co-founder and CEO Andrew Dai previously worked at Google Brain and DeepMind.

ParticipantAccuracy (%)
GPT-6 Astra (top)53.6
Median of 25 models30.9
20 human participants93.1

Limitations: Answers are graded by model judges. Scale says it regraded five models using judges from three vendors, and the rankings held with more than 95% agreement. RuntimeWire notes the benchmark also publicly tests Elorian's own "visual thinking" thesis: that today's vision-language models convert images into language before reasoning, which makes them fragile. The results also suggest that more reasoning time alone may not fix failures to spot cues, spatial relationships and social context.

Source: RuntimeWire · Scale AI on X · Hugging Face dataset

Global AI news

04/10

Architect launches Liquid Inference, a live auction for every LLM request

Inference providers bid on each request, and the buyer pays the lowest offer that meets its rules.

Architect Financial Technologies launched Liquid Inference, an exchange-style router for LLM inference. Providers post offers for specific models. Each request is auctioned across every provider quoting that model, and the cheapest offer that meets the buyer's rules wins. Developers only need to change the base URL; the rest of their code stays the same.

Limitations: Architect is a trading firm, not an AI lab. It runs the AX perpetual futures exchange. In May it acquired a US Designated Contract Market to list GPU compute futures, which is still pending regulatory review. There is no independent data yet on actual clearing prices or latency, and MarkTechPost is the only outlet reporting it.

Source: MarkTechPost

05/10

Microsoft gives Copilot local file access and OS-wide actions under "Hybrid Intelligence"

Copilot on Windows will be able to read local files, take actions across the OS, and mix local and cloud models.

At its October 7 Windows and Surface event, Microsoft showed an upgraded Copilot that can read local files on your PC and take actions across the OS. Microsoft calls the approach Hybrid Intelligence: apps use a mix of local and cloud models. Copilot EVP Jacob Andreou walked through a video demo of the Autopilot tool helping with taxes.

Limitations: The tax demo was a pre-recorded video, not a live run. Andreou said features powered by Hybrid Intelligence will arrive in Copilot over the next couple of months. Windows and Surface chief Pavan Davuluri said the new Windows search experience comes to Windows 11 PCs starting this fall. The Verge is the only outlet cited here.

Source: The Verge

06/10

NVIDIA-led PivotOPD trains multi-turn agents to avoid and recover from pivotal mistakes

An on-policy distillation method aimed at multi-turn tasks where one early wrong step can derail everything after it.

Researchers from NVIDIA, Princeton University and the University of Maryland introduced PivotOPD, which uses on-policy distillation to train multi-turn LLM agents. The goal is twofold: avoid early pivotal mistakes, and recover when they happen. On ALFWorld, WebShop and search-based QA (with Qwen3-1.7B/8B student models), it posted the best average score against 13 baseline methods.

  • ALFWorld: 73.7% with a 1.7B student model, 93.0% with an 8B student
  • Pivotal-mistake recovery rate: 72.7%
  • SWE-Bench Verified (separate experiment): with a Nemotron-3.5-SFT student, up from 62.8% to 66.0%, versus 63.0% for standard OPD and 73.0% for the teacher

Limitations: This is a research method, not a product you can call. The SWE-Bench Verified experiment was compared only with standard OPD and the teacher model, not with the 13 baselines. The results come from the authors' own experiments and haven't been independently reproduced.

Source: MarkTechPost

07/10

AWS publishes an auto-remediation pattern for after DevOps Agent investigations

The agent only diagnoses and never changes production resources; a separate, controlled workflow applies the fix.

AWS DevOps Agent triages incidents using correlated metrics, logs and application topology, then returns a root cause analysis and recommended fixes. For safety, organizations usually keep it in observe-and-report mode. This AWS post shows how to build an automated remediation workflow with Lambda Durable Functions, EventBridge and Amazon Bedrock to handle the step after diagnosis. Sample code is in an aws-samples repository.

Limitations: This is a reference implementation from a blog post, not a managed feature. The agent itself still only diagnoses. You have to design and validate the permission boundaries and approval steps for automated fixes yourself.

Source: AWS Machine Learning Blog · GitHub sample repository

08/10

Meta rolls out AI tools to find ads that secretly lead to child abuse material

Some ads and accounts look normal but send users to illegal content off-platform; Meta has launched AI tools built to catch them.

Meta said on Wednesday that it acted on 33.2 million pieces of child sexual exploitation content on Facebook and Instagram in the first half of 2026. More than 97% of it was found before any user reported it. In India, it acted on 5.3 million pieces over the same period, with more than 98% detected proactively. The new AI tools target ads and accounts that look normal but point people to harmful content elsewhere online.

Limitations: All figures are self-reported by Meta, with no third-party audit. TechCrunch is the only outlet reporting it so far.

Source: TechCrunch

Regional and early signals

09/10

Leiphone on Apsara: Qwen3.8-Flash activates only 6B parameters per inference

Alibaba's plan for its next phase is to keep scaling up while redesigning the architecture to cut costs.

In a recap of the Apsara Conference, Leiphone reports that Qwen3.8, released in August, has 2.4T parameters. That is about 33 times the 72B of the Qwen2.5 flagship from two years ago. Qwen3.8-Flash uses sparse attention, linear attention and external memory:

  • Active parameters: 6B per inference
  • Training cost: one-ninth of the previous generation
  • Input price: one-quarter of the previous generation per million tokens
  • Long context: 8.6x higher prefill throughput on very long contexts

Limitations: These figures all come from Alibaba's own talks and haven't been independently checked. The post-Qwen4 parameter roadmap was reported earlier, and this piece adds no new release timeline. Chinese-language source.

Source: Leiphone (Chinese-language source)

10/10

STEPX Neo agent phone from 阶跃终端 launches October 13

The device maker 阶跃终端 has announced its first phone built natively around a large model and agents.

阶跃终端 will hold a "Ready Builder One" launch in Shanghai on October 13 to unveil STEPX Neo, its first phone built natively around a large model and agents. The company says the model, system and hardware were all designed for agents, and it will also share updates on ecosystem partnerships.

Limitations: This is a pre-announcement. Specs, price and the model version have not been disclosed. The article was supplied by the company and republished by QbitAI, so it reflects the company's own claims. Chinese-language source.

Source: QbitAI (Chinese-language source)

Generate one yourself with JIUFENG

Write a prompt in your browser and get an image — 1K, 2K or 4K output, up to 15 reference images. 50 free generations on sign-up, no credit card.

Generate free