In this article
AI Highlights

DeepSeek V4 Flash Hosting Starts at $0.15/M Tokens

Key Takeaways
  • DeepSeek gets low-cost hosting, Qwen powers local Junie, and GPT‑5.6 reaches Kiro, alongside new agent infrastructure and robotics signals.
jiufeng
August 25, 2026
21 min read
DeepSeek V4 Flash Hosting Starts at $0.15/M Tokens

Overview

10 stories in this issue. The first 3 are today's priorities.

Popular Model Updates

  1. Top · DeepSeek V4 Flash gets $0.15-per-million hosting
  2. Top · JetBrains builds local Junie on Qwen3.6-27B
  3. Top · GPT‑5.6 arrives in Kiro
  4. Mistral-HUMAIN partnership targets localized models

Global AI News 5. Managed Ray lands on SageMaker HyperPod 6. ARD opens a specification for agent discovery 7. Agent Lightning 1.0.1 packages agent optimization 8. Fake Codex installer uses ClickFix on macOS

Regional and Early Signals 9. GEN-1.5 learns tasks from seconds-long demos 10. JD builds a 7,000-square-meter robot service hub

AI signal map for 2026-08-25
AI signal map for 2026-08-25

Jiufeng graphic based on the sources cited in this issue.

DeepSeek V4 Flash gets $0.15-per-million hosting

OneTriangle has launched hosted DeepSeek V4 Flash, emphasizing low prices and low-latency inference.

The service costs $0.15 per million input tokens and $0.35 per million output tokens, running on an eight-H100 stack. The reported reserved configuration delivers about 10,039 output tokens per second at 128 concurrent requests, with median time to first token of roughly 0.35 seconds. DeepSeek’s official repository lists the V4 Flash 0731 weights under the MIT license.

OneTriangle’s “fastest” claim lacks an independent comparison within the report, and its rates exceed DeepSeek’s direct API pricing.

deepseek-ai/DeepSeek-V4-Flash-0731 · Hugging Face
deepseek-ai/DeepSeek-V4-Flash-0731 · Hugging Face

Image source: huggingface; mirrored on Jiufeng R2.

Source: RuntimeWire · Launch thread · Model repository

JetBrains builds local Junie on Qwen3.6-27B

Junie Local bundles a 4-bit Qwen3.6-27B model for running a coding agent entirely on a Mac.

Entering /local in Junie downloads roughly 20GB and starts the local server without a separate runtime or endpoint configuration. It uses no token quota and keeps code on the device, but requires an M5 Mac with 64GB of memory.

JetBrains says Qwen3.6-27B scored on par with Sonnet 4.5 in its private tests, while GPT-5 scored slightly higher. It also says reasoning-enabled Qwen3.8 took roughly four times longer to complete tasks; the full benchmark setup was not disclosed.

Source: JetBrains

GPT‑5.6 arrives in Kiro

Kiro now offers GPT‑5.6 for planning, building, reviewing, and testing software.

OpenAI says GPT‑5.6 is now available in Kiro to help developers plan, build, review, and test software, positioning it as a price-performance option for these tasks.

The supplied announcement gives no price, benchmark score, context length, or rollout scope. The available account comes solely from OpenAI.

Source: OpenAI

Mistral-HUMAIN partnership targets localized models

The companies plan model development, localization, and deployment for Saudi Arabia and the wider Middle East.

The partnership is valued at hundreds of millions of euros and covers AI infrastructure, advanced model development, and solution deployment, with cybersecurity among its initial focus areas. Mistral frames the project as part of a regional sovereign-AI effort.

The announcement says the parties will “pursue” development but names no model, parameter count, schedule, or access point.

Source: Mistral AI

Global AI News

Managed Ray lands on SageMaker HyperPod

AWS has integrated Ray cluster creation, monitoring, and development connections into SageMaker HyperPod.

The capabilities run on Amazon EKS, letting users create and monitor Ray clusters and attach JupyterLab or Code Editor notebooks to live clusters. Ray Dashboard and Amazon Managed Grafana observability are included, while the open-source KubeRay operator manages cluster lifecycles.

AWS disclosed no pricing, supported-region list, or performance comparison with self-managed KubeRay, leaving the operational gains unquantified.

Source: AWS · KubeRay

ARD opens a specification for agent discovery

ARD aims to standardize discovery of agents, tools, skills, and MCP servers distributed across environments.

AWS Agent Registry provides a centralized, searchable catalog for agents, MCP servers, tools, skills, and custom resources. The related Agentic Resource Discovery specification is available on GitHub under the Apache License 2.0.

The material describes the specification and registry structure but reports no cross-client compatibility tests, external adoption figures, or large-catalog search benchmarks.

Source: AWS · ARD GitHub

Agent Lightning 1.0.1 packages agent optimization

Microsoft has released a skill that helps coding agents optimize other editable AI agents against benchmarks.

Given an editable agent and a benchmark, the skill guides changes to prompts, tools, workflows, models, and reasoning settings. Its measured iterations balance accuracy, cost, latency, and reliability; v1.0.1 is the first official release.

The release material provides no standardized improvement or operating-cost results. The Hacker News post had 27 points and one comment at capture time, so community evidence remains limited.

Source: GitHub · Hacker News

Fake Codex installer uses ClickFix on macOS

Attackers are using search ads and a counterfeit Codex download page to persuade Mac users to run malware.

Cato researchers found that sponsored results for searches such as “codex macos download” could appear above OpenAI’s official listing. A Google Sites page copied the download portal, loaded attacker-controlled content through an iframe, and instructed victims to paste a command into Terminal; researchers mapped three infrastructure sets.

Payload delivery was observed only on macOS and still required the victim to execute the command. The report did not disclose the number of affected users.

Source: SiliconANGLE

Regional and Early Signals

GEN-1.5 learns tasks from seconds-long demos

GEN-1.5 can learn short-horizon manipulation tasks from one 3–12 second demonstration, but remains a research release.

The model consumes sensorimotor data within a 30-second context window and averaged 59% success with a ±10% standard deviation across 10 tasks using in-context learning. Ten gradient steps on five minutes of data per task raised the average to 83% with a ±9% standard deviation.

The team describes the tasks as simple and short-horizon. Generalist AI reported the results, and no public weights or API are currently available, so the system is not yet deployable.

Source: MarkTechPost

JD builds a 7,000-square-meter robot service hub

The “Robot Home” supports 2,056 competition robots with storage, repair, charging, and digital check-in services.

The event drew 666 teams from 16 countries. JD’s facility includes storage for more than 1,500 robots, 24 repair stations, nearly 200 charging cabinets, and 100 independent charging positions; its robot-and-battery identification system claims 30-second check-in.

This item relies on a single Chinese-language regional report. The available material provides no service failure rate, completed-repair count, or independent verification of the reported figures.

Source: Leiphone — Chinese-language source