FORSMILE
Issue #7Published October 4, 2026 / Covering Sep 28-Oct 4日本語で読む

This Week in AI, 13 Items — Three New Models Land at $2/$10, and the Contest Moves to Tokens per Task and Who Can Use Them

The week in one place

The week’s thread is that three new models were priced at $2 and $10 per million tokens. OpenAI’s GPT-6.1 Sol (29 September) and Anthropic’s Claude Sonnet 5.5 (28 September) share a list price of $2 input and $10 output, and Google’s Gemini 4 Argon (30 September) carries the same introductory price, but it rises to $4/$20 once the introductory period ends, and for now only trusted cyber defenders can use it. OpenAI and Anthropic lead with cost per task and Google with a 95% discount on cached input, so with equal unit prices the difference can only be measured on your own workload. The second movement is that always-on, team-shared agents became the centre of the product, as with OpenAI’s dots and SpaceXAI’s Team Bots (OpenAI DevDay was on 29 September, with more than 20 announcements). In the same week OpenAI disclosed that it had disrupted an organised distillation campaign, and Anthropic put classifiers against reasoning extraction into Sonnet 5.5, so keeping outputs and thinking from being carried away is becoming standard equipment.

Models & Products

OpenAI releases GPT-6.1 Sol: near-Astra performance at a fifth of the price, with cached input at $0.10

An update to GPT-6 Sol: the list price stays at $2/$10 while performance moves closer to Astra.

Released on 29 September, the day of DevDay 2026. OpenAI says it delivers "near-Astra intelligence at one-fifth of Astra’s standard input and output token prices". Per million tokens it is $2 input, $10 output and $0.10 cached input, against $10, $50 and $1.00 for Astra and $0.10, $0.50 and $0.01 for Luna. On DeepSWE v1.1 it matches Astra at roughly one-fifth of the cost and beats GPT-6 Sol’s best score by 6.4 points; on AutomationBench 1.0.6 it scores 2.2 points above Claude Opus 5.5 (medium) at roughly a third of the cost. The share of answers containing a factual error fell from 11.4% to 7.7% at low reasoning effort, on deliberately hard prompts that OpenAI says are not representative of typical use. It is available in ChatGPT Work and Codex (Plus, Pro, Business, Enterprise, Edu) and in the API as `gpt-6.1-sol`, and it is not yet in Chat. AWS made it generally available on Amazon Bedrock the same day.

So What

For work you run through the API, you can try near-Astra results at a fifth of the unit price. But OpenAI chose the benchmarks and reasoning settings, and it states that Astra still scores highest on scientific research at 68.1%. The split of Astra for the hardest reasoning and Sol for complex coding and computer use also appears in the 2 October model guide. If you try it in ChatGPT, pick it in "Work", not in Chat.

Models & Products

Anthropic releases Claude Sonnet 5.5: the same $2/$10 as Sonnet 5, and 70.6% on Terminal-Bench 4.0

A full refresh of Sonnet: the unit price is unchanged, it uses fewer tokens for the same job, and it is 30%+ faster.

Released on 28 September as the second model in the Claude 5.5 family, below Opus 5.5 (22 September). Per million tokens it costs $2 input and $10 output, with cache reads at $0.20 and cache writes at $2.50 (Opus 5.5 is $4, $20, $0.20 and $5). It scores 70.6% on Terminal-Bench 4.0 (Sonnet 5: 10.3%; Opus 5.5: 66.4%, which is Opus 5.5 at its highest effort, Xhigh) and 80.1% on OSWorld 2.1 (partial credit). Anthropic says it runs 30%+ faster than Sonnet 5 and costs up to 30% less per task. Haiku 5.5 is said to join "in the coming weeks". It is available on all platforms including AWS, Google Cloud and Azure, as `claude-sonnet-5-5`.

So What

If you run Sonnet with thinking off, check before migrating. Anthropic’s migration guide says you must switch to the new `between_tools` setting before moving to Sonnet 5.5. On cyber, it is the first Sonnet to ship with Opus 5.5-style safeguards, so higher-risk cybersecurity tasks visibly fall back to Sonnet 5 (routine bug fixing is unaffected). It is also the first Sonnet with classifiers that block reasoning extraction, and its thinking can no longer be decoupled from the account that created it, which affects workflows that move conversations between accounts, including switching accounts mid-session in Claude Code.

Models & Products

Google announces Gemini 4 Argon: trusted cyber defenders first, with an introductory $2/$10

A new frontier model with a 1M-token output limit, entering a phased rollout before general availability.

Announced on 30 September. Google is starting with cyber defenders through its Fairwind Program and says it will widen access in phases while taking part in the U.S. government’s voluntary pre-release process. Developers, enterprises and consumers come "as soon as possible", starting with paid API customers and Google AI Ultra subscribers. The introductory price is $2 per million input tokens and $10 per million output tokens, with cached input at 95% off, and it becomes $4/$20 after the introductory period. The output limit grows from 64K to 1M tokens. Google reports 77.9% on DeepSWE v1.1 (which it calls state of the art), first place at 51.3% on Zapier’s AutomationBench, and a first-place tie at 68% on CWE-bench v1. Trusted defenders and Google’s own teams get it without cyber guardrails.

So What

You probably cannot use this model yet even if you apply. Access is limited to trusted cyber defenders and no developer date is given, so $2/$10 is not a price you can buy today but a limited-time preview price, doubling afterwards (the end date of the introductory period is not stated). The scores are on benchmarks Google chose, with no direct comparison to other vendors’ latest models. Putting frontier capability behind a vetted tier and then widening in phases is the same pattern as the three vendors in issue 3 (6 September).

Models & Products

OpenAI launches "dots", always-on agents that work 24/7 on their own cloud computer

You can now keep one persistent agent of your own inside ChatGPT.

Announced on 29 September. Built on GPT-6 Astra, each dot has its own cloud computer and browser and reaches more than 4,000 apps through plugins. You can talk to it from ChatGPT, Slack and Teams, and in "proactive research", where it looks into things before being asked, it uses only read-only tools on the apps you have connected. It is rolling out to Pro and Business Premium in eligible markets; Enterprise, Edu and Healthcare can try the beta when a workspace admin turns it on (it is off by default). Your first dot is included at no extra cost, and conversations with a dot do not count toward ChatGPT usage limits, though tasks it starts in Codex or ChatGPT Work do. OpenAI also says it is working with Microsoft to let businesses manage specialist dots through Agent 365’s governance controls.

So What

The biggest caution is that costs cannot be forecast. The first dot is free, and the post only describes a future where you add more dots or scale their speed and monthly workload, so no price is published. On privacy, content from Business, Enterprise and Edu workspaces is not used to improve models by default, and personal plans can choose. A dot’s notes to itself and its proactive research are not trained on directly, but may inform an eligible conversation or task depending on settings. Sensitive actions such as changing a password always stay with the person.

Models & Products

SpaceXAI announces "Team Bots", shared Grok Bots that learn as a team and answer in Slack

One Bot is shared across a team, while each person’s conversations stay separate, on common knowledge and tools.

Announced on 28 September (SpaceXAI is the name used on x.ai’s news page). A Team Bot bundles context (files, instructions, skills), plugins (Salesforce, Notion, GitHub and others), credentials and memories around a role or workflow. The Bot is shared, but each person’s conversations with it stay private, and it keeps separate context and memories per user. In Slack it has its own handle. SpaceXAI’s own examples include an account Bot that posts a briefing every morning and an engineering Bot that steered Cursor cloud agents so that a five-person team shipped more than 100 PRs a day (the company’s own claim).

So What

Where OpenAI’s dots are one persistent agent per person, this is "one agent shared by a team". The reason it is here is that in the same week vendors put always-on, shared agents at the centre of their products. The figures are the company’s own case studies, with no evidence of reproducibility in other environments.

Models & Products

OpenAI launches Ultrafast for GPT-6 Astra and a new "Pro 500" plan; AWS Bedrock followed the next day

A paid speed tier that makes Astra generate up to 8x faster. Pricing could not be confirmed for this issue.

Started on 29 September. Ultrafast is a speed tier for GPT-6 Astra that promises up to 8x faster token generation in Codex (300 tokens per second) and up to 6x in the API. It is available in the API and in ChatGPT Work and Codex on Pro 500 and Enterprise, with a GPT-6.1 Sol version "coming soon". NVIDIA said on 1 October that it runs on Blackwell GPUs, and AWS announced Amazon Bedrock support on 30 September. The new Pro 500 plan gives 25 times the ChatGPT Plus allowance and includes Ultrafast.

So What

Speed is an add-on, its pricing is in the "Ultrafast guide", and this issue did not confirm it. The monthly price of Pro 500 is also absent from the text I retrieved. OpenAI and NVIDIA say speed matters most when an agent loops through "write, call a tool, check" many times. Before you consider it, check the price in the primary sources (the Ultrafast guide and the plan pricing page).

Developer Tools

OpenAI adds computer use to the Agents API; Bedrock Managed Agents, which runs entirely inside AWS, enters preview

OpenAI’s agent runtime can now be used both through the API and inside AWS.

Announced on 29 September. The Agents API now supports computer use, and brings Codex’s multi-agent capabilities, tool search, tool calling and context compaction into your own application (OpenAI runs the infrastructure). Amazon Bedrock Managed Agents, built jointly by AWS and OpenAI on a customised version of the Agents API, lets you run agents entirely inside AWS with the identities, permissions and governance you already use. AWS announced that it entered preview on 29 September. Separately, the Decisions API, which uses Luna to pick from a finite set of defined answers for classification and routing, is in limited preview, with a broad release planned "in the coming days".

So What

Organisations that run on AWS now have a way to use OpenAI agents without the data leaving AWS. But the AWS side is in preview, and this issue did not confirm production terms such as pricing or availability guarantees. The Decisions API is only said to be "broadly released in the coming days", so I will check whether it shipped in the next issue.

Developer Tools

OpenAI makes Codex run in the cloud around the clock, adds voice to the CLI and introduces Security Cloud

Codex can keep working in the cloud even when your own computer is closed.

Announced on 29 September. Codex in the cloud lets you choose between your computer, remote control from a phone, or the cloud from any device, with reusable environments that share approved settings and permissions across a team (Plus, Pro, Business, Healthcare, Education, Enterprise). The Codex CLI can now start tasks by voice and gains an `/agents` view for tracking several tasks at once. Codex Security Cloud scans whole GitHub repositories on demand or on a schedule and keeps checking new commits; it investigates findings, removes duplicates and prepares fixes in the cloud, and includes access to models offered through Daybreak Blue without a separate Daybreak application (Pro, Business, Enterprise and Edu, on desktop and web).

So What

The change is that it keeps running with your laptop closed. On the other hand, a security scan run in the cloud means handing the reading of whole repositories and the preparing of fixes to OpenAI’s cloud, so decide the scope of repositories and permissions before you turn it on. Security Cloud is for Pro and above; Plus is not included.

Developer Tools

Claude Code 2.1.284 to 2.1.289: Sonnet 5.5 becomes the default Sonnet, and "Claude Mods" lets plugins change deeper behaviour

Besides tracking the new model, the range of what plugins are allowed to do widened.

Six releases from 28 September to 3 October. 2.1.284 made Sonnet 5.5 (`claude-sonnet-5-5`, 1M context, $2/$10 with $0.20 cache reads) the default Sonnet on the Anthropic API. 2.1.285 added an `allowedProviders` managed setting that limits which API providers a machine may use, and a `CLAUDE_CODE_DISABLE_WEB_FETCH` environment variable. 2.1.287 introduced Claude Mods, where plugins may modify deeper behaviour, along with a built-in mod, "You should know", a side agent that watches your back, enabled with `/plugin enable cc-plugin-you-should-know@builtin` (for first-party sessions with telemetry on). 2.1.289 fixed a deny or ask rule on a nested part of a compound shell command not holding over a user-installed mod’s approval on managed machines, and a user-installed plugin being able to rewrite the descriptions of an organization-managed MCP server’s sign-in tools.

So What

If you restrict permissions with rules, it is worth moving to 2.1.289, after Mods arrived. The two fixes landed in 2.1.289, so earlier versions could evidently be affected (under the conditions in the fix notes: managed machines and organization-managed MCP servers). `allowedProviders` is a way for administrators to narrow where models are reached from.

Market & Industry

OpenAI launches Marketplace to put existing commitments toward partner software, plus a way to use ChatGPT allowances in 16 partners’ tools

A contract with OpenAI can now pay for other companies’ software too.

Announced on 29 September. OpenAI Marketplace lets eligible enterprise customers apply part of their existing OpenAI commitment toward approved partner software. The first 32 partners include Figma for creative work; Adobe, Sierra, Decagon, HubSpot, Salesforce and ServiceNow for customer experience; Harvey and Legora for legal; Palo Alto Networks and CrowdStrike for cybersecurity; and Baseten for open-source models. Separately, Sign in with ChatGPT lets you use your ChatGPT allowance across 16 partners, including Cognition’s Devin, Notion, Vercel, T3, OpenClaw and Dactyl, and control how much each can use. Sharing plan usage is for Plus and Pro users.

So What

The money in an OpenAI contract can now move as a budget for neighbouring tools. For the software companies, it adds a path for OpenAI’s customers to buy their products. For buyers, the conditions for applying a commitment (which contracts qualify, any cap, how it is settled) are not stated in the announcement, so do not read it as freely usable until you confirm with your account team.

Market & Industry

Anthropic launches Claude Frontier Academy, with $100 million to train 10,000 engineers by the end of 2027

Anthropic sees talent as the bottleneck to AI adoption and will train customers’ and partners’ engineers to its own standard.

Announced on 2 October. With a $100 million commitment, it aims to train 10,000 Frontier Deployed Engineers (FDEs) by the end of 2027. The first cohorts include engineers from Accenture, Bain, Capgemini, Commonwealth Bank of Australia, Deloitte, McKinsey, Morgan Stanley and Novo Nordisk, running in San Francisco, New York and London. After a multi-day in-person programme and a graded practical, engineers earn the Claude Resident Engineer badge and move into a 12-week residency leading a real use case at their own organisation. Passing a final assessment earns the FDE badge, with the first expected in early 2027. Participation is by nomination from the organisation.

So What

This puts "the number of people who can deploy it" at the centre of competition, not model performance. It targets customers’ and partners’ engineers and goes through the Anthropic account team, so it is not a scheme an individual can apply to. How the credential is valued in the market will not be known until the first graduates in early 2027.

Regulation & Safety

OpenAI says it disrupted a coordinated campaign to extract protected reasoning, with a core cluster tied to Moonshot AI associates

It found even a method that had the model decrypt hidden reasoning in a different conversation, and frames the problem as one shared across vendors.

Published on 30 September. The earliest activity was on 1 July, with a spike on 24 and 25 July of 16,000 requests using the extraction pattern from over 4,000 users (attempted, not necessarily successful extractions), and a related cluster of more than 15,000 users was fully disrupted by 28 July. One method copied encrypted reasoning from one conversation and asked a model in another to decrypt and transcribe it. OpenAI attributes a core cluster to individuals associated with Moonshot AI, the developer of Kimi, while saying it is unclear whether all the operators came from one actor. The response included account enforcement, closing a pathway that let someone replay another user’s encrypted reasoning to recover its contents, and holding streamed output that might expose reasoning; it also shared findings through the Frontier Model Forum and government channels.

So What

Systems that make reasoning portable or replayable become an attack surface (OpenAI itself writes that such systems may face related risks). The same week, Anthropic shipped Sonnet 5.5 with classifiers against reasoning extraction. Using a vendor model’s outputs or thinking to train or reproduce another model is the kind of use now being shut down (OpenAI says it violates its terms). But the attribution to Moonshot AI is OpenAI’s assessment, and this issue could not confirm that company’s response from a primary source.

Regulation & Safety

OpenAI publishes guidelines for "safety cases" in frontier AI training, with a veto for senior leaders and public postmortems

Early guidelines saying that the argument for risk should be written down before a training run continues.

Published on 28 September. It says structured safety documentation, ideally safety cases, should be required before continuing any frontier reinforcement-learning training run, and calls them an aspirational north star, since matching aviation or nuclear rigour is hard. The technical part has three layers, alignment training, containment and monitoring; the operational part lists a dissent (pre-mortem) from another team, approval with a veto by senior leaders, accountability for the responsible leader, runbooks and SLAs for pausing, and giving auditors enough access. The third part covers investigating severe misalignment incidents, including root-cause analysis, a postmortem and public disclosure of the results. OpenAI says these are current recommendations it is implementing and that they may change within weeks.

So What

This is a proposal from OpenAI, not a standard. It has no legal force and no guarantee that others will adopt it. Even so, alongside the misalignment reports in issue 5 and Anthropic’s embedded outside evaluator, the direction of making training-time safety verifiable from outside is lining up across vendors’ official documents, which is worth noting.

Watching (no confirmed primary source yet)

  • The window starts on 28 September so it joins cleanly to issue 6, which covered 22 to 27 September, with no overlap and no gap. The item that issue’s watching list marked for this issue, "OpenAI DevDay 2026 (29 September) will be covered next", is closed here.
  • No date is given for Gemini 4 Argon’s release to developers or the general public, and the source does not say when the $2/$10 introductory price ends. I will cover it once paid API and Google AI Ultra access opens.
  • Whether OpenAI’s Decisions API (limited preview, "broad release in the coming days") and the GPT-6.1 Sol version of Ultrafast ("coming soon") have shipped will be checked next issue. This issue did not read the primary sources for Ultrafast or Pro 500 pricing (the Ultrafast guide and the plan pricing page).
  • Moonshot AI’s response could not be confirmed from a primary source. OpenAI’s attribution is its assessment that a core cluster is associated with Moonshot AI, and it says it is unclear whether all observed operators were a single actor.
  • Microsoft AI’s speech-recognition model (1 October) could not be read because its page body did not load, so it is not an item. For Meta, Apple and the Chinese labs (DeepSeek, Alibaba, Moonshot), no new announcement in the target week could be confirmed from a primary source (ai.meta.com’s listing could not be retrieved, so Meta may not be fully covered).

"Primary" links go to the announcing party's own publication (company blog, press release, official docs). "Reporting" links go to news coverage or third-party analysis. Figures and dates are as verified on the publication date.

← All Weekly AI News issues