AI models in September 2026: who leads, what it costs, what to choose
Between September 1 and 22, 2026, Anthropic, OpenAI, Google and SpaceXAI released at least seven models. This map sorts out what changed, what each tier costs and what the leaderboards really measure, with dates and sources, so you can build a shortlist on solid ground.
In September 2026 three US labs set the pace: Anthropic with Claude Fable 5.1 and Claude Opus 5.5, OpenAI with GPT-6 Astra, Sol and Luna, and Google with a new Flash model almost every month while Gemini 3.5 Pro stays delayed. Grok 4.7, Meta's closed Muse Spark and the open-weight labs complete the map.
Prices now fall into clear bands, from $10/$50 per million tokens down to $0.30/$1.20, and leadership depends on the test. Our advice: shortlist by use case, keep a fallback from a second vendor and decide with your own prompts in Spanish and French.
The closed labs and their flagships
Who sells what, since when and at what list price.
| Lab | Model | Released | Input / output | Context |
|---|---|---|---|---|
| Anthropic | Claude Fable 5.1 | Sept 1, 2026 | 10 / 50 | 1M |
| Anthropic | Claude Opus 5.5 | Sept 22, 2026 | 4 / 20 | 1M |
| Anthropic | Claude Sonnet 5 | June 30, 2026 | 2 / 10 | 1M |
| OpenAI | GPT-6 Astra | Sept 3, 2026 | 10 / 50 | 1.05M |
| OpenAI | GPT-6 Sol | Sept 22, 2026 | 2 / 10 | 1.05M |
| OpenAI | GPT-6 Luna | Sept 22, 2026 | 0.10 / 0.50 | 1.05M |
| Gemini 3.8 Flash | Sept 2, 2026 | 0.75 / 3.75 until Dec 31, 2026, then 1.50 / 7.50 | 1M | |
| Gemini 3.1 Pro (preview) | Feb 19, 2026 | 2 / 12 up to 200K tokens | 1M | |
| SpaceXAI | Grok 4.7 | Sept 21, 2026 | 2 / 6 up to 200K tokens | 500K |
Anthropic. Claude Fable 5.1 is its most capable widely released model; its twin, Mythos 5.1, is limited to verified US organizations. Anthropic's docs now recommend starting with Opus 5.5 for most workloads, and Sonnet 5.5 and Haiku 5.5 are announced for the coming weeks (details in our Opus 5.5 analysis).
OpenAI. GPT-6 Astra is its first model rated "Critical" for cybersecurity under its Preparedness Framework, so the public version refuses advanced offensive tasks. GPT-6 Sol and Luna launched at half the promotional price of their GPT-5.6 predecessors; GPT-5.6 Terra ($2/$12) remains the middle option and the o-series is now legacy.
Google. Gemini 3.5 Pro, expected in June, is late: Bloomberg reported on July 16 that its coding performance fell short. Meanwhile a new Flash model shipped almost every month (3.5 in May, 3.6 in July, 3.7 in August, 3.8 in September), and Gemini 3.1 Pro is still in preview.
SpaceXAI. SpaceX absorbed xAI in February 2026, renamed it in July and bought the code editor Cursor in August. Grok 4.7 ranks sixth on the official Terminal-Bench 4.0 board (37.6%).
Meta, Microsoft, Amazon. Muse Spark (April 8) was Meta's first model without open weights; version 1.3 scores 48 on the Artificial Analysis index. Microsoft unveiled MAI-Thinking-1, its first reasoning model, in June (price unconfirmed). Amazon moved Nova Premier and other Nova models to maintenance mode, according to Business Insider (via eWeek).
Figures checked on September 25 and 26, 2026, against the sources below. Prices, rankings and even availability change monthly: check again before signing a contract.
Open weights in brief
Open-weight models have grown huge (Kimi K3 has 2.8 trillion parameters, 104 billion active per token) but still trail the closed leaders: the best, Xiaomi's MiMo-V2.6-Pro, scores 46 on the Artificial Analysis index against 58 for Claude Opus 5.5.
DeepSeek V4.1-Flash (MIT license) is the cheapest near-frontier API we found, at $0.30/$1.20 at peak hours and half that off-peak, and licenses are drifting toward custom terms with revenue thresholds. More in our open-weight guide.
Price bands: four tiers, not one market
The list price tells you the band; your cost per task tells you the bill.
| Band | Typical price | Examples | Fits |
|---|---|---|---|
| Premium | 10 / 50 | Claude Fable 5.1, GPT-6 Astra | Hardest reasoning, long-horizon agents |
| Upper mid | 4 / 20 | Claude Opus 5.5, GPT-5.6 Sol (promotional) | Default for agents and knowledge work |
| Mid | about 2 / 10 | Claude Sonnet 5, GPT-6 Sol, GPT-5.6 Terra (2 / 12), Grok 4.7 (2 / 6) | Volume with solid quality |
| Budget | 1 / 5 or less | Claude Haiku 4.5, Gemini 3.8 Flash (0.75 / 3.75), GPT-6 Luna (0.10 / 0.50) | Classification, extraction, routing |
| Cheapest near-frontier | 0.30 / 1.20 | DeepSeek V4.1-Flash, half price off-peak | Batch work on data that may leave your perimeter |
The $10/$50 tier is a 2026 novelty (Claude Fable 5 in June, then Fable 5.1 and GPT-6 Astra), while the middle got cheaper: GPT-6 Sol costs half its predecessor's promotional price, Opus 5.5 is 20% cheaper per token than Opus 5, and Sonnet 5 kept $2/$10 instead of rising to $3/$15.
Three details move a real bill more than the list price: long prompts (OpenAI doubles the input rate above 272K tokens; Claude has no surcharge), cache reads (5% of the input price on Opus 5.5, 10% on GPT-6) and scheduled rises (Gemini 3.6 to 3.8 Flash double on January 1, 2027). Method in our AI cost guide.
Leaderboards: what they measure, who leads
Each ranking answers a narrower question than "which model is best".
Vendor tables are run or chosen by the lab itself, usually at its highest effort and on benchmark versions a few weeks old. Independent boards apply one method to everyone but lag behind: on September 26, Anthropic's reported 66.4% for Opus 5.5 on Terminal-Bench 4.0 was not yet on the official board.
Versions also break comparisons: OpenAI's Astra post quotes Artificial Analysis index v4.1.1 values of up to 65.7, while v4.3.2 tops out at 58, a new scale rather than weaker models. Arena, for its part, ranks by human preference votes, not by tasks completed.
| Board | What it measures | Leader | Worth knowing |
|---|---|---|---|
| Arena, text | Human preference votes | Claude Opus 5.5 (1,509) | Top five are all Claude; GPT-6 Astra is 26th |
| Artificial Analysis index v4.3.2 | Ten agentic and knowledge-work evaluations | Claude Opus 5.5, max effort (58) | Fable 5.1 and GPT-6 Astra tie at 53 |
| Terminal-Bench 4.0, official | Agentic tasks in a terminal | GPT-6 Astra (58.2%) | Opus 5.5 not listed yet |
| ARC-AGI-2 | Abstract reasoning, with cost per task | GPT-6 Astra (95.0%, $1.12 per task) | Opus 5.5 at high effort: 93.3% for $0.408 |
| Humanity's Last Exam | Expert-level questions | Claude Opus 5.5 (61.4%, measured by Artificial Analysis) | Anthropic reports 67.7% with tools |
None of these tests uses business Spanish or French: they filter a shortlist, they don't decide it. Our method is in the evals guide.
How to choose by use case
A shortlist to test, not a verdict.
| Use case | Shortlist to test | Why |
|---|---|---|
| Customer service in Spanish | Claude Sonnet 5, GPT-6 Sol, Gemini 3.8 Flash; GPT-6 Luna for routing | Volume pricing; no public benchmark measures Latin American service Spanish, so test tone, accuracy and refusals |
| Coding agents | Claude Opus 5.5, GPT-6 Astra, Claude Fable 5.1; GLM-5.3 if you need open weights | Leaders on Terminal-Bench 4.0; GLM-5.3 is the best open-weights entry on the official board |
| Long documents | Claude Opus 5.5, GPT-6 Astra, Gemini 3.8 Flash | 1M-token windows; OpenAI reports 96.3% for Astra on MRCR v2 at 512K to 1M tokens; no long-context surcharge on Claude |
| Voice | Gemini 3.8 Live, Amazon Nova 2 Sonic, Mistral Voxtral for transcription | Live audio models; Voxtral Mini Transcribe Realtime has Apache 2.0 weights (see our voice guide) |
| Regulated or on-premise data | Self-hosted Gemma 4, Muse Glimmer, Qwen3.8-27B or Mistral Small 4 | Apache 2.0 licenses, and no data leaves your infrastructure |
| High-volume classification | GPT-6 Luna, Gemini 3.1 Flash-Lite, DeepSeek V4.1-Flash, Claude Haiku 4.5 | From $0.10 per million input tokens, half that in batch; Haiku 4.5 is only committed until October 15, 2026 |
Include a second vendor and an open-weight option in every shortlist. Our method (30 fixed prompts per client, five engines per run, three languages) is described in how we choose models.
Seven verified trends of 2026
- A premium tier, yet cheaper capability. On ARC-AGI-2, a score of about 61 to 65% cost $2.25 per task with Claude Opus 4.6 in February and $0.042 with DeepSeek V4 Flash in July.
- Reasoning effort is a setting. Claude, GPT-5.6, GPT-6 and Gemini 3.x expose effort levels; on ARC-AGI-2, Opus 5.5 costs $0.241 per task at low effort and $1.85 at max (see our reasoning guide).
- 1M tokens of context is the norm for flagships; Grok 4.7 stops at 500K.
- Offensive cyber capability is gated. Mythos 5.1 is limited to verified US organizations, Gemini 3.8 Flash Cyber to trusted defenders in Google's Fairwind Program, and OpenAI plans wider defensive access to Astra through its Daybreak program.
- Products disappear quickly. OpenAI closed the Sora app in April, about seven months after launch, removed the Sora 2 API on September 24 and shut its Atlas browser on August 9; Nova Premier reached end of life on September 14.
- US policy can cut access. A US export-control directive made Anthropic pull Fable 5 and Mythos 5 for all users on June 12 (Fable 5 returned on July 1), and Executive Order 14409 (June 2) calls for a voluntary framework giving the government up to 30 days of pre-release access to "covered frontier models".
- Labs debate a slower pace. In We Must Pace the Frontier, Dario Amodei called for slowing capability gains and committed Anthropic to embedded third-party evaluators; Sam Altman matched that commitment on September 14. One trigger he cites: the July incident in which OpenAI agents under evaluation reached Hugging Face systems (OpenAI's report).
What it means for companies in Colombia, Latin America and Europe
Keep a tested fallback. The Fable 5 withdrawal lasted 19 days and hit every customer. A production system in Bogotá, Mexico City or Paris needs a second model from another vendor, tested on the same prompts, and an open-weight option you could host.
Budget for residency. OpenAI adds 10% for regional processing on models released since March 5, 2026; Claude costs 10% more on regional Bedrock or Google Cloud endpoints.
Put retirement dates in contracts. Anthropic commits to Opus 5.5 until at least September 22, 2027, but to Haiku 4.5 only until October 15, 2026. Ask every vendor for such dates, notice periods and a migration clause.
Test Spanish and French yourself. At Slash AI Lab every candidate goes through 30 fixed prompts per client in Spanish, English and French before deployment; no leaderboard replaces that step.
Check the reach of the EU AI Act. Article 2(1)(c) covers providers and deployers outside the EU when their system's output is used in the EU, so a Colombian firm serving French clients can be in scope. In Colombia, check how Law 1581 of 2012 applies before sending personal data abroad. More in our 2026 regulatory map.
For each route, name a default model, a fallback from another vendor and an open-weight option, then re-run your evaluation at every major release: in 2026, roughly once a month.
Key takeaways
- September 2026 brought Claude Fable 5.1 and Opus 5.5, GPT-6 Astra, Sol and Luna, Gemini 3.8 Flash and Grok 4.7; Gemini 3.5 Pro is still missing.
- Prices form bands, from $10/$50 for the premium tier down to $0.30/$1.20 for DeepSeek V4.1-Flash.
- Leadership depends on the test: Anthropic leads Arena text and the Artificial Analysis index, OpenAI the official Terminal-Bench 4.0 board and ARC-AGI-2.
- Shortlist per use case, then decide with your own Spanish and French prompts.
- Plan for access shocks and closures: a fallback from a second vendor, retirement dates in contracts and residency surcharges in the budget.
Sources
- Models overview
- Introducing Claude Opus 5.5
- Update on access to Claude Fable 5 and Claude Mythos 5
- GPT-6 Astra
- Introducing GPT-6 Sol and GPT-6 Luna
- API pricing
- Gemini API pricing
- Gemini 3.5 Pro delayed over coding, Bloomberg reports
- Models and pricing
- Text leaderboard
- AI model leaderboards (Intelligence Index)
- ARC-AGI leaderboard
Editorial note: this analysis reflects the public information available on the review date. Models, prices and rules change fast; every third-party figure links to its source, and our opinions are labeled as such. Spotted an error? Write to contact@slash-digital.io.
The questions we hear often
What is the best AI model in September 2026?
There is no single answer. On September 25 and 26, 2026, Claude Opus 5.5 led the Arena text board and the Artificial Analysis index, while GPT-6 Astra led the official Terminal-Bench 4.0 board and ARC-AGI-2. The best model for you is the one that wins on your own tasks at a cost you can defend.
How much does a frontier model cost?
$10 per million input tokens and $50 per million output tokens for Claude Fable 5.1 and GPT-6 Astra, $4/$20 for Claude Opus 5.5 and about $2/$10 for Claude Sonnet 5 and GPT-6 Sol. Budget models start at $0.10/$0.50 with GPT-6 Luna.
Has Google released Gemini 3.5 Pro?
Not as of September 26, 2026. Bloomberg reported a delay in July; Google's newest model is Gemini 3.8 Flash, released on September 2, and Gemini 3.1 Pro is still in preview.
Are open-weight models good enough for business?
For many tasks, yes, and they are the basis of any on-premise option. On independent boards the best of them still trail the closed leaders (46 against 58 on the Artificial Analysis index). Our open-weight guide covers licenses and hosting.
Can companies in Colombia and France use these models?
Yes, through the vendors' APIs and the main clouds, but plan for residency surcharges, data-protection rules and the possibility that US policy restricts a model, as happened with Claude Fable 5 in June 2026.
More analysis to read next
Claude Opus 5.5: what changes for businesses
Launched September 22, 2026: cheaper than Opus 5, Fable 5.1-level results per Anthropic. Pricing, benchmarks, breaking changes and when to use it.
Local AI and open weightsOpen-weight models in 2026: which to use, under which license
Kimi K3, Qwen3.8, DeepSeek V4.1, GLM-5.3, Gemma 4, gpt-oss, Mistral: sizes, licenses, gap to the frontier and how to host them without legal surprises.
ModelsHow we choose a model
Language, code, agents, cost and privacy: the evaluation we run before every deployment, with 30 real prompts per client.
Where to next
Let's put it in production
Tell us your challenge. We reply within 24 business hours with an honest first read: if we can help, we'll say how; if not, we'll say who can.
I reply personally. No endless forms, no canned replies.