AI Release Tracker
Release Tracker
Key milestones in AI history · updated weekly
starGPT-6 Astra
llmOpenAI's next-generation flagship, released to approved users Sep 3 with wider rollout the next day. 1.05M context (128K max output), $10/$50 per 1M. State-of-the-art on computer use, browsing, software engineering, and cybersecurity (100% on ExploitBench per OpenAI); built on its largest training run yet (100,000+ GPUs at the Stargate Texas site). OpenAI president Greg Brockman called it a "generational leap."
starClaude Fable 5.1
llmIncremental update to Fable 5 — same $10/$50 per 1M list price, cheaper prompt-cache-hit rate. Anthropic emphasized long multi-step agentic work (e.g. Terminal-Bench-Science) over headline SWE-bench gains. GA across the Claude API, AWS, Google Cloud, and Azure.
Grok 4.6
llmReleased Aug 12, 35 days after Grok 4.5. Same $2/$6 per 1M and 500K context as 4.5, adds an "xhigh" reasoning effort level; notable gains on DeepSWE (54 -> 65.9) and APEX-Agents (47.1 -> 57.5) benchmarks.
starClaude Opus 5
llmAnthropic's fourth model release in under two months. Same $5/$25 per 1M as Opus 4.8, 1M context, 96% SWE-bench Verified — took the top spot on the Artificial Analysis Intelligence Index at launch.
starGemini 3.6 Flash, 3.5 Flash-Lite & 3.5 Flash Cyber
llmReleased July 21. Gemini 3.6 Flash is the new workhorse — up to 17% fewer output tokens and fewer tool calls per task than 3.5 Flash, $1.50/$7.50 per 1M list price (intro rate $0.75/$3.75 through Dec 31, 2026). 3.5 Flash-Lite targets high-throughput/low-latency work at $0.30/$2.50. 3.5 Flash Cyber is tuned for agentic workloads. Google also teased Gemini 4.
starKimi K3
llmReleased July 16 — 2.8T-parameter open-weight MoE, 1M-token context, native vision, full weights due July 27. Moonshot claims it competes with Claude Fable 5 and beats Opus 4.8 and GPT-5.6 Sol; the largest open-weight model to date, triggering a market reaction compared to the original DeepSeek shock.
starGrok 4.5
llmFirst release under the new SpaceXAI brand (post SpaceX merger), built for coding and agentic work. Released July 8. 500K context, $2/$6 per 1M, roughly 2x more token-efficient than comparable frontier models.
GPT-Live-1
multimodalFull-duplex voice models (GPT-Live-1 + mini) that listen and speak simultaneously for natural, interruptible conversation. Released July 8, now default for ChatGPT Voice.
starGPT-5.6 (Sol, Terra, Luna)
llmThree-tier GPT-5.6 family released July 1: Sol (flagship reasoning, $5/$30 per 1M), Terra ($2.50/$15) and Luna (fast, $1/$6) — all with a 1M-token context window.
starClaude Sonnet 5
llmAnthropic's most agentic Sonnet yet, near-Opus performance at lower cost. Released June 30, $2/$10 per 1M intro pricing (through Aug 31), then $3/$15. Default model for free and Pro plans.
starClaude Fable 5
llmAnthropic's new frontier tier above Opus. 1M context, 80.3% SWE-bench, $10/$50 per 1M. Free on Pro/Max through June 22.
Gemini 3.5 Pro (announced)
multimodalAnnounced at Google I/O 2026 (May 19) targeting a June GA; slipped three times (June, July, a widely-reported July 17 date) and remains unreleased as of Sep 9, 2026, limited to a Vertex AI enterprise preview. Reported cause: DeepMind scrapped an earlier build over recursive tool-calling and SVG-generation failures. Slated to power Google AI Mode in Search once shipped.
Siri AI (Apple WWDC 2026)
llmApple rebuilt Siri with generative AI powered by Google Gemini models. Announced at WWDC June 8.
Microsoft MAI-Thinking-1
llmMicrosoft's first in-house reasoning model family, reducing strategic dependence on OpenAI.
starGemini 3.5 Flash
llmFast, multimodal, 1M context. Powers Google AI Mode in Search. Announced at I/O 2026 May 19.
starGPT-5.5
llmReleased April 23. $5/$30 per 1M tokens, ~1M context. Neck-and-neck with Claude Opus on coding.
Grok 4.3
llmxAI flagship released April 17 (full API rollout April 30). $1.25/$2.50 per 1M — frontier quality at mid-tier price. 1M context, native video input.
starDeepSeek V4 Pro
llm1.6T parameter MoE (49B active), MIT license, 1M context, $0.435/$0.87 per 1M. First OSS model at near-frontier level.
GPT-5.4
llmReleased March 5 — flagship refresh between GPT-5 and GPT-5.5.
starGemini 3.1 Pro
multimodalReleased Feb 19. Tops 13/16 benchmarks: 94.3% GPQA Diamond, 80.6% SWE-bench. $2/$12 — 7.5x cheaper than GPT-5.4.
Claude Sonnet 4.6
llmReleased Feb 17 — the workhorse Sonnet tier that preceded Sonnet 5.
Qwen 3.5 Family
llmOpen-weight family topped by a 397B (17B active) flagship with smaller variants. Strongest OSS reasoning line. 256K native context.