Between January 1 and September 23, 2026, 26 labs shipped 101 language-model releases. That is one every 2.6 days. On September 22 alone, Anthropic released Claude Opus 5.5 and OpenAI released GPT-6 Sol and Luna about an hour later.
This is the full list, in date order. Every date below is confirmed by at least two independent sources, usually the vendor’s own post plus press coverage. Items with a single source or conflicting dates are listed separately at the end.
How to read the list
- The list was built from vendor blogs, API changelogs and press coverage, then checked against a daily news archive. Minor checkpoints and snapshot refreshes are not included.
- One entry per launch. A family released on the same day (GPT-5.6 Sol/Terra/Luna) counts once. A preview and its later general release count once.
- Speech, image, video and robotics models are excluded.
- Benchmark figures are the vendors’ own unless marked (AA) for Artificial Analysis. Benchmark versions changed during the year (Terminal-Bench went from 2.0 to 4.0), so numbers from different months are not directly comparable.
- Prices are per million input/output tokens.
January
- Jan 8: Jamba2 (AI21). Hybrid SSM-Transformer, 256K context, Apache 2.0.
- Jan 20: LongCat-Flash-Thinking-2601 (Meituan). 560B MoE reasoning model, MIT licence.
- Jan 22: ERNIE 5.0 (Baidu). 2.4T-parameter omni-modal MoE, generally available after a November 2025 preview.
- Jan 27: Kimi K2.5 (Moonshot). 1T parameters, 32B active, native vision, up to 100 coordinated sub-agents. SWE-bench Verified 76.8.
February
- Feb 5: Claude Opus 4.6 (Anthropic). $5/$25. Terminal-Bench 2.0 65.4%, ARC-AGI-2 68.8%.
- Feb 5: GPT-5.3-Codex (OpenAI). In Codex Feb 5, API Feb 24. Terminal-Bench 2.0 77.3%, SWE-bench Pro 56.8%.
- Feb 11: GLM-5 (Zhipu). 744B/40B active, 200K context, MIT, $1.00/$3.20.
- Feb 12: Gemini 3 Deep Think upgrade (Google). ARC-AGI-2 84.6% (verified by ARC Prize), Humanity’s Last Exam 48.4%.
- Feb 12: MiniMax M2.5 (MiniMax). 230B/10B active, open weights. SWE-bench Verified 80.2%, $0.15/$1.20.
- Feb 12: GPT-5.3-Codex-Spark (OpenAI). Research preview on Cerebras hardware, over 1,000 tokens per second.
- Feb 14: Doubao-Seed-2.0 (ByteDance). Pro/Lite/Mini/Code, closed weights. SWE-bench Verified 76.5.
- Feb 16: Qwen3.5 (Alibaba). 397B/17B active, Apache 2.0, 262K native context. SWE-bench Verified 76.4.
- Feb 17: Claude Sonnet 4.6 (Anthropic). $3/$15, 1M context in beta. SWE-bench Verified 79.6%.
- Feb 17: Grok 4.20 (xAI). Beta, API on Mar 10. Four-agent system, 2M context.
- Feb 19: Gemini 3.1 Pro (Google). Preview. ARC-AGI-2 77.1%, more than double Gemini 3 Pro.
March
- Mar 3: GPT-5.3 Instant (OpenAI). 26.8% fewer hallucinations with web search on high-stakes evals.
- Mar 3: Gemini 3.1 Flash-Lite (Google). Preview, GA May 7. $0.25/$1.50, 1M context.
- Mar 4: Phi-4-reasoning-vision-15B (Microsoft). 15B open multimodal reasoning model.
- Mar 5: GPT-5.4 and 5.4 Pro (OpenAI). $2.50/$15, about 1M context, native computer use. OSWorld 75% against a 72.4% human baseline.
- Mar 5: Olmo Hybrid 7B (Ai2). Fully open. Matches Olmo 3 on MMLU with 49% fewer training tokens.
- Mar 11: Nemotron 3 Super (NVIDIA). 120B/12B active hybrid Mamba-Transformer, 1M context, open weights.
- Mar 16: Mistral Small 4 (Mistral). 119B/6B active, Apache 2.0, $0.15/$0.60.
- Mar 17: GPT-5.4 mini and nano (OpenAI). $0.75/$4.50 and $0.20/$1.25.
- Mar 18: MiniMax M2.7 (MiniMax). Weights on Apr 12. SWE-Pro 56.2%.
- Mar 18: MiMo-V2-Pro (Xiaomi). 1T/42B active, 1M context, $1/$3. Previously tested anonymously as “Hunter Alpha”.
April
The busiest month: 16 releases, 13 of them with two-source dates.
- Apr 2: Gemma 4 (Google). Open weights up to 31B, Apache 2.0, 256K context, 140+ languages.
- Apr 7: Claude Mythos Preview (Anthropic). Restricted to security partners in Project Glasswing.
- Apr 7: GLM-5.1 (Zhipu). Open weights, MIT. SWE-bench Pro 58.4, first place at release.
- Apr 8: Muse Spark (Meta). First model from Meta Superintelligence Labs. Closed weights. Meta shipped no Llama model this year.
- Apr 16: Claude Opus 4.7 (Anthropic). SWE-bench Pro 64.3%, up from 53.4%. New tokenizer.
- Apr 16: Qwen3.6-35B-A3B (Alibaba). 35B/3B active, Apache 2.0. SWE-bench Verified 73.4.
- Apr 17: Grok 4.3 (xAI). API on Apr 30. 1M context, video input, $1.25/$2.50.
- Apr 22: MiMo-V2.5-Pro (Xiaomi). 1.02T/42B active, 1M context, MIT.
- Apr 23: GPT-5.5 and 5.5 Pro (OpenAI). API on Apr 24. $5/$30, twice the price of GPT-5.4. Terminal-Bench 2.0 82.7%.
- Apr 23: Hy3 preview (Tencent). 295B/21B active, 256K context. Licence excluded the EU, UK and South Korea.
- Apr 24: DeepSeek-V4 Preview (DeepSeek). V4-Pro 1.6T/49B active, V4-Flash 284B/13B active. MIT, 1M context by default. Flash at $0.14/$0.28.
- Apr 28: Nemotron 3 Nano Omni (NVIDIA). 30B/3B active, text, image, video and audio input.
- Apr 29: Granite 4.1 (IBM). Dense 3B/8B/30B, up to 512K context, Apache 2.0.
May
- May 19: Gemini 3.5 Flash (Google). GA at I/O. Terminal-Bench 2.1 76.2%, $1.50/$9.
- May 19: Qwen3.7-Max (Alibaba). Proprietary. GPQA 92.4.
- May 20: Command A+ (Cohere). 218B/25B active. First Cohere flagship under Apache 2.0.
- May 28: Claude Opus 4.8 (Anthropic). 1M context. SWE-bench Verified 88.6%, GPQA 93.6%. Fast mode 3x cheaper.
- May 29: Step 3.7 Flash (StepFun). 198B vision-language MoE, Apache 2.0, $0.20/$1.15.
June
- Jun 1: MiniMax M3 (MiniMax). 1M context, image and video input. SWE-bench Pro 59.0%, $0.60/$2.40.
- Jun 2: Seven MAI models (Microsoft). Led by MAI-Thinking-1, 35B active, 256K context, trained in-house.
- Jun 4: Nemotron 3 Ultra (NVIDIA). 550B/55B active, 1M context, open weights. Index 48 (AA).
- Jun 8: Apple Foundation Models 3 (Apple). Announced at WWDC, ships with OS 27. A 3B on-device model, a 20B sparse on-device model and cloud models.
- Jun 9: Claude Fable 5 and Mythos 5 (Anthropic). A new tier above Opus, $10/$50, 1M context. Mythos 5 limited to vetted organizations. Both suspended Jun 12, restored Jul 1.
- Jun 13: GLM-5.2 (Zhipu). 1M context, MIT. SWE-bench Pro 62.1. Top open model at index 51 (AA).
- Jun 24: Seed 2.1 Pro and Turbo (ByteDance). Closed weights.
- Jun 30: Claude Sonnet 5 (Anthropic). $2/$10, 1M context, 128K output.
- Jun 30: LongCat-2.0 (Meituan). 1.6T/48B active, 1M context, trained on Chinese chips, MIT.
July
- Jul 6: Hy3 (Tencent). Official release, now Apache 2.0 with no regional exclusions.
- Jul 8: Grok 4.5 (xAI). 500K context, $2/$6. SWE-bench Pro 64.7%.
- Jul 9: GPT-5.6 Sol, Terra and Luna (OpenAI). Partner preview from Jun 26. 1M context. Sol $5/$30, Terra $2.50/$15, Luna $1/$6.
- Jul 9: Muse Spark 1.1 (Meta). 1M context, public API preview.
- Jul 15: Inkling (Thinking Machines). 975B multimodal, Apache 2.0. First model from the lab.
- Jul 16: Kimi K3 (Moonshot). 2.8T parameters, about 50B active, the largest open-weight model to date. Weights on Jul 27.
- Jul 19: Qwen3.8-Max (Alibaba). Preview, GA Aug 3. 2.4T/95B active, 1M context, $2/$6.
- Jul 21: Gemini 3.6 Flash and 3.5 Flash-Lite (Google). OSWorld-Verified 83.0% for 3.6 Flash.
- Jul 24: Claude Opus 5 (Anthropic). $5/$25, half the price of Fable 5.
- Jul 31: DeepSeek V4-Flash-0731 (DeepSeek). Final V4-Flash checkpoint after the preview.
August
- Aug 4: LFM2.5-2.6B (Liquid AI). On-device agent model.
- Aug 5: Muse Spark 1.2 and Muse Code (Meta). Muse Code is a terminal coding agent.
- Aug 10: Muse Glimmer (Meta). 30B, Apache 2.0, runs on one 24 GB GPU.
- Aug 12: Grok 4.6 (xAI). $2/$6. Index 61, level with GPT-5.6 Sol (AA).
- Aug 12: Qwen3.8 open weights (Alibaba). Open version of 3.8-Max. Commercial licence required above $50M revenue.
- Aug 13: Gemini 3.7 Flash (Google). $0.75/$3.75 introductory price until Dec 31.
- Aug 13: DeepSeek V4-Pro GA (DeepSeek). Low/high/max reasoning effort, 50% off-peak discount.
- Aug 14: GLM-5.3 (Zhipu). $1.40/$4.40. Terminal-Bench 3.0 28.3%, up from 4.6%.
- Aug 14: Qwen3.8-27B (Alibaba). Dense 27B, Apache 2.0. Index 52 (AA).
- Aug 26: Qwen3.8-Flash (Alibaba). Fast proprietary tier.
- Aug 26: GLM-5.3-Flash (Zhipu). 320B/18B active, open weights. The stealth model “Ox Alpha”.
- Aug 28: Hy4 preview (Tencent). 770B/49B active, 1M context, $0.83/$2.50.
September
- Sep 1: Claude Fable 5.1 and Mythos 5.1 (Anthropic). List price unchanged at $10/$50, cache reads 75% cheaper.
- Sep 2: Gemini 3.8 Flash (Google). Third Flash in six weeks. Terminal-Bench 2.1 90.8%, $0.75/$3.75.
- Sep 3: GPT-6 Astra (OpenAI). Vetted users first, GA Sep 4. $10/$50, 1.05M context. First model OpenAI rates “Critical” for cybersecurity.
- Sep 10: DeepSeek V4.1-Flash (DeepSeek). 552B MoE, new encoder-decoder design, KV cache at a quarter of the memory.
- Sep 21: Grok 4.7 (xAI). New base model, 500K context, $2/$6.
- Sep 22: Claude Opus 5.5 (Anthropic). $4/$20, down from $5/$25. Terminal-Bench 4.0 66.4% against 52.3% for Opus 5.
- Sep 22: GPT-6 Sol and Luna (OpenAI). Cheaper models derived from Astra.
- Sep 22: MiMo-V2.6 (Xiaomi). Omni-modal, 1M context, MIT. Top open model at index 46 (AA).
One source only, or conflicting dates
These releases happened, but only one source gives the exact date, or the sources disagree:
- Jan 5, LFM2.5-1.2B (Liquid AI). Jan 14, GPT-5.2-Codex in the API (OpenAI). Jan 19, GLM-4.7-Flash (Zhipu).
- Step 3.5 Flash (StepFun): Jan 29, Feb 2 or Feb 12.
- Qwen3.6-Plus (Alibaba): Mar 31 or April.
- Apr 1, GLM-5V-Turbo (Zhipu). Apr 20, Kimi K2.6 (Moonshot).
- Mistral Medium 3.5: Apr 29 or 30. ERNIE 5.1 (Baidu): Apr 29 or May 8.
- May 28, LFM2.5-8B-A1B (Liquid AI). Early June, Gemma 4 12B (Google). Jun 11, DiffusionGemma (Google). Jun 14, Kimi K2.7-Code (Moonshot). June, Qwen3.7-Plus (Alibaba).
- Aug 10, GPT-5.6-Cyber (OpenAI). Aug 21, V4-Flash-Vision-Exp (DeepSeek). Sep 2, Muse Spark 1.3 (Meta). Sep 17, Qwen3.8-Omni-Flash (Alibaba). Sep 18, GLM-5.3-FlashX (Zhipu).
Counts
All 101 releases, including the single-source ones:
- By vendor: OpenAI 11. Anthropic, Google and Alibaba 10 each. Zhipu 8. xAI, Meta and DeepSeek 5 each.
- By month: Jan 7, Feb 12, Mar 11, Apr 16, May 7, Jun 13, Jul 10, Aug 14, Sep (to the 23rd) 11.
- 57 releases came from US and European labs, 44 from Chinese labs.
What changed over the year
Flagship prices moved in opposite directions. OpenAI’s top model went from $2.50/$15 (GPT-5.4) to $10/$50 (GPT-6 Astra). Anthropic’s Opus went from $5/$25 to $4/$20. Google launched Gemini 3.7 and 3.8 Flash at $0.75/$3.75.
1M-token context became standard by mid-year: GPT-5.6 and GPT-6, every Claude model from Opus 4.8 on, Gemini Flash, DeepSeek V4, Kimi K3, GLM-5.2, MiniMax M3 and Qwen3.8-Max.
Chinese open-weight releases converged on one design: sparse MoE with 20 to 50B active parameters, 1M context, SWE-bench Pro between 57 and 62, and prices under $1.50 input.
Announced but not released by September 23
Gemini 3.5 Pro (announced at I/O in May, still in partner testing), Gemini 4, Grok 5, DeepSeek R2, OpenAI o5, a new gpt-oss, Qwen 4, any new Llama, Mistral Large 4, a new Claude Haiku.
Cover image: Anthropic, Claude Opus 5.5 announcement page, September 22, 2026.