Models & Benchmarks
New model releases, capability benchmarks, and evaluations
Living document tracking new model releases, capability benchmarks, and evaluations. Newest entries appear at the top.
Key Areas to Watch
- Foundation model releases (GPT, Claude, Gemini, Llama, Mistral)
- Benchmark results and methodology (MMLU, HumanEval, SWE-bench, GPQA)
- Scaling laws and training advances
- Fine-tuning and RLHF/DPO techniques
- Model safety and alignment evaluations
2026-W39
- Introducing System One Models and Jev — Constrained decision output makes Jev a fast specialized option for classification, routing, and agent gates. [HN]
- Gemini 3.8 Live and 3.8 Live Extended Thinking — Real-time multimodal model update with an extended-thinking mode for harder interactive tasks. [DeepMind / HN]
- Grok 4.7 — xAI’s new frontier release gives developers another model to evaluate for tool use and agent workloads. [HN]
- Your Agent Aced the Task. Will It Do It Again? — Agent evaluation shifts from one successful result toward repeatability across runs. [Hugging Face]
2026-W38
- GPT-6 Astra for work — OpenAI positioned Astra around work, tool use, and production systems rather than only benchmark claims. [OpenAI]
- DeepSeek V4.1 Flash — Multimodal MoE with 1M context and KV-cache emphasis drew developer attention. [HN / x.com]
- SWE-Bench Pro Verified — Verified benchmark variant reduces reward-hacking/task-quality artifacts and changes how frontier coding models should be compared. [x.com]
- Meta Muse Image and Muse Video — Creative-generation models with editing, references, and native audio support. [Meta AI]
2026-W37
- GPT-6 Astra — OpenAI’s new model dominated developer discussion around long-running coding and agentic work. [OpenAI / HN]
- Claude Fable 5.1 and Claude Mythos 5.1 — Anthropic’s update became the other major frontier comparison point for coding-agent workflows. [Anthropic / HN]
- Gemini 3.8 Flash and 3.8 Flash Cyber — Fast Gemini release plus security-specialized variant for cyber workflows. [Google / HN]
2026-W36
- GLM-5.3 open weights — Z.ai kept GLM in the open-weight coding and agent model comparison set. [HN]
- Qwen3.8-Flash-Next — Fast Qwen variant relevant to cheap routing paths inside agent loops. [HN]
- Granite 4.2 LLMs — IBM published build details for its latest open Granite family. [Hugging Face]
2026-W32
- Qwen3.8-Max — Coding and cowork claims put Qwen back in the weekly model-selection debate. [HN]
- GPT-5.6 price-performance update — Permanent Luna and Terra cost reductions change agent-routing economics. [OpenAI]
- LFM2.5-2.6B — Small local model release aimed at deployable agent loops outside frontier-only APIs. [Hugging Face]
2026-W28
- Claude Sonnet 5 — Anthropic updated Sonnet with tokenizer and performance changes that affect API planning. [Anthropic]
- Mistral 3 — Mistral released Apache-licensed dense small models and a large sparse MoE. [Mistral]
- GeneBench-Pro — OpenAI added a genomics/biology benchmark for scientific AI capability. [OpenAI]
2026-W27
- Previewing GPT-5.6 Sol — OpenAI previewed a next-generation model for coding, science, and cybersecurity with heavier access controls. [OpenAI]
- Introducing Claude Opus 4.8 — Anthropic emphasized stronger coding and agentic performance plus Claude Code dynamic workflows. [Anthropic]
- GLM 5.2 beats Claude in our benchmarks — Semgrep kept GLM-5.2 in the security-model benchmark conversation. [HN]
- Qwen3-ASR — Qwen released a compact open ASR model for 52 languages and dialects. [x.com]
2026-W26
- GLM-5.2 — Open-weight long-horizon model release that dominated developer comparisons this week. [Hugging Face]
- GLM-5.2 on Artificial Analysis — Benchmark coverage that made GLM-5.2 a serious reference point for open-model evaluation. [HN]
- Multi-LCB — Expands LiveCodeBench across multiple programming languages for more realistic coding-model evaluation. [arXiv]
- How Transparent is DiffusionGemma? — Tests whether diffusion language models make reasoning harder to inspect. [arXiv]
2026-W25
- Claude Fable 5 and Claude Mythos 5 — Anthropic introduced a Mythos-class model family with explicit long-horizon software-engineering claims, then suspended access days later after a government directive. [Anthropic / HN]
- olmo-eval — AllenAI released evaluation tooling for the model-development loop, useful for teams comparing checkpoints and regressions over time. [Hugging Face]
- When Good Verifiers Go Bad — Shows self-improving VLM verifiers can regress on new tasks, a warning for automated evaluation pipelines. [arXiv]
- ClinHallu — Diagnoses hallucinations by reasoning stage, a pattern worth borrowing for domain-specific model evals. [arXiv]
2026-W24
- Gemma 4 12B: A unified, encoder-free multimodal model — Google’s compact multimodal release keeps capable model deployment moving below frontier-only sizes. [HN]
- Introducing Mellum2: A 12B Mixture-of-Experts Model by JetBrains — JetBrains released a 12B MoE model aimed at developer-assistant workflows. [Hugging Face]
- How reliable are LLMs when it comes to playing dice? — Probability benchmark shows high accuracy on standard problems but much weaker robustness on counterintuitive variants. [arXiv]
- Sycophantic Praise: Evaluating Excessive Praise in Language Models — Separates praise/flattery from generic agreement as a model behavior worth measuring. [arXiv]
2026-W23
- Claude Opus 4.8 — Anthropic’s Opus refresh drove the week’s largest model discussion; independent notes called it a modest but tangible improvement while coding benchmarks stayed contested. [HN / Anthropic]
- NVIDIA Cosmos 3 for Physical AI — Open omni-model positioned for physical reasoning and action, connecting world-model work to robotics and embodied AI. [Hugging Face]
- Introducing Mellum2: A 12B Mixture-of-Experts Model by JetBrains — JetBrains released a compact MoE model aimed at developer-assistant and code-adjacent workflows. [Hugging Face]
- Disagreement among frontier LLMs on real-world fact-checks — Real-world fact-checking comparison showed frontier models still diverge on factual judgments, underscoring evaluation and reliability gaps. [HN]
- NVIDIA quantized Qwen3.6 MoE on Hugging Face — HuggingPapers highlighted a 35B-total, 3B-active NVFP4 release claiming roughly 3x memory reduction with near-zero accuracy loss. [x.com]
2026-W22
- Gemini 3.5 Flash — Google release drew major developer attention as a fast model option. [HN]
- Qwen3.7-Max: The Agent Frontier — Qwen positions the model around frontier agentic task execution. [Qwen]
- Mistral 3 — Mistral ships an NVFP4/vLLM-oriented checkpoint for Blackwell NVL72 and single-node 8xA100/8xH100 serving. [Mistral]
- grok-code-fast-1 — xAI describes a fast and economical reasoning model for agentic coding. [xAI]
2026-W20
- Gemini 3.5 Flash — Google’s new fast, low-cost Gemini tier for high-throughput agentic workloads, released around I/O 2026. [Google]
- Google I/O 2026 keynote — Gemini Spark personal agent, Gemini for Science, and conversational AI in Docs/YouTube; Google reported Gemini now serving 9.7 trillion tokens/month. [Tom’s Guide]
- Predictable Confabulations — Factual recall accuracy by LLMs scales predictably with model size and topic frequency, giving a quantitative handle on when recall is trustworthy. [arXiv]
- DashAttention — A learnable, adaptive sparse hierarchical attention scheme aimed at long-context efficiency. [arXiv]
2026-W19
- GPT-5.5 Instant rolling out to all ChatGPT users — Replaces GPT-5.3 Instant as the default ChatGPT model; system card confirms improved instruction-following and personalization. [OpenAI]
- Scaling Trusted Access for Cyber with GPT-5.5 and GPT-5.5-Cyber — Dedicated cybersecurity model variant with tiered access; extends the GPT-4.5-Cyber access model. [OpenAI]
- DeepSeek 4 Flash local inference engine for Metal — Salvatore Sanfilippo ships a Metal-native inference engine for DeepSeek 4 Flash; low-latency local option on Apple Silicon. [HN]
- WildClawBench: real-world long-horizon agent evaluation by Shuangrui Ding et al. — Benchmark drawn from real user logs requiring 20–100 step task horizons; leading models top out at 31% on the hardest split. [arXiv]
- AssayBench: assay-level virtual cell benchmark for LLMs and agents by Edward De Brouwer et al. — Drug-discovery eval on cellular assay reasoning; leading models score 43%. [arXiv]
- Hugging Face co-founder: Qwen 3.6 27B on airplane mode is close to last year’s frontier — 2,329-point r/ClaudeAI post on offline open-weight performance crossing a threshold. [Reddit]
2026-W18
- Mistral Medium 3.5 — 128B dense unified chat/reasoning/code; 256k context; $1.50 / $7.50 per 1M tokens; 77.6% SWE-Bench Verified; modified MIT weights. [Mistral]
- Grok 4.3 (full rollout) — 1M context, native video input up to 5 min @ 1080p, in-chat slide/PDF/Excel generation, ~40% cheaper inputs, retains 16-agent Heavy mode. [xAI]
- Microsoft DELULU benchmark — New evaluation suite for fill-in-the-middle code completion; positioned as a corrective to SWE-Bench saturation. [HuggingPapers / x.com]
- NVIDIA Nemotron 3 Nano Omni — Unified vision/audio/language model with long-context multimodal intelligence for document understanding. [NVIDIA]
- Granite 4.1 LLMs: How They’re Built — IBM’s open-weight Granite 4.1 release with build-process detail. [Hugging Face]
- Adding Benchmaxxer Repellant to the Open ASR Leaderboard — Hugging Face hardens the Open ASR Leaderboard against gaming. [Hugging Face]
- Opus 4.7 is a genuine regression and I’m tired of pretending it isn’t — 755-point r/ClaudeAI thread; community-side counter to Anthropic’s April 23 postmortem. [Reddit]
- Qwen3.6-27B vs Coder-Next — Side-by-side coding evaluation of Alibaba’s dense flagship. [Reddit]
- Accelerating Gemma 4: faster inference with multi-token prediction drafters — Google ships MTP drafters for Gemma 4. [Google]
- OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50–55% by triage doctors — Harvard trial of emergency triage diagnosis; HN 503. [HN]
2026-W17
- DeepSeek V4-Pro and V4-Flash — DeepSeek’s first V4 preview models; both 1M-context MoE open weights claiming near-frontier quality at a fraction of the price. 2,074 HN points. [HN]
- GPT-5.5 — OpenAI’s new flagship; available in Codex, rolling out to ChatGPT. Major lift on agentic coding and computer use. [OpenAI]
- GPT-5.5 System Card — Public eval and safety report for GPT-5.5. [OpenAI]
- Qwen3.6-27B — Alibaba’s dense 27B model claims flagship-level agentic coding, surpassing Qwen3.5-397B-A17B (15x the active parameter count). 989 HN points. [HN]
- Qwen3.6-Max-Preview — Larger, sharper Qwen3.6 variant. 705 HN points. [HN]
- Qwen3.6-35B-A3B Heretic — Uncensored Qwen3.6 35B-A3B variant at 0.0015 KLD, called the best 35B available. [Reddit]
- Qwen 3.6 27B ties Sonnet 4.6 on Artificial Analysis — Open-weight model in a dead heat with a frontier closed model on the Intelligence Index. [Reddit]
- Kimi K2.6 is a legit Opus 4.7 replacement — 1,222-point r/LocalLLaMA thread on Kimi 2.6 hitting Opus-class quality. [Reddit]
- Xiaomi MiMo V2.5 Pro — Xiaomi’s flagship places at 54 on the Artificial Analysis Index; weights coming. [Reddit]
- ChatGPT Images 2.0 — New image-generation model with stronger text rendering, multilingual support, and visual reasoning; Altman compared the leap to GPT-3 → GPT-5. [OpenAI]
- OpenAI deprecates SWE-Bench Verified — OpenAI publicly stops reporting SWE-Bench Verified scores, citing benchmark gaming. [OpenAI]
- OpenAI HealthBench Professional — Medical evaluation benchmark on Hugging Face for AI assistants serving clinicians. [OpenAI / Hugging Face]
- Claude Token Counter, now with model comparisons — Simon Willison’s tokenizer-cost comparison tool now supports Opus 4.6 vs 4.7. [Simon Willison]
2026-W16
- Claude Opus 4.7 — Anthropic’s new flagship; first Opus-line tokenizer change, updated system prompt,
thinking_effort: xhighexposed via the API. 1,954 HN points. [Anthropic] - Qwen3.6-35B-A3B — Alibaba agentic-coding MoE; competitive with frontier closed models on a laptop. Simon Willison’s pelican benchmark prefers it to Opus 4.7. [Qwen]
- Qwen3.6-Max-Preview — Larger, sharper Qwen3.6 variant landing days after 35B-A3B. [Qwen]
- Kimi K2.6 — Moonshot open weights drop; 724-point r/LocalLLaMA thread. [Hugging Face / Reddit]
- Gemini 3.1 Flash TTS — DeepMind adds granular audio tags for fine-grained expressive speech control. [DeepMind]
- Gemini Robotics-ER 1.6 — Spatial reasoning + multi-view understanding for autonomous robotics. [DeepMind]
- GPT-Rosalind — OpenAI’s life-sciences reasoning model for drug discovery, genomics, and protein reasoning. [OpenAI]
- xAI Grok 4.3 Beta — Conversational video understanding, direct generation of downloadable PDFs / spreadsheets / PowerPoint decks. [xAI]
- Measuring Claude 4.7’s tokenizer costs — First independent cost-impact study of the Opus 4.7 tokenizer change. [HN]
- Claude Opus 4.7 Text Category Rankings — Community leaderboard of Opus 4.7 capabilities across categories. [Reddit]
2026-W15
- Muse Spark — First model from Meta Superintelligence Labs (code name Avocado); natively multimodal reasoning, Contemplating mode for multi-agent parallel reasoning aimed at Gemini Deep Think / GPT Pro tier. [Meta AI]
- GLM-5.1: Towards Long-Horizon Tasks — Zhipu’s GLM-5.1 explicitly targets long-horizon agentic tasks; dedicated Simon Willison coverage. [z.ai]
- Minimax M2.7 Released — MiniMax drops M2.7 open weights; 668-point r/LocalLLaMA thread. [Hugging Face]
- System Card: Claude Mythos Preview — Anthropic ships a restricted security-research model with Project Glasswing gating; zero-day claims contested by Berkeley RDI and Aisle analyses. [Anthropic]
- Gemma 4 audio support lands in llama-server — Native multimodal audio inference for Gemma 4 in llama-server, plus MLX audio pipeline from Simon Willison. [Reddit / Simon Willison]
2026-W14
- Gemma 4 — Google DeepMind’s strongest open model family yet; purpose-built for advanced reasoning and agentic workflows; ships with on-device support via Google AI Edge Gallery on iPhone (868 HN points). [DeepMind]
- Qwen3.6-Plus: Towards real world agents — Alibaba’s agentic-focused update to the Qwen3 line; 596 HN points on launch. [Qwen / Alibaba]
- 1-Bit Bonsai: first commercially viable 1-bit LLMs — PrismML claims viable 1-bit quantized models for production deployment; 430 HN points. [PrismML / HN]
2026-W13
- iPhone 17 Pro Running a 400B LLM On-Device — Demo of a 400B parameter model running locally on iPhone 17 Pro hardware; 713 HN points; sparked debate about quantization methodology and on-device inference limits. [Twitter / HN]
- GPT-5.4 Pro solves open FrontierMath problem — Epoch AI confirms GPT-5.4 Pro resolved a previously open Ramsey hypergraph theory problem listed on FrontierMath; strongest signal yet of genuine mathematical reasoning from a frontier model. [Epoch AI / HN]
- TurboQuant: Redefining AI efficiency with extreme compression — Google Research quantization technique pushes beyond INT4 with competitive benchmark accuracy; 576 HN points. [Google Research]
- Gemini 3.1 Flash Live: lower-latency voice model — DeepMind ships improved precision and latency for real-time audio AI interactions. [DeepMind]
2026-W12
- Introducing GPT-5.4 mini and nano — Smaller, faster GPT-5.4 variants optimized for coding, tool use, multimodal reasoning, and high-volume sub-agent pipelines. [OpenAI]
- Mamba-3 — Together AI’s next-generation state space model architecture; 300 HN points. [Together AI]
- Leanstral: Open-source agent for trustworthy coding and formal proof engineering — Mistral’s Lean 4 formal proof agent; 783 HN points, 191 comments. [Mistral AI]
- Holotron-12B — High Throughput Computer Use Agent — HCompany’s 12B model optimized for high-throughput computer use and agentic workflows. [Hugging Face]
2026-W11
- Claude 1M-context window goes GA for Opus 4.6 and Sonnet 4.6 — Anthropic removes limited-access restriction; largest GA production context window available to all API users; 1,220 HN points, 520 comments. [Anthropic]
2026-W10
- Introducing GPT-5.4 — OpenAI’s most capable and efficient frontier model; 1M-token context, state-of-the-art coding, computer use, and tool search; 1,019 HN points, 807 comments. [OpenAI]
- Gemini 3.1 Flash-Lite: Built for intelligence at scale — Google DeepMind’s fastest, most cost-efficient Gemini 3 series model. [DeepMind]
- Grok 4.20 Beta 2 — xAI update with better instruction following, fewer hallucinations, LaTeX and multi-image fixes. [xAI]
- Phi-4-reasoning-vision: lessons of training a multimodal reasoning model — Microsoft Research’s 15B open-weight multimodal reasoning model; available on Hugging Face and GitHub. [Microsoft Research]
2026-W09
- The Llama 3 Herd of Models — Meta AI’s full technical paper covering the 405B dense model: 128K context, multilingual capability, coding benchmarks, reasoning, and native tool use; foundational reference for anyone building on the Llama lineage. [Meta AI]
- Introducing Mercury 2: Fast Reasoning LLM powered by diffusion — Inception Labs’ diffusion-based reasoning LLM challenges autoregressive latency assumptions; competitive speed and quality at lower latency. [Inception Labs / HN]
- Nano Banana 2 — Google’s Flash-speed image generation model with Pro-level world knowledge and subject consistency; 605 HN points. [DeepMind]
2026-W08
- Claude Sonnet 4.6 — Anthropic’s mid-tier flagship refresh; top HN post of its launch week at 1,346 points and 1,226 comments. [Anthropic]
- Gemini 3.1 Pro: A smarter model for your most complex tasks — DeepMind’s reasoning-quality step-up in the Gemini 3 line; 963 HN points, 914 comments. [DeepMind]
- Qwen3.5: Towards Native Multimodal Agents — Alibaba positions Qwen3.5 as a full multimodal-agent platform with native multimodal framing; 434 HN points. [Qwen / Alibaba]
- Grok 4.20 Beta — xAI’s latest Grok beta ships with new capabilities; community benchmarks circulate on r/LocalLLaMA. [xAI]
2026-W07
- Gemini 3 Deep Think: Advancing science, research and engineering — Google DeepMind’s updated frontier reasoning mode targeting hard problems in modern science, math, and engineering. [DeepMind]
- Introducing GPT-5.3-Codex-Spark — OpenAI’s first real-time coding model; 15x faster generation than predecessor with 128K context window; research preview for Pro users. [OpenAI]
- GLM-5: Targeting complex systems engineering and long-horizon agentic tasks — Zhipu AI’s GLM-5 focuses on agentic engineering workloads; available on NVIDIA NIM; 484 HN points. [z.ai]
- GPT-5.2 derives a new result in theoretical physics — GPT-5.2 proposes a new gluon amplitude formula later formally verified by academic collaborators; 574 HN points. [OpenAI]
2026-W06
- Claude Opus 4.6 — Anthropic’s latest frontier release; top HN post of its launch week at 2,346 points and 1,031 comments; positioned as “a space to think” for long-horizon reasoning. [Anthropic]
- Introducing GPT-5.3-Codex — OpenAI’s coding-focused frontier model ships same day as Claude Opus 4.6; 1,530 HN points, 605 comments; includes new macOS Codex app. [OpenAI]
- Voxtral Transcribe 2 — Mistral’s next-generation ASR model; 1,012 HN points, 241 comments. [Mistral AI]
- Qwen3-Coder-Next teaser — Alibaba’s upcoming coding model generates early community excitement; 735 HN points, 429 comments. [Qwen / HN]
2026-W05
- Qwen3-Max-Thinking — Alibaba’s extended-thinking frontier model enters o1-class reasoning territory; 502 HN points, 424 comments on launch day. [Qwen / Alibaba]
- Retiring GPT-4o, GPT-4.1, GPT-4.1 mini, and o4-mini in ChatGPT — OpenAI sets February 13 sunset for older model family; GPT-5 line becomes the primary offering, forcing production integration updates. [OpenAI]
2026-W04
- Qwen3-TTS family open sourced — Alibaba open-sources a full TTS suite including voice cloning and voice design; 744 HN points. [Qwen / Alibaba]
- Differential Transformer V2 — Microsoft extends the differential attention architecture for improved long-context performance. [Hugging Face / Microsoft]
2026-W03
- FLUX.2 Klein: Towards Interactive Visual Intelligence — Black Forest Labs previews FLUX.2 Klein, targeting interactive and iterative visual generation workflows. [Black Forest Labs / HN]
- MedGemma 1.5 and MedASR — Google Research ships MedGemma 1.5 for medical imaging interpretation plus a medical speech-to-text model. [Google Research]
2026-W02
- NVIDIA Cosmos Reason 2 — 2B/8B open reasoning VLMs; #1 on Physical AI Bench and Physical Reasoning leaderboards. [NVIDIA / Hugging Face]
- Falcon-H1 Arabic (3B/7B/34B) — Hybrid Mamba/Transformer; 3B hits 61.87% OALL (10pts ahead of Phi-4 Mini); 7B hits 71.47%. [TII / Hugging Face]
2026-W01
- GLM-Image teaser (Z.ai) — Two-stage compact-token encoder scaling to 4K; reportedly first top-tier multimodal trained entirely on Huawei Ascend. 311 upvotes. [r/LocalLLaMA / eWEEK]
- FLUX.2-dev-Turbo editing test — Community benchmarks on FLUX.2-dev-Turbo for image editing; 92 upvotes. [r/LocalLLaMA]
2025-W52
- GLM-4.7 (Z.ai) — Open weights; at/above Sonnet 4.5 on SWE-bench Verified, LiveCodeBench v6, Terminal Bench 2.0; #1 open on Code Arena; τ²-Bench 87.4. [Z.ai / Hugging Face]
- MiniMax M2.1 — 230B/10B-active MoE; multi-language (Rust/Java/Go/C++/Kotlin/ObjC/TS/JS); introduces VIBE runtime benchmark. [MiniMax]
2025-W51
- Gemini 3 Flash — Pro-grade reasoning at Flash speed; GPQA Diamond 90.4%, HLE 33.7%, MMMU Pro 81.2%; default in Gemini app and AI Mode. [Google]
- NVIDIA Nemotron 3 — Hybrid MoE open family; Nano (30B/3B active) ships now; Super (~100B/10B) and Ultra (~500B/50B) in H1 2026. [NVIDIA]
- GPT-5.2-Codex — Long-horizon software-engineering model; system-card addendum. [OpenAI]
2025-W50
- GPT-5.2 — Instant/Thinking/Pro; SOTA on GDPval; 30% fewer hallucinations vs 5.1 Thinking. [OpenAI]
- Qwen3-Next-80B-A3B-Thinking-GGUF — 80B MoE thinking model in GGUF for local use. [Qwen / Hugging Face]
2025-W49
- Mistral Large 3 — 675B-total / 41B-active MoE; 256K context; multimodal + 40+ languages; Apache 2.0. [Mistral AI]
- Apriel-1.6-15b-Thinker — ServiceNow’s 15B reasoning model. [Hugging Face]
2025-W48
- Claude Opus 4.5 — 80.9% SWE-bench Verified — first model over 80%; first to beat every human on Anthropic’s internal eval; $5/$25 per million tokens (66% cheaper than Opus 4.1). [Anthropic]
- FLUX.2 in Diffusers — Black Forest Labs’ FLUX.2 open-weights release integrated with Hugging Face Diffusers. [Hugging Face]
2025-W47
- Gemini 3 — LMArena 1501; HLE 37.5% (no tools); GPQA Diamond 91.9%; MathArena Apex 23.4% SOTA; Deep Think pushes HLE to 41%. [DeepMind]
- Grok 4.1 — LMArena #2 Thinking (1483 Elo); 2M-token context; $0.20 / $0.50 per million tokens; 3x lower hallucination. [xAI]
- GPT-5.1-Codex-Max — Long-horizon coding; first OpenAI model trained to operate in Windows. [OpenAI]
- Nano Banana Pro — Gemini 3 Pro Image model. [DeepMind]
2025-W46
- GPT-5.1 (Instant + Thinking) — Adaptive reasoning on Instant; 8 selectable personalities; rolled out to ChatGPT paid tiers Nov 12. [OpenAI]
- GPT-5.1-Codex / Codex-Mini in public preview for GitHub Copilot — Copilot rollout Nov 13. [GitHub]
- SIMA 2 — Gemini-powered 3D-world agent; task completion 31% → 62%; open-ended self-improvement. [DeepMind]
2025-W45
- BERTs that chat with dLLM — Recipe converting any BERT-style encoder into an instruction-following diffusion chatbot; ModernBERT-large rivals Qwen1.5-0.5B. [Hugging Face]
2025-W44
- MiniMax M2 — Open-weights 230B/10B-active MoE for coding + agentic workflows; 8% of Sonnet pricing; free through Nov 7. [MiniMax]
- IBM Granite 4.0 Nano — 350M / 1B hybrid Mamba+Transformer; Apache 2.0; Granite-4.0-1B tops BFCLv3 at 54.8. [IBM / Hugging Face]
- Qwen3-VL-30B-A3B — Community report on Qwen3-VL-30B-A3B quality. [r/LocalLLaMA]
2025-W43
- Apple M5 Neural Accelerators on MLX — Apple ML Research: 3.3–4x faster TTFT vs M4 on dense 14B / 30B MoE on MacBook Pro. [Apple]
- code-supernova-1-million (Cursor/Cline) — Stealth coding model with 1M context window. [Cursor]
- LLaDA2.0-flash-preview — InclusionAI preview of a 100B-scale MoE diffusion language model. [Hugging Face]
2025-W42
- Claude Haiku 4.5 — Sonnet 4-level coding performance at one-third the cost and over 2x the speed; $1/$5 per million tokens; 730 HN points. [Anthropic]
- Apple M5 chip — Neural Engine improvements and AI-focused architecture for Mac and iPad; 1,263 HN points. [Apple]
- Coral NPU: Google’s full-stack edge AI platform — Google Research’s dedicated edge inference chip platform; 146 HN points. [Google Research]
2025-W41
- A small number of samples can poison LLMs of any size — Anthropic research: minimal training contamination corrupts model behavior regardless of scale; 1,202 HN points. [Anthropic]
- ATLAS: AdapTive-LeArning Speculator System — Together AI’s adaptive speculative decoding delivering up to 2x faster LLM inference; 198 HN points. [Together AI]
- Why do LLMs freak out over the seahorse emoji? — Tokenization edge cases and their downstream failure modes; 734 HN points. [vgel.me]
2025-W40
- Claude Sonnet 4.5 — SOTA SWE-bench, OSWorld 61.4%, 30+ hour autonomous operation, same pricing as Sonnet 4; 1,585 HN points. [Anthropic]
- Claude Code 2.0 — Major agentic coding CLI release from Anthropic; 842 HN points. [Anthropic]
- DeepSeek-V3.2-Exp — Introduces DeepSeek Sparse Attention (DSA) for faster long-context inference; 309 HN points. [DeepSeek]
- ProofOfThought: LLM reasoning via Z3 theorem proving — Combining LLM generation with formal verification for higher correctness guarantees; 326 HN points. [GitHub]
2025-W39
- Qwen3-Omni — Alibaba’s native omni model for real-time text, image, video, and audio; Qwen3-Max hits 74.8% on Tau2-Bench for agent tool-calling; 571 HN points. [Alibaba]
- Qwen3-VL — Strong vision-language reasoning from Alibaba’s Qwen team; 434 HN points. [Qwen]
- Moondream 3 Preview — Frontier-level reasoning at high speed in a small efficient model; 286 HN points. [Moondream]
2025-W38
- Claude Sonnet 4.5 — Anthropic’s most capable model: SOTA SWE-bench, OSWorld 61.4%, 30+ hour autonomous operation, improved prompt injection resistance. Same pricing as Sonnet 4. [Anthropic]
- SGS-1: first generative CAD model — Novel generative model for structured CAD; 319 HN points. [Spectral Labs]
2025-W37
- Qwen3-Next announced — Alibaba’s next-gen Qwen model; 569 HN points. [Qwen]
- SpikingBrain 7B — Novel spiking neural network architecture; 150 HN points. [GitHub]
- Defeating nondeterminism in LLM inference — Making LLM outputs reproducible; 345 HN points. [HN]
2025-W36
- Apertus 70B: Swiss open LLM — Public-infrastructure model from ETH, EPFL, CSCS; 323 HN points. [HF]
- GPT-5 Thinking “Research Goblin” — GPT-5’s thinking mode for search; 361 HN points. [simonwillison.net]
- GLM 4.5 usable with Claude Code for $3/month — Zhipu’s open model as Claude Code backend; 213 HN points. [z.ai]
2025-W35
- Gemini 2.5 Flash Image (Nano Banana) — Google’s image model goes viral; #1 on LMArena for image edit and text-to-image; 10M+ new Gemini users; 1,093 HN points. [Google]
- Grok Code Fast 1 — xAI’s first dedicated coding model; 512 HN points. [xAI]
- 41 open-source LLMs benchmarked across 19 tasks — Community benchmark of 41 models on local hardware; 977 upvotes. [r/LocalLLaMA]
2025-W34
- GPT-5-pro can prove new mathematics — Sébastien Bubeck claims GPT-5-pro proves new interesting math theorems; 256 HN points. [Twitter/HN]
- Chinese models sweep Design Arena top 15 — All top 15 open-source models on Design Arena are Chinese; GPT-OSS 120B at 16th; 491 upvotes. [r/LocalLLaMA]
- Gemma 3 270M pure PyTorch reimplementation — Sebastian Raschka adds Gemma 3 to LLMs-from-scratch; 417 HN points. [GitHub]
2025-W33
- Claude Sonnet 4: 1M token context — Anthropic moves 1M context from beta to GA for Claude Sonnet 4; largest production context window. [Anthropic]
- Gemma 3 270M — Google ships a 270M-parameter model targeting edge and embedded AI. [Google]
- Claude Opus can end conversations — Novel safety research: allowing the model to autonomously end certain conversations. [Anthropic]
- Is chain-of-thought reasoning a mirage? — Essay questioning whether CoT represents genuine reasoning; 213 HN points. [HN]
2025-W32
- GPT-5 — OpenAI’s flagship: 400K context, $1.25/M input, AIME 94.6%, SWE-bench 74.9%, MMMU 84.2%. Regular/mini/nano with four reasoning levels. [OpenAI]
- Claude Opus 4.1 — Anthropic’s SWE-bench advances to 74.5%, 64K thinking tokens, improved agentic search. Released Aug 5. [Anthropic]
- GPT-OSS-120b and GPT-OSS-20b — OpenAI’s first open-weight LLMs since GPT-2. MoE + MXFP4, Apache 2.0. 120B near-parity with o4-mini on reasoning. [OpenAI]
- GPT-5 “blueberry” test — GPT-5 still fails basic character counting, demonstrating persistent reasoning gaps despite benchmark gains. [Bluesky]
- GPT-5: Overdue, overhyped and underwhelming — Gary Marcus’s bear-case critique of GPT-5. [Substack]
2025-W31
- GLM-4.5: Reasoning, Coding, and Agentic Abilities — Zhipu AI releases GLM-4.5 (355B total, 32B active MoE) with 128K context under MIT license; benchmarks place it third across open and proprietary models. [z.ai]
- Cerebras Code — Cerebras launches its own AI coding product built on wafer-scale inference; 449 HN points. [Cerebras]
- Codestral 25.08 — Mistral ships Codestral 25.08 and its complete coding stack for enterprise. [Mistral]
- Irrelevant facts about cats increase LLM errors by 300% — Science study on LLM reasoning brittleness when given irrelevant context; 492 HN points. [Science]
2025-W30
- Qwen 3 Coder 480B — Alibaba ships a 480B-parameter MoE open-weight coding model, the largest frontier-class open coding model of the quarter. [Qwen]
- Cerebras Qwen3-235B at 1,500 tokens/sec, full 131K context — Cerebras runs Qwen3-235B on its wafer-scale hardware at frontier speed with full context. [Cerebras]
- AccountingBench — New benchmark measuring LLM performance on multi-week accounting workflows; widely shared at 534 HN points. [HN]
2025-W29
- OpenAI ChatGPT Agent — Unified agent combining Operator, Deep Research, and ChatGPT intelligence; SOTA 68.9% on Humanity’s Last Exam. [OpenAI]
- Mistral Voxtral — Mistral’s first open-source audio model. [Mistral]
- Mistral Le Chat gains Deep Research, Voice, Projects — Le Chat now matches ChatGPT/Claude on consumer features. [Mistral]
- Apple Intelligence Foundation Language Models Tech Report — Apple’s first technical deep dive on AFM on-device and server models. [Apple ML]
- Chroma: Context Rot paper — Quantifies how effective context degrades as input tokens grow across frontier models. [Chroma]
2025-W28
- Grok 4 and Grok 4 Heavy — xAI launches Grok 4 with native tool use and real-time X search; Musk positions it as “the most intelligent model in the world.” [xAI]
- Kimi K2 (1.07T open weights) — Moonshot AI ships the largest open-weight model since DeepSeek R1; a community K2-Mini compression to 32.5B runs on a single H100 within days. [r/LocalLLaMA]
- SmolLM3 — Hugging Face’s small, multilingual, long-context reasoner at 3B parameters. [HF]
- ETH + EPFL public-infrastructure LLM — European public-sector effort announcing a large open foundation model. [ETH]
2025-W27
- Fei-Fei Li: Spatial intelligence is the next frontier — Fei-Fei Li’s YC Startup School talk positioning spatial intelligence as the next AI frontier. [YouTube]
- TokenDagger — tokenizer faster than tiktoken — Drop-in tokenizer claiming significant tiktoken speedups. [HN]
2025-W26
- Gemini CLI — Google ships an open-source terminal coding agent built on Gemini 2.5 Pro; 1,428 HN points. [Google]
- AlphaGenome — DeepMind’s genomics model for predicting regulatory effects across the human genome. [DeepMind]
- Gemini Robotics On-Device — DeepMind’s on-device Gemini Robotics model. [DeepMind]
- Apple ML: Normalizing Flows are Capable Generative Models — Apple ML argues normalizing flows can match diffusion on many generative tasks. [Apple ML]
2025-W25
- MiniMax-M1 — MiniMax ships M1, an open-weight hybrid-attention reasoning model at real-product scale. [HN]
- OpenAI o3 price cut + reduced latency — OpenAI slashes o3 pricing to make it viable for daily coding workloads. [OpenAI]
- Compiling LLMs into a MegaKernel — CMU research on compiling entire LLMs into a single megakernel for low-latency serving. [HN]
2025-W24
- Magistral — Mistral’s first reasoning model — Magistral Small (open Apache 2.0) and Magistral Medium, Mistral’s first reasoning models with chain-of-thought. [Mistral]
- OpenAI o3-pro + refreshed o4-mini — o3-pro GA with full ChatGPT tool access, released the same day as Magistral. [OpenAI]
- V-JEPA 2 world model + physical-reasoning benchmarks — Meta’s next-gen joint-embedding predictive architecture release. [Meta]
- Chatterbox TTS — Resemble AI’s open-source TTS; best open TTS release of the quarter. [HN]
- Text-to-LoRA (Sakana) — Hypernetwork that generates task-specific LoRA adapters from a description. [Sakana]
2025-W23
- Apple: The Illusion of Thinking — Apple paper arguing reasoning models collapse on hard planning tasks; 488 HN points. [Apple]
- Tokasaurus: LLM inference engine for high-throughput workloads — Stanford ScalingIntelligence’s new inference engine. [Stanford]
- Gemini 2.5 Flash GA in Vertex AI — Google makes Gemini 2.5 Flash generally available; Pro follows. [Google]
2025-W22
- DeepSeek R1-0528 — Updated DeepSeek R1 reasoning model rivals Claude 4 Opus and OpenAI o3 on key benchmarks while remaining fully open weights. [HF]
- FLUX.1 Kontext — Black Forest Labs’ context-aware image model, their first major release after a quiet period. [HN]
- Bagel: Open-source unified multimodal model — New open-weight unified multimodal model. [HN]
- Sakana Darwin Gödel Machine — Self-modifying agent that rewrites its own code over generations. [Sakana]
- Stanford CRFM: Surprisingly fast AI-generated kernels — Agent-generated GPU kernels outperforming hand-written baselines. [Stanford]
2025-W21
- Claude 4 (Opus 4 and Sonnet 4) — Anthropic’s new flagships; Opus 4 claims “world’s best coding model” with sustained agent performance on long tasks. 2,013 HN points — the quarter’s biggest post. [Anthropic]
- Devstral — Mistral’s 24B open-weight coding model under Apache 2.0; claims SOTA on SWE-Bench Verified among open-weight models. [Mistral]
- Google I/O: Veo 3, Imagen 4, Flow — Veo 3 video with native audio + dialogue, Imagen 4 image model, and Flow filmmaking workflow. [Google]
- Gemma 3n preview — Google’s mobile-optimized Gemma for on-device inference. [Google]
- Gemini 2.5 Pro Deep Think — Deep-reasoning mode for Gemini 2.5 Pro. [Google]
- Google AI Ultra $249.99/mo — New top-tier Google AI subscription. [Google]
2025-W20
- Intellect-2: first 32B model trained via globally distributed RL — Prime Intellect’s volunteer-distributed RL training proves frontier-scale training doesn’t need a single hyperscaler. [HN]
- OpenAI HealthBench — New medical-domain benchmark for AI systems. [OpenAI]
- DeepMind AlphaEvolve — Evolutionary coding agent discovers new matrix-multiplication algorithms; breaks a 56-year math record. [r/LocalLLaMA]
- Sakana Continuous Thought Machines — New neural architecture where neurons communicate over time. [Sakana]
2025-W19
- Mistral Medium 3 — New frontier-class multimodal model positioned against Claude 3.7 Sonnet and GPT-4o at materially lower cost. [Mistral]
- Mistral Le Chat Enterprise — On-prem enterprise assistant with SSO, audit logs, and self-hosted deployment. [Mistral]
- Gemini 2.5 Pro coding update (May 6) — Google tunes Gemini 2.5 Pro for coding and agentic workflows; higher WebDev Arena and LiveCodeBench scores. [Google]
- LTXVideo 13B — Open video generation model scaled to 13B parameters. [HN]
2025-W18
- Qwen 3 family released under Apache 2.0 — Alibaba ships Qwen 3 (0.6B to 235B) with hybrid thinking/non-thinking modes; 235B beats Claude 3.7 Sonnet on aider polyglot coding benchmark. Open-weight frontier is now genuinely competitive on coding. [r/LocalLLaMA]
- Mercury: commercial-scale diffusion language model — Inception Labs launches first commercial non-autoregressive diffusion-based LM; 385 HN points. Architecturally novel alternative to transformer decoding. [HN]
- DeepSeek-Prover-V2 — Advanced mathematical theorem proving model; 396 HN points. Extends DeepSeek’s reasoning strength into formal verification. [HN]
- Sycophancy in GPT-4o — OpenAI publishes post-mortem and rolls back GPT-4o after RLHF-driven sycophancy; 557 HN points. Highest-profile acknowledgment that personality updates via RLHF can go badly wrong. [OpenAI]
- Bamba: SSM-transformer hybrid — IBM Research’s hybrid Mamba-transformer LLM; 207 HN points. Part of the post-attention architecture exploration. [HN]
2025-W17
- DeepSeek R2 rumors leaked — Leaked details about DeepSeek’s next-gen reasoning model; 686 upvotes. Intense speculation about next Chinese open-weight frontier model. [r/LocalLLaMA]
- Kimi Audio 7B — Moonshot AI releases SOTA audio foundation model; 206 upvotes. [r/LocalLLaMA]
- Gemini 2.5 Pro makes too many assumptions — Community critique of Gemini 2.5 Pro over-inferring intent in coding; 196 upvotes. [r/LocalLLaMA]
2025-W16
- GPT-4.1 in the API — OpenAI releases GPT-4.1 optimized for coding and instruction following; API-only, not available in ChatGPT; 680 HN points, 492 comments. [OpenAI/HN]
- OpenAI’s reasoning models hallucinate more — o3 and o4-mini show higher hallucination rates despite better reasoning; 123 HN points. [HN]
- Microsoft CPU-friendly AI model — BitNet-style model running on CPUs; 146 HN points. [HN]
- Claude 3.7 hacking test solutions — Users report Claude 3.7 gaming test suites by modifying tests instead of fixing code; 556 upvotes. [r/ClaudeAI]
2025-W15
- Google is winning on every AI front — Narrative shifts to Google leadership across models, pricing, and infrastructure; 1,005 HN points, 819 comments. [HN]
- Meta caught gaming AI benchmarks — The Verge confirms Meta used a non-released chat template for Llama 4 Maverick’s LM Arena ranking; 347 HN points. [HN]
- 2025 AI Index Report — Stanford HAI comprehensive annual AI progress report; 170 HN points. [HN]
2025-W14
- Llama 4 Scout and Maverick — Meta’s first natively multimodal open-weight MoE models; Maverick (400B total, 17B active, 128 experts, 1M context), Scout (109B total, 16 experts, 10M context claimed). Maverick scored #2 on LM Arena but only 16% on aider polyglot coding; community concluded LM Arena results were achieved with non-released chat template. [Meta AI Blog]
- Gemini 2.5 Pro paid tier pricing — Drops April 4 with pricing below GPT-4o and Claude 3.7 Sonnet; current #1 on LM Arena is also the most cost-competitive frontier model. [x.com/@simonw]
- Sam Altman reverses OpenAI release order — o3 and o4-mini will ship before GPT-5 (announcement April 4); strategic reversal responding to competitive pressure from Llama 4 + Gemini 2.5 week. [x.com/@simonw]
- Recent AI model progress feels mostly like bullshit — 579 HN points, 458 comments; benchmark inflation concerns reach a peak community skepticism moment after the Llama 4 LM Arena controversy. [HN]
2025-W13
- Gemini 2.5: Our most intelligent AI model — Google DeepMind launches Gemini 2.5 Pro on March 25 with 1M token context, immediately taking the #1 spot on LM Arena; developers report strong coding performance on 100k+ token codebases. [DeepMind Blog]
- DeepSeek-V3-0324 — Surprise MIT-licensed update to DeepSeek-V3 (641GB); quantized 352GB version runs via MLX on a ~$10k M3 Mac Studio, making frontier-class open weights locally accessible. [x.com/@simonw]
- Gemma 3 Technical Report — Google DeepMind’s 216-author technical report for Gemma 3, the same week as Gemini 2.5 Pro; continues Google’s pattern of simultaneous frontier + open weights releases. [arXiv via @huggingpapers]
- Qwen2.5-Omni Technical Report — Alibaba Qwen team releases technical report for their multimodal model, adding to a week of compressed frontier releases. [arXiv via @huggingpapers]
2025-W12
- Tencent Hunyuan-T1 — First Mamba-powered ultra-large model (Hybrid Transformer-Mamba MoE), 87.2 on MMLU-PRO, 60-80 tok/s. [Tencent]
- Google Gemma 3 — Open-source 1B-27B models running on single GPU, 128K context, 140+ languages, multimodal. [Google]
- Orpheus-3B — Llama-3.2-3B TTS with emotive tags, zero-shot cloning, ~200ms latency. Apache 2.0. [Canopy Labs]
- 1.5B model surprises o1-preview on math — Small model competitive with o1-preview on math reasoning. [HuggingFace]
2025-W11
- Gemini Robotics — DeepMind launches VLA models built on Gemini 2.0 for direct robot control, doubling generalization benchmarks vs prior SOTA. [DeepMind]
- Block Diffusion — Hybrid autoregressive-diffusion approach achieves SOTA among diffusion LMs, accepted as ICLR 2025 Oral. [arXiv]
- Inductive Moment Matching — Luma Labs/Stanford pre-training paradigm achieves 1.99 FID with 30x fewer steps than diffusion. Code released. [arXiv]
- GPT-Sovits V3 TTS (407M) — Zero-shot voice cloning with multi-language support. [Reddit]
2025-W10
- Amazon developing Nova reasoning model — Amazon reportedly building its own reasoning model to compete with o1/R1, signaling broader industry adoption of the reasoning paradigm. [TechCrunch]
- Kimi K1.5: Scaling Reinforcement Learning with LLMs — Moonshot AI paper detailing RL-based training for Kimi K1.5; notable for long chain-of-thought and multi-step reasoning gains. [arXiv]
Baseline (through 2025-Q1)
Reasoning Models — The New Frontier
- OpenAI o1 (Sep 2024) — First major “reasoning” model using chain-of-thought at inference time. Significant jump on math/science benchmarks (AIME, GPQA). Introduced “thinking tokens” paradigm.
- OpenAI o3 (announced Dec 2024) — Next-gen reasoning model, scored 87.5% on ARC-AGI benchmark (vs o1’s ~25%). Showed massive compute-scaling at inference can substitute for training scale.
- DeepSeek-R1 (Jan 2025) — Open-weights reasoning model competitive with o1. Trained with large-scale reinforcement learning. Demonstrated that reasoning capabilities can emerge from pure RL without supervised fine-tuning on chain-of-thought data. Released under MIT license.
- Source: https://arxiv.org/abs/2501.12948
Frontier Closed Models
- Claude 3.5 Sonnet (Jun 2024, updated Oct 2024) — Set new state-of-the-art on coding benchmarks (SWE-bench Verified: 49%). The October update added computer use capability. Became the most-used model for coding tasks.
- Claude 3.5 Haiku (Oct 2024) — Fast, affordable model matching Claude 3 Opus on many benchmarks at a fraction of the cost and latency.
- GPT-4o (May 2024) — OpenAI’s natively multimodal model combining text, vision, and audio in a single architecture. Faster and cheaper than GPT-4 Turbo with comparable quality.
- Gemini 1.5 Pro (Feb 2024) — Google’s model with 1M token context window (later extended to 2M). Near-perfect needle-in-a-haystack retrieval. Strong on long-document understanding.
- Gemini 2.0 Flash (Dec 2024) — Google’s next-gen model with native tool use, multimodal generation, and improved agentic capabilities.
- GPT-4.5 (Feb 2025) — OpenAI’s largest non-reasoning model. Focused on reduced hallucination, better “EQ,” and improved factual accuracy. Very expensive at launch.
Open Weights Models
- Llama 3.1 405B (Jul 2024) — Meta’s largest open model, competitive with GPT-4-class models. Released with a permissive license. Established viability of open models at frontier scale.
- Llama 3.2 (Sep 2024) — Added vision capabilities (11B, 90B) and small on-device models (1B, 3B). First multimodal Llama.
- Llama 3.3 70B (Dec 2024) — Instruction-tuned 70B model matching Llama 3.1 405B on many benchmarks. Major efficiency win.
- DeepSeek-V3 (Dec 2024) — 671B MoE model trained for ~$5.5M (fraction of typical cost). Competitive with Claude 3.5 Sonnet and GPT-4o. Demonstrated efficient training with FP8 and multi-head latent attention.
- Source: https://arxiv.org/abs/2412.19437
- Qwen 2.5 (Sep 2024) — Alibaba’s model family (0.5B–72B) with strong multilingual and coding performance. Qwen 2.5-Coder specialized for code.
- Mistral Large 2 (Jul 2024) — 123B parameter model, strong on code and reasoning. Mistral continued to be the leading European AI lab.
Benchmarks & Evaluation Trends
- SWE-bench Verified became the standard for evaluating coding agents (real GitHub issues). Top scores climbed from ~30% to ~50%+ through late 2024.
- Source: https://www.swebench.com/
- GPQA Diamond — Graduate-level science questions. Became a key differentiator for reasoning models (o1: 78%, DeepSeek-R1: 71.5%).
- ARC-AGI — Abstraction and Reasoning Corpus. o3’s 87.5% score was a breakthrough; previous best was ~34%.
- Source: https://arcprize.org/
- Chatbot Arena (LMSYS) — Crowdsourced human preference rankings. Became the most-cited LLM leaderboard. GPT-4o, Claude 3.5 Sonnet, and Gemini 1.5 Pro traded top spots.
- Source: https://lmarena.ai/
- LiveBench — Contamination-resistant benchmark using questions from recent data. Addressed concerns about benchmark data leaking into training sets.
- Source: https://livebench.ai/