Home / State of AI / Models & Benchmarks

Models & Benchmarks

New model releases, capability benchmarks, and evaluations

Living document tracking new model releases, capability benchmarks, and evaluations. Newest entries appear at the top.

Key Areas to Watch

  • Foundation model releases (GPT, Claude, Gemini, Llama, Mistral)
  • Benchmark results and methodology (MMLU, HumanEval, SWE-bench, GPQA)
  • Scaling laws and training advances
  • Fine-tuning and RLHF/DPO techniques
  • Model safety and alignment evaluations

2026-W39

2026-W38

  • GPT-6 Astra for work — OpenAI positioned Astra around work, tool use, and production systems rather than only benchmark claims. [OpenAI]
  • DeepSeek V4.1 Flash — Multimodal MoE with 1M context and KV-cache emphasis drew developer attention. [HN / x.com]
  • SWE-Bench Pro Verified — Verified benchmark variant reduces reward-hacking/task-quality artifacts and changes how frontier coding models should be compared. [x.com]
  • Meta Muse Image and Muse Video — Creative-generation models with editing, references, and native audio support. [Meta AI]

2026-W37

2026-W36

  • GLM-5.3 open weights — Z.ai kept GLM in the open-weight coding and agent model comparison set. [HN]
  • Qwen3.8-Flash-Next — Fast Qwen variant relevant to cheap routing paths inside agent loops. [HN]
  • Granite 4.2 LLMs — IBM published build details for its latest open Granite family. [Hugging Face]

2026-W32

  • Qwen3.8-Max — Coding and cowork claims put Qwen back in the weekly model-selection debate. [HN]
  • GPT-5.6 price-performance update — Permanent Luna and Terra cost reductions change agent-routing economics. [OpenAI]
  • LFM2.5-2.6B — Small local model release aimed at deployable agent loops outside frontier-only APIs. [Hugging Face]

2026-W28

  • Claude Sonnet 5 — Anthropic updated Sonnet with tokenizer and performance changes that affect API planning. [Anthropic]
  • Mistral 3 — Mistral released Apache-licensed dense small models and a large sparse MoE. [Mistral]
  • GeneBench-Pro — OpenAI added a genomics/biology benchmark for scientific AI capability. [OpenAI]

2026-W27

  • Previewing GPT-5.6 Sol — OpenAI previewed a next-generation model for coding, science, and cybersecurity with heavier access controls. [OpenAI]
  • Introducing Claude Opus 4.8 — Anthropic emphasized stronger coding and agentic performance plus Claude Code dynamic workflows. [Anthropic]
  • GLM 5.2 beats Claude in our benchmarks — Semgrep kept GLM-5.2 in the security-model benchmark conversation. [HN]
  • Qwen3-ASR — Qwen released a compact open ASR model for 52 languages and dialects. [x.com]

2026-W26

  • GLM-5.2 — Open-weight long-horizon model release that dominated developer comparisons this week. [Hugging Face]
  • GLM-5.2 on Artificial Analysis — Benchmark coverage that made GLM-5.2 a serious reference point for open-model evaluation. [HN]
  • Multi-LCB — Expands LiveCodeBench across multiple programming languages for more realistic coding-model evaluation. [arXiv]
  • How Transparent is DiffusionGemma? — Tests whether diffusion language models make reasoning harder to inspect. [arXiv]

2026-W25

  • Claude Fable 5 and Claude Mythos 5 — Anthropic introduced a Mythos-class model family with explicit long-horizon software-engineering claims, then suspended access days later after a government directive. [Anthropic / HN]
  • olmo-eval — AllenAI released evaluation tooling for the model-development loop, useful for teams comparing checkpoints and regressions over time. [Hugging Face]
  • When Good Verifiers Go Bad — Shows self-improving VLM verifiers can regress on new tasks, a warning for automated evaluation pipelines. [arXiv]
  • ClinHallu — Diagnoses hallucinations by reasoning stage, a pattern worth borrowing for domain-specific model evals. [arXiv]

2026-W24

2026-W23

2026-W22

  • Gemini 3.5 Flash — Google release drew major developer attention as a fast model option. [HN]
  • Qwen3.7-Max: The Agent Frontier — Qwen positions the model around frontier agentic task execution. [Qwen]
  • Mistral 3 — Mistral ships an NVFP4/vLLM-oriented checkpoint for Blackwell NVL72 and single-node 8xA100/8xH100 serving. [Mistral]
  • grok-code-fast-1 — xAI describes a fast and economical reasoning model for agentic coding. [xAI]

2026-W20

  • Gemini 3.5 Flash — Google’s new fast, low-cost Gemini tier for high-throughput agentic workloads, released around I/O 2026. [Google]
  • Google I/O 2026 keynote — Gemini Spark personal agent, Gemini for Science, and conversational AI in Docs/YouTube; Google reported Gemini now serving 9.7 trillion tokens/month. [Tom’s Guide]
  • Predictable Confabulations — Factual recall accuracy by LLMs scales predictably with model size and topic frequency, giving a quantitative handle on when recall is trustworthy. [arXiv]
  • DashAttention — A learnable, adaptive sparse hierarchical attention scheme aimed at long-context efficiency. [arXiv]

2026-W19


2026-W18


2026-W17

  • DeepSeek V4-Pro and V4-Flash — DeepSeek’s first V4 preview models; both 1M-context MoE open weights claiming near-frontier quality at a fraction of the price. 2,074 HN points. [HN]
  • GPT-5.5 — OpenAI’s new flagship; available in Codex, rolling out to ChatGPT. Major lift on agentic coding and computer use. [OpenAI]
  • GPT-5.5 System Card — Public eval and safety report for GPT-5.5. [OpenAI]
  • Qwen3.6-27B — Alibaba’s dense 27B model claims flagship-level agentic coding, surpassing Qwen3.5-397B-A17B (15x the active parameter count). 989 HN points. [HN]
  • Qwen3.6-Max-Preview — Larger, sharper Qwen3.6 variant. 705 HN points. [HN]
  • Qwen3.6-35B-A3B Heretic — Uncensored Qwen3.6 35B-A3B variant at 0.0015 KLD, called the best 35B available. [Reddit]
  • Qwen 3.6 27B ties Sonnet 4.6 on Artificial Analysis — Open-weight model in a dead heat with a frontier closed model on the Intelligence Index. [Reddit]
  • Kimi K2.6 is a legit Opus 4.7 replacement — 1,222-point r/LocalLLaMA thread on Kimi 2.6 hitting Opus-class quality. [Reddit]
  • Xiaomi MiMo V2.5 Pro — Xiaomi’s flagship places at 54 on the Artificial Analysis Index; weights coming. [Reddit]
  • ChatGPT Images 2.0 — New image-generation model with stronger text rendering, multilingual support, and visual reasoning; Altman compared the leap to GPT-3 → GPT-5. [OpenAI]
  • OpenAI deprecates SWE-Bench Verified — OpenAI publicly stops reporting SWE-Bench Verified scores, citing benchmark gaming. [OpenAI]
  • OpenAI HealthBench Professional — Medical evaluation benchmark on Hugging Face for AI assistants serving clinicians. [OpenAI / Hugging Face]
  • Claude Token Counter, now with model comparisons — Simon Willison’s tokenizer-cost comparison tool now supports Opus 4.6 vs 4.7. [Simon Willison]

2026-W16

  • Claude Opus 4.7 — Anthropic’s new flagship; first Opus-line tokenizer change, updated system prompt, thinking_effort: xhigh exposed via the API. 1,954 HN points. [Anthropic]
  • Qwen3.6-35B-A3B — Alibaba agentic-coding MoE; competitive with frontier closed models on a laptop. Simon Willison’s pelican benchmark prefers it to Opus 4.7. [Qwen]
  • Qwen3.6-Max-Preview — Larger, sharper Qwen3.6 variant landing days after 35B-A3B. [Qwen]
  • Kimi K2.6 — Moonshot open weights drop; 724-point r/LocalLLaMA thread. [Hugging Face / Reddit]
  • Gemini 3.1 Flash TTS — DeepMind adds granular audio tags for fine-grained expressive speech control. [DeepMind]
  • Gemini Robotics-ER 1.6 — Spatial reasoning + multi-view understanding for autonomous robotics. [DeepMind]
  • GPT-Rosalind — OpenAI’s life-sciences reasoning model for drug discovery, genomics, and protein reasoning. [OpenAI]
  • xAI Grok 4.3 Beta — Conversational video understanding, direct generation of downloadable PDFs / spreadsheets / PowerPoint decks. [xAI]
  • Measuring Claude 4.7’s tokenizer costs — First independent cost-impact study of the Opus 4.7 tokenizer change. [HN]
  • Claude Opus 4.7 Text Category Rankings — Community leaderboard of Opus 4.7 capabilities across categories. [Reddit]

2026-W15

  • Muse Spark — First model from Meta Superintelligence Labs (code name Avocado); natively multimodal reasoning, Contemplating mode for multi-agent parallel reasoning aimed at Gemini Deep Think / GPT Pro tier. [Meta AI]
  • GLM-5.1: Towards Long-Horizon Tasks — Zhipu’s GLM-5.1 explicitly targets long-horizon agentic tasks; dedicated Simon Willison coverage. [z.ai]
  • Minimax M2.7 Released — MiniMax drops M2.7 open weights; 668-point r/LocalLLaMA thread. [Hugging Face]
  • System Card: Claude Mythos Preview — Anthropic ships a restricted security-research model with Project Glasswing gating; zero-day claims contested by Berkeley RDI and Aisle analyses. [Anthropic]
  • Gemma 4 audio support lands in llama-server — Native multimodal audio inference for Gemma 4 in llama-server, plus MLX audio pipeline from Simon Willison. [Reddit / Simon Willison]

2026-W14

  • Gemma 4 — Google DeepMind’s strongest open model family yet; purpose-built for advanced reasoning and agentic workflows; ships with on-device support via Google AI Edge Gallery on iPhone (868 HN points). [DeepMind]
  • Qwen3.6-Plus: Towards real world agents — Alibaba’s agentic-focused update to the Qwen3 line; 596 HN points on launch. [Qwen / Alibaba]
  • 1-Bit Bonsai: first commercially viable 1-bit LLMs — PrismML claims viable 1-bit quantized models for production deployment; 430 HN points. [PrismML / HN]

2026-W13


2026-W12


2026-W11


2026-W10


2026-W09

  • The Llama 3 Herd of Models — Meta AI’s full technical paper covering the 405B dense model: 128K context, multilingual capability, coding benchmarks, reasoning, and native tool use; foundational reference for anyone building on the Llama lineage. [Meta AI]
  • Introducing Mercury 2: Fast Reasoning LLM powered by diffusion — Inception Labs’ diffusion-based reasoning LLM challenges autoregressive latency assumptions; competitive speed and quality at lower latency. [Inception Labs / HN]
  • Nano Banana 2 — Google’s Flash-speed image generation model with Pro-level world knowledge and subject consistency; 605 HN points. [DeepMind]

2026-W08


2026-W07


2026-W06

  • Claude Opus 4.6 — Anthropic’s latest frontier release; top HN post of its launch week at 2,346 points and 1,031 comments; positioned as “a space to think” for long-horizon reasoning. [Anthropic]
  • Introducing GPT-5.3-Codex — OpenAI’s coding-focused frontier model ships same day as Claude Opus 4.6; 1,530 HN points, 605 comments; includes new macOS Codex app. [OpenAI]
  • Voxtral Transcribe 2 — Mistral’s next-generation ASR model; 1,012 HN points, 241 comments. [Mistral AI]
  • Qwen3-Coder-Next teaser — Alibaba’s upcoming coding model generates early community excitement; 735 HN points, 429 comments. [Qwen / HN]

2026-W05


2026-W04

  • Qwen3-TTS family open sourced — Alibaba open-sources a full TTS suite including voice cloning and voice design; 744 HN points. [Qwen / Alibaba]
  • Differential Transformer V2 — Microsoft extends the differential attention architecture for improved long-context performance. [Hugging Face / Microsoft]

2026-W03


2026-W02

  • NVIDIA Cosmos Reason 2 — 2B/8B open reasoning VLMs; #1 on Physical AI Bench and Physical Reasoning leaderboards. [NVIDIA / Hugging Face]
  • Falcon-H1 Arabic (3B/7B/34B) — Hybrid Mamba/Transformer; 3B hits 61.87% OALL (10pts ahead of Phi-4 Mini); 7B hits 71.47%. [TII / Hugging Face]

2026-W01

  • GLM-Image teaser (Z.ai) — Two-stage compact-token encoder scaling to 4K; reportedly first top-tier multimodal trained entirely on Huawei Ascend. 311 upvotes. [r/LocalLLaMA / eWEEK]
  • FLUX.2-dev-Turbo editing test — Community benchmarks on FLUX.2-dev-Turbo for image editing; 92 upvotes. [r/LocalLLaMA]

2025-W52

  • GLM-4.7 (Z.ai) — Open weights; at/above Sonnet 4.5 on SWE-bench Verified, LiveCodeBench v6, Terminal Bench 2.0; #1 open on Code Arena; τ²-Bench 87.4. [Z.ai / Hugging Face]
  • MiniMax M2.1 — 230B/10B-active MoE; multi-language (Rust/Java/Go/C++/Kotlin/ObjC/TS/JS); introduces VIBE runtime benchmark. [MiniMax]

2025-W51

  • Gemini 3 Flash — Pro-grade reasoning at Flash speed; GPQA Diamond 90.4%, HLE 33.7%, MMMU Pro 81.2%; default in Gemini app and AI Mode. [Google]
  • NVIDIA Nemotron 3 — Hybrid MoE open family; Nano (30B/3B active) ships now; Super (~100B/10B) and Ultra (~500B/50B) in H1 2026. [NVIDIA]
  • GPT-5.2-Codex — Long-horizon software-engineering model; system-card addendum. [OpenAI]

2025-W50

  • GPT-5.2 — Instant/Thinking/Pro; SOTA on GDPval; 30% fewer hallucinations vs 5.1 Thinking. [OpenAI]
  • Qwen3-Next-80B-A3B-Thinking-GGUF — 80B MoE thinking model in GGUF for local use. [Qwen / Hugging Face]

2025-W49

  • Mistral Large 3 — 675B-total / 41B-active MoE; 256K context; multimodal + 40+ languages; Apache 2.0. [Mistral AI]
  • Apriel-1.6-15b-Thinker — ServiceNow’s 15B reasoning model. [Hugging Face]

2025-W48

  • Claude Opus 4.5 — 80.9% SWE-bench Verified — first model over 80%; first to beat every human on Anthropic’s internal eval; $5/$25 per million tokens (66% cheaper than Opus 4.1). [Anthropic]
  • FLUX.2 in Diffusers — Black Forest Labs’ FLUX.2 open-weights release integrated with Hugging Face Diffusers. [Hugging Face]

2025-W47

  • Gemini 3 — LMArena 1501; HLE 37.5% (no tools); GPQA Diamond 91.9%; MathArena Apex 23.4% SOTA; Deep Think pushes HLE to 41%. [DeepMind]
  • Grok 4.1 — LMArena #2 Thinking (1483 Elo); 2M-token context; $0.20 / $0.50 per million tokens; 3x lower hallucination. [xAI]
  • GPT-5.1-Codex-Max — Long-horizon coding; first OpenAI model trained to operate in Windows. [OpenAI]
  • Nano Banana Pro — Gemini 3 Pro Image model. [DeepMind]

2025-W46


2025-W45

  • BERTs that chat with dLLM — Recipe converting any BERT-style encoder into an instruction-following diffusion chatbot; ModernBERT-large rivals Qwen1.5-0.5B. [Hugging Face]

2025-W44

  • MiniMax M2 — Open-weights 230B/10B-active MoE for coding + agentic workflows; 8% of Sonnet pricing; free through Nov 7. [MiniMax]
  • IBM Granite 4.0 Nano — 350M / 1B hybrid Mamba+Transformer; Apache 2.0; Granite-4.0-1B tops BFCLv3 at 54.8. [IBM / Hugging Face]
  • Qwen3-VL-30B-A3B — Community report on Qwen3-VL-30B-A3B quality. [r/LocalLLaMA]

2025-W43


2025-W42

  • Claude Haiku 4.5 — Sonnet 4-level coding performance at one-third the cost and over 2x the speed; $1/$5 per million tokens; 730 HN points. [Anthropic]
  • Apple M5 chip — Neural Engine improvements and AI-focused architecture for Mac and iPad; 1,263 HN points. [Apple]
  • Coral NPU: Google’s full-stack edge AI platform — Google Research’s dedicated edge inference chip platform; 146 HN points. [Google Research]

2025-W41


2025-W40

  • Claude Sonnet 4.5 — SOTA SWE-bench, OSWorld 61.4%, 30+ hour autonomous operation, same pricing as Sonnet 4; 1,585 HN points. [Anthropic]
  • Claude Code 2.0 — Major agentic coding CLI release from Anthropic; 842 HN points. [Anthropic]
  • DeepSeek-V3.2-Exp — Introduces DeepSeek Sparse Attention (DSA) for faster long-context inference; 309 HN points. [DeepSeek]
  • ProofOfThought: LLM reasoning via Z3 theorem proving — Combining LLM generation with formal verification for higher correctness guarantees; 326 HN points. [GitHub]

2025-W39

  • Qwen3-Omni — Alibaba’s native omni model for real-time text, image, video, and audio; Qwen3-Max hits 74.8% on Tau2-Bench for agent tool-calling; 571 HN points. [Alibaba]
  • Qwen3-VL — Strong vision-language reasoning from Alibaba’s Qwen team; 434 HN points. [Qwen]
  • Moondream 3 Preview — Frontier-level reasoning at high speed in a small efficient model; 286 HN points. [Moondream]

2025-W38

  • Claude Sonnet 4.5 — Anthropic’s most capable model: SOTA SWE-bench, OSWorld 61.4%, 30+ hour autonomous operation, improved prompt injection resistance. Same pricing as Sonnet 4. [Anthropic]
  • SGS-1: first generative CAD model — Novel generative model for structured CAD; 319 HN points. [Spectral Labs]

2025-W37


2025-W36


2025-W35


2025-W34


2025-W33


2025-W32

  • GPT-5 — OpenAI’s flagship: 400K context, $1.25/M input, AIME 94.6%, SWE-bench 74.9%, MMMU 84.2%. Regular/mini/nano with four reasoning levels. [OpenAI]
  • Claude Opus 4.1 — Anthropic’s SWE-bench advances to 74.5%, 64K thinking tokens, improved agentic search. Released Aug 5. [Anthropic]
  • GPT-OSS-120b and GPT-OSS-20b — OpenAI’s first open-weight LLMs since GPT-2. MoE + MXFP4, Apache 2.0. 120B near-parity with o4-mini on reasoning. [OpenAI]
  • GPT-5 “blueberry” test — GPT-5 still fails basic character counting, demonstrating persistent reasoning gaps despite benchmark gains. [Bluesky]
  • GPT-5: Overdue, overhyped and underwhelming — Gary Marcus’s bear-case critique of GPT-5. [Substack]

2025-W31


2025-W30


2025-W29


2025-W28

  • Grok 4 and Grok 4 Heavy — xAI launches Grok 4 with native tool use and real-time X search; Musk positions it as “the most intelligent model in the world.” [xAI]
  • Kimi K2 (1.07T open weights) — Moonshot AI ships the largest open-weight model since DeepSeek R1; a community K2-Mini compression to 32.5B runs on a single H100 within days. [r/LocalLLaMA]
  • SmolLM3 — Hugging Face’s small, multilingual, long-context reasoner at 3B parameters. [HF]
  • ETH + EPFL public-infrastructure LLM — European public-sector effort announcing a large open foundation model. [ETH]

2025-W27


2025-W26


2025-W25


2025-W24


2025-W23


2025-W22


2025-W21

  • Claude 4 (Opus 4 and Sonnet 4) — Anthropic’s new flagships; Opus 4 claims “world’s best coding model” with sustained agent performance on long tasks. 2,013 HN points — the quarter’s biggest post. [Anthropic]
  • Devstral — Mistral’s 24B open-weight coding model under Apache 2.0; claims SOTA on SWE-Bench Verified among open-weight models. [Mistral]
  • Google I/O: Veo 3, Imagen 4, Flow — Veo 3 video with native audio + dialogue, Imagen 4 image model, and Flow filmmaking workflow. [Google]
  • Gemma 3n preview — Google’s mobile-optimized Gemma for on-device inference. [Google]
  • Gemini 2.5 Pro Deep Think — Deep-reasoning mode for Gemini 2.5 Pro. [Google]
  • Google AI Ultra $249.99/mo — New top-tier Google AI subscription. [Google]

2025-W20


2025-W19

  • Mistral Medium 3 — New frontier-class multimodal model positioned against Claude 3.7 Sonnet and GPT-4o at materially lower cost. [Mistral]
  • Mistral Le Chat Enterprise — On-prem enterprise assistant with SSO, audit logs, and self-hosted deployment. [Mistral]
  • Gemini 2.5 Pro coding update (May 6) — Google tunes Gemini 2.5 Pro for coding and agentic workflows; higher WebDev Arena and LiveCodeBench scores. [Google]
  • LTXVideo 13B — Open video generation model scaled to 13B parameters. [HN]

2025-W18

  • Qwen 3 family released under Apache 2.0 — Alibaba ships Qwen 3 (0.6B to 235B) with hybrid thinking/non-thinking modes; 235B beats Claude 3.7 Sonnet on aider polyglot coding benchmark. Open-weight frontier is now genuinely competitive on coding. [r/LocalLLaMA]
  • Mercury: commercial-scale diffusion language model — Inception Labs launches first commercial non-autoregressive diffusion-based LM; 385 HN points. Architecturally novel alternative to transformer decoding. [HN]
  • DeepSeek-Prover-V2 — Advanced mathematical theorem proving model; 396 HN points. Extends DeepSeek’s reasoning strength into formal verification. [HN]
  • Sycophancy in GPT-4o — OpenAI publishes post-mortem and rolls back GPT-4o after RLHF-driven sycophancy; 557 HN points. Highest-profile acknowledgment that personality updates via RLHF can go badly wrong. [OpenAI]
  • Bamba: SSM-transformer hybrid — IBM Research’s hybrid Mamba-transformer LLM; 207 HN points. Part of the post-attention architecture exploration. [HN]

2025-W17

  • DeepSeek R2 rumors leaked — Leaked details about DeepSeek’s next-gen reasoning model; 686 upvotes. Intense speculation about next Chinese open-weight frontier model. [r/LocalLLaMA]
  • Kimi Audio 7B — Moonshot AI releases SOTA audio foundation model; 206 upvotes. [r/LocalLLaMA]
  • Gemini 2.5 Pro makes too many assumptions — Community critique of Gemini 2.5 Pro over-inferring intent in coding; 196 upvotes. [r/LocalLLaMA]

2025-W16


2025-W15


2025-W14

  • Llama 4 Scout and Maverick — Meta’s first natively multimodal open-weight MoE models; Maverick (400B total, 17B active, 128 experts, 1M context), Scout (109B total, 16 experts, 10M context claimed). Maverick scored #2 on LM Arena but only 16% on aider polyglot coding; community concluded LM Arena results were achieved with non-released chat template. [Meta AI Blog]
  • Gemini 2.5 Pro paid tier pricing — Drops April 4 with pricing below GPT-4o and Claude 3.7 Sonnet; current #1 on LM Arena is also the most cost-competitive frontier model. [x.com/@simonw]
  • Sam Altman reverses OpenAI release order — o3 and o4-mini will ship before GPT-5 (announcement April 4); strategic reversal responding to competitive pressure from Llama 4 + Gemini 2.5 week. [x.com/@simonw]
  • Recent AI model progress feels mostly like bullshit — 579 HN points, 458 comments; benchmark inflation concerns reach a peak community skepticism moment after the Llama 4 LM Arena controversy. [HN]

2025-W13

  • Gemini 2.5: Our most intelligent AI model — Google DeepMind launches Gemini 2.5 Pro on March 25 with 1M token context, immediately taking the #1 spot on LM Arena; developers report strong coding performance on 100k+ token codebases. [DeepMind Blog]
  • DeepSeek-V3-0324 — Surprise MIT-licensed update to DeepSeek-V3 (641GB); quantized 352GB version runs via MLX on a ~$10k M3 Mac Studio, making frontier-class open weights locally accessible. [x.com/@simonw]
  • Gemma 3 Technical Report — Google DeepMind’s 216-author technical report for Gemma 3, the same week as Gemini 2.5 Pro; continues Google’s pattern of simultaneous frontier + open weights releases. [arXiv via @huggingpapers]
  • Qwen2.5-Omni Technical Report — Alibaba Qwen team releases technical report for their multimodal model, adding to a week of compressed frontier releases. [arXiv via @huggingpapers]

2025-W12

  • Tencent Hunyuan-T1 — First Mamba-powered ultra-large model (Hybrid Transformer-Mamba MoE), 87.2 on MMLU-PRO, 60-80 tok/s. [Tencent]
  • Google Gemma 3 — Open-source 1B-27B models running on single GPU, 128K context, 140+ languages, multimodal. [Google]
  • Orpheus-3B — Llama-3.2-3B TTS with emotive tags, zero-shot cloning, ~200ms latency. Apache 2.0. [Canopy Labs]
  • 1.5B model surprises o1-preview on math — Small model competitive with o1-preview on math reasoning. [HuggingFace]

2025-W11

  • Gemini Robotics — DeepMind launches VLA models built on Gemini 2.0 for direct robot control, doubling generalization benchmarks vs prior SOTA. [DeepMind]
  • Block Diffusion — Hybrid autoregressive-diffusion approach achieves SOTA among diffusion LMs, accepted as ICLR 2025 Oral. [arXiv]
  • Inductive Moment Matching — Luma Labs/Stanford pre-training paradigm achieves 1.99 FID with 30x fewer steps than diffusion. Code released. [arXiv]
  • GPT-Sovits V3 TTS (407M) — Zero-shot voice cloning with multi-language support. [Reddit]

2025-W10


Baseline (through 2025-Q1)

Reasoning Models — The New Frontier

  • OpenAI o1 (Sep 2024) — First major “reasoning” model using chain-of-thought at inference time. Significant jump on math/science benchmarks (AIME, GPQA). Introduced “thinking tokens” paradigm.
  • OpenAI o3 (announced Dec 2024) — Next-gen reasoning model, scored 87.5% on ARC-AGI benchmark (vs o1’s ~25%). Showed massive compute-scaling at inference can substitute for training scale.
  • DeepSeek-R1 (Jan 2025) — Open-weights reasoning model competitive with o1. Trained with large-scale reinforcement learning. Demonstrated that reasoning capabilities can emerge from pure RL without supervised fine-tuning on chain-of-thought data. Released under MIT license.

Frontier Closed Models

Open Weights Models

  • SWE-bench Verified became the standard for evaluating coding agents (real GitHub issues). Top scores climbed from ~30% to ~50%+ through late 2024.
  • GPQA Diamond — Graduate-level science questions. Became a key differentiator for reasoning models (o1: 78%, DeepSeek-R1: 71.5%).
  • ARC-AGI — Abstraction and Reasoning Corpus. o3’s 87.5% score was a breakthrough; previous best was ~34%.
  • Chatbot Arena (LMSYS) — Crowdsourced human preference rankings. Became the most-cited LLM leaderboard. GPT-4o, Claude 3.5 Sonnet, and Gemini 1.5 Pro traded top spots.
  • LiveBench — Contamination-resistant benchmark using questions from recent data. Addressed concerns about benchmark data leaking into training sets.