Home / State of AI / Multimodal AI

Multimodal AI

Vision, audio, video, and cross-modal AI models

Living document tracking vision, audio, video, and cross-modal AI models. Newest entries appear at the top.

Key Areas to Watch

  • Vision-language models (GPT-4V, Gemini, Claude vision)
  • Image and video generation (DALL-E, Midjourney, Sora, Runway)
  • Speech and audio models (Whisper, ElevenLabs, NotebookLM)
  • Document understanding and OCR
  • Real-time multimodal interaction

2026-W39

2026-W38

2026-W37

  • Gemini agentic video understanding — Gemini work on agentic video understanding points at richer multimodal tool use. [DeepMind]
  • NeoMME — Efficient multimodal-native and multilingual encoder for retrieval and cross-modal workflows. [Hugging Face]
  • Muse Spark 1.3 — Meta model API page drew attention for creative generation workflows. [Meta / HN]

2026-W36

  • Gemini 3.5 Transcribe — Google updated its transcription model line for audio-heavy AI workflows. [HN]
  • Gemini Omni 1.1 Flash — Developer-facing multimodal update for builders working across audio, vision, and text. [HN]
  • When Robots Mishear Us — Embodied-AI safety paper on voice interfaces controlling physical action. [arXiv]

2026-W32

  • Robostral Navigate — 8B RGB-only navigation model points at simpler embodied-AI deployment assumptions. [Mistral]
  • Gemini Robotics 2 — DeepMind continued pushing robotics toward whole-body intelligence and task generalization. [DeepMind/HN]
  • Mistral OCR 4 — Document AI release with multilingual OCR, bounding boxes, and self-hosting in the web-source fallback. [Mistral]

2026-W28

  • Brain2Qwerty — Meta continued non-invasive brain-to-text decoding research. [Meta AI]
  • LeRobot v0.6.0 — Open robotics tooling added imagine/evaluate/improve loops. [Hugging Face]
  • Embodied.cpp — Portable C++ runtime work for embodied AI models surfaced in developer research feeds. [x.com]

2026-W27

2026-W24

2026-W23

  • NVIDIA Cosmos 3 for Physical AI — Open omni-model for physical reasoning and action, tying multimodal world modeling to embodied systems. [Hugging Face]
  • Lumos-Nexus — Training-efficient unified video generation framework that bridges high-fidelity generation with instruction-grounded video synthesis. [arXiv]
  • TunerDiT — Training-free progressive steering method for multi-event text-to-video generation over longer horizons. [arXiv]
  • UniAudio-Token — Universal audio tokenizer aimed at combining general audio perception with linguistic alignment for Audio-LLM integration. [arXiv]
  • Reachy Mini goes fully local — Local conversation stack for a small robot, reinforcing the trend toward on-device multimodal interaction. [Hugging Face]

2026-W19


2026-W18


2026-W17


2026-W16


2026-W15


2026-W14

  • Gemma 4: Byte for byte, the most capable open models — Google’s strongest open multimodal model family; built for agentic and on-device workflows; runs on iPhone via AI Edge Gallery (868 HN points). [DeepMind]
  • Falcon Perception — TII’s Falcon Perception multimodal model; adds vision to the Falcon family. [Hugging Face / TII]
  • Granite 4.0 3B Vision — IBM’s compact 3B VLM targeting enterprise document understanding workflows. [Hugging Face / IBM]

2026-W13

  • Gemini 3.1 Flash Live — Google DeepMind ships lower-latency voice model with improved precision for real-time voice interactions. [DeepMind]
  • Lyria 3 Pro — Google’s music generation model gains long-form structural coherence; distributed via broader surfaces. [DeepMind]

2026-W12


2026-W11

  • Grok Imagine update (Elon-flagged) — xAI ships image generation update to Grok; Musk publicly criticizes output quality, drawing attention to prompt-following gaps. [xAI]

2026-W10

  • Phi-4-reasoning-vision-15B — Microsoft Research’s 15B open-weight multimodal reasoning model; available on HuggingFace and GitHub; strong reasoning-while-seeing benchmark results. [Microsoft Research]

2026-W09

  • Nano Banana 2 — Google DeepMind’s Flash-speed image generation with Pro-level world knowledge and subject consistency; 605 HN points. [DeepMind / Google]
  • Moonshine open-weights STT models — Open-weights speech-to-text models beating Whisper Large v3 on accuracy; 316 HN points. [moonshine-ai / HN]

2026-W08


2026-W07

  • Qwen-Image-2.0 — Alibaba’s new image generation model targets professional infographics and photorealistic output; 422 HN points. [Qwen / HN]

2026-W06

  • Voxtral Transcribe 2 — Mistral’s next-generation ASR model; 1012 HN points, 241 comments; strongest open transcription release of Q1 2026. [Mistral / HN]
  • Grok Imagine v1.0 — xAI ships 10-second 720p video generation model described as their biggest leap in prompt-following; released same week as SpaceX acquisition of xAI. [xAI]

2026-W05

  • Grok Imagine API — xAI launches text-to-video, image-to-video, and video editing API at $0.05/sec; positions Grok directly against Sora and Runway in programmatic video generation. [xAI]

2026-W04


2026-W03


2026-W02

  • NVIDIA Cosmos Reason 2 — Open reasoning vision-language model for physical AI; 2B and 8B; top of Physical AI Bench. [NVIDIA / Hugging Face]

2026-W01

  • GLM-Image (Z.ai, teaser) — Two-stage compact-token text-to-image encoder scaling to 4K; reportedly first top-tier multimodal trained on Huawei Ascend. [r/LocalLLaMA / eWEEK]
  • FLUX.2-dev-Turbo editing — Community reports FLUX.2-dev-Turbo is strong at image editing. [r/LocalLLaMA]

2025-W52


2025-W51

  • Gemini 3 Flash — Pro-grade multimodal reasoning at Flash speed; MMMU Pro 81.2%; default in Gemini app and AI Mode. [Google]

2025-W50


2025-W48


2025-W47


2025-W46

  • SIMA 2 — Gemini-powered 3D-world agent handles instructions given through both language and images. [DeepMind]
  • Faster Maya1 TTS model — Community TTS speedup. [r/LocalLLaMA]

2025-W44


2025-W43


2025-W35


2025-W33


2025-W32

  • GPT-5 MMMU 84.2% — Sets new SOTA on multimodal understanding benchmark; visual perception highlighted as a key GPT-5 strength. [OpenAI]
  • MicroLlaVA — Community builds a vision-language model on a single NVIDIA 4090. [r/LocalLLaMA]

2025-W31


2025-W29


2025-W28


2025-W26


2025-W24


2025-W22


2025-W21


2025-W19

  • LTXVideo 13B — Open video generation model scaled to 13B parameters. [HN]

2025-W17


2025-W16


2025-W14


2025-W13

  • Introducing 4o Image Generation — OpenAI ships native in-ChatGPT image generation on March 25; system card addendum published simultaneously. GPT-4o can now generate images inline without a separate product. [OpenAI Blog]
  • Qwen2.5-Omni Technical Report — Alibaba Qwen team’s multimodal model technical report; strong open-weights multimodal release in a week dominated by Gemini 2.5. [arXiv via @huggingpapers]
  • Video-R1: Reinforcing Video Reasoning in MLLMs — Applies reinforcement learning to improve video reasoning in multimodal LLMs; extends R1-style RL training to video understanding. [arXiv via @huggingpapers]
  • Wan: Open and Advanced Large-Scale Video Generative Models — Open-source large-scale video generation model (62 authors); adds to the field of accessible video gen alternatives. [arXiv via @huggingpapers]

2025-W12

  • MoshiVis — First open-source real-time speech + vision model. [Kyutai]
  • Orpheus-3B — Emotive TTS with tags and zero-shot cloning. Apache 2.0. [Canopy Labs]

2025-W11

  • Gemini Robotics — VLA models adding physical actions as new output modality for Gemini 2.0. [DeepMind]
  • GPT-Sovits V3 TTS — Zero-shot voice cloning, multi-language, 407M params. [Reddit]

2025-W10

  • Aya Vision: multilingual multimodal AI — Cohere for AI’s open multilingual vision-language model covering 23+ languages; addresses underrepresentation of non-English languages in multimodal models. [Hugging Face]
  • Amazon Alexa+ multimodal interactions — Upgraded Alexa gains richer multimodal capabilities on Echo Show devices, combining vision input with conversational AI. [ZDNet]

Baseline (through 2025-Q1)

Image Generation

  • DALL-E 3 (OpenAI, Oct 2023) — Integrated into ChatGPT. Strong text rendering and prompt adherence. Uses ChatGPT to rewrite prompts for better results.
  • Midjourney v6/v6.1 (2024) — Continued dominance in artistic image generation. Improved prompt following and text rendering. Added web interface beyond Discord.
  • Stable Diffusion 3 (Stability AI, Feb 2024 announcement, Jun 2024 release) — Multimodal Diffusion Transformer (MMDiT) architecture. Improved text rendering and prompt adherence. Open weights for smaller variants.
  • FLUX (Black Forest Labs, Aug 2024) — Founded by ex-Stability AI researchers. FLUX.1 models (schnell, dev, pro) quickly gained traction as high-quality open alternatives. Rectified flow transformers.
  • Ideogram — Strong text-in-image generation. Raised significant funding. Canvas feature for editing.
  • Google Imagen 3 (2024) — Google’s latest image model, available in Gemini. High photorealism.

Video Generation

  • Sora (OpenAI, Feb 2024 preview, Dec 2024 launch) — Text-to-video model generating up to 20-second clips at 1080p. Stunning quality but limited availability. ChatGPT Plus/Pro at launch.
  • Runway Gen-3 Alpha (Jun 2024) — Professional-grade video generation. Used in film/TV production. Strong motion and consistency.
  • Kling (Kuaishou) — Chinese video model with impressive quality, especially for motion. Available internationally.
  • Google Veo 2 (Dec 2024) — Google’s video generation model. High-quality 4K output. Available in VideoFX.
  • Pika — Video generation startup. Focus on stylized and motion-controlled video editing.

Audio and Speech

  • ElevenLabs — Leading voice synthesis platform. Realistic text-to-speech, voice cloning, and dubbing. Raised $180M (Jan 2025).
  • NotebookLM Audio Overviews (Google, Sep 2024) — Generates podcast-style audio discussions from documents. Viral adoption. Surprisingly natural-sounding two-host conversations.
  • OpenAI Advanced Voice Mode (Sep 2024) — Real-time conversational voice in ChatGPT using GPT-4o’s native audio capabilities. Low latency, natural turn-taking, emotional expression.
  • Whisper v3 (OpenAI) — Open-source speech recognition. Remains the standard for transcription. Runs locally via faster-whisper and whisper.cpp.

Music Generation

  • Suno — AI music generation from text prompts. Full songs with vocals, instruments, and lyrics. v3/v3.5 improved quality significantly. Copyright concerns remain.
  • Udio — Competitor to Suno with similar text-to-music capabilities. Strong on genre diversity.

Vision-Language Models

  • GPT-4o — Natively multimodal (text, vision, audio in one model). Strong visual understanding, document parsing, and image-based reasoning.
  • Claude 3.5 Sonnet — Vision capabilities for screenshot analysis, document understanding, chart reading. Used extensively in coding workflows for UI development.
  • Gemini 1.5 Pro/Flash — Strong vision with long-context advantage (can process hour-long videos). Native PDF understanding.
  • Llama 3.2 Vision (Sep 2024) — First multimodal Llama models (11B, 90B). Open-weights vision-language models.
  • Natively multimodal models (GPT-4o, Gemini) outperform adapter-based approaches. Single model handling multiple modalities is becoming the norm.
  • Real-time interaction — Voice and vision in real-time conversations. Google’s Project Astra and GPT-4o voice set new expectations.
  • Copyright litigation — Multiple lawsuits from artists, photographers, and publishers against image/video generation companies. Legal landscape still evolving.