Multimodal AI
Vision, audio, video, and cross-modal AI models
Living document tracking vision, audio, video, and cross-modal AI models. Newest entries appear at the top.
Key Areas to Watch
- Vision-language models (GPT-4V, Gemini, Claude vision)
- Image and video generation (DALL-E, Midjourney, Sora, Runway)
- Speech and audio models (Whisper, ElevenLabs, NotebookLM)
- Document understanding and OCR
- Real-time multimodal interaction
2026-W39
- Qwen Image 2.1 — Qwen updated its image-generation model for creative application workflows. [HN]
- Qwen 3.8 Omni Flash — Fast multimodal release extends Qwen’s developer model stack beyond text. [HN]
- NemotronLabs VoiceChat — Open full-duplex speech architecture combines streaming transcription, reasoning, speech, and structured tool calls. [arXiv]
- Gemini 3.8 Live and 3.8 Live Extended Thinking — Real-time multimodal interaction gains an extended-thinking path. [DeepMind]
2026-W38
- GPT-Live-1 in the API — Full-duplex voice and tool calling open better interfaces for real-time agent applications. [OpenAI]
- ChatGPT Images 2.5 — Image-generation update relevant to multimodal product workflows. [OpenAI]
- Meta Muse Image and Muse Video — Image/video generation with reference composition, editing precision, and native audio. [Meta AI]
- Gander realtime omni interaction agent — Open omni model unifying perception, full-duplex speech, and agentic execution. [x.com]
2026-W37
- Gemini agentic video understanding — Gemini work on agentic video understanding points at richer multimodal tool use. [DeepMind]
- NeoMME — Efficient multimodal-native and multilingual encoder for retrieval and cross-modal workflows. [Hugging Face]
- Muse Spark 1.3 — Meta model API page drew attention for creative generation workflows. [Meta / HN]
2026-W36
- Gemini 3.5 Transcribe — Google updated its transcription model line for audio-heavy AI workflows. [HN]
- Gemini Omni 1.1 Flash — Developer-facing multimodal update for builders working across audio, vision, and text. [HN]
- When Robots Mishear Us — Embodied-AI safety paper on voice interfaces controlling physical action. [arXiv]
2026-W32
- Robostral Navigate — 8B RGB-only navigation model points at simpler embodied-AI deployment assumptions. [Mistral]
- Gemini Robotics 2 — DeepMind continued pushing robotics toward whole-body intelligence and task generalization. [DeepMind/HN]
- Mistral OCR 4 — Document AI release with multilingual OCR, bounding boxes, and self-hosting in the web-source fallback. [Mistral]
2026-W28
- Brain2Qwerty — Meta continued non-invasive brain-to-text decoding research. [Meta AI]
- LeRobot v0.6.0 — Open robotics tooling added imagine/evaluate/improve loops. [Hugging Face]
- Embodied.cpp — Portable C++ runtime work for embodied AI models surfaced in developer research feeds. [x.com]
2026-W27
- Mistral OCR 4 — Adds document-intelligence OCR with bounding boxes and block classification. [Mistral]
- Unlimited OCR: One-shot long-horizon parsing — Baidu released an OCR system aimed at long-document parsing. [HN]
- Brain2Qwerty — Meta described non-invasive brain-signal-to-text research. [Meta AI]
2026-W24
- Gemma 4 12B: A unified, encoder-free multimodal model — Compact multimodal release aimed at developer tooling and deployable inference. [HN]
- Introducing Muse Spark: Scaling Towards Personal AI — Meta’s search-surfaced release describes native multimodal reasoning, visual chain-of-thought, tool use, and multi-agent capabilities. [Meta AI]
- MemDreamer — Hierarchical graph memory plus agentic retrieval targets hours-long video understanding without token explosion. [arXiv]
- Whisper Hallucination Detection and Mitigation — Sparse-autoencoder steering study targets hallucinations in ASR systems. [arXiv]
- MMAE: A Massive Multitask Audio Editing Benchmark — Large instruction-based audio-editing benchmark shows current models still fail complex edits. [arXiv]
- NVIDIA Jetson Brings Agentic AI to the Physical World — JetPack and NemoClaw updates connect multimodal agents to edge robotics hardware. [NVIDIA]
2026-W23
- NVIDIA Cosmos 3 for Physical AI — Open omni-model for physical reasoning and action, tying multimodal world modeling to embodied systems. [Hugging Face]
- Lumos-Nexus — Training-efficient unified video generation framework that bridges high-fidelity generation with instruction-grounded video synthesis. [arXiv]
- TunerDiT — Training-free progressive steering method for multi-event text-to-video generation over longer horizons. [arXiv]
- UniAudio-Token — Universal audio tokenizer aimed at combining general audio perception with linguistic alignment for Audio-LLM integration. [arXiv]
- Reachy Mini goes fully local — Local conversation stack for a small robot, reinforcing the trend toward on-device multimodal interaction. [Hugging Face]
2026-W19
- GPT-Realtime-Translate: live speech translation across 70+ languages — Translates live speech from 70+ input languages into 13 output languages while keeping pace with the speaker; built on the same stack as GPT-Realtime-Whisper. [OpenAI]
- GPT-Realtime-Whisper: streaming speech-to-text — Live streaming STT that transcribes as the speaker talks; released alongside GPT-5.5 Instant. [OpenAI]
- Grok voice models and audio APIs — grok-voice-think-fast-1.0 for complex multi-step voice workflows; standalone Grok STT and Grok TTS APIs built on the Tesla/Starlink voice stack. [xAI]
2026-W18
- Grok 4.3 native video input — Up to 5 minutes at 1080p in MP4/MOV/WebM, plus in-chat generation of slides, PDFs, spreadsheets, and PowerPoint decks. [xAI]
- NVIDIA Nemotron 3 Nano Omni — Unified vision/audio/language model with long-context multimodal intelligence for document understanding. [NVIDIA]
- Advancing voice intelligence with new models in the API — OpenAI rolls out new voice models in the API (May 7 — early W19). [OpenAI]
- How OpenAI delivers low-latency voice AI at scale — Engineering deep-dive on the realtime voice stack; HN 507. [HN / OpenAI]
- Adding Benchmaxxer Repellant to the Open ASR Leaderboard — Hugging Face hardens the speech-recognition leaderboard against gaming. [Hugging Face]
2026-W17
- ChatGPT Images 2.0 — OpenAI’s new image-generation model with stronger text rendering, multilingual support, and visual reasoning; Altman compared the leap to GPT-3 → GPT-5. [OpenAI]
- Where’s the raccoon with the ham radio? — Simon Willison’s hands-on stress test of ChatGPT Images 2.0. [Simon Willison]
- Qwen3 TTS running locally in real-time — One of the most expressive open-weight TTS models, demoed in real time. [Reddit]
- Gemma 4 VLA Demo on Jetson Orin Nano Super — On-device vision-language-action robotics demo. [Hugging Face]
- It’s all about the angle: photos, re-composed — Google Research generative re-composition of existing photographs. [Google Research]
- How to Use Transformers.js in a Chrome Extension — Browser-resident multimodal models inside Chrome extensions. [Hugging Face]
- xAI Speech APIs (STT + TTS) — Standalone Grok speech-to-text and text-to-speech APIs with diarization, expressive tags, and real-time + batch endpoints. [xAI]
- Meta Sapiens2 — High-resolution vision transformers pretrained on 1 billion human images for human-centric perception (pose, segmentation, depth). [Meta / Hugging Face]
- Apple DFNDR-12M — Multi-modal training dataset with synthetic captions, embeddings, and metadata for 12.8M images. [Apple / Hugging Face]
2026-W16
- Gemini 3.1 Flash TTS — Granular audio tags let callers direct expressive TTS output; Simon Willison’s hands-on found the tag control meaningfully changes delivery style. [DeepMind]
- Gemini Robotics-ER 1.6 — Spatial reasoning and multi-view understanding for embodied robotics. [DeepMind]
- Grok Speech APIs (STT + TTS) — Standalone speech-to-text and text-to-speech APIs with diarization, expressive tags, real-time and batch endpoints. [xAI]
- Grok 4.3 conversational video understanding — Grok 4.3 can discuss video content conversationally and generate fully populated PDFs, spreadsheets, and decks directly from chat. [xAI]
- HoloTab by HCompany — AI browser companion for multimodal web tasks. [Hugging Face]
- Building a Fast Multilingual OCR Model with Synthetic Data (Nemotron-OCR v2) — NVIDIA’s multilingual OCR with synthetic training data. [Hugging Face]
- Training and Finetuning Multimodal Embedding & Reranker Models with Sentence Transformers — First-class multimodal embedding + reranker recipes in Sentence Transformers. [Hugging Face]
- VEFX-Bench: A Holistic Benchmark for Generic Video Editing and VFX — Large-scale human-annotated benchmark for instruction-guided video editing systems. [arXiv]
2026-W15
- Muse Spark: natively multimodal reasoning — Meta’s first MSL model ships with visual chain-of-thought and tool use as first-class capabilities, not an afterthought on a text-native base. [Meta AI]
- Audio processing landed in llama-server with Gemma-4 — Native audio inference for Gemma 4 in the reference open-source server. [Reddit]
- Gemma 4 audio with MLX — Simon Willison’s hands-on of the MLX audio pipeline for Gemma 4 on Apple silicon. [Simon Willison]
- Waypoint-1.5: Higher-Fidelity Interactive Worlds for Everyday GPUs — Interactive-worlds model that runs on consumer GPUs; part of the emerging “generative world model” category. [Hugging Face]
- Multimodal Embedding & Reranker Models with Sentence Transformers — New recipe for training multimodal embedding + reranker models in the sentence-transformers stack. [Hugging Face]
2026-W14
- Gemma 4: Byte for byte, the most capable open models — Google’s strongest open multimodal model family; built for agentic and on-device workflows; runs on iPhone via AI Edge Gallery (868 HN points). [DeepMind]
- Falcon Perception — TII’s Falcon Perception multimodal model; adds vision to the Falcon family. [Hugging Face / TII]
- Granite 4.0 3B Vision — IBM’s compact 3B VLM targeting enterprise document understanding workflows. [Hugging Face / IBM]
2026-W13
- Gemini 3.1 Flash Live — Google DeepMind ships lower-latency voice model with improved precision for real-time voice interactions. [DeepMind]
- Lyria 3 Pro — Google’s music generation model gains long-form structural coherence; distributed via broader surfaces. [DeepMind]
2026-W12
- Three new Kitten TTS models — smallest under 25MB — New TTS model family with a sub-25MB variant for edge deployment; 561 HN points. [GitHub / HN]
2026-W11
- Grok Imagine update (Elon-flagged) — xAI ships image generation update to Grok; Musk publicly criticizes output quality, drawing attention to prompt-following gaps. [xAI]
2026-W10
- Phi-4-reasoning-vision-15B — Microsoft Research’s 15B open-weight multimodal reasoning model; available on HuggingFace and GitHub; strong reasoning-while-seeing benchmark results. [Microsoft Research]
2026-W09
- Nano Banana 2 — Google DeepMind’s Flash-speed image generation with Pro-level world knowledge and subject consistency; 605 HN points. [DeepMind / Google]
- Moonshine open-weights STT models — Open-weights speech-to-text models beating Whisper Large v3 on accuracy; 316 HN points. [moonshine-ai / HN]
2026-W08
- Qwen3.5: Towards Native Multimodal Agents — Alibaba positions Qwen3.5 as a full multimodal-agent platform with native vision integration; 434 HN points. [Qwen / Alibaba]
2026-W07
- Qwen-Image-2.0 — Alibaba’s new image generation model targets professional infographics and photorealistic output; 422 HN points. [Qwen / HN]
2026-W06
- Voxtral Transcribe 2 — Mistral’s next-generation ASR model; 1012 HN points, 241 comments; strongest open transcription release of Q1 2026. [Mistral / HN]
- Grok Imagine v1.0 — xAI ships 10-second 720p video generation model described as their biggest leap in prompt-following; released same week as SpaceX acquisition of xAI. [xAI]
2026-W05
- Grok Imagine API — xAI launches text-to-video, image-to-video, and video editing API at $0.05/sec; positions Grok directly against Sora and Runway in programmatic video generation. [xAI]
2026-W04
- Qwen3-TTS family open sourced — Full TTS suite with voice cloning and voice design capabilities from Alibaba; 744 HN points. [Qwen / Alibaba]
- Waypoint-1: Real-time interactive video diffusion — Real-time interactive video diffusion model from Overworld; published Jan 20. [Hugging Face]
2026-W03
- FLUX.2 Klein: Towards Interactive Visual Intelligence — Black Forest Labs previews FLUX.2 Klein; targets iterative, interactive visual generation workflows. [Black Forest Labs / HN]
- MedGemma 1.5 and MedASR — Google ships MedGemma 1.5 for medical imaging and a dedicated medical speech-to-text model. [Google Research]
- Veo 3.1 Ingredients to Video — DeepMind’s Veo update adds consistency and creativity controls; supports vertical video. [DeepMind]
2026-W02
- NVIDIA Cosmos Reason 2 — Open reasoning vision-language model for physical AI; 2B and 8B; top of Physical AI Bench. [NVIDIA / Hugging Face]
2026-W01
- GLM-Image (Z.ai, teaser) — Two-stage compact-token text-to-image encoder scaling to 4K; reportedly first top-tier multimodal trained on Huawei Ascend. [r/LocalLLaMA / eWEEK]
- FLUX.2-dev-Turbo editing — Community reports FLUX.2-dev-Turbo is strong at image editing. [r/LocalLLaMA]
2025-W52
- MiniMax M2.1 multimodal + VIBE benchmark — Multi-language MoE also introduces VIBE (Visual & Interactive Benchmark for Execution). [MiniMax]
2025-W51
- Gemini 3 Flash — Pro-grade multimodal reasoning at Flash speed; MMMU Pro 81.2%; default in Gemini app and AI Mode. [Google]
2025-W50
- Improved Gemini audio models for powerful voice experiences — Upgraded voice-mode audio stack. [DeepMind]
2025-W48
- Diffusers welcomes FLUX.2 — Black Forest Labs’ FLUX.2 open image model integrated with HF Diffusers. [Hugging Face]
2025-W47
- Nano Banana Pro — Gemini 3 Pro Image model with build tooling. [DeepMind]
- Gemini 3 Pro multimodal benchmarks — 81% MMMU-Pro and 87.6% Video-MMMU. [DeepMind]
- Open ASR Leaderboard: Multilingual & Long-Form Tracks — Expanded ASR benchmarking. [Hugging Face]
2025-W46
- SIMA 2 — Gemini-powered 3D-world agent handles instructions given through both language and images. [DeepMind]
- Faster Maya1 TTS model — Community TTS speedup. [r/LocalLLaMA]
2025-W44
- Qwen3-VL-30B-A3B — Community report on Qwen3-VL-30B-A3B practical quality; 264 upvotes. [r/LocalLLaMA]
- Vision = Language: Decoding VLM tokens — Community research into how VLMs encode visual tokens. [r/LocalLLaMA]
2025-W43
- Qwen’s VLM is strong — r/LocalLLaMA thread on Qwen3-VL practical quality; 131 upvotes. [r/LocalLLaMA]
- Supercharge your OCR Pipelines with Open Models — Practitioner’s guide to OCR on open OCR/VLM stacks. [Hugging Face]
2025-W35
- Gemini 2.5 Flash Image (Nano Banana) — Google’s image model goes viral; #1 on LMArena; 200M+ edits and 10M+ new users; 1,093 HN points. [Google]
- Chatterbox TTS faster on CUDA — Community optimization reaches 155 it/s on 3090. [r/LocalLLaMA]
- Hunyuan-MT-7B machine translation — Tencent releases multilingual translation models. [r/LocalLLaMA]
2025-W33
- OWhisper — Ollama for real-time speech-to-text — Local real-time speech recognition tool; 289 HN points. [HN]
- NVIDIA open multilingual speech dataset and models — Open dataset and models for multilingual speech-to-text. [NVIDIA]
2025-W32
- GPT-5 MMMU 84.2% — Sets new SOTA on multimodal understanding benchmark; visual perception highlighted as a key GPT-5 strength. [OpenAI]
- MicroLlaVA — Community builds a vision-language model on a single NVIDIA 4090. [r/LocalLLaMA]
2025-W31
- FLUX.1 Krea weights released — Krea open-sources FLUX.1 image generation weights; 369 HN points. [Krea]
- LLM Embeddings Explained: visual guide — Interactive HuggingFace Space for understanding embeddings visually; 451 HN points. [HF]
2025-W29
- Voxtral — Mistral’s first open-source audio model. [Mistral]
- Mistral Le Chat gains voice I/O — Le Chat adds voice as part of the Deep Research / Projects update. [Mistral]
- Chroma: Context Rot — Research on long-context degradation; relevant to multimodal contexts too. [Chroma]
- Show HN: Shoggoth Mini — soft tentacle robot powered by GPT-4o + RL — DIY soft-robotics project controlled by a multimodal model. [HN]
2025-W28
- Kimi K2 multimodal variants — 1.07T Kimi K2 release spawns quick vision-enabled fine-tunes. [r/LocalLLaMA]
2025-W26
- DeepMind Gemini Robotics On-Device — DeepMind’s on-device Gemini Robotics model for local control. [DeepMind]
- DeepMind AlphaGenome — Genomics model; cross-modal genome understanding. [DeepMind]
2025-W24
- Meta V-JEPA 2 world model + physical-reasoning benchmarks — Meta’s next-gen joint-embedding predictive architecture. [Meta]
- Chatterbox TTS — Resemble AI’s open-source TTS; best open TTS release of the quarter. [HN]
- Magistral Small vision fine-tune — Community vision-enabled fine-tune of Mistral Magistral Small, shipped within days of release. [r/LocalLLaMA]
- Launch HN: Vassar Robotics — $219 robot arm that learns new skills — Cheap learning robot arm signaling a home-robotics wave. [HN]
2025-W22
- FLUX.1 Kontext — Black Forest Labs’ new context-aware image model; first major BFL release after a rumored quiet period. [HN]
- Bagel: open-source unified multimodal model — New open-weight unified multimodal model. [HN]
2025-W21
- Google I/O: Veo 3, Imagen 4, Flow — Veo 3 video with native audio + dialogue, Imagen 4 image model, and Flow AI filmmaking workflow. [Google]
- Gemini Live camera + screen sharing free on Android and iOS — Real-time multimodal interaction goes free. [Google]
2025-W19
- LTXVideo 13B — Open video generation model scaled to 13B parameters. [HN]
2025-W17
- DeepMind Lyria 2 music generation — High-quality music generation with expanded access; 300 HN points, 426 comments. Significant for creative industry implications. [DeepMind/HN]
- Kimi Audio 7B: SOTA audio foundation model — Moonshot AI’s SOTA audio understanding and generation model; 206 upvotes. [r/LocalLLaMA]
- LLMs can see and hear without any training — Facebook Research shows LLMs process vision and audio with zero additional training; 210 HN points. [HN]
- NotebookLM-style Dia — Open-source alternative to NotebookLM podcast-style audio generation; 94 upvotes. [r/LocalLLaMA]
- Dia-1.6B in Jax — Jax implementation of Dia text-to-audio model; 80 upvotes. [r/LocalLLaMA]
2025-W16
- FramePack: next-frame video generation — Progressive video generation via next-frame-section prediction; local video gen model; 164 upvotes. [r/LocalLLaMA]
2025-W14
- Llama 4 Scout and Maverick — natively multimodal MoE — Meta’s first open-weight natively multimodal models (text + vision, no adapters); Scout has 10M context; Behemoth (288B active) still training. Despite disappointing coding scores, native multimodality in an open-weight MoE is architecturally significant. [Meta AI Blog]
- MoCha: Towards Movie-Grade Talking Character Synthesis — Cinematic-quality talking character synthesis; step change in avatar and video AI quality. [arXiv via @huggingpapers]
2025-W13
- Introducing 4o Image Generation — OpenAI ships native in-ChatGPT image generation on March 25; system card addendum published simultaneously. GPT-4o can now generate images inline without a separate product. [OpenAI Blog]
- Qwen2.5-Omni Technical Report — Alibaba Qwen team’s multimodal model technical report; strong open-weights multimodal release in a week dominated by Gemini 2.5. [arXiv via @huggingpapers]
- Video-R1: Reinforcing Video Reasoning in MLLMs — Applies reinforcement learning to improve video reasoning in multimodal LLMs; extends R1-style RL training to video understanding. [arXiv via @huggingpapers]
- Wan: Open and Advanced Large-Scale Video Generative Models — Open-source large-scale video generation model (62 authors); adds to the field of accessible video gen alternatives. [arXiv via @huggingpapers]
2025-W12
- MoshiVis — First open-source real-time speech + vision model. [Kyutai]
- Orpheus-3B — Emotive TTS with tags and zero-shot cloning. Apache 2.0. [Canopy Labs]
2025-W11
- Gemini Robotics — VLA models adding physical actions as new output modality for Gemini 2.0. [DeepMind]
- GPT-Sovits V3 TTS — Zero-shot voice cloning, multi-language, 407M params. [Reddit]
2025-W10
- Aya Vision: multilingual multimodal AI — Cohere for AI’s open multilingual vision-language model covering 23+ languages; addresses underrepresentation of non-English languages in multimodal models. [Hugging Face]
- Amazon Alexa+ multimodal interactions — Upgraded Alexa gains richer multimodal capabilities on Echo Show devices, combining vision input with conversational AI. [ZDNet]
Baseline (through 2025-Q1)
Image Generation
- DALL-E 3 (OpenAI, Oct 2023) — Integrated into ChatGPT. Strong text rendering and prompt adherence. Uses ChatGPT to rewrite prompts for better results.
- Midjourney v6/v6.1 (2024) — Continued dominance in artistic image generation. Improved prompt following and text rendering. Added web interface beyond Discord.
- Source: https://www.midjourney.com/
- Stable Diffusion 3 (Stability AI, Feb 2024 announcement, Jun 2024 release) — Multimodal Diffusion Transformer (MMDiT) architecture. Improved text rendering and prompt adherence. Open weights for smaller variants.
- FLUX (Black Forest Labs, Aug 2024) — Founded by ex-Stability AI researchers. FLUX.1 models (schnell, dev, pro) quickly gained traction as high-quality open alternatives. Rectified flow transformers.
- Source: https://blackforestlabs.ai/
- Ideogram — Strong text-in-image generation. Raised significant funding. Canvas feature for editing.
- Source: https://ideogram.ai/
- Google Imagen 3 (2024) — Google’s latest image model, available in Gemini. High photorealism.
Video Generation
- Sora (OpenAI, Feb 2024 preview, Dec 2024 launch) — Text-to-video model generating up to 20-second clips at 1080p. Stunning quality but limited availability. ChatGPT Plus/Pro at launch.
- Source: https://openai.com/index/sora/
- Runway Gen-3 Alpha (Jun 2024) — Professional-grade video generation. Used in film/TV production. Strong motion and consistency.
- Source: https://runwayml.com/
- Kling (Kuaishou) — Chinese video model with impressive quality, especially for motion. Available internationally.
- Google Veo 2 (Dec 2024) — Google’s video generation model. High-quality 4K output. Available in VideoFX.
- Pika — Video generation startup. Focus on stylized and motion-controlled video editing.
- Source: https://pika.art/
Audio and Speech
- ElevenLabs — Leading voice synthesis platform. Realistic text-to-speech, voice cloning, and dubbing. Raised $180M (Jan 2025).
- Source: https://elevenlabs.io/
- NotebookLM Audio Overviews (Google, Sep 2024) — Generates podcast-style audio discussions from documents. Viral adoption. Surprisingly natural-sounding two-host conversations.
- Source: https://notebooklm.google.com/
- OpenAI Advanced Voice Mode (Sep 2024) — Real-time conversational voice in ChatGPT using GPT-4o’s native audio capabilities. Low latency, natural turn-taking, emotional expression.
- Whisper v3 (OpenAI) — Open-source speech recognition. Remains the standard for transcription. Runs locally via faster-whisper and whisper.cpp.
Music Generation
- Suno — AI music generation from text prompts. Full songs with vocals, instruments, and lyrics. v3/v3.5 improved quality significantly. Copyright concerns remain.
- Source: https://suno.com/
- Udio — Competitor to Suno with similar text-to-music capabilities. Strong on genre diversity.
- Source: https://www.udio.com/
Vision-Language Models
- GPT-4o — Natively multimodal (text, vision, audio in one model). Strong visual understanding, document parsing, and image-based reasoning.
- Claude 3.5 Sonnet — Vision capabilities for screenshot analysis, document understanding, chart reading. Used extensively in coding workflows for UI development.
- Gemini 1.5 Pro/Flash — Strong vision with long-context advantage (can process hour-long videos). Native PDF understanding.
- Llama 3.2 Vision (Sep 2024) — First multimodal Llama models (11B, 90B). Open-weights vision-language models.
Key Trends
- Natively multimodal models (GPT-4o, Gemini) outperform adapter-based approaches. Single model handling multiple modalities is becoming the norm.
- Real-time interaction — Voice and vision in real-time conversations. Google’s Project Astra and GPT-4o voice set new expectations.
- Copyright litigation — Multiple lawsuits from artists, photographers, and publishers against image/video generation companies. Legal landscape still evolving.