Every notable new AI model from the major labs, newest first and curated from official announcements. Filter by provider and model type to see who shipped what, and when.
698 releases·Page 8 of 35
Provider
Model
Released
Link
Google
Gemini Robotics-ER 1.6
Gemini Robotics-ER 1.6 is Google DeepMind's reasoning-first model for embodied reasoning, built to help robots understand visual and spatial context, plan tasks, and detect success. It serves as a high-level reasoning layer that natively calls tools such as Google Search and VLA models. It improves on ER 1.5 in pointing, counting, and success detection, adds instrument reading through agentic vision, and offers stronger safety-policy compliance. Developers can access it via the Gemini API and Google AI Studio.
April 13, 2026
Vidu
viduq3
viduq3 is the standard model in Vidu's Q3 video generation series, added to the reference-to-video API on April 13, 2026. It supports smart shot switching, generates audio and video together, and delivers stronger multi-camera consistency. In reference-to-video mode it accepts up to 7 reference images and produces 3 to 16 second clips at up to 1080p with audio on by default, suited to narrative work such as short dramas and manga-style series.
April 13, 2026
Vidu
viduq3-mix
viduq3-mix is a ViduQ3-series video generation model that Vidu added to its reference-to-video API in April 2026. It generates subject-consistent video from up to seven reference images, and the official docs describe it as the most balanced tier for general-purpose scenes, with audio and visuals produced together and smart shot transitions. Output runs up to 16 seconds (5 by default) at up to 1080p and 24fps, in 16:9, 9:16, or 1:1 aspect ratios.
April 13, 2026
ByteDance
Seeduplex
Seeduplex is a native full-duplex speech LLM from ByteDance's Seed team. Moving beyond the half-duplex turn-taking paradigm, it truly listens while speaking, combining large-scale speech pre-training with reinforcement learning (RL) and joint speech-semantic modeling. The model delivers precise interference suppression and adaptive endpoint detection in complex acoustic scenarios, cutting endpoint latency by roughly 250ms and halving false response and false interruption rates. It has been fully rolled out in the Doubao App, bringing natural real-time voice interaction to over 100 million users.
April 9, 2026
Meta
Muse Spark
Muse Spark is the first model in the Muse family from Meta Superintelligence Labs, a natively multimodal reasoning model with tool-use, visual chain of thought, and multi-agent orchestration. Its Contemplating mode runs parallel reasoning agents and scores 58% on Humanity's Last Exam and 38% on FrontierScience Research. Its rebuilt pre-training stack reaches Llama 4 Maverick level capability with over an order of magnitude less compute. It is available on meta.ai and the Meta AI app, with a private API preview for select users.
April 8, 2026
Zhipu AI
GLM-5.1
GLM-5.1 is Zhipu AI's next-generation flagship model for agentic engineering, with significantly stronger coding capabilities than its predecessor. It achieves state-of-the-art performance on SWE-Bench Pro and leads GLM-5 by a wide margin on NL2Repo and Terminal-Bench 2.0. The model is built for long-horizon tasks, sustaining productive optimization across hundreds of iterations and thousands of tool calls. It is released as open source under the MIT License and works with coding agents such as Claude Code.
April 7, 2026
Grok
grok-imagine-image-quality
grok-imagine-image-quality is the Quality Mode of Grok Imagine, xAI's image generation and editing model, available through the Grok Imagine API for enterprise developers and teams. It focuses on higher realism, stronger text rendering, and finer creative control, and ranks among the strongest models on independent leaderboards. Typical uses include photorealistic product renders, marketing assets, UGC-style content, and visuals from cinematic scenes to UI icons.
April 3, 2026
Grok
grok-imagine-image-quality-20260403
grok-imagine-image-quality-20260403 is the dated snapshot of Quality Mode for xAI's Grok Imagine image model, available to enterprise developers and teams through the Grok Imagine API for both image generation and editing. xAI highlights higher realism, stronger text rendering, and better creative control, with the model ranking among the strongest entries on independent leaderboards. Typical uses include product visualization, marketing assets, and UGC-style content, and it pairs with Grok Imagine's video generation capabilities.
April 3, 2026
Qwen
Wan2.7-Image-To-Video
Wan2.7-Image-To-Video is the image-to-video model in Alibaba's Wan series, released on Alibaba Cloud Model Studio on April 3, 2026. It generates video from a first frame or from a pair of first and last frames, and supports video continuation with end-frame control. Output runs 2 to 15 seconds at 720P or 1080P, 30fps MP4, with synchronized audio. Alibaba highlights stronger performances, nuanced in dramatic scenes and punchy in action, with more dramatic and well-paced shot transitions.
April 3, 2026
Qwen
Wan2.7-Reference-To-Video
Wan2.7-Reference-To-Video (wan2.7-r2v) is Alibaba Cloud Model Studio's reference-to-video model. It accepts up to five mixed image and video references to keep characters, props, and scenes consistent, and each subject can be assigned its own voice through an audio reference. The model can also turn a single storyboard grid into a coherent multi-shot video with sound, outputting 720P or 1080P clips suited to solo performances and multi-character scenes.
April 3, 2026
Qwen
Wan2.7-Text-To-Video
Wan2.7-Text-To-Video is Alibaba Cloud's text-to-video model that turns text prompts into fluid video. It supports 720P and 1080P output with durations of 2 to 15 seconds, and can automatically generate matching background music or sound effects. The model delivers more expressive acting, with nuanced dramatic scenes and hard-hitting action, and supports multi-shot storytelling controlled through natural language prompts.
April 3, 2026
Qwen
Wan2.7-Video-Edit
Wan2.7-Video-Edit is Alibaba Cloud's Wan video editing model on Model Studio, released in April 2026. It applies both localized and global edits to existing footage from natural language prompts, covering element addition, removal, and replacement, background and lighting changes, style transfer, and rewrites of character actions or dialogue. Reference images can be used to swap in clothing and props, and the model replicates motion, camera movement, and effects from a reference video, with output at 720P or 1080P.
April 3, 2026
Google
Gemma 4
Gemma 4 is the fourth generation of Google's open model family, built on the same research as Gemini 3 for advanced reasoning and agentic workflows. The lineup spans edge-focused E2B and E4B, a 26B MoE that activates 3.8B parameters, and a 31B dense model with up to 256K context. The 31B ranks third among open models on the Arena AI text leaderboard, with Google citing wins over models 20x its size. Everything ships under the commercially permissive Apache 2.0 license, supporting 140+ languages for on-device inference, coding assistants, and sovereign deployments.
April 2, 2026
Zhipu AI
GLM-5V-Turbo
GLM-5V-Turbo is Z.AI's multimodal vision coding model, a Turbo variant of the GLM-5V series that natively understands images, video, design mockups, and document layouts at a smaller model size. It provides a 200K context window with up to 128K output tokens, plus thinking mode and function calling. Zhipu reports leading performance on core benchmarks for multimodal coding, tool use, and GUI agents, with strong results on AndroidWorld and WebVoyager. It works with agents such as Claude Code and OpenClaw for frontend recreation, GUI automation, and code debugging.
April 2, 2026
Qwen
Qwen3.6-Plus
Qwen3.6-Plus is the Plus tier of Alibaba Cloud's Qwen3.6 natively multimodal vision-language series. Per the official release, it delivers performance comparable to leading frontier models and improves markedly over the 3.5 series. The model accepts image, text, and video input with a 1M-token context window, with stronger capability in Agentic coding, front-end work, and Vibe coding as well as multimodal recognition, OCR, and object localization. Pricing starts at CNY 2 per million input tokens in the Beijing region.
April 1, 2026
Qwen
Qwen3.6-Plus-2026-04-02
Qwen3.6-Plus-2026-04-02 is the April 2, 2026 snapshot of Qwen3.6 Plus, a native vision-language model in Alibaba Cloud's Qwen3.6 family served on Model Studio. Alibaba describes it as rivaling top frontier models with clear improvements over the 3.5 series, and notes stronger Agentic coding and frontend coding alongside multimodal perception such as object recognition, OCR, and object grounding. It accepts text, image, and video input with a 1M token context window, and the pinned snapshot suits production workloads that need stable, reproducible behavior.
April 1, 2026
Qwen
Wan2.7-Image-Generator-Edit
Wan2.7-Image-Generator-Edit is the standard tier of Wan 2.7, Alibaba Cloud's image generation and editing model on Model Studio, released on April 1, 2026. It covers text-to-image, image set generation, image editing, multi-image reference generation, and interactive editing, with output at 1K or 2K resolution. Relative to the flagship Pro tier it runs faster while delivering stronger text rendering, subject consistency, and adherence to complex instructions.
April 1, 2026
Qwen
Wan2.7-Image-Generator-Edit-Pro
Wan2.7-Image-Generator-Edit-Pro is the flagship image generation and editing model in Alibaba's Wan family, released in April 2026. It covers text-to-image, multi-image generation, image editing, and interactive editing, with a brand color palette and up to 4096x4096 resolution for text-to-image output. Editing accepts up to nine reference images, and the model delivers stronger text rendering and subject consistency, suited to brand design and consistent multi-image workflows.
April 1, 2026
Google
Gemini 3.1 Flash Live
Official release — detailed description coming soon.
March 26, 2026
PixVerse
PixVerse V6
Official release — detailed description coming soon.
Four ways to read the board. Each one turns the stream of new AI models into a clear answer. Pick the view that fits your question.
01
Read the newest launches first
The board lists new AI models newest first, so the top row is always the latest launch from any lab. Every row links to the official announcement, so one click gets you to the source. Dates reflect the initial public release.
02
Follow one provider
Pick a lab in the Provider filter. OpenAI, Anthropic, Google, Meta, DeepSeek, Qwen and 20+ more each get their own view. Use it to read a lab's release rhythm: how often it ships, and which model types it bets on.
03
Compare model types
Filter by type: text/LLM, multimodal, image, video, or audio. The mix shows where the industry is pushing: text models still lead the count, while image and video generators ship in waves. Watch the balance shift month by month.
04
Catch up on what you missed
Page back through the timeline. Some months see dozens of launches, and the archive reaches back years. The board is a full release history, not just a feed of today's new AI models. Reconstruct any lab's year in a few scrolls.
New AI Models FAQ
What the board tracks, where the data comes from, and how to read it.
The board sorts releases newest first, so the top row is always the most recent launch from any provider. Recent months brought new AI models from labs like OpenAI, Anthropic, Google, Qwen and Zhipu AI. Every row links to the official announcement with its exact release date.
Twenty-nine providers and counting: OpenAI, Anthropic, Google, Meta, Microsoft, Amazon, NVIDIA, Mistral AI, DeepSeek, Qwen, Moonshot AI, Zhipu AI, ByteDance, xAI and more. Western and Chinese labs share one board, so a launch abroad still shows up the same day.
Five kinds: text/LLM, multimodal, image generation, video generation, and audio/speech. Text models carry the largest share of launches, but image and video generators ship in fast waves. The Type filter isolates any one stream.
Constantly. Major labs now ship new AI models on a weekly cadence, and busy weeks bring more than one frontier launch. Since 2024 the pace has kept climbing, and the board adds each release as its official announcement goes live.
Each row uses the date of the initial public release: the day the lab announced the model or opened it to users. Betas, rumors and paper preprints do not count. Every entry links to the official announcement, so you can verify the date yourself.
Yes. The board doubles as an archive reaching back to 2017. Page back to see how the release cadence exploded, from a handful of new AI models a year to hundreds. It is a searchable release history, not just a live feed.
Yes. The Provider filter narrows the board to one lab or any set of labs. Combine it with the Type filter to answer narrow questions, like every video model a single lab has shipped.
Yes, completely. Every release, filter and official link is free to browse, no signup needed. The board updates as new models are announced, so it pays to check back after every launch event.
Curated from official provider announcements; dates reflect the initial public release.