Skip to main content

New AI Models

Every notable new AI model from the major labs, newest first and curated from official announcements. Filter by provider and model type to see who shipped what, and when.

Provider
Type

698 releasesPage 3 of 35

ProviderModelReleasedLink
Qwen

Qwen-Audio-3.0-ASR-Flash-Streaming

Qwen-Audio-3.0-ASR-Flash-Streaming is a flagship speech recognition model from Tongyi Lab, built for low-latency, high-concurrency real-time interaction and returning text while the speaker is still talking. It supports 30 languages including Chinese, English, Japanese and Korean, plus seven major Chinese dialect groups, and combines context enhancement, hotwords, and specialized training for industry vocabulary. Typical uses include real-time meeting interpretation, contact center agent assistance, and voice assistants.

July 30, 2026
Google

Lyria 3.5

Lyria 3.5 is Google's latest AI music generation model, announced on July 29, 2026. It improves on musicality, lyrics, and vocals: richer and more natural melodic structures, higher-quality lyrics with better prompt adherence and structural awareness, and more realistic, emotionally nuanced vocals with clearer pronunciation. Users also get easier control over the tempo and duration of generated tracks. Lyria 3.5 is available in Google Flow Music.

July 29, 2026
Anthropic

Claude Opus 5

Claude Opus 5 is a text model from Anthropic that comes close to the frontier intelligence of Claude Fable 5 at half the price, and is now the default model on Claude Max. It leads all models on Frontier-Bench v0.1, more than doubling Opus 4.8's score at lower cost per task, and scores three times the next-best model on ARC-AGI 3. Built for long-running, multi-step agent work, it verifies its own output and iterates until tasks succeed, making it well suited to software engineering, research, and knowledge work.

July 24, 2026
Midjourney

V8.2

V8.2 is Midjourney's image model released in July 2026, with the update focused on aesthetics, image quality, and personalization. Compared with V8, it generates more creative, sophisticated results and sharply reduces the random low-quality images users occasionally saw before. Personalization also reads individual taste more accurately, especially for profiles with many ratings, and draws on a larger, improved image pool. It is now live on midjourney.com.

July 24, 2026
FLUX

FLUX 3

FLUX 3 is Black Forest Labs' multimodal foundation model that learns from images, videos, and audio within a unified architecture, with all capabilities built on a single multimodal flow matching model. For image work it synthesizes and edits across a wide range of styles, aspect ratios, and resolutions, and renders high-accuracy text in multiple languages, with marked gains over earlier FLUX versions on complex prompts and text generation. Access rolls out through Early Access, followed by APIs and private weights, plus an open-weight FLUX 3 Dev backbone.

July 23, 2026
Google

3.5 Flash Cyber

Gemini 3.5 Flash Cyber is Google's lightweight cybersecurity model, fine-tuned from Gemini 3.5 Flash to find, validate, and patch code vulnerabilities more cost-efficiently than large security models. It scores competitive pass@1 results against much larger models on the CyberGym benchmark and surpasses mainline 3.5 Flash and 3.6 Flash on Google's Big Sleep evaluation. Access is limited to governments and trusted partners through the CodeMender pilot.

July 21, 2026
Google

3.5 Flash-Lite

Gemini 3.5 Flash-Lite is a low-latency, cost-effective multimodal model from Google, released as a stable version for production use on July 21, 2026. It accepts text, image, video, audio, and PDF inputs, with an input token limit of 1,048,576. Optimized for subagent tasks and document parsing, it suits high-volume agentic workflows and simple data extraction where latency and API cost are the primary constraints.

July 21, 2026
Google

Gemini 3.6 Flash

Gemini 3.6 Flash is Google's fast, cost-efficient multimodal model, generally available as of July 21, 2026. It accepts text, image, video, audio, and PDF input within a 1M token context window. Compared with 3.5 Flash, it delivers higher token efficiency and stronger code and agent planning at a lower price, while addressing developer feedback on output verbosity. It excels at code generation, agentic execution, and spatial reasoning, making it a fit for rapid, iterative agent workflows.

July 21, 2026
ByteDance

Seed Audio 1.0

Seed Audio 1.0 is ByteDance's audio creation model for full-scene audio generation. It jointly models voice, sound effects, and ambience in a unified framework for end-to-end film-grade audio production, with prompt-level timing control at 100 ms intervals. The model supports zero-shot voice generation from text or reference samples, produces up to two minutes of audio in a single pass, and transfers the same voice naturally across more than 20 languages. Official evaluations show an audio Availability Rate above 90% in most scenarios such as film, short dramas, and podcast dialogue.

July 20, 2026
Qwen

Qwen-Image-3.0-Pro

Qwen-Image-3.0-Pro is an image generation model available on Alibaba Cloud Model Studio, positioned as a practical productivity tool rather than a purely aesthetic one. It accepts up to 4.5k tokens of input, generating complex layouts such as newspapers, storyboards, menus, and exam papers in a single pass. It renders text as small as 10px with precision and reproduces fine details like micro-expressions and hair strands with near-photographic realism, while natively rendering text in 12 languages and simulating web, game, and livestream interfaces.

July 20, 2026
Qwen

Qwen3.7-Flash

Qwen3.7-Flash is the fast, cost-effective tier of Alibaba's Qwen3.7 family, a natively vision-language model that accepts image, text, and video input. Compared with Qwen3.6-Flash, it delivers stronger multimodal understanding and object recognition, improved real-world perception and spatial intelligence, and more stable end-to-end execution in agentic scenarios such as Search Agent and CI Agent. Multimodal coding and the vibe coding experience are also refined. It offers a 1M context window, with input from RMB 0.2 per million tokens in the Beijing region.

July 20, 2026
Qwen

Qwen3.7-Flash-2026-07-15

Qwen3.7-Flash-2026-07-15 is the 2026-07-15 snapshot of Qwen3.7 Flash, the fast and cost-efficient tier of Alibaba Cloud's natively multimodal Qwen3.7 vision-language series on Model Studio, suited to production deployments that need a pinned version. Compared with 3.6-Flash it delivers stronger multimodal understanding and agent execution, better object recognition and spatial intelligence, and accepts image, text, and video input with a 1M-token context. Typical workloads include Search Agent and CI Agent scenarios as well as multimodal coding and vibe coding.

July 20, 2026
Google

Gemini 3.5 Flash Cyber

Gemini 3.5 Flash Cyber is Google's lightweight cybersecurity model, fine-tuned from Gemini 3.5 Flash to find, validate, and patch vulnerabilities more effectively than the mainline Flash models. Its speed and affordability suit scanning large codebases across many code paths. It proved competitive with significantly larger models on the CyberGym benchmark and surpassed mainline 3.5 Flash and 3.6 Flash on Google's Big Sleep evaluation. It is offered to governments and trusted partners via CodeMender through a limited-access pilot.

July 17, 2026
Moonshot AI

Kimi K3

Kimi K3 is Moonshot AI's open-weight, native multimodal agentic model and its most capable release to date. Built on a 2.8T-parameter MoE architecture with 104B activated parameters and a 1M-token context window, it handles text, image, and video in one model. Moonshot positions it as the world's first open 3T-class model, competitive with GPT-5.6 Sol and Claude Opus 4.8 on benchmarks such as GPQA Diamond, SWE-Marathon, and BrowseComp. It ships under the Kimi K3 License for long-horizon coding and agentic knowledge work.

July 16, 2026
Qwen

Qwen-Audio-3.0-TTS-Flash

Qwen-Audio-3.0-TTS-Flash is a high-performance text-to-speech model from Alibaba's Qwen team, optimized for real-time interaction. It supports more low-resource languages and Chinese dialects than its predecessor, with free-style instruction following and fine-grained tags for controlling emotion, tone, character, speech rate, and volume. The model is also more robust in noisy and reverberant acoustic conditions. The Flash tier prioritizes real-time synthesis, making it a fit for voice assistants, real-time conversation, and intelligent customer service.

July 14, 2026
Qwen

Qwen-Audio-3.0-TTS-Plus

Qwen-Audio-3.0-TTS-Plus is a high-performance text-to-speech model from Alibaba's Qwen team built for high-quality speech generation. It supports more low-resource languages and Chinese dialects, with free-style instructions and fine-grained tags for controlling emotion, tone, speaking rate, and speaking style. It is also more robust under noise and reverberation, with further improved audio quality, clarity, and expressiveness. The Plus tier prioritizes output quality and detail, making it a fit for content creation, audiobooks, film dubbing, and brand voice design.

July 14, 2026
Meta

Muse Spark 1.1

Muse Spark 1.1 is a multimodal reasoning model from Meta Superintelligence Labs, built for agentic tasks with major gains in tool and computer use, coding, and multimodal understanding. It actively manages a 1 million token context window, orchestrates multi-agent systems to cut end-to-end latency, and posts substantial coding gains over its predecessor on Meta Internal Coding Bench. Developers can access it through the Meta Model API public preview, or use it in Thinking mode in the Meta AI app and on meta.ai.

July 9, 2026
OpenAI

GPT-5.6

OpenAI's GPT-5.6 family of frontier models ships in three capability tiers: Sol, the new flagship; Terra, balanced for everyday work; and Luna, a fast, cost-efficient option, available across ChatGPT, Codex, and the OpenAI API. Sol delivers state-of-the-art results in coding, knowledge work, cybersecurity, and science, outperforming GPT-5.5 and Claude Fable 5 with fewer tokens and lower estimated cost. A new ultra setting coordinates multiple agents in parallel to finish demanding tasks faster.

July 9, 2026
ByteDance

Seedream 5.0 Pro

Seedream 5.0 Pro is a multimodal image generation model from ByteDance's Seed team, built for advanced reasoning and professional production. It handles high-density infographics with rich, accurate text rendering and supports interactive editing that follows spatial annotations and sketches, including layer separation. With authentic lighting and skin textures plus native prompting in a dozen widely used languages, it fits education graphics, poster design, and multilingual content workflows.

July 8, 2026
OpenAI

GPT-Live

GPT-Live is OpenAI's new real-time voice model built on a full-duplex architecture, so it listens and speaks at the same time and handles natural turn-taking, interruptions, and live translation. It is OpenAI's smartest voice model to date, delegating search and deeper reasoning to a frontier model in the background (GPT-5.5 at launch) while the conversation keeps flowing. In human evaluations it was clearly preferred over Advanced Voice Mode. GPT-Live-1 and GPT-Live-1 mini are rolling out to ChatGPT users worldwide on iOS, Android, and the web, with API access planned.

July 8, 2026

How to track new AI models

Four ways to read the board. Each one turns the stream of new AI models into a clear answer. Pick the view that fits your question.

01

Read the newest launches first

The board lists new AI models newest first, so the top row is always the latest launch from any lab. Every row links to the official announcement, so one click gets you to the source. Dates reflect the initial public release.

02

Follow one provider

Pick a lab in the Provider filter. OpenAI, Anthropic, Google, Meta, DeepSeek, Qwen and 20+ more each get their own view. Use it to read a lab's release rhythm: how often it ships, and which model types it bets on.

03

Compare model types

Filter by type: text/LLM, multimodal, image, video, or audio. The mix shows where the industry is pushing: text models still lead the count, while image and video generators ship in waves. Watch the balance shift month by month.

04

Catch up on what you missed

Page back through the timeline. Some months see dozens of launches, and the archive reaches back years. The board is a full release history, not just a feed of today's new AI models. Reconstruct any lab's year in a few scrolls.

New AI Models FAQ

What the board tracks, where the data comes from, and how to read it.

The board sorts releases newest first, so the top row is always the most recent launch from any provider. Recent months brought new AI models from labs like OpenAI, Anthropic, Google, Qwen and Zhipu AI. Every row links to the official announcement with its exact release date.

Twenty-nine providers and counting: OpenAI, Anthropic, Google, Meta, Microsoft, Amazon, NVIDIA, Mistral AI, DeepSeek, Qwen, Moonshot AI, Zhipu AI, ByteDance, xAI and more. Western and Chinese labs share one board, so a launch abroad still shows up the same day.

Five kinds: text/LLM, multimodal, image generation, video generation, and audio/speech. Text models carry the largest share of launches, but image and video generators ship in fast waves. The Type filter isolates any one stream.

Constantly. Major labs now ship new AI models on a weekly cadence, and busy weeks bring more than one frontier launch. Since 2024 the pace has kept climbing, and the board adds each release as its official announcement goes live.

Each row uses the date of the initial public release: the day the lab announced the model or opened it to users. Betas, rumors and paper preprints do not count. Every entry links to the official announcement, so you can verify the date yourself.

Yes. The board doubles as an archive reaching back to 2017. Page back to see how the release cadence exploded, from a handful of new AI models a year to hundreds. It is a searchable release history, not just a live feed.

Yes. The Provider filter narrows the board to one lab or any set of labs. Combine it with the Type filter to answer narrow questions, like every video model a single lab has shipped.

Yes, completely. Every release, filter and official link is free to browse, no signup needed. The board updates as new models are announced, so it pays to check back after every launch event.

Curated from official provider announcements; dates reflect the initial public release.