Skip to main content

New AI Models

Every notable new AI model from the major labs, newest first and curated from official announcements. Filter by provider and model type to see who shipped what, and when.

Provider
Type

698 releasesPage 2 of 35

ProviderModelReleasedLink
OpenAI

GPT-5.6-Cyber

GPT-5.6-Cyber is OpenAI's purpose-trained cybersecurity model, built on GPT-5.6 Sol to strengthen specialized work such as zero-day vulnerability discovery and exploit-chain development while reducing refusals on dual-use security tasks. It completes 95.0% of requests on OpenAI's internal Advanced Cybersecurity Completion Rate evaluation, up from 1.5% for GPT-5.6 Sol. Access is limited to approved defenders through the Daybreak Red tier, with identity verification and authorized-use restrictions.

August 10, 2026
Qwen

Qwen Audio Realtime3.0

qwen-audio-3.0-realtime-plus is the standard edition of Alibaba Cloud Model Studio's next-generation real-time full-duplex speech model, ranked first overall in the Artificial Analysis Speech-to-Speech benchmark. It balances model intelligence with natural duplex turn-taking while keeping speech reasoning intact, and keeps end-to-end latency low through parallel inference and omnidirectional streaming. It prioritizes high-quality responses, making it well suited to voice assistants, intelligent customer service, and AI companion scenarios.

August 10, 2026
Qwen

Qwen Audio Realtime3.0 Fast

Qwen Audio Realtime3.0 Fast is the speed-focused tier of Qwen's real-time duplex speech model on Alibaba Cloud Model Studio. The series ranked first overall in Artificial Analysis's Speech-to-Speech benchmark and keeps full reasoning ability over natural duplex turn-taking, using parallel inference and omnidirectional streaming to hold end-to-end latency down. Compared with the standard tier, this version prioritizes top response speed for latency-sensitive voice interaction.

August 10, 2026
Grok

grok-imagine-image-2.0

Imagine Image 2.0 is xAI's image generation model, generally available as the Quality Mode on grok.com/imagine, in the iOS and Android apps, and through the API. It follows detailed instructions with sharp typography and consistent layout, and supports precise editing with the magic wand, background removal, multi-ref editing with up to 5 input images, and smart resize for any aspect ratio. Templates turn common workflows such as product shots, headshots, and icons into ready-made starting points.

August 8, 2026
Grok

grok-4.6

Grok 4.6 is xAI's latest text model, built on Grok 4.5 with a focus on long-running agents and ambitious interactive and visual work. It matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index and posts frontier-level results across several agentic coding and knowledge work benchmarks. It is especially strong at turning a broad product idea into a working first version, and is available via the API starting at $2 per million input tokens and $6 per million output tokens.

August 6, 2026
Qwen

Wan3.0

Wan3.0 is Alibaba's all-in-one video generation model, unifying text-to-video, image-to-video (first and last frame control), reference-based generation, and video editing in one model. It accepts multimodal inputs such as text, images, video, audio, files, and webpages, and generates clips up to 30 seconds long with native speech, BGM, and sound effects. Output covers 480P, 720P, and 1080P at 30fps, with production-grade character consistency.

August 6, 2026
ByteDance

SeedRealtime

SeedRealtime is ByteDance Seed's native audio-visual full-duplex LLM that unifies audio, video, and text in a single architecture for real-time interaction over continuous multimodal streams. Its end-to-end design handles perception, understanding, decision-making, and response generation in one model, replacing multi-stage cascaded systems. In human evaluations it cut conversational pacing issues by half versus cascaded models, with fewer interruptions, lower latency, and fewer false triggers. Now fully rolled out, it suits noisy multi-speaker settings, device-operation guidance, and study support.

August 5, 2026
Meta

Muse Spark 1.2

Muse Spark 1.2 is Meta's coding-focused LLM, an update to Muse Spark 1.1 that improves code generation, complex debugging, and codebase understanding while keeping its strength in general agent tasks. It was co-trained with the Muse Code terminal coding agent, and it targets long-horizon work such as whole-repository generation and large end-to-end projects, progressively improving GPU kernel performance over 1,000+ tool calls in Meta's optimization case study. It is available today through Muse Code and the Meta Model API with expanded global access.

August 5, 2026
OpenAI

GPT-Daybreak

OpenAI's Daybreak family of cybersecurity models is built for defensive workflows, helping security teams find, validate, and fix vulnerabilities across large codebases. Daybreak Red, the security-specialized model in the family, targets authorized vulnerability research, exploit validation, penetration testing, and red teaming. OpenAI researchers used it to identify two previously unknown V8 flaws that could be chained to escape the heap sandbox, and verified defenders can request elevated access through Daybreak Access.

August 5, 2026
Mistral AI

Shieldstral 1.0 (3B)

Shieldstral 1.0 (3B) is an open-weights multimodal safety classifier from Mistral AI. It frames content moderation as plain-language question answering: policies are supplied as text queries at inference time, and a single forward pass returns a calibrated safety score, letting one checkpoint adapt to new policies without retraining. It covers text, image, and text+image content, unifying prompt classification, response moderation, refusal detection, and toxicity detection, and matches open guard models up to 7x its size on text safety. Weights ship under Apache 2.0 and run on a single 16GB GPU.

August 4, 2026
Qwen

Qwen-Image-3.0

Qwen-Image-3.0 is the Standard-tier image generation model in Alibaba's Qwen series, aimed at everyday creative work. It accepts inputs of up to 4.5k tokens, handles complex text-and-image instructions in a single pass, and renders text as small as 10px clearly across 12 languages and 20+ fonts. Posters, web pages, and interface designs can be generated in batches at lower cost, making it a practical choice for sustained content production.

August 4, 2026
Qwen

Qwen3.8-Max

Qwen3.8-Max is Alibaba Qwen's flagship MoE model with 2.4 trillion parameters. It lifts coding and office productivity, autonomously coding for days to deliver complete projects, and handles hundreds of professional tasks in law, finance, and design, delivering production-grade results end to end in a single conversation. Native vision understanding spans planning, execution, and verification, with deep semantic parsing of long documents and videos, and the model plans and iterates on its own across long-horizon tasks.

August 2, 2026
ByteDance

Seedance 2.5

Seedance 2.5 is ByteDance Seed's next-generation audio-video joint generation model, producing videos up to 30 seconds in a single pass with the option to extend twice for longer narratives. It reads reference videos with greater precision, capturing intent, framing, and cinematic language, and handles a wide range of audio and visual editing requests. Capabilities such as white-model control, green-screen editing, professional camera movement, and performance blocking support complex video production and professional creative workflows.

July 31, 2026
DeepSeek

DeepSeek-V4-Flash-0731

DeepSeek-V4-Flash-0731 is the official public beta of the DeepSeek-V4-Flash API, released on July 31, 2026 under the model name deepseek-v4-flash. It delivers significantly stronger agent performance, scoring well above V4-Pro-Preview on benchmarks such as Terminal Bench 2.1 (82.7), DeepSWE, and Cybergym. The model keeps the same architecture and size as V4-Flash-Preview, with only re-post-training applied, and it natively supports the Responses API format with specific adaptation for Codex.

July 31, 2026
Hunyuan

HY-ASR-3-NoStream

Hunyuan HY-ASR-3-NoStream is the non-streaming ASR model from Tencent Hunyuan, the file transcription form of Hy ASR 3.0 preview. Built on the Hy3 foundation model, it combines high-precision recognition with deep semantic understanding and supports Mandarin, English and 20 Chinese dialects. It uses context to correct homophones and resolve ambiguity, stays robust in noisy and whispered speech, and suits batch transcription of meetings, interviews and media recordings.

July 31, 2026
Hunyuan

HY-ASR-3-Stream

HY-ASR-3-Stream is Tencent Hunyuan's streaming ASR model, released in preview as Hy-ASR-3.0-preview and built on the Hy3 LLM base. It transcribes Mandarin, English, and 20 Chinese dialects in a single engine and returns results over WebSocket in real time, producing text as the speaker talks. Deep semantic understanding improves general recognition, context awareness, and dialect coverage, suiting low-latency scenarios such as instant messaging transcription, live call captions, and voice assistant command parsing.

July 31, 2026
Google

Gemini Robotics ER 2

Gemini Robotics ER 2 is Google DeepMind's most capable embodied reasoning model, a vision language model that acts as the robot's high-level brain. It observes the environment, plans multi-step tasks that last several minutes, and coordinates with VLA models to execute them, with support for multi-robot collaboration. Google describes it as its safest robotics model to date in safety constraint following and human proximity benchmarks. It is available in Google AI Studio and in private preview on the Gemini Enterprise Agent Platform.

July 30, 2026
HaiLuo

MiniMax H3

MiniMax H3 is a general-purpose multimodal generation model that understands unified context across text, images, video, and audio, generating video with native stereo sound at up to 15 seconds and native 2K resolution. Built on components such as H3-VAE and the H3-Omni Transformer, it handles instruction following, accurate text and brand rendering, and V2V motion transfer. Per-second pricing at 2K is less than a third of mainstream models, and open weights are planned. It targets advertising, e-commerce, product design, UI/UX, and gaming.

July 30, 2026
Qwen

Qwen-Audio-3.0-ASR-Flash

Qwen-Audio-3.0-ASR-Flash is a non-realtime speech recognition model from Alibaba's Qwen team, available on the Bailian platform for transcribing audio files over HTTP, with up to 5 minutes and 2GB per request. It supports 30 languages including Chinese, English, Japanese and Korean, covers the seven major Chinese dialect groups, and offers hotwords and context prompting to improve accuracy on specific terms. Output adds punctuation prediction and text normalization, with optimized classical poetry recognition for use cases such as call-center recordings, podcasts and audiobooks.

July 30, 2026
Qwen

qwen-audio-3.0-asr-flash-filetrans

qwen-audio-3.0-asr-flash-filetrans is a non-real-time ASR model from Alibaba Cloud Bailian built for transcribing complete audio and video files, with asynchronous processing for recordings up to 12 hours and 2GB each. It integrates Context augmentation and strengthens recognition of industry-specific hotwords, covering 30 languages including Chinese, English, Japanese, and Korean. Typical uses include short-video subtitles, customer service quality inspection, and financial dual recording, billed at 0.00022 CNY per second of audio.

July 30, 2026

How to track new AI models

Four ways to read the board. Each one turns the stream of new AI models into a clear answer. Pick the view that fits your question.

01

Read the newest launches first

The board lists new AI models newest first, so the top row is always the latest launch from any lab. Every row links to the official announcement, so one click gets you to the source. Dates reflect the initial public release.

02

Follow one provider

Pick a lab in the Provider filter. OpenAI, Anthropic, Google, Meta, DeepSeek, Qwen and 20+ more each get their own view. Use it to read a lab's release rhythm: how often it ships, and which model types it bets on.

03

Compare model types

Filter by type: text/LLM, multimodal, image, video, or audio. The mix shows where the industry is pushing: text models still lead the count, while image and video generators ship in waves. Watch the balance shift month by month.

04

Catch up on what you missed

Page back through the timeline. Some months see dozens of launches, and the archive reaches back years. The board is a full release history, not just a feed of today's new AI models. Reconstruct any lab's year in a few scrolls.

New AI Models FAQ

What the board tracks, where the data comes from, and how to read it.

The board sorts releases newest first, so the top row is always the most recent launch from any provider. Recent months brought new AI models from labs like OpenAI, Anthropic, Google, Qwen and Zhipu AI. Every row links to the official announcement with its exact release date.

Twenty-nine providers and counting: OpenAI, Anthropic, Google, Meta, Microsoft, Amazon, NVIDIA, Mistral AI, DeepSeek, Qwen, Moonshot AI, Zhipu AI, ByteDance, xAI and more. Western and Chinese labs share one board, so a launch abroad still shows up the same day.

Five kinds: text/LLM, multimodal, image generation, video generation, and audio/speech. Text models carry the largest share of launches, but image and video generators ship in fast waves. The Type filter isolates any one stream.

Constantly. Major labs now ship new AI models on a weekly cadence, and busy weeks bring more than one frontier launch. Since 2024 the pace has kept climbing, and the board adds each release as its official announcement goes live.

Each row uses the date of the initial public release: the day the lab announced the model or opened it to users. Betas, rumors and paper preprints do not count. Every entry links to the official announcement, so you can verify the date yourself.

Yes. The board doubles as an archive reaching back to 2017. Page back to see how the release cadence exploded, from a handful of new AI models a year to hundreds. It is a searchable release history, not just a live feed.

Yes. The Provider filter narrows the board to one lab or any set of labs. Combine it with the Type filter to answer narrow questions, like every video model a single lab has shipped.

Yes, completely. Every release, filter and official link is free to browse, no signup needed. The board updates as new models are announced, so it pays to check back after every launch event.

Curated from official provider announcements; dates reflect the initial public release.