Skip to main content

New AI Models

Every notable new AI model from the major labs, newest first and curated from official announcements. Filter by provider and model type to see who shipped what, and when.

Provider
Type

698 releasesPage 1 of 35

ProviderModelReleasedLink
Google

Gemini Omni 1.1 Flash

Gemini Omni 1.1 Flash is Google DeepMind's generative video model, available in production through the Gemini API in Google AI Studio. Scene extension draws on up to 10 seconds of prior context, extends clips in 10-second increments to 40 seconds total, and accepts up to three seconds of video as reference input for visual consistency. Its 360p draft mode renders up to 60% faster at about a third of the cost of standard 720p, with 1080p and 4K upscaling available.

August 27, 2026
Midjourney

V8.2 Edit Model

Released on August 27, 2026 for open community testing, this is Midjourney's first V8.2 image edit model. It edits images from text instructions, generates new images from up to 4 image references at once (replacing omni-reference), and supports inpainting and outpainting for local changes and canvas expansion. Personalization, moodboards, and srefs are also supported. It is available on midjourney.com and alpha.midjourney.com, or via the --edit command on Discord.

August 27, 2026
Midjourney

V8.2 image edit model

Midjourney's first V8.2 image edit model is now open for everyone to test. It edits images from text instructions, generates new images from up to four references (replacing omni-reference), and supports inpainting and outpainting, along with personalization, moodboards, and srefs. You can use it by attaching an image to the prompt bar, clicking edit in the lightbox, or typing --edit on Discord.

August 27, 2026
Zhipu AI

GLM-5.3-Flash

GLM-5.3-Flash is the first natively multimodal model in Zhipu AI's GLM-5 series, with 320B total and 18B activated parameters, a hybrid sparse-plus-linear attention architecture, and a context window of up to 1M tokens. It outperforms GLM-5.2 across coding and agentic benchmarks at one-tenth the price and approaches Claude Opus 4.8, scoring 57 on the Artificial Analysis Intelligence Index at $0.045 per task. Suited to high-volume coding, agent workflows, and visual document understanding, its weights are available on Hugging Face.

August 27, 2026
Hunyuan

hy4-preview

Hy4 preview is Tencent Hunyuan's new flagship open-source LLM built on a Mixture-of-Experts (MoE) architecture, with 770B total parameters, 49B activated per token, and a 1M-token context window. A native MTP layer enables speculative decoding for faster inference. Tuned for real productivity work in software engineering, office and data analysis, game development, and scientific research, it edged out GLM 5.3 and Kimi K3 in an internal blind evaluation across 203 engineering tasks. Weights ship under Apache 2.0, including an FP8 variant, and serve via vLLM or SGLang.

August 26, 2026
Qwen

Qwen3.8-Flash

Qwen3.8-Flash is a multimodal LLM from Alibaba's Qwen team that pairs strong understanding and generation capability with fast response times. It natively supports a million-token context window, handling long documents, code repositories, and complex conversations in a single pass. The model targets coding assistance, agent workflows, and vision-language tasks, works with OpenAI and Anthropic API protocols in tools such as Claude Code and Codex, and offers competitive inference costs for high-concurrency applications.

August 25, 2026
OpenAI

GPT‑5.6

The GPT‑5.6 model family is OpenAI's latest flagship series, spanning Sol, Terra, and Luna, and is now available in Kiro, a software development agent for AI-native coding. It delivers stronger performance per dollar and on-demand capability for complex tasks, applied to long-running development work grounded in requirements, codebases, and team standards. In testing by OpenAI and AWS, GPT‑5.6 Terra completed Terminal-Bench 2.1 tasks in Kiro at roughly 82% cost reduction.

August 24, 2026
DeepSeek

DeepSeek-V4-Flash-Vision-Exp

DeepSeek-V4-Flash-Vision-Exp is an experimental multimodal model released on the DeepSeek API Platform on August 21, 2026. It matches DeepSeek-V4-Flash on text tasks, including agents, reasoning, and world knowledge, while showing a major step up on multimodal agent benchmarks that brings it close to Opus-4.8. Images are billed at up to 384 tokens each at V4-Flash pricing and can be passed via base64, external URLs, or the Files API, making it a fit for agent workflows that pair visual understanding with tool calls.

August 21, 2026
ElevenLabs

Eleven v3 Conversational

Eleven v3 Conversational is ElevenLabs' most expressive realtime speech synthesis model, producing natural, emotionally rich speech at low latency of roughly 280ms. It supports 70+ languages and provides audio tags for fine-grained control over delivery. Typical uses include realtime support voice agents, AI assistants, and interactive characters, accessed via the new Text to Dialogue WebSocket.

August 20, 2026
Qwen

Wan3.0-Video-Prime

Wan3.0-Video-Prime is the accelerated version of Alibaba Cloud's Wan video generation model, delivering significantly faster generation speed while maintaining high-quality output. It accepts text, image, video, and audio inputs in a single model, covering text-to-video, image-to-video (first frame and first-last frame), and reference-based video generation. It outputs video up to 30 seconds long at 480P, 720P, or 1080P with a 30fps frame rate.

August 20, 2026
OpenAI

GPT-5.6 Luna

GPT-5.6 Luna is the fastest, most cost-efficient tier in OpenAI's GPT-5.6 lineup, alongside the flagship Sol and the balanced Terra. It approaches GPT-5.5's overall capability at less than half the estimated cost and outscores Claude Opus 4.8 on the Artificial Analysis Coding Agent Index in about a third of the time. Priced at $1 per million input tokens and $6 per million output tokens, with an 80% price cut announced on July 30, it targets cost-sensitive workloads.

August 19, 2026
Zhipu AI

GLM-5.3

GLM-5.3 is Zhipu AI's flagship coding model, sharing the GLM-5.2 base model, with all gains from scaled post-training. Z.ai calls it the most capable open-weights model for coding, reaching open-source SOTA on Terminal Bench 3.0 and Agents' Last Exam and improving 50% over GLM-5.2 on the in-house Z.ai Code Bench. Its emergent cyber capability posts a state-of-the-art 84.5% on CyberGym for vulnerability discovery. Weights follow two weeks after launch, and it fits agentic coding and long-horizon work in agents like ZCode and Claude Code.

August 18, 2026
Qwen

Qwen3.8-27B

Qwen3.8-27B is a 27B dense vision-language model in Alibaba's Qwen3.8 open-model family, with native image and video understanding. It strengthens coding and office-productivity capabilities in both text and vision modalities over Qwen3.6-27B, posting large gains on benchmarks such as SWE-bench Pro and OSWorld. The model offers a native 256K context window, extensible to 1M tokens, and is suited to coding, professional work, and long-horizon agentic tasks.

August 17, 2026
DeepSeek

DeepSeek-V4-Pro-0813

DeepSeek-V4-Pro-0813 is the general-availability release of DeepSeek's V4 Pro text model, bringing major Agent upgrades with strong production gains. It supports three reasoning effort levels (low, high, max), native OpenAI Responses API compatibility, and one-click setup for Codex. Agentic and coding benchmarks such as Terminal Bench and DeepSWE show clear gains over the V4-Pro preview, and API pricing now uses peak and off-peak rates, with off-peak at half the peak price.

August 13, 2026
Google

Gemini 3.7 Flash

Gemini 3.7 Flash is Google's workhorse model for coding and agents, released on August 13, 2026, three weeks after Gemini 3.6 Flash. It posts clear gains over its predecessor on benchmarks such as FrontierCode 1.1 Main (43.6% vs 34.4%) and DeepSWE v1.1 (65.3% vs 49.0%), and can build web UI that matches reference screenshots or design systems. Introductory pricing runs $0.75 per million input tokens and $3.75 per million output tokens, and it is available through the Gemini API, Google Antigravity, and AI Studio.

August 13, 2026
Google

sign-language-to-text (SL2T)

SL2T is a massively multilingual sign-language-to-text translation model from Google DeepMind and Android that converts signing into streaming text, so users can sign anywhere they would normally type. Trained on more than 100,000 hours of data spanning over 50 sign languages, it launches with American Sign Language (ASL) to English and scores 70 BLEURT zero-shot on the FLEURS-ASL benchmark. It is available at no cost in Gboard and Live Transcribe on Pixel 11, with more devices and languages to follow.

August 12, 2026
Qwen

Qwen3.8-2.4T-A95B

Qwen3.8-2.4T-A95B is the open-source release of Qwen's latest flagship series, launched in August 2026. It uses a sparse MoE architecture with 2.4 trillion total parameters and about 95 billion active parameters per step, combining hybrid attention with a 1 million token context window. Key benchmark results include GPQA Diamond 92.6, PaperBench 93.0, OSWorld 86.1, and BabyVision 82.0, with a fourth-place global ranking on CodeArena.

August 12, 2026
Microsoft

MAI-Code-1.1-Flash

MAI-Code-1.1-Flash is a small, efficient coding model from Microsoft, now in production in GitHub Copilot. Compared with the 1.0 release announced at Microsoft Build in June, it produces higher quality code at 25% greater token efficiency and one quarter of the cost, with a 22% gain on Terminal-Bench 2.1 and 15% on .NET tasks. Tokens also stream 25% faster, making it a practical fit for everyday coding assistance across IDE and CLI workflows.

August 11, 2026
NVIDIA

Nemotron 3.5 Lightning (30B A3B)

NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model with 30B total parameters (3B active, A3B) built for long-running agentic AI workloads. NVIDIA reports up to 4x faster output speed and 30% faster agentic task completion than other models in its class, with frontier-level accuracy on PinchBench. Offered with NVFP4 quantization, it runs locally on RTX PCs, DGX Spark, DGX Station and Jetson, and can be post-trained with NeMo for high-volume tasks such as code review and tool use.

August 11, 2026
Meta

Muse Glimmer-30B

Muse Glimmer-30B is a 30-billion-parameter agentic model from Meta Superintelligence Labs, open sourced under Apache 2.0 and optimized for always-on local agent workflows on a single consumer GPU. It handles precise tool calling, long-horizon multi-step reasoning, and interleaved text and image input across more than 100 languages. Meta reports strong results for its size class on agentic benchmarks such as SWE-Bench and tau-Bench versus Gemma4-31B and Qwen3.6-27B, with quantization bringing the model under 20 GB for local coding, function calling, and LLM-as-a-judge evaluation.

August 10, 2026

How to track new AI models

Four ways to read the board. Each one turns the stream of new AI models into a clear answer. Pick the view that fits your question.

01

Read the newest launches first

The board lists new AI models newest first, so the top row is always the latest launch from any lab. Every row links to the official announcement, so one click gets you to the source. Dates reflect the initial public release.

02

Follow one provider

Pick a lab in the Provider filter. OpenAI, Anthropic, Google, Meta, DeepSeek, Qwen and 20+ more each get their own view. Use it to read a lab's release rhythm: how often it ships, and which model types it bets on.

03

Compare model types

Filter by type: text/LLM, multimodal, image, video, or audio. The mix shows where the industry is pushing: text models still lead the count, while image and video generators ship in waves. Watch the balance shift month by month.

04

Catch up on what you missed

Page back through the timeline. Some months see dozens of launches, and the archive reaches back years. The board is a full release history, not just a feed of today's new AI models. Reconstruct any lab's year in a few scrolls.

New AI Models FAQ

What the board tracks, where the data comes from, and how to read it.

The board sorts releases newest first, so the top row is always the most recent launch from any provider. Recent months brought new AI models from labs like OpenAI, Anthropic, Google, Qwen and Zhipu AI. Every row links to the official announcement with its exact release date.

Twenty-nine providers and counting: OpenAI, Anthropic, Google, Meta, Microsoft, Amazon, NVIDIA, Mistral AI, DeepSeek, Qwen, Moonshot AI, Zhipu AI, ByteDance, xAI and more. Western and Chinese labs share one board, so a launch abroad still shows up the same day.

Five kinds: text/LLM, multimodal, image generation, video generation, and audio/speech. Text models carry the largest share of launches, but image and video generators ship in fast waves. The Type filter isolates any one stream.

Constantly. Major labs now ship new AI models on a weekly cadence, and busy weeks bring more than one frontier launch. Since 2024 the pace has kept climbing, and the board adds each release as its official announcement goes live.

Each row uses the date of the initial public release: the day the lab announced the model or opened it to users. Betas, rumors and paper preprints do not count. Every entry links to the official announcement, so you can verify the date yourself.

Yes. The board doubles as an archive reaching back to 2017. Page back to see how the release cadence exploded, from a handful of new AI models a year to hundreds. It is a searchable release history, not just a live feed.

Yes. The Provider filter narrows the board to one lab or any set of labs. Combine it with the Type filter to answer narrow questions, like every video model a single lab has shipped.

Yes, completely. Every release, filter and official link is free to browse, no signup needed. The board updates as new models are announced, so it pays to check back after every launch event.

Curated from official provider announcements; dates reflect the initial public release.