Skip to main content

New AI Models

Every notable new AI model from the major labs, newest first and curated from official announcements. Filter by provider and model type to see who shipped what, and when.

Provider
Type

698 releasesPage 5 of 35

ProviderModelReleasedLink
Qwen

Qwen3.7-Max-2026-06-08

Qwen3.7-Max-2026-06-08 is a dated snapshot of the largest and most capable Max model in the Qwen3.7 series, suited to production workloads that need stable, reproducible behavior. Compared with the May 20 snapshot, it adds visual understanding to perceive real-world scenes, accepts text, image, and video input, and offers a 1M context window with multimodal interaction and hybrid agentic capabilities. It targets agent workloads such as coding, office productivity, and long-horizon autonomous execution.

June 10, 2026
Anthropic

Claude Fable 5

Claude Fable 5 is Anthropic's Mythos-class model, sharing its base with Claude Mythos 5 and made safe for general use. Anthropic reports state-of-the-art results on nearly all tested benchmarks, with its lead widening as tasks get longer and more complex. It excels at software engineering, knowledge work, vision, and scientific research, and is more token-efficient than earlier Claude models. It is available via the Claude API as claude-fable-5, priced at $10 per million input tokens and $50 per million output tokens.

June 9, 2026
Anthropic

Claude Mythos 5

Claude Mythos 5 shares its underlying model with Claude Fable 5, with safeguards lifted in select areas. Deployed through Project Glasswing in collaboration with the US government, it serves a small group of cyberdefenders and infrastructure providers, and Anthropic calls its cybersecurity capabilities the strongest of any model in the world. In internal testing it accelerated parts of drug design by about 10 times, and it is priced at $10 per million input and $50 per million output tokens, less than half the price of Mythos Preview.

June 9, 2026
Google

Gemma 4 12B

Gemma 4 12B is Google DeepMind's open multimodal model that brings agentic AI to consumer laptops. Built on an encoder-free architecture, it natively handles text, image, and audio inputs and runs locally on machines with 16GB of VRAM or unified memory. Its performance approaches the larger 26B MoE sibling at less than half the memory footprint, with a 256K context window, configurable thinking modes, and native function calling. Weights are available under Apache 2.0.

June 9, 2026
NVIDIA

Nemotron 3 Ultra (550B A55B)

NVIDIA Nemotron 3 Ultra is an open Mixture-of-Experts model built for reasoning and orchestration in long-running agents, with 550B total parameters and 55B active per token in a hybrid Mamba-Transformer design that supports a 1M-token context. NVIDIA reports up to 5x the throughput of comparable open models and up to 30% lower cost per completed task on SWE-bench and Terminal-Bench 2.0. Released under the OpenMDW-1.1 license, a single NVFP4 checkpoint runs across Ampere, Hopper, and Blackwell GPUs.

June 4, 2026
Microsoft

MAI-Code-1-Flash

MAI-Code-1-Flash is Microsoft's lightweight agentic coding model for fast, everyday developer workflows, trained from the ground up on clean, traceable, enterprise-grade data with no distillation from third-party models. In Microsoft's production-harness evaluations it outperforms Claude Haiku 4.5 on SWE-Bench Verified, SWE-Bench Pro, SWE-Bench Multilingual, and Terminal Bench 2, leading 51.2% to 35.2% on SWE-Bench Pro, and it solves harder problems with up to 60% fewer tokens. It adapts its reasoning effort to task complexity and is rolling out to GitHub Copilot individual users in VS Code.

June 2, 2026
Microsoft

MAI-Thinking-1

MAI-Thinking-1 is Microsoft AI's reasoning model, built as a sparse Mixture of Experts with 35B active and roughly 1T total parameters. It matches Claude Opus 4.6 on SWE-Bench Pro, scores 97.0% on AIME 2025, and was preferred over Claude Sonnet 4.6 in blind human side-by-side evaluations spanning 1,276 tasks. Trained without distilling from third-party models, it offers a 256k token context window, function calling, and Chat Completions API compatibility. It is available in public preview on Microsoft Foundry.

June 2, 2026
MiniMax

MiniMax M3

MiniMax M3 is an open-weight model from MiniMax built for coding and agentic work, with native image and video input and a context window of up to 1M tokens. It uses MSA (MiniMax Sparse Attention), which cuts per-token compute at 1M context to about 1/20 of the previous generation. It scores 59.0% on SWE-Bench Pro and 66.0% on Terminal-Bench 2.1, with clear coding gains over M2, and is available through MiniMax Code, the Token Plan, and the API.

June 1, 2026
Qwen

Qwen3.7-Plus

Qwen3.7-Plus is the cost-effective Plus model in Alibaba's Qwen3.7 family, built to balance capability and cost. It upgrades vision-language abilities on a strong text foundation, offers a 1M context window, and keeps full agentic abilities in coding, tool use, and productivity workflows. It can read screens and operate GUIs, generate code from visual references, and navigate mobile apps end to end, making it a fit for AI coding, agent development, chatbots, and document processing.

June 1, 2026
Qwen

Qwen3.7-Plus-2026-05-26

Qwen3.7-Plus-2026-05-26 is a dated snapshot of qwen3.7-plus, the cost-effective Plus model in Alibaba's Qwen3.7 series. It upgrades vision-language capability on top of strong text ability, accepts image, text, and video input, and retains complete agentic capabilities for coding, tool use, and productivity workflows, from reading screens and operating GUIs to end-to-end mobile app navigation. With a 1M-token context window, the pinned snapshot fits production workloads that need stable behavior.

June 1, 2026
Anthropic

Claude Opus 4.8

Claude Opus 4.8 is the newest version of Anthropic's flagship Opus line, building on Opus 4.7 with improvements across coding, agentic, and reasoning benchmarks. It launches with effort control, dynamic workflows in Claude Code that run hundreds of parallel subagents, and a fast mode at 2.5x speed. Anthropic reports it is around four times less likely than its predecessor to let flaws in code it has written pass unremarked. Pricing is unchanged at $5 per million input tokens and $25 per million output tokens.

May 28, 2026
Grok

grok-imagine-video-1.5

grok-imagine-video-1.5 is xAI's video generation model in the Grok Imagine family, now generally available in the xAI API after its preview. It improves on its predecessor in audio and physics: sound effects, ambience, and dialogue are generated in one pass and synced to the action, while motion stays consistent across longer clips. The API supports 1 to 15 second outputs up to 1080p, and the model is also available through grok.com/imagine and the Grok iOS and Android apps.

May 27, 2026
Grok

grok-imagine-video-1.5-preview

grok-imagine-video-1.5-preview is the preview release of Grok Imagine Video 1.5, xAI's image-to-video model with improved motion, physics, and audio over the previous generation. Sound effects, ambience, and dialogue are generated in the same pass, with clearer and better-synced speech. Clips run up to 15 seconds at up to 1080p, and the Fast variant produces 6-second 720p videos in about 25 seconds. The model is available through the xAI API and Grok Imagine on grok.com and mobile apps.

May 27, 2026
Hunyuan

hunyuan-mt2-1.8b-chat

Hy-MT2-1.8B is the lightweight member of Tencent's Hy-MT2 family of fast-thinking multilingual translation models. It supports translation across 33 languages with an 8192-token context and follows instructions such as terminology constraints and style control, surpassing mainstream commercial APIs like Microsoft and Doubao overall in the official evaluation. With AngelSlim 1.25-bit quantization it shrinks to just 440 MB for on-device deployment, and it is released under the Apache 2.0 license.

May 21, 2026
Hunyuan

hunyuan-mt2-7b-chat

Hy-MT2-7B is Tencent Hunyuan's open source multilingual translation model, the 7B member of a fast-thinking translation family built for complex real-world scenarios. It translates across 33 languages with an 8192-token context window and, in Hunyuan's evaluations covering general, business, and domain-specific translation, outperforms open source models such as DeepSeek-V4-Pro and Kimi K2.6. It supports instruction-following translation features like terminology constraints and style control, and is released under Apache 2.0 on HuggingFace and ModelScope.

May 21, 2026
Hunyuan

hunyuan-translation-pro-chat

hunyuan-translation-pro-chat is Tencent Hunyuan's flagship commercial translation API, the top tier of the Hy-MT2 family built on the 30B-A3B (MoE) model for professional domains that demand high translation quality. The fast-thinking model supports translation across 33 languages with an 8K context window and follows instructions for terminology constraints, style control, and structured data. Official evaluations report results ahead of open-source models such as DeepSeek-V4-Pro and Kimi K2.6 in fast-thinking mode.

May 21, 2026
Qwen

Qwen3.7-Max

Qwen3.7-Max is the largest and most capable model in Alibaba's Qwen3.7 series, built for the agent era and currently offered as a text-only model on Alibaba Cloud Model Studio. Its core strength lies in the breadth and depth of its agentic capabilities, covering coding, office productivity, and long-horizon autonomous execution. It supports a 1M token context window and thinking mode, with pricing at 12 CNY per million input tokens and 36 CNY per million output tokens in the Beijing region.

May 21, 2026
Qwen

Qwen3.7-Max-2026-05-20

Qwen3.7-Max-2026-05-20 is the May 20, 2026 snapshot of Qwen3.7-Max, the largest and most capable model in Alibaba's Qwen3.7 series, currently open for text-only capabilities with support for text generation and deep thinking. Qwen3.7 is a new flagship generation built for the agentic era, and its core strength lies in the breadth and depth of its agent capabilities, spanning coding, office and productivity work, and long-horizon autonomous execution.

May 21, 2026
Qwen

Qwen3.5-LiveTranslate-Flash-Realtime

Qwen3.5-LiveTranslate-Flash-Realtime is Alibaba's multilingual model for real-time audio and video translation, the realtime version of Qwen3.5-LiveTranslate-Flash built on the Qwen3.5-Omni foundation. It translates across 60 languages, 29 of them with both audio and text output, at simultaneous-translation latency as low as 2.8 seconds, with quality close to offline results. By combining audio and image input, it uses visual context to stay accurate in noisy or ambiguous settings, and works with live video streams and local video files.

May 19, 2026
Qwen

Qwen3.5-LiveTranslate-Flash-Realtime-2026-05-19

Qwen3.5-LiveTranslate-Flash-Realtime-2026-05-19 is a vision-enhanced real-time speech and video interpretation model from Alibaba's Qwen team, the realtime edition of Qwen3.5-LiveTranslate-Flash built on the Qwen3.5-Omni foundation. It translates between 60 languages, 29 of them with speech output, at simultaneous interpretation latency as low as 2.8 seconds, and uses visual cues such as lip movement and on-screen text to improve accuracy in noisy settings. Suited to live video streams, cross-language talks, and video translation, this dated snapshot matches the current stable release for production use.

May 19, 2026

How to track new AI models

Four ways to read the board. Each one turns the stream of new AI models into a clear answer. Pick the view that fits your question.

01

Read the newest launches first

The board lists new AI models newest first, so the top row is always the latest launch from any lab. Every row links to the official announcement, so one click gets you to the source. Dates reflect the initial public release.

02

Follow one provider

Pick a lab in the Provider filter. OpenAI, Anthropic, Google, Meta, DeepSeek, Qwen and 20+ more each get their own view. Use it to read a lab's release rhythm: how often it ships, and which model types it bets on.

03

Compare model types

Filter by type: text/LLM, multimodal, image, video, or audio. The mix shows where the industry is pushing: text models still lead the count, while image and video generators ship in waves. Watch the balance shift month by month.

04

Catch up on what you missed

Page back through the timeline. Some months see dozens of launches, and the archive reaches back years. The board is a full release history, not just a feed of today's new AI models. Reconstruct any lab's year in a few scrolls.

New AI Models FAQ

What the board tracks, where the data comes from, and how to read it.

The board sorts releases newest first, so the top row is always the most recent launch from any provider. Recent months brought new AI models from labs like OpenAI, Anthropic, Google, Qwen and Zhipu AI. Every row links to the official announcement with its exact release date.

Twenty-nine providers and counting: OpenAI, Anthropic, Google, Meta, Microsoft, Amazon, NVIDIA, Mistral AI, DeepSeek, Qwen, Moonshot AI, Zhipu AI, ByteDance, xAI and more. Western and Chinese labs share one board, so a launch abroad still shows up the same day.

Five kinds: text/LLM, multimodal, image generation, video generation, and audio/speech. Text models carry the largest share of launches, but image and video generators ship in fast waves. The Type filter isolates any one stream.

Constantly. Major labs now ship new AI models on a weekly cadence, and busy weeks bring more than one frontier launch. Since 2024 the pace has kept climbing, and the board adds each release as its official announcement goes live.

Each row uses the date of the initial public release: the day the lab announced the model or opened it to users. Betas, rumors and paper preprints do not count. Every entry links to the official announcement, so you can verify the date yourself.

Yes. The board doubles as an archive reaching back to 2017. Page back to see how the release cadence exploded, from a handful of new AI models a year to hundreds. It is a searchable release history, not just a live feed.

Yes. The Provider filter narrows the board to one lab or any set of labs. Combine it with the Type filter to answer narrow questions, like every video model a single lab has shipped.

Yes, completely. Every release, filter and official link is free to browse, no signup needed. The board updates as new models are announced, so it pays to check back after every launch event.

Curated from official provider announcements; dates reflect the initial public release.