GPT-5.6 Sol
OpenAI · Reasoning LLM
OpenAI’s highest-capability GPT-5.6 model for difficult professional, analytical, and coding work.
- Family
- GPT-5.6
- Context
- 1,050,000 tokens
- API
- Yes
GPT-5.6 Terra
OpenAI · Reasoning LLM
A balanced GPT-5.6 model designed to deliver strong quality at a lower cost than Sol.
- Family
- GPT-5.6
- Context
- 1,050,000 tokens
- API
- Yes
GPT-5.6 Luna
OpenAI · General LLM
The fast, lower-cost GPT-5.6 option for high-volume applications and responsive user experiences.
- Family
- GPT-5.6
- Context
- 1,050,000 tokens
- API
- Yes
GPT-5
OpenAI · Reasoning LLM
Original GPT-5 flagship for reasoning, coding, and multimodal work.
- Family
- GPT-5
- Context
- 400,000 tokens
- API
- Yes
GPT-5 mini
OpenAI · Reasoning LLM
Smaller GPT-5 model for cost-sensitive production workloads.
- Family
- GPT-5
- Context
- 400,000 tokens
- API
- Yes
GPT-5 nano
OpenAI · Reasoning LLM
Very small GPT-5 model for classification, extraction, and routing.
- Family
- GPT-5
- Context
- 400,000 tokens
- API
- Yes
text-embedding-3-large
OpenAI · Embedding
OpenAI’s higher-quality embedding model for semantic search, recommendations, and retrieval.
- Family
- Embedding 3
- Context
- 8,191 tokens
- API
- Yes
text-embedding-3-small
OpenAI · Embedding
A compact, inexpensive embedding model for search and retrieval at scale.
- Family
- Embedding 3
- Context
- 8,191 tokens
- API
- Yes
Claude Fable 5
Anthropic · Reasoning LLM
Anthropic’s premium everyday work model, tuned for high-quality documents, spreadsheets, analysis, and agent workflows.
- Family
- Claude 5
- Context
- 1,000,000 tokens
- API
- Yes
Claude Opus 4.8
Anthropic · Reasoning LLM
Anthropic’s strongest broadly available model for advanced coding, agents, computer use, and complex professional work.
- Family
- Claude 4
- Context
- 1,000,000 tokens
- API
- Yes
Claude Sonnet 5
Anthropic · Reasoning LLM
A high-performance Claude model that balances frontier coding and agent capability with production-friendly cost.
- Family
- Claude 5
- Context
- 1,000,000 tokens
- API
- Yes
Claude Haiku 4.5
Anthropic · General LLM
Anthropic’s fast, lower-cost Claude model for responsive assistants and high-volume processing.
- Family
- Claude 4
- Context
- 200,000 tokens
- API
- Yes
Claude Mythos 5
Anthropic · Reasoning LLM
An invitation-only Claude 5 model for specialized, high-end professional and agentic workloads.
- Family
- Claude 5
- Context
- 1,000,000 tokens
- API
- Limited
Gemini 3.5 Flash
Google · Multimodal LLM
Google’s current fast flagship for agentic coding, multimodal understanding, and sustained intelligent workflows.
- Family
- Gemini 3.5
- Context
- 1,048,576 tokens
- API
- Yes
Gemini 3.1 Pro Preview
Google · Multimodal LLM
A premium Gemini model for complex reasoning, advanced software engineering, and sophisticated agent workflows.
- Family
- Gemini 3.1
- Context
- 1,048,576 tokens
- API
- Yes
Gemini 3.1 Flash-Lite
Google · Multimodal LLM
A cost-efficient Gemini model for high-volume multimodal applications.
- Family
- Gemini 3.1
- Context
- 1,048,576 tokens
- API
- Yes
Gemini 3 Flash Preview
Google · Multimodal LLM
A fast Gemini 3 model for responsive multimodal assistants and tool-using applications.
- Family
- Gemini 3
- Context
- 1,048,576 tokens
- API
- Yes
Gemini 3.1 Flash Live Preview
Google · Multimodal LLM
A live conversational model that can listen, see, speak, and respond in real time.
- Family
- Gemini Live
- Context
- 131,072 tokens
- API
- Yes
Gemini 3.1 Flash TTS Preview
Google · Audio / Speech
Google’s Gemini-based speech model for creating natural spoken audio from text.
- Family
- Gemini TTS
- Context
- Varies
- API
- Yes
Gemini Omni Flash
Google · Video
A conversational video model for generating and editing video through natural-language interaction.
- Family
- Gemini Omni
- Context
- Varies
- API
- Yes
Nano Banana 2
Google · Image
A production image model for generating and editing images while following detailed instructions.
- Family
- Nano Banana
- Context
- 131,072 tokens
- API
- Yes
Nano Banana 2 Lite
Google · Image
A faster, lower-cost version of Nano Banana 2 for high-volume image creation and edits.
- Family
- Nano Banana
- Context
- 131,072 tokens
- API
- Yes
Nano Banana Pro
Google · Image
Google’s premium Nano Banana image tier for higher-quality generation and professional creative work.
- Family
- Nano Banana
- Context
- Varies
- API
- Yes
Veo 3.1 Preview
Google · Video
Google’s high-quality text- and image-to-video model for cinematic clips.
- Family
- Veo
- Context
- Varies
- API
- Yes
Lyria 3 Pro Preview
Google · Audio / Speech
Google’s premium model for generating music from text and creative direction.
- Family
- Lyria
- Context
- Varies
- API
- Yes
Gemini Embedding 2
Google · Embedding
A Gemini embedding model for converting content into vectors for search, recommendations, and RAG.
- Family
- Gemini Embedding
- Context
- Varies
- API
- Yes
Grok 4.5
xAI / SpaceXAI · Reasoning LLM
xAI’s flagship reasoning model for coding, research, current-information workflows, and tool-using agents.
- Family
- Grok 4
- Context
- 1,000,000 tokens
- API
- Yes
Grok 4.3
xAI / SpaceXAI · Reasoning LLM
A high-capability Grok model with long context and configurable reasoning for production applications.
- Family
- Grok 4
- Context
- 1,000,000 tokens
- API
- Yes
DeepSeek-V4-Pro
DeepSeek · Reasoning LLM
DeepSeek’s larger V4 model for demanding reasoning, coding, and long-context agent workloads.
- Family
- DeepSeek V4
- Context
- 1,000,000 tokens
- API
- Yes
DeepSeek-V4-Flash
DeepSeek · Reasoning LLM
A smaller, faster DeepSeek V4 model designed for efficient reasoning and high-volume production.
- Family
- DeepSeek V4
- Context
- 1,000,000 tokens
- API
- Yes
DeepSeek-V3.2
DeepSeek · Reasoning LLM
An open DeepSeek model that combines general chat, reasoning, and tool use.
- Family
- DeepSeek V3
- Context
- 128,000 tokens
- API
- Yes
DeepSeek-R1
DeepSeek · Reasoning LLM
A widely used open reasoning model designed to show strong step-by-step performance in math, coding, and logic.
- Family
- DeepSeek R1
- Context
- 128,000 tokens
- API
- Yes
Llama 4 Maverick
Meta · Multimodal LLM
Meta’s larger Llama 4 model for strong multimodal assistants, coding, and enterprise applications.
- Family
- Llama 4
- Context
- 1,000,000 tokens
- API
- Via Meta and partners
Llama 4 Scout
Meta · Multimodal LLM
A more deployable Llama 4 model notable for its extremely long context window.
- Family
- Llama 4
- Context
- 10,000,000 tokens
- API
- Via Meta and partners
Llama 3.3 70B Instruct
Meta · General LLM
A strong open 70B text model for chat, coding, and enterprise self-hosting.
- Family
- Llama 3
- Context
- 131,072 tokens
- API
- Via Meta and partners
Llama 3.2 90B Vision Instruct
Meta · Multimodal LLM
Meta’s large open vision-language model for understanding images and documents.
- Family
- Llama 3.2
- Context
- 131,072 tokens
- API
- Via Meta and partners
Llama 3.2 3B Instruct
Meta · General LLM
A compact Llama model suitable for on-device and edge text applications.
- Family
- Llama 3.2
- Context
- 131,072 tokens
- API
- Via Meta and partners
Llama 3.1 405B Instruct
Meta · General LLM
One of Meta’s largest open-weight dense language models for advanced text generation and research.
- Family
- Llama 3.1
- Context
- 131,072 tokens
- API
- Via Meta and partners
Qwen3-235B-A22B
Alibaba Cloud / Qwen · Reasoning LLM
Qwen’s flagship open MoE model for reasoning, coding, multilingual work, and agents.
- Family
- Qwen3
- Context
- 131,072 tokens
- API
- Yes and self-hosted
Qwen3-30B-A3B
Alibaba Cloud / Qwen · Reasoning LLM
A highly efficient open Qwen3 MoE model that activates only a small portion of its parameters.
- Family
- Qwen3
- Context
- 131,072 tokens
- API
- Yes and self-hosted
Qwen3-32B
Alibaba Cloud / Qwen · Reasoning LLM
A dense 32B Qwen model balancing deployability with strong reasoning and multilingual capability.
- Family
- Qwen3
- Context
- 131,072 tokens
- API
- Yes and self-hosted
Qwen3-Coder-480B-A35B
Alibaba Cloud / Qwen · Coding LLM
A large open coding model for software agents, repository work, and long autonomous development tasks.
- Family
- Qwen3-Coder
- Context
- 262,144 tokens
- API
- Yes and self-hosted
Qwen3-Embedding-8B
Alibaba Cloud / Qwen · Embedding
A large open embedding model for multilingual semantic search and retrieval.
- Family
- Qwen3 Embedding
- Context
- 32,768 tokens
- API
- Yes and self-hosted
Qwen VLo
Alibaba Cloud / Qwen · General LLM
A unified Qwen model for understanding images and creating or editing visual content.
- Family
- Qwen VLo
- Context
- Varies
- API
- Preview
Mistral Large 3
Mistral AI · Multimodal LLM
Mistral’s large open-weight multimodal model for enterprise assistants, agents, and long-context work.
- Family
- Mistral Large
- Context
- 262,144 tokens
- API
- Yes and self-hosted
Mistral Medium 3.5
Mistral AI · Multimodal LLM
A production model balancing strong agentic and coding performance with lower cost than Mistral Large.
- Family
- Mistral Medium
- Context
- 262,144 tokens
- API
- Yes and self-hosted
Mistral Small 4
Mistral AI · Multimodal LLM
A compact hybrid model that combines general chat, coding, agents, and reasoning in one deployable package.
- Family
- Mistral Small
- Context
- 262,144 tokens
- API
- Yes and self-hosted
Ministral 3 14B
Mistral AI · Multimodal LLM
A compact 14B Mistral model designed for local, edge, and private multimodal applications.
- Family
- Ministral 3
- Context
- 262,144 tokens
- API
- Yes and self-hosted
Ministral 3 8B
Mistral AI · Multimodal LLM
A compact 8B Mistral model designed for local, edge, and private multimodal applications.
- Family
- Ministral 3
- Context
- 262,144 tokens
- API
- Yes and self-hosted
Ministral 3 3B
Mistral AI · Multimodal LLM
A compact 3B Mistral model designed for local, edge, and private multimodal applications.
- Family
- Ministral 3
- Context
- 262,144 tokens
- API
- Yes and self-hosted
Devstral 2
Mistral AI · Coding LLM
A Mistral coding model optimized for software engineering agents and repository-scale work.
- Family
- Devstral
- Context
- 262,144 tokens
- API
- Yes
Magistral Medium 1.2
Mistral AI · Reasoning LLM
A Mistral model focused on transparent, multi-step reasoning for difficult analytical tasks.
- Family
- Magistral
- Context
- 131,072 tokens
- API
- Yes
Command A+
Cohere · Multimodal LLM
Cohere’s most capable enterprise model for multilingual agents, vision, reasoning, and retrieval workflows.
- Family
- Command
- Context
- 131,072 tokens
- API
- Yes
Command A
Cohere · General LLM
An enterprise-focused language model built for tool use, retrieval, and multilingual business applications.
- Family
- Command
- Context
- 262,144 tokens
- API
- Yes
Command A Reasoning
Cohere · Reasoning LLM
A Command model specialized for complex reasoning while retaining enterprise tool and retrieval features.
- Family
- Command
- Context
- 262,144 tokens
- API
- Yes
Command A Vision
Cohere · Multimodal LLM
A Cohere model for understanding business documents, images, and mixed visual-text content.
- Family
- Command
- Context
- 131,072 tokens
- API
- Yes
Command R7B
Cohere · General LLM
A compact Cohere model for efficient enterprise retrieval, chat, and tool use.
- Family
- Command R
- Context
- 131,072 tokens
- API
- Yes
Aya Expanse 32B
Cohere For AI · General LLM
An open multilingual model designed to perform well across many languages and cultures.
- Family
- Aya Expanse
- Context
- 8,192 tokens
- API
- Via partners / self-hosted
Phi-4
Microsoft · General LLM
A compact Microsoft model designed for strong reasoning and math relative to its size.
- Family
- Phi-4
- Context
- 16,384 tokens
- API
- Yes and self-hosted
Phi-4-mini-instruct
Microsoft · General LLM
A smaller Phi-4 model for efficient local and edge text applications with long context.
- Family
- Phi-4
- Context
- 131,072 tokens
- API
- Yes and self-hosted
Phi-4-multimodal-instruct
Microsoft · Multimodal LLM
A compact Phi model that can understand text, images, and audio in private or edge deployments.
- Family
- Phi-4
- Context
- 131,072 tokens
- API
- Yes and self-hosted
Phi-4-reasoning
Microsoft · Reasoning LLM
A Phi-4 variant tuned specifically for multi-step reasoning in math and complex problem solving.
- Family
- Phi-4
- Context
- 32,768 tokens
- API
- Yes and self-hosted
Phi-4-mini-reasoning
Microsoft · Reasoning LLM
A compact reasoning model for efficient math, logic, and educational applications.
- Family
- Phi-4
- Context
- 128,000 tokens
- API
- Yes and self-hosted
Amazon Nova 2 Lite
Amazon Web Services · Multimodal LLM
AWS’s cost-efficient Nova 2 model for multimodal reasoning, agents, and enterprise applications.
- Family
- Nova 2
- Context
- 1,000,000 tokens
- API
- Yes — Amazon Bedrock
Amazon Nova 2 Sonic
Amazon Web Services · Audio / Speech
AWS’s speech-to-speech model for natural, low-latency voice conversations.
- Family
- Nova 2
- Context
- Varies
- API
- Yes — Amazon Bedrock
Amazon Nova Multimodal Embeddings
Amazon Web Services · Embedding
An AWS embedding model that places text, images, audio, video, and documents into a shared vector space.
- Family
- Nova 2
- Context
- Varies
- API
- Yes — Amazon Bedrock
Amazon Nova Premier
Amazon Web Services · Multimodal LLM
The highest-capability original Amazon Nova model for complex enterprise tasks and teacher-model use.
- Family
- Amazon Nova
- Context
- 1,000,000 tokens
- API
- Yes — Amazon Bedrock
Amazon Nova Canvas
Amazon Web Services · Image
AWS’s image generation model for creating and editing visual assets.
- Family
- Amazon Nova
- Context
- Varies
- API
- Yes — Amazon Bedrock
Amazon Nova Reel
Amazon Web Services · Video
AWS’s generative video model for creating short video clips from prompts and images.
- Family
- Amazon Nova
- Context
- Varies
- API
- Yes — Amazon Bedrock
Jamba Large
AI21 Labs · General LLM
AI21’s enterprise flagship for long documents, retrieval, and production language applications.
- Family
- Jamba
- Context
- 262,144 tokens
- API
- Yes
Jamba2 Mini
AI21 Labs · General LLM
A more efficient Jamba model for enterprise workloads that need long context and controllable behavior.
- Family
- Jamba 2
- Context
- 262,144 tokens
- API
- Yes
Jamba2 3B
AI21 Labs · General LLM
A compact Jamba model intended for on-device, private, and resource-constrained applications.
- Family
- Jamba 2
- Context
- 262,144 tokens
- API
- Model dependent
Nemotron 3 Ultra
NVIDIA · Reasoning LLM
NVIDIA’s very large open reasoning model for enterprise agents, coding, and high-end self-hosted inference.
- Family
- Nemotron 3
- Context
- Varies
- API
- NVIDIA NIM and self-hosted
Llama-3.1-Nemotron-Ultra-253B-v1
NVIDIA · Reasoning LLM
An NVIDIA-tuned Llama model for advanced reasoning and enterprise agent workflows.
- Family
- Nemotron
- Context
- 131,072 tokens
- API
- NVIDIA NIM and self-hosted
FLUX.2 [max]
Black Forest Labs · Image
The highest-quality FLUX.2 tier for photorealistic generation, strong prompt adherence, and complex image edits.
- Family
- FLUX.2
- Context
- Varies
- API
- Yes
FLUX.2 [pro]
Black Forest Labs · Image
A production-focused FLUX.2 model balancing high image quality, speed, and multi-reference editing.
- Family
- FLUX.2
- Context
- Varies
- API
- Yes
FLUX.2 [klein] 9B
Black Forest Labs · Image
A compact FLUX.2 model optimized for very fast, high-volume image generation.
- Family
- FLUX.2
- Context
- Varies
- API
- Yes
Stable Diffusion 3.5 Large
Stability AI · Image
Stability AI’s large open image model for high-quality text-to-image and image-to-image generation.
- Family
- Stable Diffusion 3.5
- Context
- Varies
- API
- Yes and self-hosted
Stable Diffusion 3.5 Large Turbo
Stability AI · Image
A faster SD3.5 Large variant designed to generate strong images in only a few sampling steps.
- Family
- Stable Diffusion 3.5
- Context
- Varies
- API
- Yes and self-hosted
Stable Diffusion 3.5 Medium
Stability AI · Image
A smaller SD3.5 model designed to run on consumer hardware while preserving good prompt following.
- Family
- Stable Diffusion 3.5
- Context
- Varies
- API
- Yes and self-hosted
Runway Gen-4.5
Runway · Video
Runway’s advanced video model for realistic motion, prompt adherence, and cinematic visual generation.
- Family
- Runway Gen-4
- Context
- Varies
- API
- Yes / platform dependent
Runway Aleph
Runway · Video
A video editing model that can add, remove, transform, or restyle elements in existing footage using text instructions.
- Family
- Runway Aleph
- Context
- Varies
- API
- Platform dependent
Runway Gen-4 Images
Runway · Image
Runway’s image generation system for consistent characters, locations, and cinematic visual styles.
- Family
- Runway Images
- Context
- Varies
- API
- Platform dependent
Eleven v3
ElevenLabs · Audio / Speech
ElevenLabs’ most expressive speech model for emotional delivery, dialogue, and multilingual narration.
- Family
- Eleven
- Context
- Varies
- API
- Yes
Eleven Flash v2.5
ElevenLabs · Audio / Speech
A very fast, lower-cost speech model designed for voice agents and interactive applications.
- Family
- Eleven
- Context
- Varies
- API
- Yes
Eleven Multilingual v2
ElevenLabs · Audio / Speech
A stable, natural-sounding speech model for long-form multilingual narration.
- Family
- Eleven
- Context
- Varies
- API
- Yes
Scribe v2
ElevenLabs · Audio / Speech
ElevenLabs’ accurate batch transcription model for meetings, subtitles, and long recordings.
- Family
- Scribe
- Context
- Varies
- API
- Yes
Scribe v2 Realtime
ElevenLabs · Audio / Speech
A live transcription model designed for voice agents, meetings, and conversational applications.
- Family
- Scribe
- Context
- Varies
- API
- Yes