Models
Search and compare to find the model you need.
Auto
warp/auto
Easy questions go to a light model, hard ones to the top model. If one host fails, it falls over to the next model in the chain.
Claude Opus 4.8 · GPT-5.6 Sol · Kimi K3 · Kimi K2.6 · DeepSeek V4 Pro +4
Constituent models
9
Output $/M
$0.40–$30.00
Code
warp/code
Hard code goes to the top model; lighter code goes to strong, low-cost coding models.
Claude Fable 5 · GPT-5.6 Sol · GPT-5.6 Terra · Kimi K3 · GLM-5.2 +4
Constituent models
9
Output $/M
$0.80–$50.00
Top speed
warp/nitro
Routes to the fastest of the curated candidate pool by measured throughput and latency. For live chat and autocomplete.
Candidate pool
48
Output $/M
$0.15–$50.00
Claude Fable 5
anthropic/claude-fable-5
Top-tier default model
Input $/M
$10.00
Output $/M
$50.00
Context
500K
Providers
1
GPT-5.6 Sol
openai/gpt-5.6-sol
GPT-5.6 top line
Input $/M
$5.00
Output $/M
$30.00
Context
1.1M
Providers
1
Claude Opus 4.8
anthropic/claude-opus-4.8
Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token...
Input $/M
$5.00
Output $/M
$25.00
Context
1M
Providers
1
GPT-5.5
openai/gpt-5.5
GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and improved token efficiency on hard tasks. It features a 1M+ token...
Input $/M
$5.00
Output $/M
$30.00
Context
1.1M
Providers
1
Claude Sonnet 5
anthropic/claude-sonnet-5
Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports adaptive thinking with selectable reasoning effort levels (low, medium, high, max,...
Input $/M
$2.00
Output $/M
$10.00
Context
1M
Providers
1
Claude Opus 4.7
anthropic/claude-opus-4-7
Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.6, it delivers stronger performance on...
Input $/M
$5.00
Output $/M
$25.00
Context
1M
Providers
1
GLM-5.2
zhipu/glm-5.2
Opus-class on coding benchmarks, 46% of the cost with caching
Input $/M
$0.85
Output $/M
$2.68
Context
1.0M
Providers
14
GPT-5.4
openai/gpt-5.4
GPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system. It features a 1M+ token context window (922K input, 128K output) with support for...
Input $/M
$2.50
Output $/M
$15.00
Context
400K
Providers
1
Grok 4.5
xai/grok-4.5
Grok 4.5 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.
Input $/M
$2.00
Output $/M
$6.00
Context
500K
Providers
1
Claude Sonnet 4.6
anthropic/claude-sonnet-4-6
Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at iterative development, complex codebase navigation, end-to-end project management with...
Input $/M
$3.00
Output $/M
$15.00
Context
1M
Providers
1
GLM 5.1
zhipu/glm-5.1
GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. Unlike previous models built around minute-level interactions, GLM-5.1 can work independently and continuously on...
Input $/M
$1.05
Output $/M
$3.50
Context
205K
Providers
11
Gemini 3.5 Flash
google/gemini-3.5-flash
Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution...
Input $/M
$1.50
Output $/M
$9.00
Context
1.0M
Providers
2
Kimi K3
moonshot/kimi-k3
Moonshot 플래그십 · 1M 컨텍스트
Input $/M
$3.00
Output $/M
$15.00
Context
1.0M
Providers
3
Gemini 3.1 Pro
google/gemini-3.1-pro
Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows. Building on the multimodal foundation...
Input $/M
$2.00
Output $/M
$12.00
Context
1.0M
Providers
2
Gemini 3.1 Pro Preview
google/gemini-3.1-pro-preview
Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows. Building on the multimodal foundation...
Input $/M
$2.00
Output $/M
$12.00
Context
1.0M
Providers
1
Grok 4.20
xai/grok-4.20
2M context
Input $/M
$1.25
Output $/M
$2.50
Context
2M
Providers
1
Gemini 3 Flash
google/gemini-3-flash
Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It delivers near Pro level reasoning and tool...
Input $/M
$0.50
Output $/M
$3.00
Context
1.0M
Providers
2
Kimi K2.7 Code
moonshot/kimi-k2.7-code
Coding-specialized
Input $/M
$0.74
Output $/M
$3.50
Context
262K
Providers
8
MiMo V2.5 Pro
xiaomimimo/mimo-v2.5-pro
MiMo-V2.5-Pro is purpose-built to push the boundaries of complex software engineering and extreme long-horizon tasks. Compared to its predecessor, it achieves a comprehensive leap in general agentic capabilities, advancing the human-AI collaboration paradigm toward true "autonomous delivery." Without human intervention, it stably orchestrates massive workflows requiring up to a thousand tool calls in a single session, not only precisely capturing implicit requirements within ultra-long contexts but also demonstrating exceptional global architectural planning and self-correction discipline. In core agentic scenarios and long-horizon complexities, MiMo-V2.5-Pro is fully equipped to go head-to-head with top-tier global models like Claude Opus 4.6 and GPT-5.4. Backed by this exceptionally high execution confidence and long-term logical consistency, it completely sheds the "co-pilot" label, ready to take on truly serious, professional-grade workloads in real-world business environments.
Input $/M
$1.00
Output $/M
$3.00
Context
1.0M
Providers
4
Kimi K2.6
moonshot/kimi-k2.6
Kimi K2.6 is Moonshot AI's next-generation multimodal model, designed for long-horizon coding, coding-driven UI/UX generation, and multi-agent orchestration. It handles complex end-to-end coding tasks across Python, Rust, and Go, and...
Input $/M
$0.66
Output $/M
$3.41
Context
262K
Providers
11
Qwen3.7 Plus
qwen/qwen3.7-plus
Qwen3.7-Plus is a cost-effective model in Alibaba's Qwen3.7 series. It supports text and image input with text output, building on the series' text capabilities with a comprehensive upgrade to its...
Input $/M
$0.32
Output $/M
$1.28
Context
1M
Providers
1
Qwen3.6 Max Preview
qwen/qwen3.6-max-preview
Qwen3.6-Max-Preview is a proprietary frontier model from Alibaba Cloud built on a sparse mixture-of-experts architecture with approximately 1 trillion total parameters. It is optimized for agentic coding, tool use, and...
Input $/M
$1.30
Output $/M
$7.80
Context
262K
Providers
1
DeepSeek V4 Pro
deepseek/deepseek-v4-pro
DeepSeek flagship · 1M context
Input $/M
$1.30
Output $/M
$2.60
Context
1.0M
Providers
8
GLM-5
zhipu/glm-5
GLM flagship
Input $/M
$0.60
Output $/M
$2.08
Context
1.0M
Providers
7
GPT-5.4 mini
openai/gpt-5.4-mini
GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image inputs with strong performance across reasoning, coding,...
Input $/M
$0.75
Output $/M
$4.50
Context
400K
Providers
1
MiMo V2 Pro
xiaomimimo/mimo-v2-pro
Xiaomi’s MiMo-V2-Pro is a flagship, >1T parameter foundation model with a 1M context window, specifically optimized for advanced agentic workflows. Highly compatible with frameworks like OpenClaw, it secures top-tier global rankings on PinchBench and ClawBench, with a perceived capability approaching Opus 4.6. Built to act as the cognitive core of intelligent systems, MiMo-V2-Pro seamlessly orchestrates complex tasks, drives engineering automation, and consistently delivers dependable real-world outcomes.
Input $/M
$2.00
Output $/M
$6.00
Context
1.0M
Providers
1
Gemini 2.5 Pro
google/gemini-2.5-pro
Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy...
Input $/M
$1.25
Output $/M
$10.00
Context
1M
Providers
1
MiniMax M3
minimax/minimax-m3
MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...
Input $/M
$0.30
Output $/M
$1.20
Context
1.0M
Providers
6
Qwen3.6 Plus
qwen/qwen3.6-plus
Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong scalability and high-performance inference. Compared to the 3.5 series, it delivers...
Input $/M
$0.50
Output $/M
$3.00
Context
1M
Providers
3
Grok 4.3
xai/grok-4.3
Grok 4.3 is a reasoning model from xAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual...
Input $/M
$1.25
Output $/M
$2.50
Context
1M
Providers
1
Qwen3.5 397B A17B
qwen/qwen3.5-397b-a17b
The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. It delivers...
Input $/M
$0.45
Output $/M
$3.00
Context
262K
Providers
6
GLM 4.7
zhipu/glm-4.7
GLM-4.7 is Z.ai’s latest flagship model, featuring upgrades in two key areas: enhanced programming capabilities and more stable multi-step reasoning/execution. It demonstrates significant improvements in executing complex agent tasks while...
Input $/M
$0.40
Output $/M
$1.75
Context
205K
Providers
6
Inkling
thinkingmachines/inkling
Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...
Input $/M
$1.00
Output $/M
$4.05
Context
524K
Providers
3
GLM 5v Turbo
zhipu/glm-5v-turbo
GLM-5V-Turbo is Z.ai’s first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks. It natively handles image, video, and text inputs, excels at long-horizon planning, complex coding,...
Input $/M
$1.20
Output $/M
$4.00
Context
205K
Providers
1
GPT-5.1
openai/gpt-5.1
GPT-5.1 is the latest frontier-grade model in the GPT-5 series, offering stronger general-purpose reasoning, improved instruction adherence, and a more natural conversational style compared to GPT-5. It uses adaptive reasoning...
Input $/M
$1.25
Output $/M
$10.00
Context
400K
Providers
1
DeepSeek V4 Flash
deepseek/deepseek-v4-flash
DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and...
Input $/M
$0.09
Output $/M
$0.18
Context
1M
Providers
6
GPT-5.2
openai/gpt-5.2
Top-tier fallback model
Input $/M
$1.75
Output $/M
$14.00
Context
400K
Providers
1
Qwen3 Max
qwen/qwen3-max
Qwen flagship
Input $/M
$1.20
Output $/M
$6.00
Context
262K
Providers
2
Gemini 3.1 Flash Lite
google/gemini-3.1-flash-lite
Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image, video, audio, and PDF inputs, and is designed for lightweight agentic...
Input $/M
$0.25
Output $/M
$1.50
Context
1.0M
Providers
2
Gemini 3.1 Flash Lite Preview
google/gemini-3.1-flash-lite-preview
Gemini 3.1 Flash Lite Preview is Google's high-efficiency model optimized for high-volume use cases. It outperforms Gemini 2.5 Flash Lite on overall quality and approaches Gemini 2.5 Flash performance across...
Input $/M
$0.25
Output $/M
$1.50
Context
1.0M
Providers
1
MiMo V2.5
xiaomimimo/mimo-v2.5
Xiaomi's native omni-modal model (text / image / video / audio understanding) with Pro-level agentic performance at ~half the inference cost and a 1M-token context — built for cost-efficient, perception-rich agent workflows.
Input $/M
$0.40
Output $/M
$2.00
Context
1.0M
Providers
5
Kimi K2 Thinking
moonshot/kimi-k2-thinking
Reasoning (Thinking) specialized
Input $/M
$0.47
Output $/M
$2.00
Context
262K
Providers
3
DeepSeek V3.2
deepseek/deepseek-v3.2
DeepSeek-V3.2 is a large language model designed to harmonize high computational efficiency with strong reasoning and agentic tool-use performance. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism...
Input $/M
$0.27
Output $/M
$0.40
Context
164K
Providers
6
GLM 4.6
zhipu/glm-4.6
Compared with GLM-4.5, this generation brings several key improvements: Longer context window: The context window has been expanded from 128K to 200K tokens, enabling the model to handle more complex...
Input $/M
$0.50
Output $/M
$2.00
Context
205K
Providers
4
DeepSeek V3.2 Exp
deepseek/deepseek-v3.2-exp
DeepSeek-V3.2-Exp is an experimental large language model released by DeepSeek as an intermediate step between V3.1 and future architectures. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism...
Input $/M
$0.27
Output $/M
$0.41
Context
164K
Providers
1
Qwen3 235B
qwen/qwen3-235b
Qwen3-235B-A22B-Instruct-2507 is a multilingual, instruction-tuned mixture-of-experts language model based on the Qwen3-235B architecture, with 22B active parameters per forward pass. It is optimized for general-purpose text generation, including instruction following,...
Input $/M
$0.09
Output $/M
$0.60
Context
262K
Providers
6
DeepSeek R1 0528
deepseek/deepseek-r1-0528
May 28th update to the [original DeepSeek R1](/deepseek/deepseek-r1) Performance on par with [OpenAI o1](/openai/o1), but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active...
Input $/M
$0.50
Output $/M
$2.15
Context
164K
Providers
3
DeepSeek V3.1
deepseek/deepseek-v3.1
DeepSeek-V3.1 is a hybrid model that supports both thinking mode and non-thinking mode.DeepSeek-V3.1 is post-trained on the top of DeepSeek-V3.1-Base, which is built upon the original V3 base checkpoint through a two-phase long context extension approach, following the methodology outlined in the original DeepSeek-V3 report. We have expanded our dataset by collecting additional long documents and substantially extending both training phases. The 32K extension phase has been increased 10-fold to 630B tokens, while the 128K extension phase has been extended by 3.3x to 209B tokens.
Input $/M
$0.25
Output $/M
$0.95
Context
164K
Providers
3
MiniMax M2.7
minimax/minimax-m2.7
MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built to actively participate in its own evolution, M2.7 integrates advanced agentic capabilities through multi-agent...
Input $/M
$0.25
Output $/M
$1.00
Context
205K
Providers
6
Qwen3.5 122B A10B
qwen/qwen3.5-122b-a10b
The Qwen3.5 122B-A10B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. In terms of...
Input $/M
$0.29
Output $/M
$2.40
Context
262K
Providers
3
DeepSeek V3.1 Terminus
deepseek/deepseek-v3.1-terminus
DeepSeek-V3.1 Terminus is an update to [DeepSeek V3.1](/deepseek/deepseek-chat-v3.1) that maintains the model's original capabilities while addressing issues reported by users, including language consistency and agent capabilities, further optimizing the model's...
Input $/M
$0.27
Output $/M
$0.95
Context
164K
Providers
2
Claude Haiku 4.5
anthropic/claude-haiku-4.5
Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance...
Input $/M
$1.00
Output $/M
$5.00
Context
200K
Providers
1
GLM 4.5
zhipu/glm-4.5
GLM-4.5 is our latest flagship foundation model, purpose-built for agent-based applications. It leverages a Mixture-of-Experts (MoE) architecture and supports a context length of up to 128k tokens. GLM-4.5 delivers significantly...
Input $/M
$0.60
Output $/M
$2.20
Context
131K
Providers
2
Gemini 2.5 Flash
google/gemini-2.5-flash
Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks. It includes built-in "thinking" capabilities, enabling it to provide responses with greater...
Input $/M
$0.30
Output $/M
$2.50
Context
1.0M
Providers
2
Qwen3.5 27B
qwen/qwen3.5-27b
The Qwen3.5 27B native vision-language Dense model incorporates a linear attention mechanism, delivering fast response times while balancing inference speed and performance. Its overall capabilities are comparable to those of...
Input $/M
$0.26
Output $/M
$2.60
Context
262K
Providers
4
GPT-5.3 Codex
openai/gpt-5.3-codex
Coding-specialized (Codex)
Input $/M
$1.75
Output $/M
$14.00
Context
400K
Providers
1
GPT-5.4 nano
openai/gpt-5.4-nano
GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks. It supports text and image inputs and is designed for low-latency...
Input $/M
$0.20
Output $/M
$1.25
Context
400K
Providers
1
Qwen3 Next 80B A3B Instruct
qwen/qwen3-next-80b-a3b-instruct
Qwen3-Next-80B-A3B-Instruct is an instruction-tuned chat model in the Qwen3-Next series optimized for fast, stable responses without “thinking” traces. It targets complex tasks across reasoning, code generation, knowledge QA, and multilingual...
Input $/M
$0.09
Output $/M
$1.10
Context
262K
Providers
3
Qwen3 235B A22B Thinking 2507
qwen/qwen3-235b-a22b-thinking-2507
Qwen3-235B-A22B-Thinking-2507 is a high-performance, open-weight Mixture-of-Experts (MoE) language model optimized for complex reasoning tasks. It activates 22B of its 235B parameters per forward pass and natively supports up to 262,144...
Input $/M
$0.23
Output $/M
$2.30
Context
262K
Providers
5
DeepSeek R1
deepseek/deepseek-r1
DeepSeek R1 is here: Performance on par with [OpenAI o1](/openai/o1), but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active in an inference pass....
Input $/M
$4.00
Output $/M
$4.00
Context
66K
Providers
1
Qwen3.5 35B A3B
qwen/qwen3.5-35b-a3b
The Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture that integrates linear attention mechanisms and a sparse mixture-of-experts model, achieving higher inference efficiency. Its overall...
Input $/M
$0.14
Output $/M
$1.00
Context
262K
Providers
5
DeepSeek V3 0324
deepseek/deepseek-v3-0324
DeepSeek V3, a 685B-parameter, mixture-of-experts model, is the latest iteration of the flagship chat model family from the DeepSeek team.
Input $/M
$0.24
Output $/M
$0.90
Context
164K
Providers
3
Qwen3 VL 235B A22B Thinking
qwen/qwen3-vl-235b-a22b-thinking
Qwen3-VL-235B-A22B Thinking is a multimodal model that unifies strong text generation with visual understanding across images and video. The Thinking model is optimized for multimodal reasoning in STEM and math....
Input $/M
$0.98
Output $/M
$3.95
Context
131K
Providers
1
MiniMax M2.5
minimax/minimax-m2.5
MiniMax-M2.5 is a SOTA large language model designed for real-world productivity. Trained in a diverse range of complex real-world digital working environments, M2.5 builds upon the coding expertise of M2.1...
Input $/M
$0.30
Output $/M
$1.20
Context
205K
Providers
6
GPT-5 mini
openai/gpt-5-mini
GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following and safety-tuning benefits as GPT-5, but with reduced latency and cost....
Input $/M
$0.25
Output $/M
$2.00
Context
400K
Providers
1
Qwen3 Coder 480B A35B Instruct
qwen/qwen3-coder-480b-a35b-instruct
Qwen3-Coder-480B-A35B-Instruct is a cutting-edge open coding model from Qwen, matching Claude Sonnet’s performance in agentic programming, browser automation, and core development tasks. With native 256K context (extendable to 1M tokens via YaRN), it excels at repository-scale analysis and features specialized function-call support for platforms like Qwen Code and CLINE—making it ideal for complex, real-world development workflows.
Input $/M
$0.38
Output $/M
$1.55
Context
262K
Providers
3
GLM 4.6v
zhipu/glm-4.6v
GLM-4.6V is a large multimodal model designed for high-fidelity visual understanding and long-context reasoning across images, documents, and mixed media. It supports up to 128K tokens, processes complex page layouts...
Input $/M
$0.30
Output $/M
$0.90
Context
131K
Providers
1
GLM 4.5 Air
zhipu/glm-4.5-air
GLM-4.5-Air is the lightweight variant of our latest flagship model family, also purpose-built for agent-centric applications. Like GLM-4.5, it adopts the Mixture-of-Experts (MoE) architecture but with a more compact parameter...
Input $/M
$0.13
Output $/M
$0.85
Context
131K
Providers
2
Qwen3 Next 80B A3B Thinking
qwen/qwen3-next-80b-a3b-thinking
Qwen3-Next-80B-A3B-Thinking is a reasoning-first chat model in the Qwen3-Next line that outputs structured “thinking” traces by default. It’s designed for hard multi-step problems; math proofs, code synthesis/debugging, logic, and agentic...
Input $/M
$0.15
Output $/M
$1.50
Context
262K
Providers
2
GLM 4.7 Flash
zhipu/glm-4.7-flash
As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agentic coding use cases, strengthening coding capabilities, long-horizon task planning,...
Input $/M
$0.06
Output $/M
$0.40
Context
200K
Providers
3
Gemma 3 27B It
google/gemma-3-27b-it
Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities,...
Input $/M
$0.08
Output $/M
$0.16
Context
131K
Providers
3
Nvidia Nemotron 3 Super 120B A12B
nvidia/nvidia-nemotron-3-super-120b-a12b
NVIDIA Nemotron 3 Super is a hybrid Mixture-of-Experts (MoE) model engineered for highest compute efficiency and accuracy in multi-agent applications and specialized agentic systems. It is optimized to run many collaborating agents per application on a single GPU, delivering high accuracy for reasoning, tool use, and instruction following.
Input $/M
$0.09
Output $/M
$0.40
Context
262K
Providers
1
DeepSeek V3
deepseek/deepseek_v3
DeepSeek-V3 is the latest model from the DeepSeek team, building upon the instruction following and coding abilities of the previous versions. Pre-trained on nearly 15 trillion tokens, the reported evaluations reveal that the model outperforms other open-source models and rivals leading closed-source models.
Input $/M
$0.89
Output $/M
$0.89
Context
64K
Providers
1
DeepSeek V3
deepseek/deepseek-v3
DeepSeek-V3, a strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token. To achieve efficient inference and cost-effective training, DeepSeek-V3 adopts Multi-head Latent Attention (MLA) and DeepSeekMoE architectures, which were thoroughly validated in DeepSeek-V2.
Input $/M
$0.32
Output $/M
$0.89
Context
66K
Providers
1
GLM 4.5v
zhipu/glm-4.5v
GLM-4.5V is a vision-language foundation model for multimodal agent applications. Built on a Mixture-of-Experts (MoE) architecture with 106B parameters and 12B activated parameters, it achieves state-of-the-art results in video understanding,...
Input $/M
$0.60
Output $/M
$1.80
Context
66K
Providers
1
GPT OSS 120B
openai/gpt-oss-120b
gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. It activates 5.1B parameters per forward pass and is optimized...
Input $/M
$0.04
Output $/M
$0.17
Context
131K
Providers
8
o3 Mini
openai/o3-mini
OpenAI o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and coding. This model supports the `reasoning_effort` parameter, which can be set to...
Input $/M
$1.10
Output $/M
$4.40
Context
200K
Providers
1
Qwen3 32B
qwen/qwen3-32b
Qwen3-32B is a dense 32.8B parameter causal language model from the Qwen3 series, optimized for both complex reasoning and efficient dialogue. It supports seamless switching between a "thinking" mode for...
Input $/M
$0.08
Output $/M
$0.28
Context
41K
Providers
2
MiniMax M2
minimax/minimax-m2
MiniMax-M2 is a compact, high-efficiency large language model optimized for end-to-end coding and agentic workflows. With 10 billion activated parameters (230 billion total), it delivers near-frontier intelligence across general reasoning,...
Input $/M
$0.30
Output $/M
$1.20
Context
205K
Providers
1
Gemma 3 12B It
google/gemma-3-12b-it
Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities,...
Input $/M
$0.05
Output $/M
$0.15
Context
131K
Providers
2
GPT-5 nano
openai/gpt-5-nano
GPT-5-Nano is the smallest and fastest variant in the GPT-5 system, optimized for developer tools, rapid interactions, and ultra-low latency environments. While limited in reasoning depth compared to its larger...
Input $/M
$0.05
Output $/M
$0.40
Context
400K
Providers
1
GPT 5.2 Codex
openai/gpt-5.2-codex
GPT-5.2-Codex is an upgraded version of GPT-5.1-Codex optimized for software engineering and coding workflows. It is designed for both interactive development sessions and long, independent execution of complex engineering tasks....
Input $/M
$1.75
Output $/M
$14.00
Context
272K
Providers
1
GPT 5.1 Codex
openai/gpt-5.1-codex
GPT-5.1-Codex is a specialized version of GPT-5.1 optimized for software engineering and coding workflows. It is designed for both interactive development sessions and long, independent execution of complex engineering tasks....
Input $/M
$1.25
Output $/M
$10.00
Context
272K
Providers
1
GPT 5.1 Codex Mini
openai/gpt-5.1-codex-mini
GPT-5.1-Codex-Mini is a smaller and faster version of GPT-5.1-Codex
Input $/M
$0.25
Output $/M
$2.00
Context
272K
Providers
1
Llama 4 Maverick
meta-llama/llama-4-maverick
Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architecture with 128 experts and 17 billion active parameters per forward...
Input $/M
$0.15
Output $/M
$0.60
Context
1.0M
Providers
3
Llama 4 Scout
meta-llama/llama-4-scout
Llama 4 Scout 17B Instruct (16E) is a mixture-of-experts (MoE) language model developed by Meta, activating 17 billion parameters out of a total of 109B. It supports native multimodal input...
Input $/M
$0.08
Output $/M
$0.30
Context
328K
Providers
2
Qwen2.5 VL 72B Instruct
qwen/qwen2.5-vl-72b-instruct
Qwen2.5-VL is proficient in recognizing common objects such as flowers, birds, fish, and insects. It is also highly capable of analyzing texts, charts, icons, graphics, and layouts within images.
Input $/M
$0.80
Output $/M
$0.80
Context
33K
Providers
1
DALL·E 3
openai/dall-e-3
은퇴 — OpenAI가 2026-05-12 서비스 종료 (gpt-image-2로 대체)
Price (per image)
$0.04/image
Context
—
Providers
1
GPT-5.6 Terra
openai/gpt-5.6-terra
GPT-5.6 standard line
Input $/M
$2.50
Output $/M
$15.00
Context
1.1M
Providers
1
GPT-5.6 Luna
openai/gpt-5.6-luna
GPT-5.6 lightweight line
Input $/M
$1.00
Output $/M
$6.00
Context
1.1M
Providers
1
Tencent Hunyuan 3
tencent/hy3
Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active, 192 experts with top-8 routing) built for reasoning, agentic workflows, and real-world production use. It supports a configurable reasoning effort:...
Input $/M
$0.20
Output $/M
$0.80
Context
262K
Providers
3
GLM 5.2 Fast
zhipu/glm-5.2-fast
GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering,...
Input $/M
$1.20
Output $/M
$4.10
Context
1.0M
Providers
1
GLM 5.2 Batch
zhipu/glm-5.2-batch
GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering,...
Input $/M
$0.75
Output $/M
$2.50
Context
—
Providers
1
Kimi K2.7 Code HighSpeed
moonshot/kimi-k2.7-code-highspeed
코딩 특화 · 고속 서빙(~180 tok/s)
Input $/M
$1.90
Output $/M
$8.00
Context
262K
Providers
2
MiniMax M3 Preview
minimax/minimax-m3-preview
MiniMax-M3 preview is a 1.4T-parameter frontier model from MiniMax for coding, agentic workflows, and complex reasoning, served at fp8 with a 512K context window.
Input $/M
$0.30
Output $/M
$1.20
Context
524K
Providers
1
Nemotron 3 Ultra 550B A55B
nvidia/nemotron-3-ultra-550b-a55b
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
Input $/M
$0.60
Output $/M
$3.60
Context
512K
Providers
3
Nemotron 3 Ultra
nvidia/nemotron-3-ultra
NVIDIA 550B open weights
Input $/M
$0.50
Output $/M
$2.20
Context
1M
Providers
2
Nemotron Content Safety 3.5
nvidia/nemotron-content-safety-3.5
Nemotron Content Safety 3.5 is a multimodal safety classifier developed by NVIDIA. A compact safety model that handles text, images, and custom policies. It outputs a safe/unsafe classification plus a reasoning trace, and can be used as an inference-time guardrail, as a judge for LLM safety testing and evaluation, or with the accompanying training dataset to post-train models for safer behavior.
Input $/M
$0.20
Output $/M
$0.20
Context
131K
Providers
1
Qwen3.7 Max
qwen/qwen3.7-max
Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series. It supports text input and output and is designed for agent-centric workloads, with particular strengths in coding, office and productivity tasks,...
Input $/M
$2.50
Output $/M
$7.50
Context
1M
Providers
4
Grok Build 0.1
xai/grok-build-0.1
Grok Build 0.1 is xAI’s fast coding model trained specifically for agentic software engineering workflows. It supports text and image inputs with text output, and is optimized for interactive coding...
Input $/M
$1.00
Output $/M
$2.00
Context
256K
Providers
1
Qwen3.6 27B
qwen/qwen3.6-27b
Qwen3.6 27B is a dense 27-billion-parameter language model from the Qwen Team at Alibaba, released in April 2026. It features hybrid multimodal capabilities — accepting text, image, and video inputs...
Input $/M
$0.32
Output $/M
$3.20
Context
262K
Providers
5
Qwen3.6 35B A3B
qwen/qwen3.6-35b-a3b
Qwen3.6-35B-A3B is an open-weight multimodal model from Alibaba Cloud with 35 billion total parameters and 3 billion active parameters per token. It uses a hybrid sparse mixture-of-experts architecture combining Gated...
Input $/M
$0.15
Output $/M
$0.95
Context
262K
Providers
4
GPT 5.5 Pro
openai/gpt-5.5-pro
GPT-5.5 Pro is OpenAI’s high-capability model optimized for deep reasoning and accuracy on complex, high-stakes workloads. It features a 1M+ token context window (922K input, 128K output) with support for...
Input $/M
$30.00
Output $/M
$180.00
Context
1.1M
Providers
1
Hy3 Preview
tencent/hy3-preview
Hy3 preview is a high-efficiency Mixture-of-Experts model from Tencent designed for agentic workflows and production use. It supports configurable reasoning levels across disabled, low, and high modes, allowing it to...
Input $/M
$0.18
Output $/M
$0.60
Context
262K
Providers
1
GPT Image 2
openai/gpt-image-2
토큰 과금 · 1024² 기준 low $0.006 / medium $0.053 / high $0.211
Price (per image)
~$0.05/image
Context
—
Providers
1
Qwen3.6 Plus 2026 04.02
qwen/qwen3.6-plus-2026-04-02
Input $/M
$0.50
Output $/M
$3.00
Context
262K
Providers
1
Gemma 4 26B A4B It
google/gemma-4-26b-a4b-it
Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at...
Input $/M
$0.07
Output $/M
$0.34
Context
262K
Providers
4
Gemma 4 31B It
google/gemma-4-31b-it
Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...
Input $/M
$0.13
Output $/M
$0.38
Context
262K
Providers
8
Gemma 4 31B It Turbo
google/gemma-4-31b-it-turbo
Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...
Input $/M
$0.12
Output $/M
$0.37
Context
262K
Providers
1
Grok 4.20 Multi-Agent
xai/grok-4.20-multi-agent
멀티 에이전트 오케스트레이션 모드
Input $/M
$1.25
Output $/M
$2.50
Context
2M
Providers
1
Nemotron Cascade 2 30B A3B
nvidia/nemotron-cascade-2-30b-a3b
Nemotron Cascade 2 30B A3B is a reasoning-optimized language model from NVIDIA, designed for efficient inference with strong reasoning capabilities across complex tasks.
Input $/M
$0.14
Output $/M
$0.80
Context
256K
Providers
1
MiniMax M2.7 Highspeed
minimax/minimax-m2.7-highspeed
MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built to actively participate in its own evolution, M2.7 integrates advanced agentic capabilities through multi-agent...
Input $/M
$0.60
Output $/M
$2.40
Context
205K
Providers
1
MiniMax M2.7 Turbo
minimax/minimax-m2.7-turbo
MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built to actively participate in its own evolution, M2.7 integrates advanced agentic capabilities through multi-agent...
Input $/M
$0.38
Output $/M
$1.70
Context
197K
Providers
1
Mistral Small 2603
mistral/mistral-small-2603
Mistral Small 4 unifies instruction following, reasoning, coding, and vision in a single 119B MoE model with 256K context and configurable reasoning effort.
Input $/M
$0.19
Output $/M
$0.75
Context
256K
Providers
1
GLM 5 Turbo
zhipu/glm-5-turbo
GLM-5 Turbo is a new model from Z.ai designed for fast inference and strong performance in agent-driven environments such as OpenClaw scenarios. It is deeply optimized for real-world agent workflows...
Input $/M
$1.20
Output $/M
$4.00
Context
203K
Providers
2
Nemotron 120B A12B
nvidia/nemotron-120b-a12b
Flagship LLM for both reasoning and non-reasoning tasks. It excels in code generation and agentic execution.
Input $/M
$0.30
Output $/M
$0.75
Context
203K
Providers
1
Qwen3.5 9B
qwen/qwen3.5-9b
Qwen3.5-9B is a multimodal foundation model from the Qwen3.5 family, designed to deliver strong reasoning, coding, and visual understanding in an efficient 9B-parameter architecture. It uses a unified vision-language design...
Input $/M
$0.10
Output $/M
$0.15
Context
262K
Providers
3
Grok 4.20 Non-Reasoning
xai/grok-4.20-non-reasoning
추론 생략 모드 · 저지연
Input $/M
$1.25
Output $/M
$2.50
Context
2M
Providers
1
GPT 5.4 Pro
openai/gpt-5.4-pro
GPT-5.4 Pro is OpenAI's most advanced model, building on GPT-5.4's unified architecture with enhanced reasoning capabilities for complex, high-stakes tasks. It features a 1M+ token context window (922K input, 128K...
Input $/M
$30.00
Output $/M
$180.00
Context
1.1M
Providers
1
Gemma 4 E4b It
google/gemma-4-e4b-it
Input $/M
$0.02
Output $/M
$0.10
Context
131K
Providers
1
Gemini 3.1 Pro Preview Customtools
google/gemini-3.1-pro-preview-customtools
Gemini 3.1 Pro Preview Custom Tools is a variant of Gemini 3.1 Pro that improves tool selection behavior by preventing overuse of a general bash tool when more efficient third-party...
Input $/M
$2.00
Output $/M
$12.00
Context
1.0M
Providers
1
Qwen3.5 Plus
qwen/qwen3.5-plus
The Qwen3.5 native vision-language series Plus models are based on a hybrid architecture design that integrates linear attention mechanisms with sparse Mixture-of-Experts (MoE), achieving higher inference efficiency. Across various task evaluations, the 3.5 series demonstrates exceptional performance comparable to current top-tier frontier models, marking a leap forward in both plain text and multimodal capabilities compared to the 3 series.
Input $/M
$0.40
Output $/M
$2.40
Context
992K
Providers
1
MiniMax M2.5 Highspeed
minimax/minimax-m2.5-highspeed
MiniMax-M2.5 is a SOTA large language model designed for real-world productivity. Trained in a diverse range of complex real-world digital working environments, M2.5 builds upon the coding expertise of M2.1...
Input $/M
$0.60
Output $/M
$2.40
Context
205K
Providers
1
Qwen3 Max Thinking
qwen/qwen3-max-thinking
Qwen3-Max-Thinking is the flagship reasoning model in the Qwen3 series, designed for high-stakes cognitive tasks that require deep, multi-step reasoning. By significantly scaling model capacity and reinforcement learning compute, it...
Input $/M
$1.20
Output $/M
$6.00
Context
262K
Providers
1
Qwen3 Coder Next
qwen/qwen3-coder-next
Qwen3-Coder-Next is an open-weight causal language model optimized for coding agents and local development workflows. It uses a sparse MoE design with 80B total parameters and only 3B activated per...
Input $/M
$0.20
Output $/M
$1.50
Context
262K
Providers
1
DeepSeek OCR 2
deepseek/deepseek-ocr-2
DeepSeek-OCR 2 is a multimodal document recognition model released by DeepSeek AI, serving as an upgrade to DeepSeek-OCR. By introducing the DeepEncoder V2 architecture, it achieves a paradigm shift in visual encoding from "fixed scanning" to "semantic reasoning." The model replaces the original CLIP encoder with a lightweight language model (Qwen2-0.5B) and incorporates a causal flow query mechanism, while retaining the DeepSeek-3B-MoE decoder. The model requires only 256 to 1120 visual tokens to cover complex document pages. On the OmniDocBench v1.5 benchmark, it achieves an overall score of 91.09%, a 3.73% improvement over its predecessor, with reading order recognition edit distance reduced from 0.085 to 0.057.
Input $/M
$0.03
Output $/M
$0.03
Context
8K
Providers
1
Kimi K2.5
moonshot/kimi-k2.5
Kimi K2.5 is Moonshot AI's native multimodal model, delivering state-of-the-art visual coding capability and a self-directed agent swarm paradigm. Built on Kimi K2 with continued pretraining over approximately 15T mixed...
Input $/M
$0.45
Output $/M
$2.25
Context
262K
Providers
7
Solar Pro 3
upstage/solar-pro3
Korean-optimized
Input $/M
$0.15
Output $/M
$0.60
Context
128K
Providers
1
GLM 4.7 H
zhipu/glm-4.7-h
GLM-4.7 is Z.AI's latest flagship model, with major upgrades focused on advanced coding capabilities and more reliable multi-step reasoning and execution. It shows clear gains in complex agent workflows, while delivering a more natural conversational experience and stronger front-end design sensibility.
Input $/M
$0.60
Output $/M
$2.20
Context
205K
Providers
1
Qwen3 VL 235B A22B
qwen/qwen3-vl-235b-a22b
Qwen3-VL 235B vision-language model with MoE architecture. The most powerful VL model in the Qwen series with superior visual perception, OCR, and multimodal reasoning.
Input $/M
$0.21
Output $/M
$1.90
Context
128K
Providers
1
Mistral Small 3.2 24B Instruct
mistral/mistral-small-3.2-24b-instruct
Mistral Small 3.2 is a 24B parameter model optimized for efficiency and performance. Ideal for general-purpose tasks with balanced speed and capability.
Input $/M
$0.09
Output $/M
$0.25
Context
256K
Providers
1
MiniMax M2.1
minimax/minimax-m2.1
MiniMax-M2.1 is a lightweight, state-of-the-art large language model optimized for coding, agentic workflows, and modern application development. With only 10 billion activated parameters, it delivers a major jump in real-world...
Input $/M
$0.30
Output $/M
$1.20
Context
205K
Providers
1
Gemini 3 Flash Preview
google/gemini-3-flash-preview
Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It delivers near Pro level reasoning and tool...
Input $/M
$0.50
Output $/M
$3.00
Context
1.0M
Providers
1
Nemotron 3 Nano 30B A3B
nvidia/nemotron-3-nano-30b-a3b
NVIDIA Nemotron 3 Nano 30B A3B is a small language MoE model with highest compute efficiency and accuracy for developers to build specialized agentic AI systems. The model is fully...
Input $/M
$0.05
Output $/M
$0.20
Context
262K
Providers
3
GPT 5.2 Pro
openai/gpt-5.2-pro
GPT-5.2 Pro is OpenAI’s most advanced model, offering major improvements in agentic coding and long context performance over GPT-5 Pro. It is optimized for complex tasks that require step-by-step reasoning,...
Input $/M
$21.00
Output $/M
$168.00
Context
272K
Providers
1
Autoglm Phone 9B Multilingual
zhipu/autoglm-phone-9b-multilingual
Phone Agent is a mobile intelligent assistant framework built on AutoGLM, capable of understanding smartphone screens through multimodal perception and executing automated operations to complete tasks. The system controls devices via ADB (Android Debug Bridge), uses a vision-language model for screen understanding, and leverages intelligent planning to generate and execute action sequences. Users can simply describe tasks in natural language—for example, “Open Xiaohongshu and search for food recommendations.” Phone Agent will automatically parse the intent, understand the current UI, plan the next steps, and carry out the entire workflow. The system also includes: Sensitive action confirmation mechanisms Human-in-the-loop fallback for login or verification code scenarios Remote ADB debugging, allowing device connection via WiFi or network for flexible remote control and development
Input $/M
$0.04
Output $/M
$0.14
Context
66K
Providers
1
GPT 5.1 Codex Max
openai/gpt-5.1-codex-max
GPT-5.1-Codex-Max is OpenAI’s latest agentic coding model, designed for long-running, high-context software development tasks. It is based on an updated version of the 5.1 reasoning stack and trained on agentic...
Input $/M
$1.25
Output $/M
$10.00
Context
272K
Providers
1
DeepSeek OCR
deepseek/deepseek-ocr
DeepSeek-OCR as an initial investigation into the feasibility of compressing long contexts via optical 2D mapping. DeepSeek-OCR consists of two components: DeepEncoder and DeepSeek3B-MoE-A570M as the decoder. Specifically, DeepEncoder serves as the core engine, designed to maintain low activations under high-resolution input while achieving high compression ratios to ensure an optimal and manageable number of vision tokens. Experiments show that when the number of text tokens is within 10 times that of vision tokens (i.e., a compression ratio < 10x), the model can achieve decoding (OCR) precision of 97%. Even at a compression ratio of 20x, the OCR accuracy still remains at about 60%. This shows considerable promise for research areas such as historical long-context compression and memory forgetting mechanisms in LLMs.
Input $/M
$0.03
Output $/M
$0.03
Context
8K
Providers
1
Qwen3 VL 8B Instruct
qwen/qwen3-vl-8b-instruct
Qwen3-VL-8B-Instruct is a multimodal vision-language model from the Qwen3-VL series, built for high-fidelity understanding and reasoning across text, images, and video. It features improved multimodal fusion with Interleaved-MRoPE for long-horizon...
Input $/M
$0.08
Output $/M
$0.50
Context
131K
Providers
1
GPT 5 Pro
openai/gpt-5-pro
GPT-5 Pro is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that require step-by-step reasoning, instruction following, and...
Input $/M
$15.00
Output $/M
$120.00
Context
400K
Providers
1
Qwen3 VL 30B A3B Thinking
qwen/qwen3-vl-30b-a3b-thinking
Qwen3-VL-30B-A3B-Thinking is a multimodal model that unifies strong text generation with visual understanding for images and videos. Its Thinking variant enhances reasoning in STEM, math, and complex tasks. It excels...
Input $/M
$0.20
Output $/M
$1.00
Context
131K
Providers
1
Qwen3 VL 30B A3B Instruct
qwen/qwen3-vl-30b-a3b-instruct
Qwen3-VL-30B-A3B-Instruct is a multimodal model that unifies strong text generation with visual understanding for images and videos. Its Instruct variant optimizes instruction-following for general multimodal tasks. It excels in perception...
Input $/M
$0.15
Output $/M
$0.60
Context
131K
Providers
2
GPT 5 Codex
openai/gpt-5-codex
GPT-5-Codex is a specialized version of GPT-5 optimized for software engineering and coding workflows. It is designed for both interactive development sessions and long, independent execution of complex engineering tasks....
Input $/M
$1.25
Output $/M
$10.00
Context
272K
Providers
1
Qwen3 VL 235B
qwen/qwen3-vl-235b
Open-weight vision model
Input $/M
$0.20
Output $/M
$0.88
Context
131K
Providers
2
Qwen3 Omni 30B A3B Instruct
qwen/qwen3-omni-30b-a3b-instruct
The Qwen-Omni model accepts combined inputs of text and a single additional modality (image, audio, or video) to generate responses in text or speech. It offers a variety of human-like voices, supports speech output in multiple languages and dialects, and is suitable for applications such as text creation, visual recognition, and voice assistants
Input $/M
$0.25
Output $/M
$0.97
Context
66K
Providers
1
Qwen3 Omni 30B A3B Thinking
qwen/qwen3-omni-30b-a3b-thinking
The Qwen-Omni model accepts combined inputs of text and a single additional modality (image, audio, or video) to generate responses in text or speech. It offers a variety of human-like voices, supports speech output in multiple languages and dialects, and is suitable for applications such as text creation, visual recognition, and voice assistants
Input $/M
$0.25
Output $/M
$0.97
Context
66K
Providers
1
Kimi K2 0905
moonshot/kimi-k2-0905
Kimi K2 0905 is the September update of [Kimi K2 0711](moonshotai/kimi-k2). It is a large-scale Mixture-of-Experts (MoE) language model developed by Moonshot AI, featuring 1 trillion total parameters with 32...
Input $/M
$0.60
Output $/M
$2.50
Context
262K
Providers
1
Qwen Mt Plus
qwen/qwen-mt-plus
Qwen-MT is a large language model optimized for machine translation, built upon the foundation of the Tongyi Qianwen model. It supports translation across 92 languages — including Chinese, English, Japanese, Korean, French, Spanish, German, Thai, Indonesian, Vietnamese, Arabic, and more — enabling seamless multilingual communication.
Input $/M
$0.25
Output $/M
$0.75
Context
16K
Providers
1
Kimi K2 Instruct 0905
moonshot/kimi-k2-instruct-0905
Kimi K2 0905 is the September update of Kimi K2 0711. It is a large-scale Mixture-of-Experts (MoE) language model developed by Moonshot AI, featuring 1 trillion total parameters with 32 billion active per forward pass. It supports long-context inference up to 256k tokens, extended from the previous 128k. This update improves agentic coding with higher accuracy and better generalization across scaffolds, and enhances frontend coding with more aesthetic and functional outputs for web, 3D, and related tasks. Kimi K2 is optimized for agentic capabilities, including advanced tool use, reasoning, and code synthesis. It excels across coding (LiveCodeBench, SWE-bench), reasoning (ZebraLogic, GPQA), and tool-use (Tau2, AceBench) benchmarks. The model is trained with a novel stack incorporating the MuonClip optimizer for stable large-scale MoE training.
Input $/M
$0.57
Output $/M
$2.29
Context
262K
Providers
1
GPT 5
openai/gpt-5
GPT-5 is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that require step-by-step reasoning, instruction following, and accuracy...
Input $/M
$1.25
Output $/M
$10.00
Context
272K
Providers
1
GPT OSS 120B Turbo
openai/gpt-oss-120b-turbo
gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. It activates 5.1B parameters per forward pass and is optimized...
Input $/M
$0.15
Output $/M
$0.60
Context
131K
Providers
1
GPT OSS 20B
openai/gpt-oss-20b
gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for...
Input $/M
$0.03
Output $/M
$0.14
Context
131K
Providers
5
Qwen3 Coder 30B A3B Instruct
qwen/qwen3-coder-30b-a3b-instruct
Qwen3-Coder-30B-A3B-Instruct is a 30.5B parameter Mixture-of-Experts (MoE) model with 128 experts (8 active per forward pass), designed for advanced code generation, repository-scale understanding, and agentic tool use. Built on the...
Input $/M
$0.07
Output $/M
$0.27
Context
160K
Providers
1
GLM 4.5 Air FP8
zhipu/glm-4.5-air-fp8
GLM-4.5-Air is the lightweight variant of our latest flagship model family, also purpose-built for agent-centric applications. Like GLM-4.5, it adopts the Mixture-of-Experts (MoE) architecture but with a more compact parameter...
Input $/M
$0.20
Output $/M
$1.10
Context
131K
Providers
1
Qwen3 Coder 480B A35B Instruct Turbo
qwen/qwen3-coder-480b-a35b-instruct-turbo
Qwen3-Coder-480B-A35B-Instruct is the Qwen3's most agentic code model, featuring Significant Performance on Agentic Coding, Agentic Browser-Use and other foundational coding tasks, achieving results comparable to Claude Sonnet.
Input $/M
$0.30
Output $/M
$1.00
Context
262K
Providers
2
Kimi K2 Instruct
moonshot/kimi-k2-instruct
Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model with 32 billion activated parameters and 1 trillion total parameters. Trained with the Muon optimizer, Kimi K2 achieves exceptional performance across frontier knowledge, reasoning, and coding tasks while being meticulously optimized for agentic capabilities.Specifically designed for tool use, reasoning, and autonomous problem-solving.
Input $/M
$0.57
Output $/M
$2.30
Context
131K
Providers
1
Mistral Small 3.2 24B Instruct 2506
mistral/mistral-small-3.2-24b-instruct-2506
Mistral-Small-3.2-24B-Instruct is a drop-in upgrade over the 3.1 release, with markedly better instruction following, roughly half the infinite-generation errors, and a more robust function-calling interface—while otherwise matching or slightly improving on all previous text and vision benchmarks.
Input $/M
$0.08
Output $/M
$0.20
Context
128K
Providers
1
MiniMax M1 80K
minimax/minimax-m1-80k
MiniMax-M1: The World's First Open-Weight, Large-Scale Hybrid Attention Inference Model MiniMax-M1 adopts a Mixture of Experts (MoE) architecture and integrates the Flash Attention mechanism. The model contains a total of 456 billion parameters, with 45.9 billion parameters activated per token. Natively, the M1 model supports a context length of 1 million tokens—8 times that of DeepSeek R1. Additionally, by combining the CISPO algorithm with an efficient hybrid attention design for reinforcement learning training, MiniMax-M1 achieves industry-leading performance in long-context reasoning and real-world software engineering scenarios.
Input $/M
$0.55
Output $/M
$2.20
Context
1M
Providers
1
o3 Pro
openai/o3-pro
The o-series of models are trained with reinforcement learning to think before they answer and perform complex reasoning. The o3-pro model uses more compute to think harder and provide consistently...
Input $/M
$20.00
Output $/M
$80.00
Context
200K
Providers
1
DeepSeek R1 0528 Qwen3 8B
deepseek/deepseek-r1-0528-qwen3-8b
DeepSeek-R1-0528-Qwen3-8B is a high-performance reasoning model based on the Qwen3 8B Base model, enhanced through the integration of DeepSeek-R1-0528's Chain-of-Thought (CoT) optimization. In the AIME 2024 evaluation, this open-source model achieved state-of-the-art (SOTA) performance, delivering a 10% improvement over the original Qwen3 8B while matching the reasoning capabilities of the much larger 235-billion-parameter Qwen3-235B-thinking.
Input $/M
$0.06
Output $/M
$0.09
Context
128K
Providers
1
Gemma 3n E4b It
google/gemma-3n-e4b-it
Gemma 3n E4B-it is optimized for efficient execution on mobile and low-resource devices, such as phones, laptops, and tablets. It supports multimodal inputs—including text, visual data, and audio—enabling diverse tasks...
Input $/M
$0.06
Output $/M
$0.12
Context
33K
Providers
1
Llama Guard 4 12B
meta-llama/llama-guard-4-12b
Llama Guard 4 is a Llama 4 Scout-derived multimodal pretrained model, fine-tuned for content safety classification. Similar to previous versions, it can be used to classify content in both LLM...
Input $/M
$0.18
Output $/M
$0.18
Context
164K
Providers
1
DeepSeek Prover V2 671B
deepseek/deepseek-prover-v2-671b
DeepSeek Launches Open-Source Model DeepSeek-Prover-V2-671B, Specializing in Mathematical Theorem Proving The new model employs a Mixture of Experts (MoE) architecture and is trained using the Lean 4 framework for formal reasoning. With 671 billion parameters, it leverages reinforcement learning and large-scale synthetic data to significantly enhance automated theorem-proving capabilities.
Input $/M
$0.70
Output $/M
$2.50
Context
160K
Providers
1
Qwen3 Next 80B
qwen/qwen3-next-80b
Optimized for speed and efficiency.
Input $/M
$0.35
Output $/M
$1.90
Context
256K
Providers
1
Qwen3 235B A22B FP8
qwen/qwen3-235b-a22b-fp8
Qwen3-235B-A22B is a 235B parameter mixture-of-experts (MoE) model developed by Qwen, activating 22B parameters per forward pass. It supports seamless switching between a "thinking" mode for complex reasoning, math, and...
Input $/M
$0.20
Output $/M
$0.80
Context
41K
Providers
1
Qwen3 30B A3B
qwen/qwen3-30b-a3b
Qwen3, the latest generation in the Qwen large language model series, features both dense and mixture-of-experts (MoE) architectures to excel in reasoning, multilingual support, and advanced agent tasks. Its unique...
Input $/M
$0.12
Output $/M
$0.50
Context
41K
Providers
1
Qwen3 30B A3B FP8
qwen/qwen3-30b-a3b-fp8
Qwen3, the latest generation in the Qwen large language model series, features both dense and mixture-of-experts (MoE) architectures to excel in reasoning, multilingual support, and advanced agent tasks. Its unique...
Input $/M
$0.09
Output $/M
$0.45
Context
41K
Providers
1
Qwen3 32B FP8
qwen/qwen3-32b-fp8
Qwen3-32B is a dense 32.8B parameter causal language model from the Qwen3 series, optimized for both complex reasoning and efficient dialogue. It supports seamless switching between a "thinking" mode for...
Input $/M
$0.10
Output $/M
$0.45
Context
41K
Providers
1
Qwen3 14B
qwen/qwen3-14b
Qwen3-14B is a dense 14.8B parameter causal language model from the Qwen3 series, designed for both complex reasoning and efficient dialogue. It supports seamless switching between a "thinking" mode for...
Input $/M
$0.12
Output $/M
$0.24
Context
41K
Providers
1
Qwen3 8B FP8
qwen/qwen3-8b-fp8
Qwen3-8B is a dense 8.2B parameter causal language model from the Qwen3 series, designed for both reasoning-heavy tasks and efficient dialogue. It supports seamless switching between "thinking" mode for math,...
Input $/M
$0.04
Output $/M
$0.14
Context
128K
Providers
1
Qwen3 4B FP8
qwen/qwen3-4b-fp8
Achieves effective integration of reasoning and non-reasoning modes, allowing seamless switching during conversations. The model delivers state-of-the-art (SOTA) reasoning performance among models of the same scale, with significantly enhanced human preference alignment. Notable improvements are seen in creative writing, role-playing, multi-turn dialogue, and instruction following, leading to a clearly improved user experience.
Input $/M
$0.03
Output $/M
$0.03
Context
128K
Providers
1
o4 Mini
openai/o4-mini
OpenAI o4-mini is a compact reasoning model in the o-series, optimized for fast, cost-efficient performance while retaining strong multimodal and agentic capabilities. It supports tool use and demonstrates competitive reasoning...
Input $/M
$1.10
Output $/M
$4.40
Context
200K
Providers
1
o3
openai/o3
o3 is a well-rounded and powerful model across domains. It sets a new standard for math, science, coding, and visual reasoning tasks. It also excels at technical writing and instruction-following....
Input $/M
$2.00
Output $/M
$8.00
Context
200K
Providers
1
GLM 4 32B 0414
zhipu/glm-4-32b-0414
GLM-4-32B-0414 is the latest open-source model in the GLM series, featuring 32 billion parameters. Its performance is comparable to OpenAI's GPT series and DeepSeek's V3/R1 series, while also supporting highly user-friendly local deployment capabilities. GLM-4-32B-Base-0414 was pre-trained on 15T of high-quality data, including a large amount of reasoning-type synthetic data, which laid a solid foundation for subsequent reinforcement learning extensions. In the post-training stage, in addition to human preference alignment for dialogue scenarios, the research team enhanced the model’s performance in instruction following, engineering code, and function calling using techniques such as rejection sampling and reinforcement learning, thereby strengthening the atomic capabilities required for agent tasks. GLM-4-32B-0414 has achieved strong results in engineering code generation, artifact creation, function calling, search-based question answering, and report generation. On several benchmarks, its performance appr
Input $/M
$0.55
Output $/M
$1.66
Context
32K
Providers
1
GPT 4.1
openai/gpt-4.1
GPT-4.1 is a flagship large language model optimized for advanced instruction following, real-world software engineering, and long-context reasoning. It supports a 1 million token context window and outperforms GPT-4o and...
Input $/M
$2.00
Output $/M
$8.00
Context
1.0M
Providers
1
GPT 4.1 Mini
openai/gpt-4.1-mini
GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost. It retains a 1 million token context window and scores 45.1% on hard...
Input $/M
$0.40
Output $/M
$1.60
Context
1.0M
Providers
1
GPT 4.1 Nano
openai/gpt-4.1-nano
For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a small size with its 1 million...
Input $/M
$0.10
Output $/M
$0.40
Context
1.0M
Providers
1
Llama 3.3 70B
meta-llama/llama-3.3-70b
Input $/M
$0.70
Output $/M
$2.80
Context
128K
Providers
1
o1 Pro
openai/o1-pro
The o1 series of models are trained with reinforcement learning to think before they answer and perform complex reasoning. The o1-pro model uses more compute to think harder and provide...
Input $/M
$150.00
Output $/M
$600.00
Context
200K
Providers
1
Gemma 3 4B It
google/gemma-3-4b-it
Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities,...
Input $/M
$0.05
Output $/M
$0.10
Context
131K
Providers
1
Sonar Pro
perplexity/sonar-pro
Search-grounding flagship (search billed separately per request)
Input $/M
$3.00
Output $/M
$15.00
Context
200K
Providers
1
Sonar Reasoning Pro
perplexity/sonar-reasoning-pro
Search + reasoning (CoT) combined (search billed separately per request)
Input $/M
$2.00
Output $/M
$8.00
Context
128K
Providers
1
Gemini 1.5 Flash
google/gemini-1.5-flash
Gemini 1.5 Flash is Google's foundation model that performs well at a variety of multimodal tasks such as visual understanding, classification, summarization, and creating content from image, audio and video. It's adept at processing visual and text inputs such as photographs, documents, infographics, and screenshots. Gemini 1.5 Flash is designed for high-volume, high-frequency tasks where cost and latency matter.
Input $/M
$0.08
Output $/M
$0.30
Context
1M
Providers
1
DeepSeek V3 Turbo
deepseek/deepseek-v3-turbo
DeepSeek-V3 is the latest model from the DeepSeek team, building upon the instruction following and coding abilities of the previous versions. Pre-trained on nearly 15 trillion tokens, the reported evaluations reveal that the model outperforms other open-source models and rivals leading closed-source models.
Input $/M
$0.40
Output $/M
$1.30
Context
64K
Providers
1
DeepSeek R1 Distill Qwen 32B
deepseek/deepseek-r1-distill-qwen-32b
DeepSeek R1 Distill Qwen 32B is a distilled large language model based on Qwen 2.5 32B, using outputs from DeepSeek R1. It outperforms OpenAI's o1-mini across various benchmarks, achieving new state-of-the-art results for dense models. Other benchmark results include: AIME 2024 pass@1: 72.6 MATH-500 pass@1: 94.3 CodeForces Rating: 1691 The model leverages fine-tuning from DeepSeek R1's outputs, enabling competitive performance comparable to larger frontier models.
Input $/M
$0.30
Output $/M
$0.30
Context
64K
Providers
1
DeepSeek R1 Distill Qwen 14B
deepseek/deepseek-r1-distill-qwen-14b
DeepSeek R1 Distill Qwen 14B is a distilled large language model based on Qwen 2.5 14B, using outputs from DeepSeek R1. It outperforms OpenAI's o1-mini across various benchmarks, achieving new state-of-the-art results for dense models. Other benchmark results include: AIME 2024 pass@1: 69.7 MATH-500 pass@1: 93.9 CodeForces Rating: 1481 The model leverages fine-tuning from DeepSeek R1's outputs, enabling competitive performance comparable to larger frontier models.
Input $/M
$0.15
Output $/M
$0.15
Context
33K
Providers
1
Mistral Small 24B Instruct 2501
mistral/mistral-small-24b-instruct-2501
Mistral Small 3 is a 24B-parameter language model optimized for low-latency performance across common AI tasks. Released under the Apache 2.0 license, it features both pre-trained and instruction-tuned versions designed for efficient local deployment. The model achieves 81% accuracy on the MMLU benchmark and performs competitively with larger models like Llama 3.3 70B and Qwen 32B, while operating at three times the speed on equivalent hardware.
Input $/M
$0.05
Output $/M
$0.08
Context
33K
Providers
1
Sonar
perplexity/sonar
Answers grounded in live web search (search billed separately per request)
Input $/M
$1.00
Output $/M
$1.00
Context
128K
Providers
1
DeepSeek R1 Distill Llama 70B
deepseek/deepseek-r1-distill-llama-70b
DeepSeek R1 Distill Llama 70B is a distilled large language model based on [Llama-3.3-70B-Instruct](/meta-llama/llama-3.3-70b-instruct), using outputs from [DeepSeek R1](/deepseek/deepseek-r1). The model combines advanced distillation techniques to achieve high performance across...
Input $/M
$0.80
Output $/M
$0.80
Context
8K
Providers
2
DeepSeek R1 Turbo
deepseek/deepseek-r1-turbo
DeepSeek R1 is here: Performance on par with [OpenAI o1](/openai/o1), but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active in an inference pass....
Input $/M
$0.70
Output $/M
$2.50
Context
64K
Providers
1
Qwen 2 VL 72B Instruct
qwen/qwen-2-vl-72b-instruct
Qwen2 VL 72B is a multimodal LLM from the Qwen Team with the following key enhancements: SoTA understanding of images of various resolution & ratio: Qwen2-VL achieves state-of-the-art performance on visual understanding benchmarks, including MathVista, DocVQA, RealWorldQA, MTVQA, etc. Understanding videos of 20min+: Qwen2-VL can understand videos over 20 minutes for high-quality video-based question answering, dialog, content creation, etc. Agent that can operate your mobiles, robots, etc.: with the abilities of complex reasoning and decision making, Qwen2-VL can be integrated with devices like mobile phones, robots, etc., for automatic operation based on visual environment and text instructions. Multilingual Support: to serve global users, besides English and Chinese, Qwen2-VL now supports the understanding of texts in different languages inside images, including most European languages, Japanese, Korean, Arabic, Vietnamese, etc.
Input $/M
$0.45
Output $/M
$0.45
Context
33K
Providers
1
o1
openai/o1
The latest and strongest model family from OpenAI, o1 is designed to spend more time thinking before responding. The o1 model series is trained with large-scale reinforcement learning to reason...
Input $/M
$15.00
Output $/M
$60.00
Context
200K
Providers
1
Llama 3.3 70B Instruct
meta-llama/llama-3.3-70b-instruct
The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model...
Input $/M
$0.14
Output $/M
$0.40
Context
6K
Providers
3
Llama 3.3 70B Instruct Turbo
meta-llama/llama-3.3-70b-instruct-turbo
The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model...
Input $/M
$0.10
Output $/M
$0.32
Context
131K
Providers
2
Qwen2.5 7B Instruct Turbo
qwen/qwen2.5-7b-instruct-turbo
Qwen2.5 7B is the latest series of Qwen large language models. Qwen2.5 brings the following improvements upon Qwen2: - Significantly more knowledge and has greatly improved capabilities in coding and...
Input $/M
$0.30
Output $/M
$0.30
Context
33K
Providers
1
Qwen2.5 7B Instruct
qwen/qwen2.5-7b-instruct
Qwen2.5 7B is the latest series of Qwen large language models. Qwen2.5 brings the following improvements upon Qwen2: - Significantly more knowledge and has greatly improved capabilities in coding and...
Input $/M
$0.07
Output $/M
$0.07
Context
32K
Providers
1
Llama 3.2 3B
meta-llama/llama-3.2-3b
Input $/M
$0.15
Output $/M
$0.60
Context
128K
Providers
1
Llama 3.2 3B Instruct
meta-llama/llama-3.2-3b-instruct
Llama 3.2 3B is a 3-billion-parameter multilingual large language model, optimized for advanced natural language processing tasks like dialogue generation, reasoning, and summarization. Designed with the latest transformer architecture, it...
Input $/M
$0.03
Output $/M
$0.05
Context
33K
Providers
1
Llama 3.2 1B Instruct
meta-llama/llama-3.2-1b-instruct
Llama 3.2 1B is a 1-billion-parameter language model focused on efficiently performing natural language tasks, such as summarization, dialogue, and multilingual text analysis. Its smaller size allows it to operate...
Input $/M
$0.02
Output $/M
$0.02
Context
131K
Providers
1
Qwen2.5 72B Instruct
qwen/qwen2.5-72b-instruct
Qwen2.5 72B is the latest series of Qwen large language models. Qwen2.5 brings the following improvements upon Qwen2: - Significantly more knowledge and has greatly improved capabilities in coding and...
Input $/M
$0.36
Output $/M
$0.40
Context
33K
Providers
1
Qwen 2 7B Instruct
qwen/qwen-2-7b-instruct
Qwen2 is the newest series in the Qwen large language model family. Qwen2 7B is a transformer-based model that demonstrates exceptional performance in language understanding, multilingual capabilities, programming, mathematics, and reasoning.
Input $/M
$0.05
Output $/M
$0.05
Context
33K
Providers
1
Mistral Nemo
mistral/mistral-nemo
A 12B parameter model with a 128k token context length built by Mistral in collaboration with NVIDIA. The model is multilingual, supporting English, French, German, Spanish, Italian, Portuguese, Chinese, Japanese, Korean, Arabic, and Hindi. It supports function calling and is released under the Apache 2.0 license.
Input $/M
$0.04
Output $/M
$0.17
Context
60K
Providers
1
Meta Llama 3.1 70B Instruct Turbo
meta-llama/meta-llama-3.1-70b-instruct-turbo
Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 70B instruct-tuned version is optimized for high quality dialogue usecases. It has demonstrated strong...
Input $/M
$0.40
Output $/M
$0.40
Context
131K
Providers
1
Llama 3.1 8B Instruct
meta-llama/llama-3.1-8b-instruct
Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 8B instruct-tuned version is fast and efficient. It has demonstrated strong performance compared to...
Input $/M
$0.02
Output $/M
$0.05
Context
16K
Providers
2
Meta Llama 3.1 8B Instruct Turbo
meta-llama/meta-llama-3.1-8b-instruct-turbo
Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 8B instruct-tuned version is fast and efficient. It has demonstrated strong performance compared to...
Input $/M
$0.02
Output $/M
$0.03
Context
131K
Providers
1
GPT 4o Mini
openai/gpt-4o-mini
GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. As their most advanced small model, it is many multiples more affordable...
Input $/M
$0.15
Output $/M
$0.60
Context
128K
Providers
1
Mistral Nemo Instruct 2407
mistral/mistral-nemo-instruct-2407
12B model trained jointly by Mistral AI and NVIDIA, it significantly outperforms existing models smaller or similar in size.
Input $/M
$0.02
Output $/M
$0.03
Context
131K
Providers
1
Solar Mini
upstage/solar-mini
Korean-optimized (lightweight)
Input $/M
$0.15
Output $/M
$0.15
Context
33K
Providers
1
GPT 4o
openai/gpt-4o
GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligence level of [GPT-4 Turbo](/models/openai/gpt-4-turbo) while being twice as...
Input $/M
$2.50
Output $/M
$10.00
Context
128K
Providers
1
Llama 3 70B Instruct
meta-llama/llama-3-70b-instruct
Meta's latest class of model (Llama 3) launched with a variety of sizes & flavors. This 70B instruct-tuned version was optimized for high quality dialogue usecases. It has demonstrated strong performance compared to leading closed-source models in human evaluations.
Input $/M
$0.51
Output $/M
$0.74
Context
8K
Providers
1
Llama 3 8B Instruct
meta-llama/llama-3-8b-instruct
Meta's latest class of model (Llama 3) launched with a variety of sizes & flavors. This 8B instruct-tuned version was optimized for high quality dialogue usecases. It has demonstrated strong performance compared to leading closed-source models in human evaluations.
Input $/M
$0.04
Output $/M
$0.04
Context
8K
Providers
1
GPT 4 Turbo
openai/gpt-4-turbo
The latest GPT-4 Turbo model with vision capabilities. Vision requests can now use JSON mode and function calling. Training data: up to December 2023.
Input $/M
$10.00
Output $/M
$30.00
Context
128K
Providers
1
text-embedding-3-large
openai/text-embedding-3-large
고정밀 임베딩 · 3072차원
Input $/M
$0.13
Context
8K
Providers
1
text-embedding-3-small
openai/text-embedding-3-small
범용 임베딩 · 1536차원
Input $/M
$0.02
Context
8K
Providers
1
GPT 3.5 Turbo 16K
openai/gpt-3.5-turbo-16k
This model offers four times the context length of gpt-3.5-turbo, allowing it to support approximately 20 pages of text in a single request at a higher cost. Training data: up...
Input $/M
$3.00
Output $/M
$4.00
Context
16K
Providers
1
GPT 4
openai/gpt-4
OpenAI's flagship model, GPT-4 is a large-scale multimodal language model capable of solving difficult problems with greater accuracy than previous models due to its broader general knowledge and advanced reasoning...
Input $/M
$30.00
Output $/M
$60.00
Context
8K
Providers
1
GPT 3.5 Turbo
openai/gpt-3.5-turbo
GPT-3.5 Turbo is OpenAI's fastest model. It can understand and generate natural language or code, and is optimized for chat and traditional completion tasks. Training data up to Sep 2021.
Input $/M
$0.50
Output $/M
$1.50
Context
16K
Providers
1
Whisper
openai/whisper-1
$0.006/min (STT)
Price (per min)
$0.006/min
Context
—
Providers
1
Whisper Large V3
openai/whisper-large-v3
Price (per min)
$0.00045/min
Context
—
Providers
2
Whisper Large V3 Turbo
openai/whisper-large-v3-turbo
Price (per min)
$0.0001998/min
Context
—
Providers
1
Gemini 3.1 Flash Image
google/gemini-3.1-flash-image
Price (per image)
~$0.05/image
Context
—
Providers
1
Gemini 3 Pro Image
google/gemini-3-pro-image
Price (per image)
~$0.13/image
Context
—
Providers
1
Embeddinggemma 300M
google/embeddinggemma-300m
Input $/M
<$0.01
Context
2K
Providers
1
Qwen3 Embedding 0.6B
qwen/qwen3-embedding-0.6b
Input $/M
$0.01
Context
33K
Providers
2
Qwen3 Embedding 4B
qwen/qwen3-embedding-4b
Input $/M
$0.02
Context
33K
Providers
1
Qwen3 Embedding 8B
qwen/qwen3-embedding-8b
Input $/M
$0.01
Context
33K
Providers
2
Voxtral Mini 3B 2507
mistral/voxtral-mini-3b-2507
Price (per min)
$0.001/min
Context
—
Providers
1
Voxtral Small 24B 2507
mistral/voxtral-small-24b-2507
Price (per min)
$0.003/min
Context
—
Providers
1
Nemotron 3.5 Asr Streaming 0.6B
nvidia/nemotron-3.5-asr-streaming-0.6b
Price (per min)
$0.0015/min
Context
—
Providers
1
Nemotron 3.5 Asr Streaming Multilingual 0.6B
nvidia/nemotron-3.5-asr-streaming-multilingual-0.6b
Price (per min)
$0.0001998/min
Context
—
Providers
1
Nemotron 3 Asr Streaming 0.6B
nvidia/nemotron-3-asr-streaming-0.6b
Price (per min)
$0.0015/min
Context
—
Providers
1
Parakeet Tdt 0.6B V3
nvidia/parakeet-tdt-0.6b-v3
Price (per min)
$0.0015/min
Context
—
Providers
1
Llama Nemotron Embed VL 1B V2
nvidia/llama-nemotron-embed-vl-1b-v2
Input $/M
$0.01
Context
10K
Providers
2
Cobuddy
baidu/cobuddy
CoBuddy is a specialized code generation model developed by Baidu, engineered with targeted optimizations for Coding and AI Agent scenarios. It delivers exceptional performance, characterized by high inference throughput and ultra-low end-to-end (E2E) latency.
Input $/M
$0.28
Output $/M
$1.13
Context
131K
Providers
1
Lfm2.5 8B A1B
liquidai/lfm2.5-8b-a1b
Input $/M
$0.03
Output $/M
$0.12
Context
128K
Providers
1
Step 3.7 Flash
stepfun/step-3.7-flash
Step 3.7 Flash is StepFun's latest high-efficiency multimodal Mixture-of-Experts model. It pairs a 196B-parameter language backbone with a vision encoder for native image and video understanding, activating roughly 11B parameters...
Input $/M
$0.20
Output $/M
$1.15
Context
262K
Providers
2
Ring 2.6 1T
inclusionai/ring-2.6-1t
Ring-2.6-1T is a 1T-parameter-scale thinking model with 63B active parameters, built for real-world agent workflows that require both strong capability and operational efficiency. It is optimized for coding agents, tool...
Input $/M
$0.30
Output $/M
$2.50
Context
262K
Providers
1
Seed 2.0 Code
bytedance/seed-2.0-code
A coding model optimized for real-world development environments, with reliable tool use in common IDEs such as Claude Code. It delivers strong front-end performance and supports Skills.
Input $/M
$0.50
Output $/M
$3.00
Context
256K
Providers
1
Ling 2.6 1T
inclusionai/ling-2.6-1t
Ling-2.6-1T is an instant (instruct) model from inclusionAI and the company’s trillion-parameter flagship, designed for real-world agents that require fast execution and high efficiency at scale. It uses a “fast...
Input $/M
$0.30
Output $/M
$2.50
Context
262K
Providers
1
Ling 2.6 Flash
inclusionai/ling-2.6-flash
Ling-2.6-flash is an instant (instruct) model from inclusionAI with 104B total parameters and 7.4B active parameters, designed for real-world agents that require fast responses, strong execution, and high token efficiency....
Input $/M
$0.10
Output $/M
$0.30
Context
262K
Providers
1
Seed 2.0 Pro
bytedance/seed-2.0-pro
Built for the Agent era, it delivers stable performance in complex reasoning and long-horizon tasks, including multi-step planning, visual-text reasoning, video understanding, and advanced analysis.
Input $/M
$0.50
Output $/M
$3.00
Context
256K
Providers
1
Kat Coder Pro V2
kwaipilot/kat-coder-pro-v2
KAT-Coder-Pro V2 is the latest high-performance model in KwaiKAT’s KAT-Coder series, designed for complex enterprise-grade software engineering and SaaS integration. It builds on the agentic coding strengths of earlier versions,...
Input $/M
$0.30
Output $/M
$1.20
Context
262K
Providers
1
Seed 2.0 Mini
bytedance/seed-2.0-mini
Built for low-latency, high-concurrency, cost-sensitive use cases, with flexible deployment, four-tier thinking, and multimodal
Input $/M
$0.10
Output $/M
$0.40
Context
262K
Providers
2
Seed 1.8
bytedance/seed-1.8
Optimized specifically for multimodal agent scenarios. It features enhanced agent capabilities, upgraded multimodal comprehension, and more flexible context management.
Input $/M
$0.25
Output $/M
$2.00
Context
256K
Providers
1
Voyage 4
voyage/voyage-4
범용 임베딩 · 1024차원(가변) · voyage-4 계열은 임베딩 공간 공유
Input $/M
$0.06
Context
32K
Providers
1
Voyage 4 Large
voyage/voyage-4-large
최고 품질·다국어 · 1024차원(가변) · text-embedding-3-large보다 저렴
Input $/M
$0.12
Context
32K
Providers
1
Voyage 4 Lite
voyage/voyage-4-lite
지연·비용 최적 · 1024차원(256/512/2048 가변)
Input $/M
$0.02
Context
32K
Providers
1
Kat Coder Pro
kwaipilot/kat-coder-pro
KAT-Coder-Pro V2 by KwaiKAT is a non-reasoning model optimized for agentic coding. It delivers strong performance on reasoning-style tasks while requiring significantly fewer output tokens than peer models. With the 1210 release, it achieved a score of 64 on the Artificial Analysis Intelligence Index, placing it in the global Top 10 and ranking first among all non-reasoning models.
Input $/M
$0.30
Output $/M
$1.20
Context
256K
Providers
1
K EXAONE 236B A23B
lgai-exaone/k-exaone-236b-a23b
Input $/M
$0.20
Output $/M
$0.80
Context
—
Providers
1
MiMo V2 Flash
xiaomimimo/mimo-v2-flash
Xiaomi MiMo-V2-Flash is a proprietary MoE model developed by Xiaomi, designed for extreme inference efficiency with 309B total parameters (15B active). By incorporating an innovative Hybrid attention architecture and multi-layer MTP inference acceleration, it ranks among the top 2 global open-source models across multiple Agent benchmarks. Its coding capabilities surpass all open-source models and rival the industry-leading closed-source model, Claude 4.5 Sonnet—yet at only 2.5% of the inference cost and with 2x the generation speed, successfully pushing the limits of both model performance and efficiency.
Input $/M
$0.11
Output $/M
$0.33
Context
262K
Providers
1
Cogito V2 1 671B
deepcogito/cogito-v2-1-671b
Cogito v2.1 671B MoE represents one of the strongest open models globally, matching performance of frontier closed and open models. This model is trained using self play with reinforcement learning...
Input $/M
$1.25
Output $/M
$1.25
Context
164K
Providers
1
ERNIE 4.5 VL 28B A3B Thinking
baidu/ernie-4.5-vl-28b-a3b-thinking
Built upon the powerful ERNIE-4.5-VL-28B-A3B architecture, the newly upgraded ERNIE-4.5-VL-28B-A3B-Thinking achieves a remarkable leap forward in multimodal reasoning capabilities. 🧠✨ Through an extensive mid-training phase, the model absorbed a vast and highly diverse corpus of premium visual-language reasoning data. This massive-scale training process dramatically boosted the model’s representation power while deepening the semantic alignment between visual and language modalities—unlocking unprecedented capabilities in nuanced visual-textual reasoning. 📊 The model leverages cutting-edge multimodal reinforcement learning techniques on verifiable tasks, integrating GSPO and IcePop strategies to stabilize MoE training combined with dynamic difficulty sampling for exceptional learning efficiency. ⚡ Responding to strong community demand, we’ve significantly strengthened the model’s grounding performance with improved instruction-following capabilities, making visual grounding functions more accessible than eve
Input $/M
$0.39
Output $/M
$0.39
Context
131K
Providers
1
ERNIE 4.5 21B A3B Thinking
baidu/ernie-4.5-21b-a3b-thinking
ERNIE-4.5-21B-A3B-Thinking is a text-based Mixture of Experts (MoE) post-training model featuring 21B total parameters with 3B active parameters per token. It delivers enhanced performance on reasoning tasks, including logical reasoning, mathematics, science, coding, text generation, and academic benchmarks that typically require human expertise. The model offers efficient tool utilization capabilities and supports up to 128K tokens for long-context understanding.
Input $/M
$0.07
Output $/M
$0.28
Context
131K
Providers
1
Morph V3 Large
morph/morph-v3-large
Morph's high-accuracy apply model for complex code edits. ~4,500 tokens/sec with 98% accuracy for precise code transformations. The model requires the prompt to be in the following format: <instruction>{instruction}</instruction> <code>{initial_code}</code>...
Input $/M
$0.90
Output $/M
$1.90
Context
16K
Providers
1
Morph V3 Fast
morph/morph-v3-fast
Morph's fastest apply model for code edits. ~10,500 tokens/sec with 96% accuracy for rapid code transformations. The model requires the prompt to be in the following format: <instruction>{instruction}</instruction> <code>{initial_code}</code> <update>{edit_snippet}</update>...
Input $/M
$0.80
Output $/M
$1.20
Context
16K
Providers
1
ERNIE 4.5 VL 424B A47B
baidu/ernie-4.5-vl-424b-a47b
ERNIE-4.5-VL-424B-A47B is a multimodal Mixture-of-Experts (MoE) model from Baidu’s ERNIE 4.5 series, featuring 424B total parameters with 47B active per token. It is trained jointly on text and image data...
Input $/M
$0.42
Output $/M
$1.25
Context
123K
Providers
1
ERNIE 4.5 21B A3B
baidu/ernie-4.5-21b-a3b
The ERNIE 4.5 series of open-source models adopts a Mixture-of-Experts (MoE) architecture, representing an innovative multimodal heterogeneous model structure. It achieves cross-modal knowledge fusion through a parameter-sharing mechanism while retaining dedicated parameter spaces for individual modalities. This architecture is particularly well-suited for the continuous pre-training paradigm from large language models to multimodal models, significantly enhancing multimodal understanding capabilities while maintaining or even improving performance in text-based tasks. The models are efficiently trained, inferred, and deployed using the PaddlePaddle deep learning framework. During the pre-training of large language models, the Model FLOPs Utilization (MFU) reaches 47%. Experimental results demonstrate that this series of models achieves state-of-the-art (SOTA) performance across multiple text and multimodal benchmarks, with particularly outstanding results in instruction following, world knowledge memorizatio
Input $/M
$0.07
Output $/M
$0.28
Context
120K
Providers
1
ERNIE 4.5 300B A47B Paddle
baidu/ernie-4.5-300b-a47b-paddle
The ERNIE 4.5 series of open-source models adopts a Mixture-of-Experts (MoE) architecture, representing an innovative multimodal heterogeneous model structure. It achieves cross-modal knowledge fusion through a parameter-sharing mechanism while retaining dedicated parameter spaces for individual modalities. This architecture is particularly well-suited for the continuous pre-training paradigm from large language models to multimodal models, significantly enhancing multimodal understanding capabilities while maintaining or even improving performance in text-based tasks. The models are efficiently trained, inferred, and deployed using the PaddlePaddle deep learning framework. During the pre-training of large language models, the Model FLOPs Utilization (MFU) reaches 47%. Experimental results demonstrate that this series of models achieves state-of-the-art (SOTA) performance across multiple text and multimodal benchmarks, with particularly outstanding results in instruction following, world knowledge memorizatio
Input $/M
$0.28
Output $/M
$1.10
Context
123K
Providers
1
ERNIE 4.5 VL 28B A3B
baidu/ernie-4.5-vl-28b-a3b
The ERNIE 4.5 series of open-source models adopts a Mixture-of-Experts (MoE) architecture, representing an innovative multimodal heterogeneous model structure. It achieves cross-modal knowledge fusion through a parameter-sharing mechanism while retaining dedicated parameter spaces for individual modalities. This architecture is particularly well-suited for the continuous pre-training paradigm from large language models to multimodal models, significantly enhancing multimodal understanding capabilities while maintaining or even improving performance in text-based tasks. The models are efficiently trained, inferred, and deployed using the PaddlePaddle deep learning framework. During the pre-training of large language models, the Model FLOPs Utilization (MFU) reaches 47%. Experimental results demonstrate that this series of models achieves state-of-the-art (SOTA) performance across multiple text and multimodal benchmarks, with particularly outstanding results in instruction following, world knowledge memorizatio
Input $/M
$0.14
Output $/M
$0.56
Context
30K
Providers
1
Phi 4
microsoft/phi-4
[Microsoft Research](/microsoft) Phi-4 is designed to perform well in complex reasoning tasks and can operate efficiently in situations with limited memory or where quick responses are needed. At 14 billion...
Input $/M
$0.07
Output $/M
$0.14
Context
16K
Providers
1
Voyage Code 3
voyage/voyage-code-3
코드 검색 특화 임베딩 · 1024차원(가변)
Input $/M
$0.18
Context
32K
Providers
1
Wizardlm 2 8x22b
microsoft/wizardlm-2-8x22b
WizardLM-2 8x22B is Microsoft AI's most advanced Wizard model. It demonstrates highly competitive performance compared to leading proprietary models, and it consistently outperforms all existing state-of-the-art opensource models. It is...
Input $/M
$0.62
Output $/M
$0.62
Context
66K
Providers
1
Bge M3
baai/bge-m3
Input $/M
$0.01
Context
8K
Providers
2
Bge En Icl
baai/bge-en-icl
Input $/M
$0.01
Context
8K
Providers
2
Nova 3 En
deepgram/nova-3-en
Price (per min)
$0.0015/min
Context
—
Providers
1
Nova 3 Multi
deepgram/nova-3-multi
Price (per min)
$0.0015/min
Context
—
Providers
1
Flux
deepgram/flux
Price (per min)
$0.0015/min
Context
—
Providers
1
Multilingual E5 Large Instruct
intfloat/multilingual-e5-large-instruct
Input $/M
$0.01
Context
1K
Providers
3

