Warp

Models

Search and compare to find the model you need.

Showing 268 of 268

Claude Fable 5

anthropic/claude-fable-5

Top-tier default model

Input $/M

$10.00

Output $/M

$50.00

Context

500K

Providers

1

GPT-5.6 Sol

openai/gpt-5.6-sol

GPT-5.6 top line

Input $/M

$5.00

Output $/M

$30.00

Context

1.1M

Providers

1

Claude Opus 4.8

anthropic/claude-opus-4.8

Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token...

Input $/M

$5.00

Output $/M

$25.00

Context

1M

Providers

1

GPT-5.5

openai/gpt-5.5

GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and improved token efficiency on hard tasks. It features a 1M+ token...

Input $/M

$5.00

Output $/M

$30.00

Context

1.1M

Providers

1

Claude Sonnet 5

anthropic/claude-sonnet-5

Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports adaptive thinking with selectable reasoning effort levels (low, medium, high, max,...

Input $/M

$2.00

Output $/M

$10.00

Context

1M

Providers

1

Claude Opus 4.7

anthropic/claude-opus-4-7

Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.6, it delivers stronger performance on...

Input $/M

$5.00

Output $/M

$25.00

Context

1M

Providers

1

GLM-5.2

zhipu/glm-5.2

Opus-class on coding benchmarks, 46% of the cost with caching

Input $/M

$0.85

Output $/M

$2.68

Context

1.0M

Providers

14

GPT-5.4

openai/gpt-5.4

GPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system. It features a 1M+ token context window (922K input, 128K output) with support for...

Input $/M

$2.50

Output $/M

$15.00

Context

400K

Providers

1

Grok 4.5

xai/grok-4.5

Grok 4.5 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.

Input $/M

$2.00

Output $/M

$6.00

Context

500K

Providers

1

Claude Sonnet 4.6

anthropic/claude-sonnet-4-6

Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at iterative development, complex codebase navigation, end-to-end project management with...

Input $/M

$3.00

Output $/M

$15.00

Context

1M

Providers

1

GLM 5.1

zhipu/glm-5.1

GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. Unlike previous models built around minute-level interactions, GLM-5.1 can work independently and continuously on...

Input $/M

$1.05

Output $/M

$3.50

Context

205K

Providers

11

Gemini 3.5 Flash

google/gemini-3.5-flash

Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution...

Input $/M

$1.50

Output $/M

$9.00

Context

1.0M

Providers

2

Kimi K3

moonshot/kimi-k3

Moonshot 플래그십 · 1M 컨텍스트

Input $/M

$3.00

Output $/M

$15.00

Context

1.0M

Providers

3

Gemini 3.1 Pro

google/gemini-3.1-pro

Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows. Building on the multimodal foundation...

Input $/M

$2.00

Output $/M

$12.00

Context

1.0M

Providers

2

Gemini 3.1 Pro Preview

google/gemini-3.1-pro-preview

Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows. Building on the multimodal foundation...

Input $/M

$2.00

Output $/M

$12.00

Context

1.0M

Providers

1

Grok 4.20

xai/grok-4.20

2M context

Input $/M

$1.25

Output $/M

$2.50

Context

2M

Providers

1

Gemini 3 Flash

google/gemini-3-flash

Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It delivers near Pro level reasoning and tool...

Input $/M

$0.50

Output $/M

$3.00

Context

1.0M

Providers

2

Kimi K2.7 Code

moonshot/kimi-k2.7-code

Coding-specialized

Input $/M

$0.74

Output $/M

$3.50

Context

262K

Providers

8

MiMo V2.5 Pro

xiaomimimo/mimo-v2.5-pro

MiMo-V2.5-Pro is purpose-built to push the boundaries of complex software engineering and extreme long-horizon tasks. Compared to its predecessor, it achieves a comprehensive leap in general agentic capabilities, advancing the human-AI collaboration paradigm toward true "autonomous delivery." Without human intervention, it stably orchestrates massive workflows requiring up to a thousand tool calls in a single session, not only precisely capturing implicit requirements within ultra-long contexts but also demonstrating exceptional global architectural planning and self-correction discipline. In core agentic scenarios and long-horizon complexities, MiMo-V2.5-Pro is fully equipped to go head-to-head with top-tier global models like Claude Opus 4.6 and GPT-5.4. Backed by this exceptionally high execution confidence and long-term logical consistency, it completely sheds the "co-pilot" label, ready to take on truly serious, professional-grade workloads in real-world business environments.

Input $/M

$1.00

Output $/M

$3.00

Context

1.0M

Providers

4

Kimi K2.6

moonshot/kimi-k2.6

Kimi K2.6 is Moonshot AI's next-generation multimodal model, designed for long-horizon coding, coding-driven UI/UX generation, and multi-agent orchestration. It handles complex end-to-end coding tasks across Python, Rust, and Go, and...

Input $/M

$0.66

Output $/M

$3.41

Context

262K

Providers

11

Qwen3.7 Plus

qwen/qwen3.7-plus

Qwen3.7-Plus is a cost-effective model in Alibaba's Qwen3.7 series. It supports text and image input with text output, building on the series' text capabilities with a comprehensive upgrade to its...

Input $/M

$0.32

Output $/M

$1.28

Context

1M

Providers

1

Qwen3.6 Max Preview

qwen/qwen3.6-max-preview

Qwen3.6-Max-Preview is a proprietary frontier model from Alibaba Cloud built on a sparse mixture-of-experts architecture with approximately 1 trillion total parameters. It is optimized for agentic coding, tool use, and...

Input $/M

$1.30

Output $/M

$7.80

Context

262K

Providers

1

DeepSeek V4 Pro

deepseek/deepseek-v4-pro

DeepSeek flagship · 1M context

Input $/M

$1.30

Output $/M

$2.60

Context

1.0M

Providers

8

GLM-5

zhipu/glm-5

GLM flagship

Input $/M

$0.60

Output $/M

$2.08

Context

1.0M

Providers

7

GPT-5.4 mini

openai/gpt-5.4-mini

GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image inputs with strong performance across reasoning, coding,...

Input $/M

$0.75

Output $/M

$4.50

Context

400K

Providers

1

MiMo V2 Pro

xiaomimimo/mimo-v2-pro

Xiaomi’s MiMo-V2-Pro is a flagship, >1T parameter foundation model with a 1M context window, specifically optimized for advanced agentic workflows. Highly compatible with frameworks like OpenClaw, it secures top-tier global rankings on PinchBench and ClawBench, with a perceived capability approaching Opus 4.6. Built to act as the cognitive core of intelligent systems, MiMo-V2-Pro seamlessly orchestrates complex tasks, drives engineering automation, and consistently delivers dependable real-world outcomes.

Input $/M

$2.00

Output $/M

$6.00

Context

1.0M

Providers

1

Gemini 2.5 Pro

google/gemini-2.5-pro

Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy...

Input $/M

$1.25

Output $/M

$10.00

Context

1M

Providers

1

MiniMax M3

minimax/minimax-m3

MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...

Input $/M

$0.30

Output $/M

$1.20

Context

1.0M

Providers

6

Qwen3.6 Plus

qwen/qwen3.6-plus

Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong scalability and high-performance inference. Compared to the 3.5 series, it delivers...

Input $/M

$0.50

Output $/M

$3.00

Context

1M

Providers

3

Grok 4.3

xai/grok-4.3

Grok 4.3 is a reasoning model from xAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual...

Input $/M

$1.25

Output $/M

$2.50

Context

1M

Providers

1

Qwen3.5 397B A17B

qwen/qwen3.5-397b-a17b

The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. It delivers...

Input $/M

$0.45

Output $/M

$3.00

Context

262K

Providers

6

GLM 4.7

zhipu/glm-4.7

GLM-4.7 is Z.ai’s latest flagship model, featuring upgrades in two key areas: enhanced programming capabilities and more stable multi-step reasoning/execution. It demonstrates significant improvements in executing complex agent tasks while...

Input $/M

$0.40

Output $/M

$1.75

Context

205K

Providers

6

Inkling

thinkingmachines/inkling

Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...

Input $/M

$1.00

Output $/M

$4.05

Context

524K

Providers

3

GLM 5v Turbo

zhipu/glm-5v-turbo

GLM-5V-Turbo is Z.ai’s first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks. It natively handles image, video, and text inputs, excels at long-horizon planning, complex coding,...

Input $/M

$1.20

Output $/M

$4.00

Context

205K

Providers

1

GPT-5.1

openai/gpt-5.1

GPT-5.1 is the latest frontier-grade model in the GPT-5 series, offering stronger general-purpose reasoning, improved instruction adherence, and a more natural conversational style compared to GPT-5. It uses adaptive reasoning...

Input $/M

$1.25

Output $/M

$10.00

Context

400K

Providers

1

DeepSeek V4 Flash

deepseek/deepseek-v4-flash

DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and...

Input $/M

$0.09

Output $/M

$0.18

Context

1M

Providers

6

GPT-5.2

openai/gpt-5.2

Top-tier fallback model

Input $/M

$1.75

Output $/M

$14.00

Context

400K

Providers

1

Qwen3 Max

qwen/qwen3-max

Qwen flagship

Input $/M

$1.20

Output $/M

$6.00

Context

262K

Providers

2

Gemini 3.1 Flash Lite

google/gemini-3.1-flash-lite

Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image, video, audio, and PDF inputs, and is designed for lightweight agentic...

Input $/M

$0.25

Output $/M

$1.50

Context

1.0M

Providers

2

Gemini 3.1 Flash Lite Preview

google/gemini-3.1-flash-lite-preview

Gemini 3.1 Flash Lite Preview is Google's high-efficiency model optimized for high-volume use cases. It outperforms Gemini 2.5 Flash Lite on overall quality and approaches Gemini 2.5 Flash performance across...

Input $/M

$0.25

Output $/M

$1.50

Context

1.0M

Providers

1

MiMo V2.5

xiaomimimo/mimo-v2.5

Xiaomi's native omni-modal model (text / image / video / audio understanding) with Pro-level agentic performance at ~half the inference cost and a 1M-token context — built for cost-efficient, perception-rich agent workflows.

Input $/M

$0.40

Output $/M

$2.00

Context

1.0M

Providers

5

Kimi K2 Thinking

moonshot/kimi-k2-thinking

Reasoning (Thinking) specialized

Input $/M

$0.47

Output $/M

$2.00

Context

262K

Providers

3

DeepSeek V3.2

deepseek/deepseek-v3.2

DeepSeek-V3.2 is a large language model designed to harmonize high computational efficiency with strong reasoning and agentic tool-use performance. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism...

Input $/M

$0.27

Output $/M

$0.40

Context

164K

Providers

6

GLM 4.6

zhipu/glm-4.6

Compared with GLM-4.5, this generation brings several key improvements: Longer context window: The context window has been expanded from 128K to 200K tokens, enabling the model to handle more complex...

Input $/M

$0.50

Output $/M

$2.00

Context

205K

Providers

4

DeepSeek V3.2 Exp

deepseek/deepseek-v3.2-exp

DeepSeek-V3.2-Exp is an experimental large language model released by DeepSeek as an intermediate step between V3.1 and future architectures. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism...

Input $/M

$0.27

Output $/M

$0.41

Context

164K

Providers

1

Qwen3 235B

qwen/qwen3-235b

Qwen3-235B-A22B-Instruct-2507 is a multilingual, instruction-tuned mixture-of-experts language model based on the Qwen3-235B architecture, with 22B active parameters per forward pass. It is optimized for general-purpose text generation, including instruction following,...

Input $/M

$0.09

Output $/M

$0.60

Context

262K

Providers

6

DeepSeek R1 0528

deepseek/deepseek-r1-0528

May 28th update to the [original DeepSeek R1](/deepseek/deepseek-r1) Performance on par with [OpenAI o1](/openai/o1), but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active...

Input $/M

$0.50

Output $/M

$2.15

Context

164K

Providers

3

DeepSeek V3.1

deepseek/deepseek-v3.1

DeepSeek-V3.1 is a hybrid model that supports both thinking mode and non-thinking mode.DeepSeek-V3.1 is post-trained on the top of DeepSeek-V3.1-Base, which is built upon the original V3 base checkpoint through a two-phase long context extension approach, following the methodology outlined in the original DeepSeek-V3 report. We have expanded our dataset by collecting additional long documents and substantially extending both training phases. The 32K extension phase has been increased 10-fold to 630B tokens, while the 128K extension phase has been extended by 3.3x to 209B tokens.

Input $/M

$0.25

Output $/M

$0.95

Context

164K

Providers

3

MiniMax M2.7

minimax/minimax-m2.7

MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built to actively participate in its own evolution, M2.7 integrates advanced agentic capabilities through multi-agent...

Input $/M

$0.25

Output $/M

$1.00

Context

205K

Providers

6

Qwen3.5 122B A10B

qwen/qwen3.5-122b-a10b

The Qwen3.5 122B-A10B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. In terms of...

Input $/M

$0.29

Output $/M

$2.40

Context

262K

Providers

3

DeepSeek V3.1 Terminus

deepseek/deepseek-v3.1-terminus

DeepSeek-V3.1 Terminus is an update to [DeepSeek V3.1](/deepseek/deepseek-chat-v3.1) that maintains the model's original capabilities while addressing issues reported by users, including language consistency and agent capabilities, further optimizing the model's...

Input $/M

$0.27

Output $/M

$0.95

Context

164K

Providers

2

Claude Haiku 4.5

anthropic/claude-haiku-4.5

Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance...

Input $/M

$1.00

Output $/M

$5.00

Context

200K

Providers

1

GLM 4.5

zhipu/glm-4.5

GLM-4.5 is our latest flagship foundation model, purpose-built for agent-based applications. It leverages a Mixture-of-Experts (MoE) architecture and supports a context length of up to 128k tokens. GLM-4.5 delivers significantly...

Input $/M

$0.60

Output $/M

$2.20

Context

131K

Providers

2

Gemini 2.5 Flash

google/gemini-2.5-flash

Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks. It includes built-in "thinking" capabilities, enabling it to provide responses with greater...

Input $/M

$0.30

Output $/M

$2.50

Context

1.0M

Providers

2

Qwen3.5 27B

qwen/qwen3.5-27b

The Qwen3.5 27B native vision-language Dense model incorporates a linear attention mechanism, delivering fast response times while balancing inference speed and performance. Its overall capabilities are comparable to those of...

Input $/M

$0.26

Output $/M

$2.60

Context

262K

Providers

4

GPT-5.3 Codex

openai/gpt-5.3-codex

Coding-specialized (Codex)

Input $/M

$1.75

Output $/M

$14.00

Context

400K

Providers

1

GPT-5.4 nano

openai/gpt-5.4-nano

GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks. It supports text and image inputs and is designed for low-latency...

Input $/M

$0.20

Output $/M

$1.25

Context

400K

Providers

1

Qwen3 Next 80B A3B Instruct

qwen/qwen3-next-80b-a3b-instruct

Qwen3-Next-80B-A3B-Instruct is an instruction-tuned chat model in the Qwen3-Next series optimized for fast, stable responses without “thinking” traces. It targets complex tasks across reasoning, code generation, knowledge QA, and multilingual...

Input $/M

$0.09

Output $/M

$1.10

Context

262K

Providers

3

Qwen3 235B A22B Thinking 2507

qwen/qwen3-235b-a22b-thinking-2507

Qwen3-235B-A22B-Thinking-2507 is a high-performance, open-weight Mixture-of-Experts (MoE) language model optimized for complex reasoning tasks. It activates 22B of its 235B parameters per forward pass and natively supports up to 262,144...

Input $/M

$0.23

Output $/M

$2.30

Context

262K

Providers

5

DeepSeek R1

deepseek/deepseek-r1

DeepSeek R1 is here: Performance on par with [OpenAI o1](/openai/o1), but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active in an inference pass....

Input $/M

$4.00

Output $/M

$4.00

Context

66K

Providers

1

Qwen3.5 35B A3B

qwen/qwen3.5-35b-a3b

The Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture that integrates linear attention mechanisms and a sparse mixture-of-experts model, achieving higher inference efficiency. Its overall...

Input $/M

$0.14

Output $/M

$1.00

Context

262K

Providers

5

DeepSeek V3 0324

deepseek/deepseek-v3-0324

DeepSeek V3, a 685B-parameter, mixture-of-experts model, is the latest iteration of the flagship chat model family from the DeepSeek team.

Input $/M

$0.24

Output $/M

$0.90

Context

164K

Providers

3

Qwen3 VL 235B A22B Thinking

qwen/qwen3-vl-235b-a22b-thinking

Qwen3-VL-235B-A22B Thinking is a multimodal model that unifies strong text generation with visual understanding across images and video. The Thinking model is optimized for multimodal reasoning in STEM and math....

Input $/M

$0.98

Output $/M

$3.95

Context

131K

Providers

1

MiniMax M2.5

minimax/minimax-m2.5

MiniMax-M2.5 is a SOTA large language model designed for real-world productivity. Trained in a diverse range of complex real-world digital working environments, M2.5 builds upon the coding expertise of M2.1...

Input $/M

$0.30

Output $/M

$1.20

Context

205K

Providers

6

GPT-5 mini

openai/gpt-5-mini

GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following and safety-tuning benefits as GPT-5, but with reduced latency and cost....

Input $/M

$0.25

Output $/M

$2.00

Context

400K

Providers

1

Qwen3 Coder 480B A35B Instruct

qwen/qwen3-coder-480b-a35b-instruct

Qwen3-Coder-480B-A35B-Instruct is a cutting-edge open coding model from Qwen, matching Claude Sonnet’s performance in agentic programming, browser automation, and core development tasks. With native 256K context (extendable to 1M tokens via YaRN), it excels at repository-scale analysis and features specialized function-call support for platforms like Qwen Code and CLINE—making it ideal for complex, real-world development workflows.

Input $/M

$0.38

Output $/M

$1.55

Context

262K

Providers

3

GLM 4.6v

zhipu/glm-4.6v

GLM-4.6V is a large multimodal model designed for high-fidelity visual understanding and long-context reasoning across images, documents, and mixed media. It supports up to 128K tokens, processes complex page layouts...

Input $/M

$0.30

Output $/M

$0.90

Context

131K

Providers

1

GLM 4.5 Air

zhipu/glm-4.5-air

GLM-4.5-Air is the lightweight variant of our latest flagship model family, also purpose-built for agent-centric applications. Like GLM-4.5, it adopts the Mixture-of-Experts (MoE) architecture but with a more compact parameter...

Input $/M

$0.13

Output $/M

$0.85

Context

131K

Providers

2

Qwen3 Next 80B A3B Thinking

qwen/qwen3-next-80b-a3b-thinking

Qwen3-Next-80B-A3B-Thinking is a reasoning-first chat model in the Qwen3-Next line that outputs structured “thinking” traces by default. It’s designed for hard multi-step problems; math proofs, code synthesis/debugging, logic, and agentic...

Input $/M

$0.15

Output $/M

$1.50

Context

262K

Providers

2

GLM 4.7 Flash

zhipu/glm-4.7-flash

As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agentic coding use cases, strengthening coding capabilities, long-horizon task planning,...

Input $/M

$0.06

Output $/M

$0.40

Context

200K

Providers

3

Gemma 3 27B It

google/gemma-3-27b-it

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities,...

Input $/M

$0.08

Output $/M

$0.16

Context

131K

Providers

3

Nvidia Nemotron 3 Super 120B A12B

nvidia/nvidia-nemotron-3-super-120b-a12b

NVIDIA Nemotron 3 Super is a hybrid Mixture-of-Experts (MoE) model engineered for highest compute efficiency and accuracy in multi-agent applications and specialized agentic systems. It is optimized to run many collaborating agents per application on a single GPU, delivering high accuracy for reasoning, tool use, and instruction following.

Input $/M

$0.09

Output $/M

$0.40

Context

262K

Providers

1

DeepSeek V3

deepseek/deepseek_v3

DeepSeek-V3 is the latest model from the DeepSeek team, building upon the instruction following and coding abilities of the previous versions. Pre-trained on nearly 15 trillion tokens, the reported evaluations reveal that the model outperforms other open-source models and rivals leading closed-source models.

Input $/M

$0.89

Output $/M

$0.89

Context

64K

Providers

1

DeepSeek V3

deepseek/deepseek-v3

DeepSeek-V3, a strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token. To achieve efficient inference and cost-effective training, DeepSeek-V3 adopts Multi-head Latent Attention (MLA) and DeepSeekMoE architectures, which were thoroughly validated in DeepSeek-V2.

Input $/M

$0.32

Output $/M

$0.89

Context

66K

Providers

1

GLM 4.5v

zhipu/glm-4.5v

GLM-4.5V is a vision-language foundation model for multimodal agent applications. Built on a Mixture-of-Experts (MoE) architecture with 106B parameters and 12B activated parameters, it achieves state-of-the-art results in video understanding,...

Input $/M

$0.60

Output $/M

$1.80

Context

66K

Providers

1

GPT OSS 120B

openai/gpt-oss-120b

gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. It activates 5.1B parameters per forward pass and is optimized...

Input $/M

$0.04

Output $/M

$0.17

Context

131K

Providers

8

o3 Mini

openai/o3-mini

OpenAI o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and coding. This model supports the `reasoning_effort` parameter, which can be set to...

Input $/M

$1.10

Output $/M

$4.40

Context

200K

Providers

1

Qwen3 32B

qwen/qwen3-32b

Qwen3-32B is a dense 32.8B parameter causal language model from the Qwen3 series, optimized for both complex reasoning and efficient dialogue. It supports seamless switching between a "thinking" mode for...

Input $/M

$0.08

Output $/M

$0.28

Context

41K

Providers

2

MiniMax M2

minimax/minimax-m2

MiniMax-M2 is a compact, high-efficiency large language model optimized for end-to-end coding and agentic workflows. With 10 billion activated parameters (230 billion total), it delivers near-frontier intelligence across general reasoning,...

Input $/M

$0.30

Output $/M

$1.20

Context

205K

Providers

1

Gemma 3 12B It

google/gemma-3-12b-it

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities,...

Input $/M

$0.05

Output $/M

$0.15

Context

131K

Providers

2

GPT-5 nano

openai/gpt-5-nano

GPT-5-Nano is the smallest and fastest variant in the GPT-5 system, optimized for developer tools, rapid interactions, and ultra-low latency environments. While limited in reasoning depth compared to its larger...

Input $/M

$0.05

Output $/M

$0.40

Context

400K

Providers

1

GPT 5.2 Codex

openai/gpt-5.2-codex

GPT-5.2-Codex is an upgraded version of GPT-5.1-Codex optimized for software engineering and coding workflows. It is designed for both interactive development sessions and long, independent execution of complex engineering tasks....

Input $/M

$1.75

Output $/M

$14.00

Context

272K

Providers

1

GPT 5.1 Codex

openai/gpt-5.1-codex

GPT-5.1-Codex is a specialized version of GPT-5.1 optimized for software engineering and coding workflows. It is designed for both interactive development sessions and long, independent execution of complex engineering tasks....

Input $/M

$1.25

Output $/M

$10.00

Context

272K

Providers

1

GPT 5.1 Codex Mini

openai/gpt-5.1-codex-mini

GPT-5.1-Codex-Mini is a smaller and faster version of GPT-5.1-Codex

Input $/M

$0.25

Output $/M

$2.00

Context

272K

Providers

1

Llama 4 Maverick

meta-llama/llama-4-maverick

Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architecture with 128 experts and 17 billion active parameters per forward...

Input $/M

$0.15

Output $/M

$0.60

Context

1.0M

Providers

3

Llama 4 Scout

meta-llama/llama-4-scout

Llama 4 Scout 17B Instruct (16E) is a mixture-of-experts (MoE) language model developed by Meta, activating 17 billion parameters out of a total of 109B. It supports native multimodal input...

Input $/M

$0.08

Output $/M

$0.30

Context

328K

Providers

2

Qwen2.5 VL 72B Instruct

qwen/qwen2.5-vl-72b-instruct

Qwen2.5-VL is proficient in recognizing common objects such as flowers, birds, fish, and insects. It is also highly capable of analyzing texts, charts, icons, graphics, and layouts within images.

Input $/M

$0.80

Output $/M

$0.80

Context

33K

Providers

1

DALL·E 3

openai/dall-e-3

은퇴 — OpenAI가 2026-05-12 서비스 종료 (gpt-image-2로 대체)

Price (per image)

$0.04/image

Context

Providers

1

GPT-5.6 Terra

openai/gpt-5.6-terra

GPT-5.6 standard line

Input $/M

$2.50

Output $/M

$15.00

Context

1.1M

Providers

1

GPT-5.6 Luna

openai/gpt-5.6-luna

GPT-5.6 lightweight line

Input $/M

$1.00

Output $/M

$6.00

Context

1.1M

Providers

1

Tencent Hunyuan 3

tencent/hy3

Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active, 192 experts with top-8 routing) built for reasoning, agentic workflows, and real-world production use. It supports a configurable reasoning effort:...

Input $/M

$0.20

Output $/M

$0.80

Context

262K

Providers

3

GLM 5.2 Fast

zhipu/glm-5.2-fast

GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering,...

Input $/M

$1.20

Output $/M

$4.10

Context

1.0M

Providers

1

GLM 5.2 Batch

zhipu/glm-5.2-batch

GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering,...

Input $/M

$0.75

Output $/M

$2.50

Context

Providers

1

Kimi K2.7 Code HighSpeed

moonshot/kimi-k2.7-code-highspeed

코딩 특화 · 고속 서빙(~180 tok/s)

Input $/M

$1.90

Output $/M

$8.00

Context

262K

Providers

2

MiniMax M3 Preview

minimax/minimax-m3-preview

MiniMax-M3 preview is a 1.4T-parameter frontier model from MiniMax for coding, agentic workflows, and complex reasoning, served at fp8 with a 512K context window.

Input $/M

$0.30

Output $/M

$1.20

Context

524K

Providers

1

Nemotron 3 Ultra 550B A55B

nvidia/nemotron-3-ultra-550b-a55b

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...

Input $/M

$0.60

Output $/M

$3.60

Context

512K

Providers

3

Nemotron 3 Ultra

nvidia/nemotron-3-ultra

NVIDIA 550B open weights

Input $/M

$0.50

Output $/M

$2.20

Context

1M

Providers

2

Nemotron Content Safety 3.5

nvidia/nemotron-content-safety-3.5

Nemotron Content Safety 3.5 is a multimodal safety classifier developed by NVIDIA. A compact safety model that handles text, images, and custom policies. It outputs a safe/unsafe classification plus a reasoning trace, and can be used as an inference-time guardrail, as a judge for LLM safety testing and evaluation, or with the accompanying training dataset to post-train models for safer behavior.

Input $/M

$0.20

Output $/M

$0.20

Context

131K

Providers

1

Qwen3.7 Max

qwen/qwen3.7-max

Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series. It supports text input and output and is designed for agent-centric workloads, with particular strengths in coding, office and productivity tasks,...

Input $/M

$2.50

Output $/M

$7.50

Context

1M

Providers

4

Grok Build 0.1

xai/grok-build-0.1

Grok Build 0.1 is xAI’s fast coding model trained specifically for agentic software engineering workflows. It supports text and image inputs with text output, and is optimized for interactive coding...

Input $/M

$1.00

Output $/M

$2.00

Context

256K

Providers

1

Qwen3.6 27B

qwen/qwen3.6-27b

Qwen3.6 27B is a dense 27-billion-parameter language model from the Qwen Team at Alibaba, released in April 2026. It features hybrid multimodal capabilities — accepting text, image, and video inputs...

Input $/M

$0.32

Output $/M

$3.20

Context

262K

Providers

5

Qwen3.6 35B A3B

qwen/qwen3.6-35b-a3b

Qwen3.6-35B-A3B is an open-weight multimodal model from Alibaba Cloud with 35 billion total parameters and 3 billion active parameters per token. It uses a hybrid sparse mixture-of-experts architecture combining Gated...

Input $/M

$0.15

Output $/M

$0.95

Context

262K

Providers

4

GPT 5.5 Pro

openai/gpt-5.5-pro

GPT-5.5 Pro is OpenAI’s high-capability model optimized for deep reasoning and accuracy on complex, high-stakes workloads. It features a 1M+ token context window (922K input, 128K output) with support for...

Input $/M

$30.00

Output $/M

$180.00

Context

1.1M

Providers

1

Hy3 Preview

tencent/hy3-preview

Hy3 preview is a high-efficiency Mixture-of-Experts model from Tencent designed for agentic workflows and production use. It supports configurable reasoning levels across disabled, low, and high modes, allowing it to...

Input $/M

$0.18

Output $/M

$0.60

Context

262K

Providers

1

GPT Image 2

openai/gpt-image-2

토큰 과금 · 1024² 기준 low $0.006 / medium $0.053 / high $0.211

Price (per image)

~$0.05/image

Context

Providers

1

Qwen3.6 Plus 2026 04.02

qwen/qwen3.6-plus-2026-04-02

Input $/M

$0.50

Output $/M

$3.00

Context

262K

Providers

1

Gemma 4 26B A4B It

google/gemma-4-26b-a4b-it

Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at...

Input $/M

$0.07

Output $/M

$0.34

Context

262K

Providers

4

Gemma 4 31B It

google/gemma-4-31b-it

Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...

Input $/M

$0.13

Output $/M

$0.38

Context

262K

Providers

8

Gemma 4 31B It Turbo

google/gemma-4-31b-it-turbo

Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...

Input $/M

$0.12

Output $/M

$0.37

Context

262K

Providers

1

Grok 4.20 Multi-Agent

xai/grok-4.20-multi-agent

멀티 에이전트 오케스트레이션 모드

Input $/M

$1.25

Output $/M

$2.50

Context

2M

Providers

1

Nemotron Cascade 2 30B A3B

nvidia/nemotron-cascade-2-30b-a3b

Nemotron Cascade 2 30B A3B is a reasoning-optimized language model from NVIDIA, designed for efficient inference with strong reasoning capabilities across complex tasks.

Input $/M

$0.14

Output $/M

$0.80

Context

256K

Providers

1

MiniMax M2.7 Highspeed

minimax/minimax-m2.7-highspeed

MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built to actively participate in its own evolution, M2.7 integrates advanced agentic capabilities through multi-agent...

Input $/M

$0.60

Output $/M

$2.40

Context

205K

Providers

1

MiniMax M2.7 Turbo

minimax/minimax-m2.7-turbo

MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built to actively participate in its own evolution, M2.7 integrates advanced agentic capabilities through multi-agent...

Input $/M

$0.38

Output $/M

$1.70

Context

197K

Providers

1

Mistral Small 2603

mistral/mistral-small-2603

Mistral Small 4 unifies instruction following, reasoning, coding, and vision in a single 119B MoE model with 256K context and configurable reasoning effort.

Input $/M

$0.19

Output $/M

$0.75

Context

256K

Providers

1

GLM 5 Turbo

zhipu/glm-5-turbo

GLM-5 Turbo is a new model from Z.ai designed for fast inference and strong performance in agent-driven environments such as OpenClaw scenarios. It is deeply optimized for real-world agent workflows...

Input $/M

$1.20

Output $/M

$4.00

Context

203K

Providers

2

Nemotron 120B A12B

nvidia/nemotron-120b-a12b

Flagship LLM for both reasoning and non-reasoning tasks. It excels in code generation and agentic execution.

Input $/M

$0.30

Output $/M

$0.75

Context

203K

Providers

1

Qwen3.5 9B

qwen/qwen3.5-9b

Qwen3.5-9B is a multimodal foundation model from the Qwen3.5 family, designed to deliver strong reasoning, coding, and visual understanding in an efficient 9B-parameter architecture. It uses a unified vision-language design...

Input $/M

$0.10

Output $/M

$0.15

Context

262K

Providers

3

Grok 4.20 Non-Reasoning

xai/grok-4.20-non-reasoning

추론 생략 모드 · 저지연

Input $/M

$1.25

Output $/M

$2.50

Context

2M

Providers

1

GPT 5.4 Pro

openai/gpt-5.4-pro

GPT-5.4 Pro is OpenAI's most advanced model, building on GPT-5.4's unified architecture with enhanced reasoning capabilities for complex, high-stakes tasks. It features a 1M+ token context window (922K input, 128K...

Input $/M

$30.00

Output $/M

$180.00

Context

1.1M

Providers

1

Gemma 4 E4b It

google/gemma-4-e4b-it

Input $/M

$0.02

Output $/M

$0.10

Context

131K

Providers

1

Gemini 3.1 Pro Preview Customtools

google/gemini-3.1-pro-preview-customtools

Gemini 3.1 Pro Preview Custom Tools is a variant of Gemini 3.1 Pro that improves tool selection behavior by preventing overuse of a general bash tool when more efficient third-party...

Input $/M

$2.00

Output $/M

$12.00

Context

1.0M

Providers

1

Qwen3.5 Plus

qwen/qwen3.5-plus

The Qwen3.5 native vision-language series Plus models are based on a hybrid architecture design that integrates linear attention mechanisms with sparse Mixture-of-Experts (MoE), achieving higher inference efficiency. Across various task evaluations, the 3.5 series demonstrates exceptional performance comparable to current top-tier frontier models, marking a leap forward in both plain text and multimodal capabilities compared to the 3 series.

Input $/M

$0.40

Output $/M

$2.40

Context

992K

Providers

1

MiniMax M2.5 Highspeed

minimax/minimax-m2.5-highspeed

MiniMax-M2.5 is a SOTA large language model designed for real-world productivity. Trained in a diverse range of complex real-world digital working environments, M2.5 builds upon the coding expertise of M2.1...

Input $/M

$0.60

Output $/M

$2.40

Context

205K

Providers

1

Qwen3 Max Thinking

qwen/qwen3-max-thinking

Qwen3-Max-Thinking is the flagship reasoning model in the Qwen3 series, designed for high-stakes cognitive tasks that require deep, multi-step reasoning. By significantly scaling model capacity and reinforcement learning compute, it...

Input $/M

$1.20

Output $/M

$6.00

Context

262K

Providers

1

Qwen3 Coder Next

qwen/qwen3-coder-next

Qwen3-Coder-Next is an open-weight causal language model optimized for coding agents and local development workflows. It uses a sparse MoE design with 80B total parameters and only 3B activated per...

Input $/M

$0.20

Output $/M

$1.50

Context

262K

Providers

1

DeepSeek OCR 2

deepseek/deepseek-ocr-2

DeepSeek-OCR 2 is a multimodal document recognition model released by DeepSeek AI, serving as an upgrade to DeepSeek-OCR. By introducing the DeepEncoder V2 architecture, it achieves a paradigm shift in visual encoding from "fixed scanning" to "semantic reasoning." The model replaces the original CLIP encoder with a lightweight language model (Qwen2-0.5B) and incorporates a causal flow query mechanism, while retaining the DeepSeek-3B-MoE decoder. The model requires only 256 to 1120 visual tokens to cover complex document pages. On the OmniDocBench v1.5 benchmark, it achieves an overall score of 91.09%, a 3.73% improvement over its predecessor, with reading order recognition edit distance reduced from 0.085 to 0.057.

Input $/M

$0.03

Output $/M

$0.03

Context

8K

Providers

1

Kimi K2.5

moonshot/kimi-k2.5

Kimi K2.5 is Moonshot AI's native multimodal model, delivering state-of-the-art visual coding capability and a self-directed agent swarm paradigm. Built on Kimi K2 with continued pretraining over approximately 15T mixed...

Input $/M

$0.45

Output $/M

$2.25

Context

262K

Providers

7

Solar Pro 3

upstage/solar-pro3

Korean-optimized

Input $/M

$0.15

Output $/M

$0.60

Context

128K

Providers

1

GLM 4.7 H

zhipu/glm-4.7-h

GLM-4.7 is Z.AI's latest flagship model, with major upgrades focused on advanced coding capabilities and more reliable multi-step reasoning and execution. It shows clear gains in complex agent workflows, while delivering a more natural conversational experience and stronger front-end design sensibility.

Input $/M

$0.60

Output $/M

$2.20

Context

205K

Providers

1

Qwen3 VL 235B A22B

qwen/qwen3-vl-235b-a22b

Qwen3-VL 235B vision-language model with MoE architecture. The most powerful VL model in the Qwen series with superior visual perception, OCR, and multimodal reasoning.

Input $/M

$0.21

Output $/M

$1.90

Context

128K

Providers

1

Mistral Small 3.2 24B Instruct

mistral/mistral-small-3.2-24b-instruct

Mistral Small 3.2 is a 24B parameter model optimized for efficiency and performance. Ideal for general-purpose tasks with balanced speed and capability.

Input $/M

$0.09

Output $/M

$0.25

Context

256K

Providers

1

MiniMax M2.1

minimax/minimax-m2.1

MiniMax-M2.1 is a lightweight, state-of-the-art large language model optimized for coding, agentic workflows, and modern application development. With only 10 billion activated parameters, it delivers a major jump in real-world...

Input $/M

$0.30

Output $/M

$1.20

Context

205K

Providers

1

Gemini 3 Flash Preview

google/gemini-3-flash-preview

Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It delivers near Pro level reasoning and tool...

Input $/M

$0.50

Output $/M

$3.00

Context

1.0M

Providers

1

Nemotron 3 Nano 30B A3B

nvidia/nemotron-3-nano-30b-a3b

NVIDIA Nemotron 3 Nano 30B A3B is a small language MoE model with highest compute efficiency and accuracy for developers to build specialized agentic AI systems. The model is fully...

Input $/M

$0.05

Output $/M

$0.20

Context

262K

Providers

3

GPT 5.2 Pro

openai/gpt-5.2-pro

GPT-5.2 Pro is OpenAI’s most advanced model, offering major improvements in agentic coding and long context performance over GPT-5 Pro. It is optimized for complex tasks that require step-by-step reasoning,...

Input $/M

$21.00

Output $/M

$168.00

Context

272K

Providers

1

Autoglm Phone 9B Multilingual

zhipu/autoglm-phone-9b-multilingual

Phone Agent is a mobile intelligent assistant framework built on AutoGLM, capable of understanding smartphone screens through multimodal perception and executing automated operations to complete tasks. The system controls devices via ADB (Android Debug Bridge), uses a vision-language model for screen understanding, and leverages intelligent planning to generate and execute action sequences. Users can simply describe tasks in natural language—for example, “Open Xiaohongshu and search for food recommendations.” Phone Agent will automatically parse the intent, understand the current UI, plan the next steps, and carry out the entire workflow. The system also includes: Sensitive action confirmation mechanisms Human-in-the-loop fallback for login or verification code scenarios Remote ADB debugging, allowing device connection via WiFi or network for flexible remote control and development

Input $/M

$0.04

Output $/M

$0.14

Context

66K

Providers

1

GPT 5.1 Codex Max

openai/gpt-5.1-codex-max

GPT-5.1-Codex-Max is OpenAI’s latest agentic coding model, designed for long-running, high-context software development tasks. It is based on an updated version of the 5.1 reasoning stack and trained on agentic...

Input $/M

$1.25

Output $/M

$10.00

Context

272K

Providers

1

DeepSeek OCR

deepseek/deepseek-ocr

DeepSeek-OCR as an initial investigation into the feasibility of compressing long contexts via optical 2D mapping. DeepSeek-OCR consists of two components: DeepEncoder and DeepSeek3B-MoE-A570M as the decoder. Specifically, DeepEncoder serves as the core engine, designed to maintain low activations under high-resolution input while achieving high compression ratios to ensure an optimal and manageable number of vision tokens. Experiments show that when the number of text tokens is within 10 times that of vision tokens (i.e., a compression ratio < 10x), the model can achieve decoding (OCR) precision of 97%. Even at a compression ratio of 20x, the OCR accuracy still remains at about 60%. This shows considerable promise for research areas such as historical long-context compression and memory forgetting mechanisms in LLMs.

Input $/M

$0.03

Output $/M

$0.03

Context

8K

Providers

1

Qwen3 VL 8B Instruct

qwen/qwen3-vl-8b-instruct

Qwen3-VL-8B-Instruct is a multimodal vision-language model from the Qwen3-VL series, built for high-fidelity understanding and reasoning across text, images, and video. It features improved multimodal fusion with Interleaved-MRoPE for long-horizon...

Input $/M

$0.08

Output $/M

$0.50

Context

131K

Providers

1

GPT 5 Pro

openai/gpt-5-pro

GPT-5 Pro is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that require step-by-step reasoning, instruction following, and...

Input $/M

$15.00

Output $/M

$120.00

Context

400K

Providers

1

Qwen3 VL 30B A3B Thinking

qwen/qwen3-vl-30b-a3b-thinking

Qwen3-VL-30B-A3B-Thinking is a multimodal model that unifies strong text generation with visual understanding for images and videos. Its Thinking variant enhances reasoning in STEM, math, and complex tasks. It excels...

Input $/M

$0.20

Output $/M

$1.00

Context

131K

Providers

1

Qwen3 VL 30B A3B Instruct

qwen/qwen3-vl-30b-a3b-instruct

Qwen3-VL-30B-A3B-Instruct is a multimodal model that unifies strong text generation with visual understanding for images and videos. Its Instruct variant optimizes instruction-following for general multimodal tasks. It excels in perception...

Input $/M

$0.15

Output $/M

$0.60

Context

131K

Providers

2

GPT 5 Codex

openai/gpt-5-codex

GPT-5-Codex is a specialized version of GPT-5 optimized for software engineering and coding workflows. It is designed for both interactive development sessions and long, independent execution of complex engineering tasks....

Input $/M

$1.25

Output $/M

$10.00

Context

272K

Providers

1

Qwen3 VL 235B

qwen/qwen3-vl-235b

Open-weight vision model

Input $/M

$0.20

Output $/M

$0.88

Context

131K

Providers

2

Qwen3 Omni 30B A3B Instruct

qwen/qwen3-omni-30b-a3b-instruct

The Qwen-Omni model accepts combined inputs of text and a single additional modality (image, audio, or video) to generate responses in text or speech. It offers a variety of human-like voices, supports speech output in multiple languages and dialects, and is suitable for applications such as text creation, visual recognition, and voice assistants

Input $/M

$0.25

Output $/M

$0.97

Context

66K

Providers

1

Qwen3 Omni 30B A3B Thinking

qwen/qwen3-omni-30b-a3b-thinking

The Qwen-Omni model accepts combined inputs of text and a single additional modality (image, audio, or video) to generate responses in text or speech. It offers a variety of human-like voices, supports speech output in multiple languages and dialects, and is suitable for applications such as text creation, visual recognition, and voice assistants

Input $/M

$0.25

Output $/M

$0.97

Context

66K

Providers

1

Kimi K2 0905

moonshot/kimi-k2-0905

Kimi K2 0905 is the September update of [Kimi K2 0711](moonshotai/kimi-k2). It is a large-scale Mixture-of-Experts (MoE) language model developed by Moonshot AI, featuring 1 trillion total parameters with 32...

Input $/M

$0.60

Output $/M

$2.50

Context

262K

Providers

1

Qwen Mt Plus

qwen/qwen-mt-plus

Qwen-MT is a large language model optimized for machine translation, built upon the foundation of the Tongyi Qianwen model. It supports translation across 92 languages — including Chinese, English, Japanese, Korean, French, Spanish, German, Thai, Indonesian, Vietnamese, Arabic, and more — enabling seamless multilingual communication.

Input $/M

$0.25

Output $/M

$0.75

Context

16K

Providers

1

Kimi K2 Instruct 0905

moonshot/kimi-k2-instruct-0905

Kimi K2 0905 is the September update of Kimi K2 0711. It is a large-scale Mixture-of-Experts (MoE) language model developed by Moonshot AI, featuring 1 trillion total parameters with 32 billion active per forward pass. It supports long-context inference up to 256k tokens, extended from the previous 128k. This update improves agentic coding with higher accuracy and better generalization across scaffolds, and enhances frontend coding with more aesthetic and functional outputs for web, 3D, and related tasks. Kimi K2 is optimized for agentic capabilities, including advanced tool use, reasoning, and code synthesis. It excels across coding (LiveCodeBench, SWE-bench), reasoning (ZebraLogic, GPQA), and tool-use (Tau2, AceBench) benchmarks. The model is trained with a novel stack incorporating the MuonClip optimizer for stable large-scale MoE training.

Input $/M

$0.57

Output $/M

$2.29

Context

262K

Providers

1

GPT 5

openai/gpt-5

GPT-5 is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that require step-by-step reasoning, instruction following, and accuracy...

Input $/M

$1.25

Output $/M

$10.00

Context

272K

Providers

1

GPT OSS 120B Turbo

openai/gpt-oss-120b-turbo

gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. It activates 5.1B parameters per forward pass and is optimized...

Input $/M

$0.15

Output $/M

$0.60

Context

131K

Providers

1

GPT OSS 20B

openai/gpt-oss-20b

gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for...

Input $/M

$0.03

Output $/M

$0.14

Context

131K

Providers

5

Qwen3 Coder 30B A3B Instruct

qwen/qwen3-coder-30b-a3b-instruct

Qwen3-Coder-30B-A3B-Instruct is a 30.5B parameter Mixture-of-Experts (MoE) model with 128 experts (8 active per forward pass), designed for advanced code generation, repository-scale understanding, and agentic tool use. Built on the...

Input $/M

$0.07

Output $/M

$0.27

Context

160K

Providers

1

GLM 4.5 Air FP8

zhipu/glm-4.5-air-fp8

GLM-4.5-Air is the lightweight variant of our latest flagship model family, also purpose-built for agent-centric applications. Like GLM-4.5, it adopts the Mixture-of-Experts (MoE) architecture but with a more compact parameter...

Input $/M

$0.20

Output $/M

$1.10

Context

131K

Providers

1

Qwen3 Coder 480B A35B Instruct Turbo

qwen/qwen3-coder-480b-a35b-instruct-turbo

Qwen3-Coder-480B-A35B-Instruct is the Qwen3's most agentic code model, featuring Significant Performance on Agentic Coding, Agentic Browser-Use and other foundational coding tasks, achieving results comparable to Claude Sonnet.

Input $/M

$0.30

Output $/M

$1.00

Context

262K

Providers

2

Kimi K2 Instruct

moonshot/kimi-k2-instruct

Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model with 32 billion activated parameters and 1 trillion total parameters. Trained with the Muon optimizer, Kimi K2 achieves exceptional performance across frontier knowledge, reasoning, and coding tasks while being meticulously optimized for agentic capabilities.Specifically designed for tool use, reasoning, and autonomous problem-solving.

Input $/M

$0.57

Output $/M

$2.30

Context

131K

Providers

1

Mistral Small 3.2 24B Instruct 2506

mistral/mistral-small-3.2-24b-instruct-2506

Mistral-Small-3.2-24B-Instruct is a drop-in upgrade over the 3.1 release, with markedly better instruction following, roughly half the infinite-generation errors, and a more robust function-calling interface—while otherwise matching or slightly improving on all previous text and vision benchmarks.

Input $/M

$0.08

Output $/M

$0.20

Context

128K

Providers

1

MiniMax M1 80K

minimax/minimax-m1-80k

MiniMax-M1: The World's First Open-Weight, Large-Scale Hybrid Attention Inference Model MiniMax-M1 adopts a Mixture of Experts (MoE) architecture and integrates the Flash Attention mechanism. The model contains a total of 456 billion parameters, with 45.9 billion parameters activated per token. Natively, the M1 model supports a context length of 1 million tokens—8 times that of DeepSeek R1. Additionally, by combining the CISPO algorithm with an efficient hybrid attention design for reinforcement learning training, MiniMax-M1 achieves industry-leading performance in long-context reasoning and real-world software engineering scenarios.

Input $/M

$0.55

Output $/M

$2.20

Context

1M

Providers

1

o3 Pro

openai/o3-pro

The o-series of models are trained with reinforcement learning to think before they answer and perform complex reasoning. The o3-pro model uses more compute to think harder and provide consistently...

Input $/M

$20.00

Output $/M

$80.00

Context

200K

Providers

1

DeepSeek R1 0528 Qwen3 8B

deepseek/deepseek-r1-0528-qwen3-8b

DeepSeek-R1-0528-Qwen3-8B is a high-performance reasoning model based on the Qwen3 8B Base model, enhanced through the integration of DeepSeek-R1-0528's Chain-of-Thought (CoT) optimization. In the AIME 2024 evaluation, this open-source model achieved state-of-the-art (SOTA) performance, delivering a 10% improvement over the original Qwen3 8B while matching the reasoning capabilities of the much larger 235-billion-parameter Qwen3-235B-thinking.

Input $/M

$0.06

Output $/M

$0.09

Context

128K

Providers

1

Gemma 3n E4b It

google/gemma-3n-e4b-it

Gemma 3n E4B-it is optimized for efficient execution on mobile and low-resource devices, such as phones, laptops, and tablets. It supports multimodal inputs—including text, visual data, and audio—enabling diverse tasks...

Input $/M

$0.06

Output $/M

$0.12

Context

33K

Providers

1

Llama Guard 4 12B

meta-llama/llama-guard-4-12b

Llama Guard 4 is a Llama 4 Scout-derived multimodal pretrained model, fine-tuned for content safety classification. Similar to previous versions, it can be used to classify content in both LLM...

Input $/M

$0.18

Output $/M

$0.18

Context

164K

Providers

1

DeepSeek Prover V2 671B

deepseek/deepseek-prover-v2-671b

DeepSeek Launches Open-Source Model DeepSeek-Prover-V2-671B, Specializing in Mathematical Theorem Proving The new model employs a Mixture of Experts (MoE) architecture and is trained using the Lean 4 framework for formal reasoning. With 671 billion parameters, it leverages reinforcement learning and large-scale synthetic data to significantly enhance automated theorem-proving capabilities.

Input $/M

$0.70

Output $/M

$2.50

Context

160K

Providers

1

Qwen3 Next 80B

qwen/qwen3-next-80b

Optimized for speed and efficiency.

Input $/M

$0.35

Output $/M

$1.90

Context

256K

Providers

1

Qwen3 235B A22B FP8

qwen/qwen3-235b-a22b-fp8

Qwen3-235B-A22B is a 235B parameter mixture-of-experts (MoE) model developed by Qwen, activating 22B parameters per forward pass. It supports seamless switching between a "thinking" mode for complex reasoning, math, and...

Input $/M

$0.20

Output $/M

$0.80

Context

41K

Providers

1

Qwen3 30B A3B

qwen/qwen3-30b-a3b

Qwen3, the latest generation in the Qwen large language model series, features both dense and mixture-of-experts (MoE) architectures to excel in reasoning, multilingual support, and advanced agent tasks. Its unique...

Input $/M

$0.12

Output $/M

$0.50

Context

41K

Providers

1

Qwen3 30B A3B FP8

qwen/qwen3-30b-a3b-fp8

Qwen3, the latest generation in the Qwen large language model series, features both dense and mixture-of-experts (MoE) architectures to excel in reasoning, multilingual support, and advanced agent tasks. Its unique...

Input $/M

$0.09

Output $/M

$0.45

Context

41K

Providers

1

Qwen3 32B FP8

qwen/qwen3-32b-fp8

Qwen3-32B is a dense 32.8B parameter causal language model from the Qwen3 series, optimized for both complex reasoning and efficient dialogue. It supports seamless switching between a "thinking" mode for...

Input $/M

$0.10

Output $/M

$0.45

Context

41K

Providers

1

Qwen3 14B

qwen/qwen3-14b

Qwen3-14B is a dense 14.8B parameter causal language model from the Qwen3 series, designed for both complex reasoning and efficient dialogue. It supports seamless switching between a "thinking" mode for...

Input $/M

$0.12

Output $/M

$0.24

Context

41K

Providers

1

Qwen3 8B FP8

qwen/qwen3-8b-fp8

Qwen3-8B is a dense 8.2B parameter causal language model from the Qwen3 series, designed for both reasoning-heavy tasks and efficient dialogue. It supports seamless switching between "thinking" mode for math,...

Input $/M

$0.04

Output $/M

$0.14

Context

128K

Providers

1

Qwen3 4B FP8

qwen/qwen3-4b-fp8

Achieves effective integration of reasoning and non-reasoning modes, allowing seamless switching during conversations. The model delivers state-of-the-art (SOTA) reasoning performance among models of the same scale, with significantly enhanced human preference alignment. Notable improvements are seen in creative writing, role-playing, multi-turn dialogue, and instruction following, leading to a clearly improved user experience.

Input $/M

$0.03

Output $/M

$0.03

Context

128K

Providers

1

o4 Mini

openai/o4-mini

OpenAI o4-mini is a compact reasoning model in the o-series, optimized for fast, cost-efficient performance while retaining strong multimodal and agentic capabilities. It supports tool use and demonstrates competitive reasoning...

Input $/M

$1.10

Output $/M

$4.40

Context

200K

Providers

1

o3

openai/o3

o3 is a well-rounded and powerful model across domains. It sets a new standard for math, science, coding, and visual reasoning tasks. It also excels at technical writing and instruction-following....

Input $/M

$2.00

Output $/M

$8.00

Context

200K

Providers

1

GLM 4 32B 0414

zhipu/glm-4-32b-0414

GLM-4-32B-0414 is the latest open-source model in the GLM series, featuring 32 billion parameters. Its performance is comparable to OpenAI's GPT series and DeepSeek's V3/R1 series, while also supporting highly user-friendly local deployment capabilities. GLM-4-32B-Base-0414 was pre-trained on 15T of high-quality data, including a large amount of reasoning-type synthetic data, which laid a solid foundation for subsequent reinforcement learning extensions. In the post-training stage, in addition to human preference alignment for dialogue scenarios, the research team enhanced the model’s performance in instruction following, engineering code, and function calling using techniques such as rejection sampling and reinforcement learning, thereby strengthening the atomic capabilities required for agent tasks. GLM-4-32B-0414 has achieved strong results in engineering code generation, artifact creation, function calling, search-based question answering, and report generation. On several benchmarks, its performance appr

Input $/M

$0.55

Output $/M

$1.66

Context

32K

Providers

1

GPT 4.1

openai/gpt-4.1

GPT-4.1 is a flagship large language model optimized for advanced instruction following, real-world software engineering, and long-context reasoning. It supports a 1 million token context window and outperforms GPT-4o and...

Input $/M

$2.00

Output $/M

$8.00

Context

1.0M

Providers

1

GPT 4.1 Mini

openai/gpt-4.1-mini

GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost. It retains a 1 million token context window and scores 45.1% on hard...

Input $/M

$0.40

Output $/M

$1.60

Context

1.0M

Providers

1

GPT 4.1 Nano

openai/gpt-4.1-nano

For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a small size with its 1 million...

Input $/M

$0.10

Output $/M

$0.40

Context

1.0M

Providers

1

Llama 3.3 70B

meta-llama/llama-3.3-70b

Input $/M

$0.70

Output $/M

$2.80

Context

128K

Providers

1

o1 Pro

openai/o1-pro

The o1 series of models are trained with reinforcement learning to think before they answer and perform complex reasoning. The o1-pro model uses more compute to think harder and provide...

Input $/M

$150.00

Output $/M

$600.00

Context

200K

Providers

1

Gemma 3 4B It

google/gemma-3-4b-it

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities,...

Input $/M

$0.05

Output $/M

$0.10

Context

131K

Providers

1

Sonar Pro

perplexity/sonar-pro

Search-grounding flagship (search billed separately per request)

Input $/M

$3.00

Output $/M

$15.00

Context

200K

Providers

1

Sonar Reasoning Pro

perplexity/sonar-reasoning-pro

Search + reasoning (CoT) combined (search billed separately per request)

Input $/M

$2.00

Output $/M

$8.00

Context

128K

Providers

1

Gemini 1.5 Flash

google/gemini-1.5-flash

Gemini 1.5 Flash is Google's foundation model that performs well at a variety of multimodal tasks such as visual understanding, classification, summarization, and creating content from image, audio and video. It's adept at processing visual and text inputs such as photographs, documents, infographics, and screenshots. Gemini 1.5 Flash is designed for high-volume, high-frequency tasks where cost and latency matter.

Input $/M

$0.08

Output $/M

$0.30

Context

1M

Providers

1

DeepSeek V3 Turbo

deepseek/deepseek-v3-turbo

DeepSeek-V3 is the latest model from the DeepSeek team, building upon the instruction following and coding abilities of the previous versions. Pre-trained on nearly 15 trillion tokens, the reported evaluations reveal that the model outperforms other open-source models and rivals leading closed-source models.

Input $/M

$0.40

Output $/M

$1.30

Context

64K

Providers

1

DeepSeek R1 Distill Qwen 32B

deepseek/deepseek-r1-distill-qwen-32b

DeepSeek R1 Distill Qwen 32B is a distilled large language model based on Qwen 2.5 32B, using outputs from DeepSeek R1. It outperforms OpenAI's o1-mini across various benchmarks, achieving new state-of-the-art results for dense models. Other benchmark results include: AIME 2024 pass@1: 72.6 MATH-500 pass@1: 94.3 CodeForces Rating: 1691 The model leverages fine-tuning from DeepSeek R1's outputs, enabling competitive performance comparable to larger frontier models.

Input $/M

$0.30

Output $/M

$0.30

Context

64K

Providers

1

DeepSeek R1 Distill Qwen 14B

deepseek/deepseek-r1-distill-qwen-14b

DeepSeek R1 Distill Qwen 14B is a distilled large language model based on Qwen 2.5 14B, using outputs from DeepSeek R1. It outperforms OpenAI's o1-mini across various benchmarks, achieving new state-of-the-art results for dense models. Other benchmark results include: AIME 2024 pass@1: 69.7 MATH-500 pass@1: 93.9 CodeForces Rating: 1481 The model leverages fine-tuning from DeepSeek R1's outputs, enabling competitive performance comparable to larger frontier models.

Input $/M

$0.15

Output $/M

$0.15

Context

33K

Providers

1

Mistral Small 24B Instruct 2501

mistral/mistral-small-24b-instruct-2501

Mistral Small 3 is a 24B-parameter language model optimized for low-latency performance across common AI tasks. Released under the Apache 2.0 license, it features both pre-trained and instruction-tuned versions designed for efficient local deployment. The model achieves 81% accuracy on the MMLU benchmark and performs competitively with larger models like Llama 3.3 70B and Qwen 32B, while operating at three times the speed on equivalent hardware.

Input $/M

$0.05

Output $/M

$0.08

Context

33K

Providers

1

Sonar

perplexity/sonar

Answers grounded in live web search (search billed separately per request)

Input $/M

$1.00

Output $/M

$1.00

Context

128K

Providers

1

DeepSeek R1 Distill Llama 70B

deepseek/deepseek-r1-distill-llama-70b

DeepSeek R1 Distill Llama 70B is a distilled large language model based on [Llama-3.3-70B-Instruct](/meta-llama/llama-3.3-70b-instruct), using outputs from [DeepSeek R1](/deepseek/deepseek-r1). The model combines advanced distillation techniques to achieve high performance across...

Input $/M

$0.80

Output $/M

$0.80

Context

8K

Providers

2

DeepSeek R1 Turbo

deepseek/deepseek-r1-turbo

DeepSeek R1 is here: Performance on par with [OpenAI o1](/openai/o1), but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active in an inference pass....

Input $/M

$0.70

Output $/M

$2.50

Context

64K

Providers

1

Qwen 2 VL 72B Instruct

qwen/qwen-2-vl-72b-instruct

Qwen2 VL 72B is a multimodal LLM from the Qwen Team with the following key enhancements: SoTA understanding of images of various resolution & ratio: Qwen2-VL achieves state-of-the-art performance on visual understanding benchmarks, including MathVista, DocVQA, RealWorldQA, MTVQA, etc. Understanding videos of 20min+: Qwen2-VL can understand videos over 20 minutes for high-quality video-based question answering, dialog, content creation, etc. Agent that can operate your mobiles, robots, etc.: with the abilities of complex reasoning and decision making, Qwen2-VL can be integrated with devices like mobile phones, robots, etc., for automatic operation based on visual environment and text instructions. Multilingual Support: to serve global users, besides English and Chinese, Qwen2-VL now supports the understanding of texts in different languages inside images, including most European languages, Japanese, Korean, Arabic, Vietnamese, etc.

Input $/M

$0.45

Output $/M

$0.45

Context

33K

Providers

1

o1

openai/o1

The latest and strongest model family from OpenAI, o1 is designed to spend more time thinking before responding. The o1 model series is trained with large-scale reinforcement learning to reason...

Input $/M

$15.00

Output $/M

$60.00

Context

200K

Providers

1

Llama 3.3 70B Instruct

meta-llama/llama-3.3-70b-instruct

The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model...

Input $/M

$0.14

Output $/M

$0.40

Context

6K

Providers

3

Llama 3.3 70B Instruct Turbo

meta-llama/llama-3.3-70b-instruct-turbo

The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model...

Input $/M

$0.10

Output $/M

$0.32

Context

131K

Providers

2

Qwen2.5 7B Instruct Turbo

qwen/qwen2.5-7b-instruct-turbo

Qwen2.5 7B is the latest series of Qwen large language models. Qwen2.5 brings the following improvements upon Qwen2: - Significantly more knowledge and has greatly improved capabilities in coding and...

Input $/M

$0.30

Output $/M

$0.30

Context

33K

Providers

1

Qwen2.5 7B Instruct

qwen/qwen2.5-7b-instruct

Qwen2.5 7B is the latest series of Qwen large language models. Qwen2.5 brings the following improvements upon Qwen2: - Significantly more knowledge and has greatly improved capabilities in coding and...

Input $/M

$0.07

Output $/M

$0.07

Context

32K

Providers

1

Llama 3.2 3B

meta-llama/llama-3.2-3b

Input $/M

$0.15

Output $/M

$0.60

Context

128K

Providers

1

Llama 3.2 3B Instruct

meta-llama/llama-3.2-3b-instruct

Llama 3.2 3B is a 3-billion-parameter multilingual large language model, optimized for advanced natural language processing tasks like dialogue generation, reasoning, and summarization. Designed with the latest transformer architecture, it...

Input $/M

$0.03

Output $/M

$0.05

Context

33K

Providers

1

Llama 3.2 1B Instruct

meta-llama/llama-3.2-1b-instruct

Llama 3.2 1B is a 1-billion-parameter language model focused on efficiently performing natural language tasks, such as summarization, dialogue, and multilingual text analysis. Its smaller size allows it to operate...

Input $/M

$0.02

Output $/M

$0.02

Context

131K

Providers

1

Qwen2.5 72B Instruct

qwen/qwen2.5-72b-instruct

Qwen2.5 72B is the latest series of Qwen large language models. Qwen2.5 brings the following improvements upon Qwen2: - Significantly more knowledge and has greatly improved capabilities in coding and...

Input $/M

$0.36

Output $/M

$0.40

Context

33K

Providers

1

Qwen 2 7B Instruct

qwen/qwen-2-7b-instruct

Qwen2 is the newest series in the Qwen large language model family. Qwen2 7B is a transformer-based model that demonstrates exceptional performance in language understanding, multilingual capabilities, programming, mathematics, and reasoning.

Input $/M

$0.05

Output $/M

$0.05

Context

33K

Providers

1

Mistral Nemo

mistral/mistral-nemo

A 12B parameter model with a 128k token context length built by Mistral in collaboration with NVIDIA. The model is multilingual, supporting English, French, German, Spanish, Italian, Portuguese, Chinese, Japanese, Korean, Arabic, and Hindi. It supports function calling and is released under the Apache 2.0 license.

Input $/M

$0.04

Output $/M

$0.17

Context

60K

Providers

1

Meta Llama 3.1 70B Instruct Turbo

meta-llama/meta-llama-3.1-70b-instruct-turbo

Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 70B instruct-tuned version is optimized for high quality dialogue usecases. It has demonstrated strong...

Input $/M

$0.40

Output $/M

$0.40

Context

131K

Providers

1

Llama 3.1 8B Instruct

meta-llama/llama-3.1-8b-instruct

Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 8B instruct-tuned version is fast and efficient. It has demonstrated strong performance compared to...

Input $/M

$0.02

Output $/M

$0.05

Context

16K

Providers

2

Meta Llama 3.1 8B Instruct Turbo

meta-llama/meta-llama-3.1-8b-instruct-turbo

Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 8B instruct-tuned version is fast and efficient. It has demonstrated strong performance compared to...

Input $/M

$0.02

Output $/M

$0.03

Context

131K

Providers

1

GPT 4o Mini

openai/gpt-4o-mini

GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. As their most advanced small model, it is many multiples more affordable...

Input $/M

$0.15

Output $/M

$0.60

Context

128K

Providers

1

Mistral Nemo Instruct 2407

mistral/mistral-nemo-instruct-2407

12B model trained jointly by Mistral AI and NVIDIA, it significantly outperforms existing models smaller or similar in size.

Input $/M

$0.02

Output $/M

$0.03

Context

131K

Providers

1

Solar Mini

upstage/solar-mini

Korean-optimized (lightweight)

Input $/M

$0.15

Output $/M

$0.15

Context

33K

Providers

1

GPT 4o

openai/gpt-4o

GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligence level of [GPT-4 Turbo](/models/openai/gpt-4-turbo) while being twice as...

Input $/M

$2.50

Output $/M

$10.00

Context

128K

Providers

1

Llama 3 70B Instruct

meta-llama/llama-3-70b-instruct

Meta's latest class of model (Llama 3) launched with a variety of sizes & flavors. This 70B instruct-tuned version was optimized for high quality dialogue usecases. It has demonstrated strong performance compared to leading closed-source models in human evaluations.

Input $/M

$0.51

Output $/M

$0.74

Context

8K

Providers

1

Llama 3 8B Instruct

meta-llama/llama-3-8b-instruct

Meta's latest class of model (Llama 3) launched with a variety of sizes & flavors. This 8B instruct-tuned version was optimized for high quality dialogue usecases. It has demonstrated strong performance compared to leading closed-source models in human evaluations.

Input $/M

$0.04

Output $/M

$0.04

Context

8K

Providers

1

GPT 4 Turbo

openai/gpt-4-turbo

The latest GPT-4 Turbo model with vision capabilities. Vision requests can now use JSON mode and function calling. Training data: up to December 2023.

Input $/M

$10.00

Output $/M

$30.00

Context

128K

Providers

1

text-embedding-3-large

openai/text-embedding-3-large

고정밀 임베딩 · 3072차원

Input $/M

$0.13

Context

8K

Providers

1

text-embedding-3-small

openai/text-embedding-3-small

범용 임베딩 · 1536차원

Input $/M

$0.02

Context

8K

Providers

1

GPT 3.5 Turbo 16K

openai/gpt-3.5-turbo-16k

This model offers four times the context length of gpt-3.5-turbo, allowing it to support approximately 20 pages of text in a single request at a higher cost. Training data: up...

Input $/M

$3.00

Output $/M

$4.00

Context

16K

Providers

1

GPT 4

openai/gpt-4

OpenAI's flagship model, GPT-4 is a large-scale multimodal language model capable of solving difficult problems with greater accuracy than previous models due to its broader general knowledge and advanced reasoning...

Input $/M

$30.00

Output $/M

$60.00

Context

8K

Providers

1

GPT 3.5 Turbo

openai/gpt-3.5-turbo

GPT-3.5 Turbo is OpenAI's fastest model. It can understand and generate natural language or code, and is optimized for chat and traditional completion tasks. Training data up to Sep 2021.

Input $/M

$0.50

Output $/M

$1.50

Context

16K

Providers

1

Whisper

openai/whisper-1

$0.006/min (STT)

Price (per min)

$0.006/min

Context

Providers

1

Whisper Large V3

openai/whisper-large-v3

Price (per min)

$0.00045/min

Context

Providers

2

Whisper Large V3 Turbo

openai/whisper-large-v3-turbo

Price (per min)

$0.0001998/min

Context

Providers

1

Gemini 3.1 Flash Image

google/gemini-3.1-flash-image

Price (per image)

~$0.05/image

Context

Providers

1

Gemini 3 Pro Image

google/gemini-3-pro-image

Price (per image)

~$0.13/image

Context

Providers

1

Embeddinggemma 300M

google/embeddinggemma-300m

Input $/M

<$0.01

Context

2K

Providers

1

Qwen3 Embedding 0.6B

qwen/qwen3-embedding-0.6b

Input $/M

$0.01

Context

33K

Providers

2

Qwen3 Embedding 4B

qwen/qwen3-embedding-4b

Input $/M

$0.02

Context

33K

Providers

1

Qwen3 Embedding 8B

qwen/qwen3-embedding-8b

Input $/M

$0.01

Context

33K

Providers

2

Voxtral Mini 3B 2507

mistral/voxtral-mini-3b-2507

Price (per min)

$0.001/min

Context

Providers

1

Voxtral Small 24B 2507

mistral/voxtral-small-24b-2507

Price (per min)

$0.003/min

Context

Providers

1

Nemotron 3.5 Asr Streaming 0.6B

nvidia/nemotron-3.5-asr-streaming-0.6b

Price (per min)

$0.0015/min

Context

Providers

1

Nemotron 3.5 Asr Streaming Multilingual 0.6B

nvidia/nemotron-3.5-asr-streaming-multilingual-0.6b

Price (per min)

$0.0001998/min

Context

Providers

1

Nemotron 3 Asr Streaming 0.6B

nvidia/nemotron-3-asr-streaming-0.6b

Price (per min)

$0.0015/min

Context

Providers

1

Parakeet Tdt 0.6B V3

nvidia/parakeet-tdt-0.6b-v3

Price (per min)

$0.0015/min

Context

Providers

1

Llama Nemotron Embed VL 1B V2

nvidia/llama-nemotron-embed-vl-1b-v2

Input $/M

$0.01

Context

10K

Providers

2

Cobuddy

baidu/cobuddy

CoBuddy is a specialized code generation model developed by Baidu, engineered with targeted optimizations for Coding and AI Agent scenarios. It delivers exceptional performance, characterized by high inference throughput and ultra-low end-to-end (E2E) latency.

Input $/M

$0.28

Output $/M

$1.13

Context

131K

Providers

1

Lfm2.5 8B A1B

liquidai/lfm2.5-8b-a1b

Input $/M

$0.03

Output $/M

$0.12

Context

128K

Providers

1

Step 3.7 Flash

stepfun/step-3.7-flash

Step 3.7 Flash is StepFun's latest high-efficiency multimodal Mixture-of-Experts model. It pairs a 196B-parameter language backbone with a vision encoder for native image and video understanding, activating roughly 11B parameters...

Input $/M

$0.20

Output $/M

$1.15

Context

262K

Providers

2

Ring 2.6 1T

inclusionai/ring-2.6-1t

Ring-2.6-1T is a 1T-parameter-scale thinking model with 63B active parameters, built for real-world agent workflows that require both strong capability and operational efficiency. It is optimized for coding agents, tool...

Input $/M

$0.30

Output $/M

$2.50

Context

262K

Providers

1

Seed 2.0 Code

bytedance/seed-2.0-code

A coding model optimized for real-world development environments, with reliable tool use in common IDEs such as Claude Code. It delivers strong front-end performance and supports Skills.

Input $/M

$0.50

Output $/M

$3.00

Context

256K

Providers

1

Ling 2.6 1T

inclusionai/ling-2.6-1t

Ling-2.6-1T is an instant (instruct) model from inclusionAI and the company’s trillion-parameter flagship, designed for real-world agents that require fast execution and high efficiency at scale. It uses a “fast...

Input $/M

$0.30

Output $/M

$2.50

Context

262K

Providers

1

Ling 2.6 Flash

inclusionai/ling-2.6-flash

Ling-2.6-flash is an instant (instruct) model from inclusionAI with 104B total parameters and 7.4B active parameters, designed for real-world agents that require fast responses, strong execution, and high token efficiency....

Input $/M

$0.10

Output $/M

$0.30

Context

262K

Providers

1

Seed 2.0 Pro

bytedance/seed-2.0-pro

Built for the Agent era, it delivers stable performance in complex reasoning and long-horizon tasks, including multi-step planning, visual-text reasoning, video understanding, and advanced analysis.

Input $/M

$0.50

Output $/M

$3.00

Context

256K

Providers

1

Kat Coder Pro V2

kwaipilot/kat-coder-pro-v2

KAT-Coder-Pro V2 is the latest high-performance model in KwaiKAT’s KAT-Coder series, designed for complex enterprise-grade software engineering and SaaS integration. It builds on the agentic coding strengths of earlier versions,...

Input $/M

$0.30

Output $/M

$1.20

Context

262K

Providers

1

Seed 2.0 Mini

bytedance/seed-2.0-mini

Built for low-latency, high-concurrency, cost-sensitive use cases, with flexible deployment, four-tier thinking, and multimodal

Input $/M

$0.10

Output $/M

$0.40

Context

262K

Providers

2

Seed 1.8

bytedance/seed-1.8

Optimized specifically for multimodal agent scenarios. It features enhanced agent capabilities, upgraded multimodal comprehension, and more flexible context management.

Input $/M

$0.25

Output $/M

$2.00

Context

256K

Providers

1

Voyage 4

voyage/voyage-4

범용 임베딩 · 1024차원(가변) · voyage-4 계열은 임베딩 공간 공유

Input $/M

$0.06

Context

32K

Providers

1

Voyage 4 Large

voyage/voyage-4-large

최고 품질·다국어 · 1024차원(가변) · text-embedding-3-large보다 저렴

Input $/M

$0.12

Context

32K

Providers

1

Voyage 4 Lite

voyage/voyage-4-lite

지연·비용 최적 · 1024차원(256/512/2048 가변)

Input $/M

$0.02

Context

32K

Providers

1

Kat Coder Pro

kwaipilot/kat-coder-pro

KAT-Coder-Pro V2 by KwaiKAT is a non-reasoning model optimized for agentic coding. It delivers strong performance on reasoning-style tasks while requiring significantly fewer output tokens than peer models. With the 1210 release, it achieved a score of 64 on the Artificial Analysis Intelligence Index, placing it in the global Top 10 and ranking first among all non-reasoning models.

Input $/M

$0.30

Output $/M

$1.20

Context

256K

Providers

1

K EXAONE 236B A23B

lgai-exaone/k-exaone-236b-a23b

Input $/M

$0.20

Output $/M

$0.80

Context

Providers

1

MiMo V2 Flash

xiaomimimo/mimo-v2-flash

Xiaomi MiMo-V2-Flash is a proprietary MoE model developed by Xiaomi, designed for extreme inference efficiency with 309B total parameters (15B active). By incorporating an innovative Hybrid attention architecture and multi-layer MTP inference acceleration, it ranks among the top 2 global open-source models across multiple Agent benchmarks. Its coding capabilities surpass all open-source models and rival the industry-leading closed-source model, Claude 4.5 Sonnet—yet at only 2.5% of the inference cost and with 2x the generation speed, successfully pushing the limits of both model performance and efficiency.

Input $/M

$0.11

Output $/M

$0.33

Context

262K

Providers

1

Cogito V2 1 671B

deepcogito/cogito-v2-1-671b

Cogito v2.1 671B MoE represents one of the strongest open models globally, matching performance of frontier closed and open models. This model is trained using self play with reinforcement learning...

Input $/M

$1.25

Output $/M

$1.25

Context

164K

Providers

1

ERNIE 4.5 VL 28B A3B Thinking

baidu/ernie-4.5-vl-28b-a3b-thinking

Built upon the powerful ERNIE-4.5-VL-28B-A3B architecture, the newly upgraded ERNIE-4.5-VL-28B-A3B-Thinking achieves a remarkable leap forward in multimodal reasoning capabilities. 🧠✨ Through an extensive mid-training phase, the model absorbed a vast and highly diverse corpus of premium visual-language reasoning data. This massive-scale training process dramatically boosted the model’s representation power while deepening the semantic alignment between visual and language modalities—unlocking unprecedented capabilities in nuanced visual-textual reasoning. 📊 The model leverages cutting-edge multimodal reinforcement learning techniques on verifiable tasks, integrating GSPO and IcePop strategies to stabilize MoE training combined with dynamic difficulty sampling for exceptional learning efficiency. ⚡ Responding to strong community demand, we’ve significantly strengthened the model’s grounding performance with improved instruction-following capabilities, making visual grounding functions more accessible than eve

Input $/M

$0.39

Output $/M

$0.39

Context

131K

Providers

1

ERNIE 4.5 21B A3B Thinking

baidu/ernie-4.5-21b-a3b-thinking

ERNIE-4.5-21B-A3B-Thinking is a text-based Mixture of Experts (MoE) post-training model featuring 21B total parameters with 3B active parameters per token. It delivers enhanced performance on reasoning tasks, including logical reasoning, mathematics, science, coding, text generation, and academic benchmarks that typically require human expertise. The model offers efficient tool utilization capabilities and supports up to 128K tokens for long-context understanding.

Input $/M

$0.07

Output $/M

$0.28

Context

131K

Providers

1

Morph V3 Large

morph/morph-v3-large

Morph's high-accuracy apply model for complex code edits. ~4,500 tokens/sec with 98% accuracy for precise code transformations. The model requires the prompt to be in the following format: <instruction>{instruction}</instruction> <code>{initial_code}</code>...

Input $/M

$0.90

Output $/M

$1.90

Context

16K

Providers

1

Morph V3 Fast

morph/morph-v3-fast

Morph's fastest apply model for code edits. ~10,500 tokens/sec with 96% accuracy for rapid code transformations. The model requires the prompt to be in the following format: <instruction>{instruction}</instruction> <code>{initial_code}</code> <update>{edit_snippet}</update>...

Input $/M

$0.80

Output $/M

$1.20

Context

16K

Providers

1

ERNIE 4.5 VL 424B A47B

baidu/ernie-4.5-vl-424b-a47b

ERNIE-4.5-VL-424B-A47B is a multimodal Mixture-of-Experts (MoE) model from Baidu’s ERNIE 4.5 series, featuring 424B total parameters with 47B active per token. It is trained jointly on text and image data...

Input $/M

$0.42

Output $/M

$1.25

Context

123K

Providers

1

ERNIE 4.5 21B A3B

baidu/ernie-4.5-21b-a3b

The ERNIE 4.5 series of open-source models adopts a Mixture-of-Experts (MoE) architecture, representing an innovative multimodal heterogeneous model structure. It achieves cross-modal knowledge fusion through a parameter-sharing mechanism while retaining dedicated parameter spaces for individual modalities. This architecture is particularly well-suited for the continuous pre-training paradigm from large language models to multimodal models, significantly enhancing multimodal understanding capabilities while maintaining or even improving performance in text-based tasks. The models are efficiently trained, inferred, and deployed using the PaddlePaddle deep learning framework. During the pre-training of large language models, the Model FLOPs Utilization (MFU) reaches 47%. Experimental results demonstrate that this series of models achieves state-of-the-art (SOTA) performance across multiple text and multimodal benchmarks, with particularly outstanding results in instruction following, world knowledge memorizatio

Input $/M

$0.07

Output $/M

$0.28

Context

120K

Providers

1

ERNIE 4.5 300B A47B Paddle

baidu/ernie-4.5-300b-a47b-paddle

The ERNIE 4.5 series of open-source models adopts a Mixture-of-Experts (MoE) architecture, representing an innovative multimodal heterogeneous model structure. It achieves cross-modal knowledge fusion through a parameter-sharing mechanism while retaining dedicated parameter spaces for individual modalities. This architecture is particularly well-suited for the continuous pre-training paradigm from large language models to multimodal models, significantly enhancing multimodal understanding capabilities while maintaining or even improving performance in text-based tasks. The models are efficiently trained, inferred, and deployed using the PaddlePaddle deep learning framework. During the pre-training of large language models, the Model FLOPs Utilization (MFU) reaches 47%. Experimental results demonstrate that this series of models achieves state-of-the-art (SOTA) performance across multiple text and multimodal benchmarks, with particularly outstanding results in instruction following, world knowledge memorizatio

Input $/M

$0.28

Output $/M

$1.10

Context

123K

Providers

1

ERNIE 4.5 VL 28B A3B

baidu/ernie-4.5-vl-28b-a3b

The ERNIE 4.5 series of open-source models adopts a Mixture-of-Experts (MoE) architecture, representing an innovative multimodal heterogeneous model structure. It achieves cross-modal knowledge fusion through a parameter-sharing mechanism while retaining dedicated parameter spaces for individual modalities. This architecture is particularly well-suited for the continuous pre-training paradigm from large language models to multimodal models, significantly enhancing multimodal understanding capabilities while maintaining or even improving performance in text-based tasks. The models are efficiently trained, inferred, and deployed using the PaddlePaddle deep learning framework. During the pre-training of large language models, the Model FLOPs Utilization (MFU) reaches 47%. Experimental results demonstrate that this series of models achieves state-of-the-art (SOTA) performance across multiple text and multimodal benchmarks, with particularly outstanding results in instruction following, world knowledge memorizatio

Input $/M

$0.14

Output $/M

$0.56

Context

30K

Providers

1

Phi 4

microsoft/phi-4

[Microsoft Research](/microsoft) Phi-4 is designed to perform well in complex reasoning tasks and can operate efficiently in situations with limited memory or where quick responses are needed. At 14 billion...

Input $/M

$0.07

Output $/M

$0.14

Context

16K

Providers

1

Voyage Code 3

voyage/voyage-code-3

코드 검색 특화 임베딩 · 1024차원(가변)

Input $/M

$0.18

Context

32K

Providers

1

Wizardlm 2 8x22b

microsoft/wizardlm-2-8x22b

WizardLM-2 8x22B is Microsoft AI's most advanced Wizard model. It demonstrates highly competitive performance compared to leading proprietary models, and it consistently outperforms all existing state-of-the-art opensource models. It is...

Input $/M

$0.62

Output $/M

$0.62

Context

66K

Providers

1

Bge M3

baai/bge-m3

Input $/M

$0.01

Context

8K

Providers

2

Bge En Icl

baai/bge-en-icl

Input $/M

$0.01

Context

8K

Providers

2

Nova 3 En

deepgram/nova-3-en

Price (per min)

$0.0015/min

Context

Providers

1

Nova 3 Multi

deepgram/nova-3-multi

Price (per min)

$0.0015/min

Context

Providers

1

Flux

deepgram/flux

Price (per min)

$0.0015/min

Context

Providers

1

Multilingual E5 Large Instruct

intfloat/multilingual-e5-large-instruct

Input $/M

$0.01

Context

1K

Providers

3