Warp
ChatInput: TextReleased Apr 7, 2026

Built by Z.ai (Zhipu AI) · China · z.ai

GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. Unlike previous models built around minute-level interactions, GLM-5.1 can work independently and continuously on...

How to use

Just add the suffix. Hard questions still come back from this model; only easy ones drop to something cheaper, and nothing pricier than it gets used.

model: "zhipu/glm-5.1:auto"

Sets no behavior — whatever your key already stores stays in effect.

Without the suffix, the bare name always goes to this model — no routing, so nothing saved.

Context

205K

Max output

128K

Input $/1M

$1.05

Output $/1M

$3.50

Features

Tool callingJSON

Providers

DeepInfraNovitaFriendliBasetenGMI CloudVeniceW&B InferenceFireworksDigitalOceanWaferZ.ai

Providers

Price, latency, and uptime per provider serving this model. Warp tries them in order of how each host has just been behaving, moving to the next on failure. Click a row for regions and data policies.

Loading provider metrics…

Performance

Benchmark quality scores, and measured speed per provider.

Quality

GPQA DiamondPhD-level science questions
85.5%
AIMECompetition mathematics
92.2%
FrontierMathResearch-level mathematics
33.5%
SWE-bench VerifiedReal GitHub issue resolution
74.2%
SciCodeScientific code generation
43.8%
SimpleQA VerifiedFactual accuracy
37.3%

At each model's best reasoning effort

Agent arenaRank on real tool-using sessions (frontier models only)#12+1.19%
Category ranksAmong all rated models
Agent#12Coding#15Korean#35Instruction following#22

Source: LMArena

Speed

Loading provider metrics…

Parameter support

What Warp actually does with each parameter when you call this model. The answer differs by serving provider and by API surface.

Provider