Provider logo

GLM 4 Air 0111

glm-4-air-0111
Provider logo

GLM 4 Air 0111

glm-4-air-0111

MiniMax's flagship model with a 1M token context window

Added Jan 11, 2025

Context Window

128.0K

Max Output

4.1K

Avg output tokens (7d)

226 tokens

32%

Input Price (Auto)

$0.14/1M

Output Price (Auto)

$0.14/1M

Cache Read (Auto)

$0.070/1M

Benchmarks

Performance metrics and benchmarks

No benchmark data is available yet for this model.

Providers

Auto routing is available for this model. Explicit provider selection is not available.

Loading provider options…

Compare GLM 4 Air 0111 with similar models from the same provider or model family.

GLM 4.5 Air

z-ai/GLM-4.5-Air

GLM-4.5-Air is a 106B total / 12B active parameter model designed to unify frontier reasoning, coding, and agentic capabilities. On the SWE-bench Verified benchmark, it delivers the best performance at its scale with a competitive performance-to-cost ratio.

GLM 4 Plus 0111

glm-4-plus-0111

GLM 4 Plus 0111 is a 1M token context window model

GLM 4.5 Air (Thinking)

z-ai/GLM-4.5-Air:thinking

GLM-4.5-Air with thinking mode enabled for enhanced reasoning capabilities. Shows step-by-step thought process.

GLM 5.3 TEE

TEE/glm-5.3

GLM-5.3 is Z.AI's open-weight reasoning model for complex software engineering, autonomous agents, vulnerability research, and long-horizon tasks. Reasoning is always on; choose low, high, or max reasoning effort (defaults to max). This text-only TEE deployment is verified through the selected provider: Redpill attestation with signed completion receipts or the official Tinfoil SDK's ATC/EHBP verification.

GLM 5.3 Flash TEE

TEE/glm-5.3-flash

GLM-5.3 Flash is Z.AI's natively multimodal 320B MoE reasoning model with 18B active parameters. This TEE deployment is verified through the selected provider: Redpill attestation with signed completion receipts or the official Tinfoil SDK's ATC/EHBP verification.

GLM 5.3 Flash

z-ai/glm-5.3-flash

ox-alpha out of stealth! GLM-5.3 Flash is Z.ai's first natively multimodal GLM-5 model, with 320B total parameters and just 18B active parameters for efficient coding, agentic work, and precise 1M-token context. Its hybrid sparse-and-linear attention architecture helps it outperform GLM-5.2 at one-tenth the price while approaching Claude Opus 4.8 on coding and agentic benchmarks.