modelgrep
z-ai logo

Z.ai: GLM 4.7 Flash

z-ai/glm-4.7-flash

116th smartest of 182Cheaper than 93% of paidReasoningToolsJSON
Use via OpenRouter ↗
Intelligence
22.9
116th of 182
Design Elo
1218
Website
Speed
tokens/sec
Latency
first token
Input price
$0.060
24th cheapest
Context
203K
16K max out

How it compares

Smarter than36%
of all ranked models
Cheaper than93%
of all ranked models

Overview

As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agentic coding use cases, strengthening coding capabilities, long-horizon task planning,...

Benchmarks

independent · Artificial Analysis & Design Arena
Artificial Analysis66th percentile
Intelligence Index
22.9
GPQA Diamond
58%
Humanity's Last Exam
7%
SciCode
34%
Tau²-Bench (agentic)
99%
Design Arena · Elo7,792 tournaments
Website
1218
3D
1178
Data Viz
1152
svg
1086

Providers & pricing (4)

ProviderIn $/MOut $/MUptime
DeepInfrabf16$0.060$0.40090%
Cloudflare$0.060$0.40099.3%
Novitabf16$0.070$0.40092.2%
Venicefp8$0.125$0.500

Specifications

Context window203K
Max output16K
Knowledge cutoff
Input modalitiestext
Output modalitiestext
Prompt caching
Cache read price$0.010/M
ModeratedNo

GLM 4.7 Flash FAQ

How much does GLM 4.7 Flash cost?

GLM 4.7 Flash costs $0.060 per million input tokens and $0.400 per million output tokens via OpenRouter, making it 24th cheapest of 332 paid models.

How smart is GLM 4.7 Flash?

GLM 4.7 Flash scores 22.9 on the Artificial Analysis Intelligence Index, ranking 116th of 182 benchmarked models, with a GPQA Diamond score of 58%.

What is GLM 4.7 Flash's context window?

GLM 4.7 Flash supports a 203K-token context window and can output up to 16K tokens. It accepts text input.

Compare head-to-head