modelgrep
google logo

Google: Gemma 4 31B

google/gemma-4-31b-it

Cheaper than 87% of paidReasoningToolsJSONVision
Use via OpenRouter ↗
Intelligence
Design Elo
Speed
tokens/sec
Latency
first token
Input price
$0.100
44th cheapest
Context
262K
262K max out

How it compares

Cheaper than87%
of all ranked models

Overview

Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...

Providers & pricing (18)

ProviderIn $/MOut $/MUptime
DeepInfrafp4$0.090$0.34097.7%
CoreWeavebf16$0.100$0.34099.8%
OpenInferencebf16$0.100$0.35097.7%
Venicebf16$0.120$0.36099.4%
Chutesfp4$0.120$0.37097.4%
DeepInfrafp8$0.130$0.38091.2%
SiliconFlowfp8$0.130$0.40097.1%
Novitabf16$0.140$0.40097.8%
Friendli$0.140$0.40099.9%
Morphfp4$0.140$0.40098.3%
Crusoe$0.140$0.40099.1%
Parasailfp8$0.150$0.40098.2%
Phala$0.150$0.46098.2%
ModelRunfp4$0.220$0.55099.8%
Together$0.280$0.86092.7%
SambaNova$0.380$1.1599.9%
Together$0.390$0.97085.8%
Cerebrasfp16$0.990$1.49100%

Specifications

Context window262K
Max output262K
Knowledge cutoff
Input modalitiesimage, text, video
Output modalitiestext
Prompt caching
Cache read price$0.100/M
ModeratedNo

Gemma 4 31B FAQ

How much does Gemma 4 31B cost?

Gemma 4 31B costs $0.100 per million input tokens and $0.340 per million output tokens via OpenRouter, making it 44th cheapest of 332 paid models.

What is Gemma 4 31B's context window?

Gemma 4 31B supports a 262K-token context window and can output up to 262K tokens. It accepts image, text, video input.

Compare head-to-head