gemma-4-26B-A4B-it

gemma

Efficient, MoE variant of Gemma 4. Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input and generating text output.

Context
262K
Input price
$0.07/1M tokens
Output price
$0.34/1M tokens

Capabilities

vision
tool call
structured output
reasoning
json mode
streaming
fine tuning
batch

Details

Provider DeepInfra
Creator Google AI
Familygemma-4
Licensegemma
Parameters
Statusactive
Input modalitiestext, image
Output modalitiestext
Architecture
Knowledge cutoff
Training data cutoff
Release date
Deprecation date
Typechat
Reasoning tokensYes
Max input
Open weightYes
Sourceofficial
Last updated

Tools

Function Callingfunction_calling
Call external functions and APIs

Endpoints

Chat CompletionsPOST
Generate chat responses with messageshttps://api.deepinfra.com/v1/openai/v1/chat/completions

Pricing

Input
$0.07
Output
$0.34
Cache write
Cache read
Batch in
Batch out

Price across providers

ProviderInputOutputSince launch
DeepInfracheapest$0.07$0.34
OpenRouter$0.09$0.3↓31%
Novita AI$0.13$0.4
Vercel AI Gateway$0.13$0.4

Family Comparison: gemma-4

ModelContextMax outPricing
gemma-4-E4B-it131K$0.02$0.1
gemma-4-26B-A4B-it262K$0.07$0.34
gemma-4-31B-it-turbo262K$0.09$0.34
gemma-4-31B-it262K$0.13$0.38
gemma-4-31B-it-Ultra131K$0.27$0.76

API

GET/v1/models/deepinfra/google/gemma-4-26B-A4B-it