gpt-oss-120b

proprietary120B

gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. The model supports configurable reasoning depth, full chain-of-thought access, and native tool use, including function calling, browsing, and structured output generation.

Intelligence
Reasoning
Speed
Context
131K
Max output
131K
Input price
$0.04/1M tokens
Output price
$0.17/1M tokens

Capabilities

vision
tool call
structured output
reasoning
json mode
streaming
fine tuning
batch

Details

Provider DeepInfra
Creator OpenAI
Familygpt-oss
Licenseproprietary
Parameters120B
Statusactive
Input modalitiestext
Output modalitiestext
Architecture—
Knowledge cutoff
Training data cutoff—
Release date—
Deprecation date—
Typechat
Reasoning tokensYes
Max input—
Open weightYes
Sourceofficial
Last updated

Tools

Function Callingfunction_calling
Call external functions and APIs

Endpoints

Chat CompletionsPOST
Generate chat responses with messageshttps://api.deepinfra.com/v1/openai/v1/chat/completions

Pricing

Input
$0.04
↓5%$0.04
Output
$0.17
↓11%$0.19
Cache write
—
Cache read
—
Batch in
—
Batch out
—
Price history · $/1M input
$0.039
$0.037
Apr 26Jul 9
DateInputOutputChange
Apr 26, 2026$0.039$0.19—
Jul 9, 202653ef452$0.037$0.17↓5%

Input is ↓5% since Apr 26, 2026, from $0.039 to $0.037 per 1M tokens.

Price across providers

ProviderInputOutputSince launch
DeepInfracheapest$0.04$0.17↓5%
OpenRoutercheapest$0.04$0.17↓5%
Cloudflare AI Gateway$0.04$0.19—
Novita AI$0.05$0.25—
Groq$0.15$0.6—
SambaNova Cloud$0.22$0.59—
Cerebras$0.35$0.75—
Vercel AI Gateway$0.35$0.75—

Family Comparison: gpt-oss

ModelContextMax outPricing
gpt-oss-20b131K131K$0.03$0.14
gpt-oss-120b131K131K$0.04$0.17
gpt-oss-120b-Turbo131K—$0.15$0.6
gpt-oss-120b-Ultra131K—$0.2$0.95

API

GET/v1/models/deepinfra/openai/gpt-oss-120b