OpenAI GPT OSS

apache-2.0120B

`gpt-oss-120b`is our most powerful open-weight model, which fits into a single H100 GPU (117B parameters with 5.1B active parameters).

Intelligence
Reasoning
Speed
Context
131K
Max output
40K
Input price
$0.35/1M tokens
Output price
$0.75/1M tokens

Capabilities

vision
tool call
structured output
reasoning
json mode
streaming
fine tuning
batch

Details

Model IDgpt-oss-120b
Provider Cerebras
Creator OpenAI
Familygpt-oss
Licenseapache-2.0
Parameters120B
Statusactive
Input modalitiestext
Output modalitiestext
Architecture—
Knowledge cutoff
Training data cutoff—
Release date—
Deprecation date—
Typechat
Reasoning tokensYes
Max input—
Open weightYes
Prompt cachingSupported
Sourceofficial
Last updated

Tools

Function Callingfunction_calling
Call external functions and APIs

Endpoints

Chat CompletionsPOST
Generate chat responses with messageshttps://api.cerebras.ai/v1/chat/completions
CompletionsPOST
Legacy text completionhttps://api.cerebras.ai/v1/completions

Pricing

Input
$0.35
Output
$0.75
Cache write
—
Cache read
—
Batch in
—
Batch out
—

Price across providers

ProviderInputOutputSince launch
DeepInfracheapest$0.04$0.17↓5%
OpenRoutercheapest$0.04$0.17↓5%
Cloudflare AI Gateway$0.04$0.19—
Novita AI$0.05$0.25—
Groq$0.15$0.6—
SambaNova Cloud$0.22$0.59—
Cerebras$0.35$0.75—
Vercel AI Gateway$0.35$0.75—

API

GET/v1/models/cerebras/gpt-oss-120b

OpenAI GPT OSS

gpt-oss-120b

`gpt-oss-120b`is our most powerful open-weight model, which fits into a single H100 GPU (117B parameters with 5.1B active parameters).

Changes · 5 entries
gpt-oss-120bupdate118451aSep 24, 2026, 07:08 AM
endpoints["chat_completions"]→["chat_completions","completions"]
gpt-oss-120bupdate1cd659eMar 24, 2026, 04:19 AM
capabilitiesjson modeyes
gpt-oss-120bupdate2b100acMar 23, 2026, 05:30 AM
parameters120 billion→120
reasoning_tokensyes
open_weightyes
tools["function_calling"]
gpt-oss-120bupdate7a4f1f6Mar 22, 2026, 05:43 AM
namegpt-oss-120b→OpenAI GPT OSS
capabilitiesvisionyesprompt cachingyes
model_typechat
statusactive
endpoints["chat_completions"]
model_card_urlhttps://openai.com/index/gpt-oss-model-card/
tokens_per_second3000
parameters120 billion
precisionFP16/FP8 (weights only)
huggingface_urlhttps://huggingface.co/openai/gpt-oss-120b
gpt-oss-120bcreate7f5c66fMar 21, 2026, 05:16 AM