← Back to registry
OpenAI

gpt-oss:20b-cloud

Low usage

gpt-oss-20b is OpenAI's smaller open-weight model, designed for lower latency and local or specialized use cases, with MXFP4 quantization enabling it to run on systems with as little as 16GB memory.

toolsthinkingcloudCommunity r/OpenAI
Context window128K
ModalitiesText
Size20B
Pulls12.3M
Tags5
Updated10 months ago
ollama run gpt-oss:20b-cloud

Feature highlights — shared with gpt-oss-120b

Quantization

The models are post-trained with MoE weights quantized to MXFP4 format (4.25 bits per parameter). The MoE weights cover 90+% of the total parameter count, so the 20B model runs on systems with as little as 16GB memory. Ollama supports MXFP4 natively without additional quantization.

Best practices

Specs sourced from ollama.com/library/gpt-oss-20b · Retrieved 2026-08-29