← Back to registry
Z.ai

glm-5.3-flash:cloud

Medium usage

GLM-5.3-Flash is the first natively multimodal model in Z.ai's GLM-5 series. With 320B total parameters and just 18B active, it outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the cost, while approaching Claude Opus 4.8 on coding and agentic benchmarks. It runs on Ollama's hosted infrastructure in the United States and Europe.

visiontoolsthinkingcloudCommunity r/ZaiGLM
Context window1M
ModalitiesText, Image
Size321B
Pulls17.5K
Tags1
Updated2 days ago
ollama run glm-5.3-flash:cloud

Cloud tags

TagParametersActiveContextModalities
glm-5.3-flash:cloud (default)320B18B1MText, Image

A single cloud tag serves the 320B MoE model (18B active) on Ollama's hosted infrastructure in the United States and Europe, with zero data retention per Ollama's privacy policy.

Benchmark highlights — GLM-5.3-Flash vs. GLM-5.2

BenchmarkGLM-5.3-FlashGLM-5.2
Terminal Bench 2.184.381.0
DeepSWE v1.163.446.2
NL2Repo56.348.9
Toolathlon Verified78.459.9
AutomationBench v1.0.648.826.2
AA Intelligence Index v4.1.157

Key features

Best practices

Specs sourced from ollama.com/library/glm-5.3-flash · Retrieved 2026-08-29