← Back to registry
NVIDIA
nemotron-3-ultra:cloud
High usage
NVIDIA Nemotron 3 Ultra is a 550 billion parameter (55B active) open model built for long-running, agentic workflows with fast and affordable performance across hundreds of tool calls.
toolsthinkingcloud
Context window256K
ModalitiesText
Size550B
Pulls62.8K
Tags1
Updated2 months ago
ollama run nemotron-3-ultra:cloud
Model highlights
- Built for long-running agents — tuned for agent orchestration, coding agents, deep research, and complex enterprise workflows that run across hundreds of steps.
- Long context — keep entire codebases, long tool histories, and research trails in context without losing the thread.
- Frontier reasoning, high efficiency — 550B total parameters with only 55B active per token, optimized for NVFP4, NVIDIA's 4-bit floating point format that packs the model into less memory and runs faster.
Benchmarks
Nemotron 3 Ultra leads on accuracy across agent productivity, instruction following, and long-context tasks among open models, while delivering leading throughput — saving up to 30% on costs compared to other leading open models.
Best practices
- Agent orchestration — designed for hundreds of steps per run; well-suited to deep research and enterprise workflows.
- Cost efficiency — NVFP4 optimization saves up to 30% on costs vs. other leading open models.
- Long tool histories — the long context keeps multi-step tool trails coherent.