Nemotron 3 Nano is trained from scratch by NVIDIA as a unified model for both reasoning and non-reasoning tasks, using a hybrid Mixture-of-Experts architecture of 23 Mamba-2 and MoE layers plus 6 Attention layers.
| Tag | Parameters | Context | Modalities |
|---|---|---|---|
| nemotron-3-nano:30b-cloud | 30B (3.5B active) | 1M | Text |
| nemotron-3-nano:4b (local, 2.8GB) | 4B | 256K | Text |
| nemotron-3-nano:30b (local, 24GB) | 30B (3.5B active) | 1M | Text |