What is NVIDIA-Nemotron-Nano-9B-v2?
NVIDIA-Nemotron-Nano-9B-v2 is an 8.9B language model that NVIDIA trained from scratch as a unified reasoning and chat model: it writes a reasoning trace first, then the final answer, and the trace can be switched off from the system prompt. The architecture is a Mamba2-Transformer hybrid, mostly Mamba-2 and MLP layers with just four attention layers. NVIDIA published the weights on Hugging Face on August 18, 2025 under the NVIDIA Open Model License, and the GGUF builds start around 5 GB, small enough for an 8 GB machine.
| Specification | NVIDIA-Nemotron-Nano-9B-v2 |
|---|---|
| Total parameters | 8.89B (BF16 checkpoint) |
| Architecture | Mamba2-Transformer hybrid, four attention layers |
| Context window | 128K tokens |
| Modalities | Text input, text output |
| Reasoning | On by default, /think and /no_think toggles, runtime thinking budget |
| Languages | English, German, Spanish, French, Italian, Japanese |
| Training data | About twenty trillion tokens, cutoff September 2024 |
| Release date | August 18, 2025 |
| License | NVIDIA Open Model License |
Reasoning control is more granular here than in most local models. Putting /think or /no_think in the system prompt, or in any user message for per-turn control, toggles the trace, and NVIDIA notes that turning it off costs a little accuracy on harder prompts. On top of that sits a runtime thinking budget: cap the trace at a token count and the model ends its reasoning near the cap, which matters when a response has a latency target. NVIDIA recommends temperature 0.6 with top_p 0.95 when reasoning is on, and greedy decoding when it is off.
NVIDIA-Nemotron-Nano-9B-v2 benchmarks
NVIDIA's launch numbers, produced with NeMo-Skills in Reasoning-On mode (RULER is the exception, measured with reasoning off), put the model against Qwen3-8B:
| Benchmark | NVIDIA-Nemotron-Nano-9B-v2 | Qwen3-8B |
|---|---|---|
AIME25 Competition math | 72.1% | 69.3% |
MATH500 Math problems | 97.8% | 96.3% |
GPQA Expert science | 64.0% | 59.6% |
LCB Competitive coding | 71.1% | 59.5% |
BFCL v3 Tool calling | 66.9% | 66.3% |
IFEval Instruction following | 90.3% | 89.4% |
HLE Expert questions | 6.5% | 4.4% |
RULER (128K) Long context | 78.9% | 74.1% |
The 9B leads Qwen3-8B on all eight rows, though the tool calling and instruction following margins are under a point. The clear gaps are coding, where LCB jumps from 59.5 to 71.1, and long context, where it holds 78.9 on RULER at 128K.
NVIDIA-Nemotron-Nano-9B-v2 hardware requirements
The system requirement to check is memory. NVIDIA publishes the original BF16 safetensors; the GGUF ladder below comes from bartowski/nvidia_NVIDIA-Nemotron-Nano-9B-v2-GGUF.
| Memory | Build to pick | File size |
|---|---|---|
| 6 GB | IQ3_M | 5.21 GB |
| 8 GB | Q4_K_M | 6.53 GB |
| 12 GB | Q6_K | 9.14 GB |
| 16 GB | Q8_0 | 9.46 GB |
| 24 GB and up | bf16 | 17.79 GB |
The low end of this ladder is unusually flat: every build from IQ2_S to Q3_K_XL lands between 4.96 and 5.78 GB, so the two-bit files barely undercut the three-bit ones and Q4_K_M is the sensible floor. When two builds both fit, take the larger one. If the format is new to you, start with what GGUF is.
How to run NVIDIA-Nemotron-Nano-9B-v2 in Atomic Chat
Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.
- Download Atomic Chat for your platform and open it.
- Search for NVIDIA-Nemotron-Nano-9B-v2 in the model browser and open Download Options.
- Pick the build that fits the memory you have, then start a chat.
For the rest of the family, see every Nemotron model you can run locally, including the larger Llama-3.3-Nemotron-Super-49B-v1.5.
NVIDIA-Nemotron-Nano-9B-v2 license
NVIDIA-Nemotron-Nano-9B-v2 is released under the NVIDIA Open Model License Agreement, NVIDIA's own open-weights license rather than a standard one like Apache 2.0. NVIDIA states the model is ready for commercial use, so you can build products on it and run it on your own hardware under those terms.
