NVIDIA-Nemotron-Nano-9B-v2

Updated
24.08.2026
Thinking
Reasoning
Code
Tools
Multilingual

A 9B hybrid Mamba2-Transformer reasoning model from NVIDIA with toggleable thinking, 128K context, and tool calling.

At a glance

  • License: NVIDIA Open Model License, commercial use allowed
  • Parameters: 8.9B
  • Context length: 128K tokens
  • Modalities: Text input, text output
  • Minimum hardware: 6 GB of memory (IQ3_M GGUF, 5.21 GB)

What is NVIDIA-Nemotron-Nano-9B-v2?

NVIDIA-Nemotron-Nano-9B-v2 is an 8.9B language model that NVIDIA trained from scratch as a unified reasoning and chat model: it writes a reasoning trace first, then the final answer, and the trace can be switched off from the system prompt. The architecture is a Mamba2-Transformer hybrid, mostly Mamba-2 and MLP layers with just four attention layers. NVIDIA published the weights on Hugging Face on August 18, 2025 under the NVIDIA Open Model License, and the GGUF builds start around 5 GB, small enough for an 8 GB machine.

SpecificationNVIDIA-Nemotron-Nano-9B-v2
Total parameters8.89B (BF16 checkpoint)
ArchitectureMamba2-Transformer hybrid, four attention layers
Context window128K tokens
ModalitiesText input, text output
ReasoningOn by default, /think and /no_think toggles, runtime thinking budget
LanguagesEnglish, German, Spanish, French, Italian, Japanese
Training dataAbout twenty trillion tokens, cutoff September 2024
Release dateAugust 18, 2025
LicenseNVIDIA Open Model License

Reasoning control is more granular here than in most local models. Putting /think or /no_think in the system prompt, or in any user message for per-turn control, toggles the trace, and NVIDIA notes that turning it off costs a little accuracy on harder prompts. On top of that sits a runtime thinking budget: cap the trace at a token count and the model ends its reasoning near the cap, which matters when a response has a latency target. NVIDIA recommends temperature 0.6 with top_p 0.95 when reasoning is on, and greedy decoding when it is off.

NVIDIA-Nemotron-Nano-9B-v2 benchmarks

NVIDIA's launch numbers, produced with NeMo-Skills in Reasoning-On mode (RULER is the exception, measured with reasoning off), put the model against Qwen3-8B:

BenchmarkNVIDIA-Nemotron-Nano-9B-v2Qwen3-8B
AIME25
Competition math
72.1%69.3%
MATH500
Math problems
97.8%96.3%
GPQA
Expert science
64.0%59.6%
LCB
Competitive coding
71.1%59.5%
BFCL v3
Tool calling
66.9%66.3%
IFEval
Instruction following
90.3%89.4%
HLE
Expert questions
6.5%4.4%
RULER (128K)
Long context
78.9%74.1%

The 9B leads Qwen3-8B on all eight rows, though the tool calling and instruction following margins are under a point. The clear gaps are coding, where LCB jumps from 59.5 to 71.1, and long context, where it holds 78.9 on RULER at 128K.

NVIDIA-Nemotron-Nano-9B-v2 hardware requirements

The system requirement to check is memory. NVIDIA publishes the original BF16 safetensors; the GGUF ladder below comes from bartowski/nvidia_NVIDIA-Nemotron-Nano-9B-v2-GGUF.

MemoryBuild to pickFile size
6 GBIQ3_M5.21 GB
8 GBQ4_K_M6.53 GB
12 GBQ6_K9.14 GB
16 GBQ8_09.46 GB
24 GB and upbf1617.79 GB

The low end of this ladder is unusually flat: every build from IQ2_S to Q3_K_XL lands between 4.96 and 5.78 GB, so the two-bit files barely undercut the three-bit ones and Q4_K_M is the sensible floor. When two builds both fit, take the larger one. If the format is new to you, start with what GGUF is.

How to run NVIDIA-Nemotron-Nano-9B-v2 in Atomic Chat

Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.

  1. Download Atomic Chat for your platform and open it.
  2. Search for NVIDIA-Nemotron-Nano-9B-v2 in the model browser and open Download Options.
  3. Pick the build that fits the memory you have, then start a chat.

For the rest of the family, see every Nemotron model you can run locally, including the larger Llama-3.3-Nemotron-Super-49B-v1.5.

NVIDIA-Nemotron-Nano-9B-v2 license

NVIDIA-Nemotron-Nano-9B-v2 is released under the NVIDIA Open Model License Agreement, NVIDIA's own open-weights license rather than a standard one like Apache 2.0. NVIDIA states the model is ready for commercial use, so you can build products on it and run it on your own hardware under those terms.

Get the weights from Hugging Face

pip install -U "transformers>=4.48" accelerate
huggingface-cli download nvidia/NVIDIA-Nemotron-Nano-9B-v2
curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "nvidia/NVIDIA-Nemotron-Nano-9B-v2",
    "messages": [{"role": "user", "content": "Solve: integrate x^2 dx"}],
    "temperature": 0.6,
    "top_p": 0.95,
    "max_tokens": 1024
  }'
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("nvidia/NVIDIA-Nemotron-Nano-9B-v2", trust_remote_code=True)
tokenizer = AutoTokenizer.from_pretrained("nvidia/NVIDIA-Nemotron-Nano-9B-v2")
messages = [{"role": "user", "content": "Explain the Mamba-2 layer in one paragraph."}]
inputs = tokenizer.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt")
print(tokenizer.decode(model.generate(inputs, max_new_tokens=1024)[0]))
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "http://localhost:8000/v1", apiKey: "local" });
const res = await client.chat.completions.create({
  model: "nvidia/NVIDIA-Nemotron-Nano-9B-v2",
  messages: [{ role: "user", content: "Write a Python function to check if a number is prime." }],
  temperature: 0.6,
  top_p: 0.95,
});
console.log(res.choices[0].message.content);
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

NVIDIA-Nemotron-Nano-9B-v2 is a 9-billion-parameter language model trained from scratch by NVIDIA and released in August 2025. It is a unified reasoning and chat model that can generate an explicit thinking trace before its final answer, with that reasoning behavior controllable through the system prompt. It uses the Nemotron-H hybrid architecture, combining Mamba-2 layers with a small number of Transformer attention layers for faster long-output inference.

NVIDIA distilled the model specifically so it can run inference at the full 128K context on a single NVIDIA A10G GPU with 22 GiB of memory in bfloat16. At 4-bit quantization the weights fit in roughly 10-12 GB of VRAM, which brings it within reach of consumer cards like an RTX 3060 12 GB or RTX 4070. It runs through Hugging Face transformers, vLLM, and NVIDIA NIM.

Yes. The model weights are published openly on Hugging Face and governed by the NVIDIA Open Model License Agreement, which NVIDIA states makes the model ready for commercial use. It is not released under a standard OSI license such as Apache 2.0 or MIT, so deployment is permitted under the terms of NVIDIA's own license rather than an unrestricted open-source one. It is also offered free to try through hosts like OpenRouter.

The model supports a context length of up to 128K tokens for both input and output. Its officially supported languages are English, German, Spanish, French, Italian, and Japanese. It is primarily intended for English and coding tasks, with the other five languages handled as secondary capabilities.

On NVIDIA's published reasoning-on benchmarks, Nemotron-Nano-9B-v2 scores at or above Qwen3-8B: 72.1% vs 69.3% on AIME25, 97.8% vs 96.3% on MATH500, 64.0% vs 59.6% on GPQA, and 71.1% vs 59.5% on LiveCodeBench. Beyond accuracy, its Mamba2-Transformer hybrid design gives it up to roughly 6x higher inference throughput than comparable Transformer models in long-output reasoning settings.