What is Llama-3.3-Nemotron-Super-49B-v1.5?
Llama-3.3-Nemotron-Super-49B-v1.5 is NVIDIA's 49B reasoning and chat model, a derivative of Meta's Llama-3.3-70B-Instruct released on Hugging Face on July 25, 2025. NVIDIA used Neural Architecture Search (NAS) to compress the 70B reference model into 49B, cutting the memory footprint so the model fits a single GPU at high workloads. It is post-trained for reasoning, human chat preferences and agentic tasks such as RAG and tool calling, and it supports a 128K context.
| Specification | Llama-3.3-Nemotron-Super-49B-v1.5 |
|---|---|
| Total parameters | 49.9B (BF16) |
| Architecture | Dense decoder-only Transformer, NAS variant of Llama 3.3 70B |
| Base model | Meta Llama-3.3-70B-Instruct |
| Context window | 131,072 tokens |
| Modalities | Text input, text output |
| Reasoning | On by default, /no_think in the system prompt switches it off |
| Languages | English and code, plus German, French, Italian, Portuguese, Hindi, Spanish and Thai |
| Release date | July 25, 2025 |
| License | NVIDIA Open Model License + Llama 3.3 Community License |
The NAS pass produces non-standard, non-repetitive blocks: in some, attention is skipped entirely or replaced with a single linear layer, and the FFN expansion ratio changes from block to block. NVIDIA then ran block-wise distillation against the reference model on 40 billion tokens of FineWeb, Buzz-V1.2 and Dolma, followed by a multi-stage post-training pipeline: supervised fine-tuning for math, code, science and tool calling, then RPO for chat, RLVR for reasoning and iterative DPO for tool calling. Reasoning is on by default; NVIDIA recommends temperature 0.6 with top-p 0.95 for reasoning on, and greedy decoding for reasoning off.
Llama-3.3-Nemotron-Super-49B-v1.5 benchmarks
NVIDIA published its own evaluation numbers in reasoning-on mode, run with NeMo-Skills at a 64k sequence length and averaged over up to 16 runs:
| Benchmark | Llama-3.3-Nemotron-Super-49B-v1.5 |
|---|---|
MATH500 Math problems | 97.4 |
AIME 2025 Competition math | 82.71 |
GPQA Expert science | 71.97 |
LiveCodeBench Competitive coding | 73.58 |
BFCL v3 Function calling | 71.75 |
IFEval Instruction following | 88.61 |
MMLU Pro Academic knowledge | 79.53 |
These are single-model scores with no competitor columns, so read them as a profile rather than a ranking. Math is the strong suit: 97.4 on MATH500 and 87.5 on AIME 2024. Science, coding and function calling land in the low 70s, and the text-only subset of Humanity's Last Exam sits at 7.64, so frontier trivia is not what this model is built for.
Llama-3.3-Nemotron-Super-49B-v1.5 hardware requirements
The system requirement to check is memory. NVIDIA ships the weights as BF16 safetensors; the community GGUF builds below, with their real file sizes, come from unsloth/Llama-3_3-Nemotron-Super-49B-v1_5-GGUF.
| Memory | Build to pick | File size |
|---|---|---|
| 16 GB | UD-IQ2_XXS | 13.95 GB |
| 24 GB | UD-IQ3_XXS | 19.69 GB |
| 32 GB | UD-Q3_K_XL | 24.81 GB |
| 48 GB | UD-Q4_K_XL | 30.36 GB |
| 64 GB and up | UD-Q6_K_XL | 43.42 GB |
Neighbouring builds differ by a few gigabytes, so when two builds both fit, take the larger one. If the format is new to you, start with what GGUF is.
How to run Llama-3.3-Nemotron-Super-49B-v1.5 in Atomic Chat
Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.
- Download Atomic Chat for your platform and open it.
- Search for Llama-3.3-Nemotron-Super-49B-v1.5 in the model browser and open Download Options.
- Pick the build that fits the memory you have, then start a chat.
For the rest of NVIDIA's local lineup, see every Nemotron model you can run locally, or the much smaller Nemotron Nano 9B v2.
Llama-3.3-Nemotron-Super-49B-v1.5 license
Use is governed by the NVIDIA Open Model License, with the Llama 3.3 Community License Agreement on top because the model is built with Llama. NVIDIA states the model is ready for commercial use, so you can download the weights, run them on your own hardware and ship products on them, subject to both sets of terms.
