Llama-3.3-Nemotron-Super-49B-v1.5

Updated
24.08.2026
Thinking
Reasoning
Code
Tools
Multilingual

A 49B reasoning and chat LLM from NVIDIA, distilled from Llama-3.3-70B via Neural Architecture Search with a 128K context.

At a glance

  • License: NVIDIA Open Model License + Llama 3.3 Community License
  • Parameters: 49B dense, distilled from Llama-3.3-70B-Instruct
  • Context length: 131,072 tokens
  • Modalities: Text input and output
  • Minimum hardware: 16 GB of memory with the 13.95 GB UD-IQ2_XXS GGUF

What is Llama-3.3-Nemotron-Super-49B-v1.5?

Llama-3.3-Nemotron-Super-49B-v1.5 is NVIDIA's 49B reasoning and chat model, a derivative of Meta's Llama-3.3-70B-Instruct released on Hugging Face on July 25, 2025. NVIDIA used Neural Architecture Search (NAS) to compress the 70B reference model into 49B, cutting the memory footprint so the model fits a single GPU at high workloads. It is post-trained for reasoning, human chat preferences and agentic tasks such as RAG and tool calling, and it supports a 128K context.

SpecificationLlama-3.3-Nemotron-Super-49B-v1.5
Total parameters49.9B (BF16)
ArchitectureDense decoder-only Transformer, NAS variant of Llama 3.3 70B
Base modelMeta Llama-3.3-70B-Instruct
Context window131,072 tokens
ModalitiesText input, text output
ReasoningOn by default, /no_think in the system prompt switches it off
LanguagesEnglish and code, plus German, French, Italian, Portuguese, Hindi, Spanish and Thai
Release dateJuly 25, 2025
LicenseNVIDIA Open Model License + Llama 3.3 Community License

The NAS pass produces non-standard, non-repetitive blocks: in some, attention is skipped entirely or replaced with a single linear layer, and the FFN expansion ratio changes from block to block. NVIDIA then ran block-wise distillation against the reference model on 40 billion tokens of FineWeb, Buzz-V1.2 and Dolma, followed by a multi-stage post-training pipeline: supervised fine-tuning for math, code, science and tool calling, then RPO for chat, RLVR for reasoning and iterative DPO for tool calling. Reasoning is on by default; NVIDIA recommends temperature 0.6 with top-p 0.95 for reasoning on, and greedy decoding for reasoning off.

Llama-3.3-Nemotron-Super-49B-v1.5 benchmarks

NVIDIA published its own evaluation numbers in reasoning-on mode, run with NeMo-Skills at a 64k sequence length and averaged over up to 16 runs:

BenchmarkLlama-3.3-Nemotron-Super-49B-v1.5
MATH500
Math problems
97.4
AIME 2025
Competition math
82.71
GPQA
Expert science
71.97
LiveCodeBench
Competitive coding
73.58
BFCL v3
Function calling
71.75
IFEval
Instruction following
88.61
MMLU Pro
Academic knowledge
79.53

These are single-model scores with no competitor columns, so read them as a profile rather than a ranking. Math is the strong suit: 97.4 on MATH500 and 87.5 on AIME 2024. Science, coding and function calling land in the low 70s, and the text-only subset of Humanity's Last Exam sits at 7.64, so frontier trivia is not what this model is built for.

Llama-3.3-Nemotron-Super-49B-v1.5 hardware requirements

The system requirement to check is memory. NVIDIA ships the weights as BF16 safetensors; the community GGUF builds below, with their real file sizes, come from unsloth/Llama-3_3-Nemotron-Super-49B-v1_5-GGUF.

MemoryBuild to pickFile size
16 GBUD-IQ2_XXS13.95 GB
24 GBUD-IQ3_XXS19.69 GB
32 GBUD-Q3_K_XL24.81 GB
48 GBUD-Q4_K_XL30.36 GB
64 GB and upUD-Q6_K_XL43.42 GB

Neighbouring builds differ by a few gigabytes, so when two builds both fit, take the larger one. If the format is new to you, start with what GGUF is.

How to run Llama-3.3-Nemotron-Super-49B-v1.5 in Atomic Chat

Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.

  1. Download Atomic Chat for your platform and open it.
  2. Search for Llama-3.3-Nemotron-Super-49B-v1.5 in the model browser and open Download Options.
  3. Pick the build that fits the memory you have, then start a chat.

For the rest of NVIDIA's local lineup, see every Nemotron model you can run locally, or the much smaller Nemotron Nano 9B v2.

Llama-3.3-Nemotron-Super-49B-v1.5 license

Use is governed by the NVIDIA Open Model License, with the Llama 3.3 Community License Agreement on top because the model is built with Llama. NVIDIA states the model is ready for commercial use, so you can download the weights, run them on your own hardware and ship products on them, subject to both sets of terms.

Get the weights from Hugging Face

pip install -U "transformers" "vllm==0.9.2"
huggingface-cli download nvidia/Llama-3_3-Nemotron-Super-49B-v1_5
python3 -m vllm.entrypoints.openai.api_server \
  --model nvidia/Llama-3_3-Nemotron-Super-49B-v1_5 \
  --trust-remote-code --max-model-len 65536 --tensor-parallel-size 2
curl http://localhost:5000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Llama-3_3-Nemotron-Super-49B-v1_5",
    "messages": [
      {"role": "system", "content": ""},
      {"role": "user", "content": "What is 18% of 100?"}
    ],
    "temperature": 0.6,
    "top_p": 0.95
  }'
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model = AutoModelForCausalLM.from_pretrained("nvidia/Llama-3_3-Nemotron-Super-49B-v1_5", trust_remote_code=True, torch_dtype=torch.bfloat16, device_map="auto")
tokenizer = AutoTokenizer.from_pretrained("nvidia/Llama-3_3-Nemotron-Super-49B-v1_5")
messages = [{"role": "system", "content": ""}, {"role": "user", "content": "Explain the NAS approach in one paragraph."}]
inputs = tokenizer.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
out = model.generate(inputs, max_new_tokens=512, temperature=0.6, top_p=0.95)
print(tokenizer.decode(out[0], skip_special_tokens=True))
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "http://localhost:5000/v1", apiKey: "dummy" });
const res = await client.chat.completions.create({
  model: "Llama-3_3-Nemotron-Super-49B-v1_5",
  messages: [
    { role: "system", content: "" },
    { role: "user", content: "Write a function to reverse a linked list." }
  ],
  temperature: 0.6,
  top_p: 0.95
});
console.log(res.choices[0].message.content);
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

Llama-3.3-Nemotron-Super-49B-v1.5 is a 49B-parameter reasoning and chat model from NVIDIA, derived from Meta's Llama-3.3-70B-Instruct. NVIDIA used Neural Architecture Search to shrink the 70B reference model to 49B while keeping accuracy high, so it fits on a single H100 or H200 GPU. It supports a 128K-token context and is post-trained for math, code, reasoning, and tool calling.

The model was optimized to fit on a single NVIDIA H100-80GB or H200 GPU at high workloads, and NVIDIA's test hardware was 2x H100-80GB or 2x A100-80GB. Running it in full precision needs roughly 100 GB of memory; with 4-bit quantization it can run on around 48 GB of VRAM. It serves well with vLLM and Transformers on Ampere or Hopper GPUs.

The weights are released openly on Hugging Face and the model is ready for commercial use. Use is governed by the NVIDIA Open Model License, with the additional Llama 3.3 Community License Agreement since the model is built on Llama. You can download and run it yourself at no cost, subject to the terms of those licenses.

By default, with an empty system prompt, the model responds in reasoning ON mode and emits a thinking trace. Adding /no_think to the system prompt switches it to reasoning OFF mode. NVIDIA recommends temperature 0.6 and top-p 0.95 for reasoning ON, and greedy decoding for reasoning OFF.

The model is primarily intended for English and coding languages. NVIDIA also lists support for German, French, Italian, Portuguese, Hindi, Spanish, and Thai. Its post-training focused heavily on English single- and multi-turn chat, so English performance is strongest.