SmolLM2-135M-Instruct

Updated
05.10.2026
Reasoning

A 135M-parameter instruction-tuned LLM from Hugging Face’s SmolLM2 family, small enough to run on CPU and on-device.

At a glance

  • License: Apache 2.0
  • Parameters: 135M
  • Context length: Not stated on the model card
  • Modalities: Text only, English
  • Minimum hardware: Runs on CPU; the F16 GGUF is 0.27 GB

What is SmolLM2-135M-Instruct?

SmolLM2-135M-Instruct is the smallest model in Hugging Face's SmolLM2 family, which spans 135M, 360M and 1.7B parameters. It is a compact transformer decoder that Hugging Face calls lightweight enough to run on-device, and the card tags the release as safetensors, onnx and transformers.js. The quickstart shows three ways in: Python transformers, the TRL chat CLI run with the cpu device flag, and a Transformers.js pipeline in JavaScript that pulls the same checkpoint from the Hub. GGUF is a separate track: the vendor README does not mention it, and the quantized files quoted further down come from the third-party repo unsloth/SmolLM2-135M-Instruct-GGUF, where the F16 build is 0.27 GB. Hugging Face published the weights on October 31, 2024 under Apache 2.0, and the data-centric recipe behind them is written up in the SmolLM2 paper, arXiv 2502.02737.

SpecificationSmolLM2-135M-Instruct
Total parameters135M (134,515,008)
ArchitectureTransformer decoder
Base modelSmolLM2-135M
Family sizes135M, 360M, 1.7B
ModalitiesText only
LanguagesEnglish
Pretraining tokens2T
Pretraining dataFineWeb-Edu, DCLM, The Stack, plus filtered sets curated by the team
Training precisionbfloat16
Training hardware64 H100 GPUs
Training frameworknanotron
Post-trainingSFT on smol-smoltalk, then DPO on UltraFeedback
Published formatssafetensors, ONNX
PaperarXiv 2502.02737
Release dateOctober 31, 2024
LicenseApache 2.0

The instruct model starts from the SmolLM2-135M base checkpoint, adds supervised fine-tuning on a mix of public and curated datasets, then Direct Preference Optimization on UltraFeedback. Hugging Face says it additionally supports text rewriting, summarization and function calling, the last of the three only for the 1.7B, thanks to datasets developed by Argilla such as Synth-APIGen-v0.1. Hugging Face published the SFT dataset as HuggingFaceTB/smol-smoltalk and the finetuning code as an alignment-handbook recipe. The preference stage uses UltraFeedback in its binarized form, released as HuggingFaceH4/ultrafeedback_binarized. The model primarily understands and generates English, and Hugging Face is direct about the limits: the output may not always be factually accurate, logically consistent or free from bias present in the training data, so treat it as an assistive tool and verify anything that matters.

SmolLM2-135M-Instruct benchmarks

Hugging Face's own numbers, run with lighteval and published on the model card, compare the instruct model with its predecessor SmolLM-135M-Instruct. Every evaluation is zero-shot unless the row says otherwise:

BenchmarkSmolLM2-135M-InstructSmolLM-135M-Instruct
IFEval
Instruction following
29.917.2
MT-Bench
Conversation quality
19.816.8
HellaSwag
Commonsense reasoning
40.938.9
ARC (Average)
Science questions
37.333.9
PIQA
Physical reasoning
66.364.0
MMLU (cloze)
Academic knowledge
29.328.3
BBH (3-shot)
Hard reasoning
28.225.2
GSM8K (5-shot)
School math
1.41.4

SmolLM2 leads every row except GSM8K, where both models score 1.4, and the widest gap is IFEval at 29.9 against 17.2. The honest read is that this model follows short, well-scoped instructions far better than its predecessor, and that math at 135M parameters is effectively absent.

The same card reports the base pre-trained checkpoint, listed there as SmolLM2-135M-8k, against SmolLM-135M: HellaSwag 42.1 against 41.2, ARC average 43.9 against 42.4, MMLU cloze 31.5 against 30.2, CommonsenseQA 33.9 against 32.7, OpenBookQA 34.6 against 34.0, GSM8K 1.4 against 1.0, PIQA tied at 68.4, Winogrande tied at 51.3, and TriviaQA 4.1 against 4.3, the single row the older model keeps. Hugging Face frames the family as models that solve a wide range of tasks while staying lightweight enough to run on-device, and puts the advance over SmolLM1 in instruction following, knowledge and reasoning. The two tables back that up: the clearest gains are on IFEval and ARC, not on math.

SmolLM2-135M-Instruct hardware requirements

The system requirement to check is memory, and here it barely registers, because every build is under 0.3 GB and the GGUF files below come from the third-party repo unsloth/SmolLM2-135M-Instruct-GGUF.

MemoryBuild to pickFile size
512 MBSmolLM2-135M-Instruct-Q2_K0.09 GB
768 MBSmolLM2-135M-Instruct-Q4_K_M0.11 GB
1 GBSmolLM2-135M-Instruct-Q5_K_M0.11 GB
2 GBSmolLM2-135M-Instruct-Q8_00.14 GB
4 GB and upSmolLM2-135M-Instruct-F160.27 GB

When two builds both fit, take the larger one, and note that the whole ladder from Q2_K to F16 spans 0.18 GB, with Q3_K_M matching Q2_K at 0.09 GB and Q6_K matching Q8_0 at 0.14 GB.

How to run SmolLM2-135M-Instruct in Atomic Chat

Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.

  1. Download Atomic Chat for your platform and open it.
  2. Search for SmolLM2-135M-Instruct in the model browser and open Download Options.
  3. Pick the build that fits the memory you have, then start a chat.

For the rest of the line, see every SmolLM model you can run locally, the larger SmolLM3-3B, or the primer on what GGUF is if the file format is new to you.

SmolLM2-135M-Instruct license

SmolLM2-135M-Instruct is released under Apache 2.0, which permits commercial use, modification and redistribution with no royalties. Hugging Face also published the SFT dataset and the fine-tuning code, so you can retrain or adapt the model on your own data and ship the result. The SmolLM2 paper, arXiv 2502.02737, documents the data-centric training recipe behind it.

Get the weights from Hugging Face

pip install -U transformers
huggingface-cli download HuggingFaceTB/SmolLM2-135M-Instruct
curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "HuggingFaceTB/SmolLM2-135M-Instruct",
    "messages": [{"role": "user", "content": "Summarize this paragraph."}]
  }'
from transformers import AutoModelForCausalLM, AutoTokenizer
checkpoint = "HuggingFaceTB/SmolLM2-135M-Instruct"
tokenizer = AutoTokenizer.from_pretrained(checkpoint)
model = AutoModelForCausalLM.from_pretrained(checkpoint)
messages = [{"role": "user", "content": "What is gravity?"}]
inputs = tokenizer.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt")
out = model.generate(inputs, max_new_tokens=50)
print(tokenizer.decode(out[0]))
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "http://localhost:8000/v1", apiKey: "local" });
const res = await client.chat.completions.create({
  model: "HuggingFaceTB/SmolLM2-135M-Instruct",
  messages: [{ role: "user", content: "Rewrite this sentence more formally." }]
});
console.log(res.choices[0].message.content);
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

SmolLM2-135M-Instruct is a 135-million-parameter instruction-tuned language model from Hugging Face's SmolLM2 family. It uses a Llama-style transformer decoder architecture and was fine-tuned with supervised fine-tuning and Direct Preference Optimization on UltraFeedback. It is the smallest of the three SmolLM2 sizes (135M, 360M, 1.7B) and targets on-device chat and text tasks.

Because it has only 135M parameters and a memory footprint of roughly 720 MB in its default precision, SmolLM2-135M-Instruct runs comfortably on CPU and on resource-constrained devices without a dedicated GPU. CPU inference works but generates tokens more slowly than a GPU; a few GB of RAM is enough. Quantized GGUF builds shrink it further for laptops, phones, and edge hardware.

Yes. SmolLM2-135M-Instruct is released under the Apache 2.0 license, which permits free commercial and private use, modification, and redistribution as long as the license and notices are preserved. The weights are openly downloadable from Hugging Face, and Hugging Face also published the training recipe and SFT dataset.

SmolLM2-135M-Instruct supports a context window of 8K tokens, extended from the original 2K during pretraining by adjusting the data mix and the RoPE base. The model primarily understands and generates English; it was not trained for broad multilingual use, so quality drops sharply in other languages.

It suits lightweight on-device tasks such as basic chat, instruction following, text rewriting, and summarization where privacy or offline operation matters. Its IFEval and MT-Bench scores improved markedly over SmolLM-135M-Instruct, but at 135M parameters it is weak on factual knowledge and math (GSM8K around 1.4), so verify its output and reserve it for simple, constrained tasks rather than open-ended reasoning. Function calling is only available on the larger 1.7B variant.