What is SmolLM2-135M-Instruct?
SmolLM2-135M-Instruct is the smallest model in Hugging Face's SmolLM2 family, which spans 135M, 360M and 1.7B parameters. It is a compact transformer decoder that Hugging Face calls lightweight enough to run on-device, and the card tags the release as safetensors, onnx and transformers.js. The quickstart shows three ways in: Python transformers, the TRL chat CLI run with the cpu device flag, and a Transformers.js pipeline in JavaScript that pulls the same checkpoint from the Hub. GGUF is a separate track: the vendor README does not mention it, and the quantized files quoted further down come from the third-party repo unsloth/SmolLM2-135M-Instruct-GGUF, where the F16 build is 0.27 GB. Hugging Face published the weights on October 31, 2024 under Apache 2.0, and the data-centric recipe behind them is written up in the SmolLM2 paper, arXiv 2502.02737.
| Specification | SmolLM2-135M-Instruct |
|---|---|
| Total parameters | 135M (134,515,008) |
| Architecture | Transformer decoder |
| Base model | SmolLM2-135M |
| Family sizes | 135M, 360M, 1.7B |
| Modalities | Text only |
| Languages | English |
| Pretraining tokens | 2T |
| Pretraining data | FineWeb-Edu, DCLM, The Stack, plus filtered sets curated by the team |
| Training precision | bfloat16 |
| Training hardware | 64 H100 GPUs |
| Training framework | nanotron |
| Post-training | SFT on smol-smoltalk, then DPO on UltraFeedback |
| Published formats | safetensors, ONNX |
| Paper | arXiv 2502.02737 |
| Release date | October 31, 2024 |
| License | Apache 2.0 |
The instruct model starts from the SmolLM2-135M base checkpoint, adds supervised fine-tuning on a mix of public and curated datasets, then Direct Preference Optimization on UltraFeedback. Hugging Face says it additionally supports text rewriting, summarization and function calling, the last of the three only for the 1.7B, thanks to datasets developed by Argilla such as Synth-APIGen-v0.1. Hugging Face published the SFT dataset as HuggingFaceTB/smol-smoltalk and the finetuning code as an alignment-handbook recipe. The preference stage uses UltraFeedback in its binarized form, released as HuggingFaceH4/ultrafeedback_binarized. The model primarily understands and generates English, and Hugging Face is direct about the limits: the output may not always be factually accurate, logically consistent or free from bias present in the training data, so treat it as an assistive tool and verify anything that matters.
SmolLM2-135M-Instruct benchmarks
Hugging Face's own numbers, run with lighteval and published on the model card, compare the instruct model with its predecessor SmolLM-135M-Instruct. Every evaluation is zero-shot unless the row says otherwise:
| Benchmark | SmolLM2-135M-Instruct | SmolLM-135M-Instruct |
|---|---|---|
IFEval Instruction following | 29.9 | 17.2 |
MT-Bench Conversation quality | 19.8 | 16.8 |
HellaSwag Commonsense reasoning | 40.9 | 38.9 |
ARC (Average) Science questions | 37.3 | 33.9 |
PIQA Physical reasoning | 66.3 | 64.0 |
MMLU (cloze) Academic knowledge | 29.3 | 28.3 |
BBH (3-shot) Hard reasoning | 28.2 | 25.2 |
GSM8K (5-shot) School math | 1.4 | 1.4 |
SmolLM2 leads every row except GSM8K, where both models score 1.4, and the widest gap is IFEval at 29.9 against 17.2. The honest read is that this model follows short, well-scoped instructions far better than its predecessor, and that math at 135M parameters is effectively absent.
The same card reports the base pre-trained checkpoint, listed there as SmolLM2-135M-8k, against SmolLM-135M: HellaSwag 42.1 against 41.2, ARC average 43.9 against 42.4, MMLU cloze 31.5 against 30.2, CommonsenseQA 33.9 against 32.7, OpenBookQA 34.6 against 34.0, GSM8K 1.4 against 1.0, PIQA tied at 68.4, Winogrande tied at 51.3, and TriviaQA 4.1 against 4.3, the single row the older model keeps. Hugging Face frames the family as models that solve a wide range of tasks while staying lightweight enough to run on-device, and puts the advance over SmolLM1 in instruction following, knowledge and reasoning. The two tables back that up: the clearest gains are on IFEval and ARC, not on math.
SmolLM2-135M-Instruct hardware requirements
The system requirement to check is memory, and here it barely registers, because every build is under 0.3 GB and the GGUF files below come from the third-party repo unsloth/SmolLM2-135M-Instruct-GGUF.
| Memory | Build to pick | File size |
|---|---|---|
| 512 MB | SmolLM2-135M-Instruct-Q2_K | 0.09 GB |
| 768 MB | SmolLM2-135M-Instruct-Q4_K_M | 0.11 GB |
| 1 GB | SmolLM2-135M-Instruct-Q5_K_M | 0.11 GB |
| 2 GB | SmolLM2-135M-Instruct-Q8_0 | 0.14 GB |
| 4 GB and up | SmolLM2-135M-Instruct-F16 | 0.27 GB |
When two builds both fit, take the larger one, and note that the whole ladder from Q2_K to F16 spans 0.18 GB, with Q3_K_M matching Q2_K at 0.09 GB and Q6_K matching Q8_0 at 0.14 GB.
How to run SmolLM2-135M-Instruct in Atomic Chat
Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.
- Download Atomic Chat for your platform and open it.
- Search for SmolLM2-135M-Instruct in the model browser and open Download Options.
- Pick the build that fits the memory you have, then start a chat.
For the rest of the line, see every SmolLM model you can run locally, the larger SmolLM3-3B, or the primer on what GGUF is if the file format is new to you.
SmolLM2-135M-Instruct license
SmolLM2-135M-Instruct is released under Apache 2.0, which permits commercial use, modification and redistribution with no royalties. Hugging Face also published the SFT dataset and the fine-tuning code, so you can retrain or adapt the model on your own data and ship the result. The SmolLM2 paper, arXiv 2502.02737, documents the data-centric training recipe behind it.
