What is Phi-3.5-mini-instruct?
Phi-3.5-mini-instruct is a 3.8B dense language model from Microsoft, released in August 2024 as an update over the June 2024 Phi-3 Mini. It was trained on 3.4 trillion tokens of rigorously filtered public documents plus synthetic, textbook-like data, with the mix deliberately skewed toward reasoning: code, math and logic. Microsoft credits its extra post-training data with substantial gains on multilingual work, multi-turn conversation quality and reasoning. The local pitch is simple: it carries a 128K token context window, and the common quantized builds are under 3 GB, so it fits in memory that almost any machine has.
| Specification | Phi-3.5-mini-instruct |
|---|---|
| Parameters | 3.8B (3,821,079,552) |
| Architecture | Dense decoder-only Transformer |
| Context window | 128K tokens |
| Attention | Flash attention by default, eager path for V100 and older GPUs |
| Modalities | Text in, text out |
| Tokenizer | Same as Phi-3 Mini, 32,064 vocabulary with placeholder tokens |
| Training data | 3.4T tokens, cutoff October 2023 |
| Training run | 512 H100-80G GPUs, 10 days, June to August 2024 |
| Post-training | Supervised fine-tuning, then PPO and DPO |
| Prompt format | Chat template with system, user and assistant turns |
| Supported languages | 23, including Arabic, Chinese, Japanese and Russian |
| Release date | August 2024 |
| License | MIT |
Microsoft is explicit about the trade behind that size. Public documents were filtered to hold what the card calls the correct level of knowledge: a Premier League result is good data for a frontier model, cut here to leave more capacity for reasoning. The consequence is stated just as plainly: at 3.8B the model cannot store much factual knowledge, users may hit factual incorrectness, and the suggested fix is augmenting it with a search engine in RAG settings. Microsoft positions the model for memory and compute constrained environments, latency bound scenarios, and strong reasoning. The card adds two limits: code training data is mostly Python, and very long sessions can turn repetitive.
Phi-3.5-mini-instruct benchmarks
The numbers below are Microsoft's own, from the model card, produced with a single internal evaluation pipeline across all models at temperature 0. The open-weight competitors are Mistral-7B-Instruct-v0.3, Mistral-Nemo-12B, Llama-3.1-8B and Gemma-2-9B:
| Benchmark | Phi-3.5-mini | Mistral-7B | Mistral-Nemo-12B | Llama-3.1-8B | Gemma-2-9B |
|---|---|---|---|---|---|
MMLU General knowledge | 69 | 60.3 | 67.2 | 68.1 | 71.3 |
BigBench Hard CoT Hard reasoning | 69 | 33.4 | 60.2 | 63.4 | 63.5 |
GPQA Expert science | 30.4 | 15.6 | 28.6 | 26.3 | 29.2 |
GSM8K School math | 86.2 | 54.4 | 84.2 | 82.4 | 84.9 |
MATH Competition math | 48.5 | 19 | 31.2 | 47.6 | 50.9 |
HumanEval Python coding | 62.8 | 35.4 | 63.4 | 66.5 | 61 |
MBPP Basic Python | 69.6 | 50.4 | 68.1 | 69.4 | 69.3 |
A 3.8B model takes four of the seven rows against models two to three times its size, with the clearest leads on reasoning and math. Gemma-2-9B keeps the knowledge-heavy rows and Llama-3.1-8B leads HumanEval; Microsoft's full table also includes Gemini 1.5 Flash and GPT-4o-mini, which sit ahead on most benchmarks.
Two more vendor tables matter locally. On RULER, a retrieval benchmark for long context, it averages 84.1 from 4K to 128K and holds 63.6 at the full 128K: behind Llama-3.1-8B-Instruct at 88.3, far ahead of Mistral-Nemo-12B, which averages 66.2 and falls to 19.0. On RepoQA, long-context code understanding across five languages, it averages 77 against 71 for Llama-3.1-8B and 62 for Mistral-7B. Multilingual is the weaker column: a 55.2 average against 47.9 for Mistral-7B but 59.6 for Gemma-2-9B, though on the Korean set in Appendix B it averages 35.62 against 29.29 for Llama-3.1-8B-Instruct.
Phi-3.5-mini-instruct hardware requirements
The system requirement to check is memory, and here almost anything qualifies. The builds below come from the community repo bartowski/Phi-3.5-mini-instruct-GGUF.
| Memory | Build to pick | File size |
|---|---|---|
| 2 GB | IQ2_M | 1.32 GB |
| 3 GB | Q3_K_M | 1.96 GB |
| 4 GB | Q4_K_M | 2.39 GB |
| 5 GB | Q5_K_M | 2.82 GB |
| 6 GB | Q6_K | 3.14 GB |
| 8 GB and up | Q8_0 | 4.06 GB |
Neighbouring files differ by a few hundred megabytes, so when two builds both fit, take the larger one. That matters most at the IQ2 end of the ladder, where quality falls fastest. Leave headroom beyond the file size if you plan to push the 128K context. If the format is new to you, start with what GGUF is, and see the best local LLMs for a 16 GB Mac for what else fits alongside it.
If you run the original weights instead, Microsoft's reference stack is transformers 4.43.0 with torch 2.3.1, accelerate 0.31.0 and flash_attn 2.5.8, and that flash attention path was tested on A100, A6000 and H100. The vendor's own sample generates greedily, temperature 0 with sampling off, the same setting behind the numbers above.
How to run Phi-3.5-mini-instruct in Atomic Chat
Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.
- Download Atomic Chat for your platform and open it.
- Search for Phi-3.5-mini-instruct in the model browser and open Download Options.
- Pick the build that fits the memory you have, then start a chat.
The rest of the family lives on our Phi models page, and the successor in the same size class is Phi-4-mini-instruct.
Phi-3.5-mini-instruct license
Phi-3.5-mini-instruct is released under the MIT license, the most permissive of the common open licenses. It allows commercial use, modification and redistribution with no royalties; the only separate ground rule in the repo is Microsoft's trademark and brand guidelines.
