Phi-3.5-mini-instruct

Updated
05.10.2026
Reasoning
Code
Multilingual

A 3.8B dense instruction-tuned LLM from Microsoft’s Phi-3.5 family with a 128K context window and multilingual support.

At a glance

  • License: MIT
  • Parameters: 3.8B dense
  • Context length: 128K tokens
  • Modalities: Text in, text out
  • Minimum hardware: 2 GB memory, smallest GGUF build is 1.32 GB

What is Phi-3.5-mini-instruct?

Phi-3.5-mini-instruct is a 3.8B dense language model from Microsoft, released in August 2024 as an update over the June 2024 Phi-3 Mini. It was trained on 3.4 trillion tokens of rigorously filtered public documents plus synthetic, textbook-like data, with the mix deliberately skewed toward reasoning: code, math and logic. Microsoft credits its extra post-training data with substantial gains on multilingual work, multi-turn conversation quality and reasoning. The local pitch is simple: it carries a 128K token context window, and the common quantized builds are under 3 GB, so it fits in memory that almost any machine has.

SpecificationPhi-3.5-mini-instruct
Parameters3.8B (3,821,079,552)
ArchitectureDense decoder-only Transformer
Context window128K tokens
AttentionFlash attention by default, eager path for V100 and older GPUs
ModalitiesText in, text out
TokenizerSame as Phi-3 Mini, 32,064 vocabulary with placeholder tokens
Training data3.4T tokens, cutoff October 2023
Training run512 H100-80G GPUs, 10 days, June to August 2024
Post-trainingSupervised fine-tuning, then PPO and DPO
Prompt formatChat template with system, user and assistant turns
Supported languages23, including Arabic, Chinese, Japanese and Russian
Release dateAugust 2024
LicenseMIT

Microsoft is explicit about the trade behind that size. Public documents were filtered to hold what the card calls the correct level of knowledge: a Premier League result is good data for a frontier model, cut here to leave more capacity for reasoning. The consequence is stated just as plainly: at 3.8B the model cannot store much factual knowledge, users may hit factual incorrectness, and the suggested fix is augmenting it with a search engine in RAG settings. Microsoft positions the model for memory and compute constrained environments, latency bound scenarios, and strong reasoning. The card adds two limits: code training data is mostly Python, and very long sessions can turn repetitive.

Phi-3.5-mini-instruct benchmarks

The numbers below are Microsoft's own, from the model card, produced with a single internal evaluation pipeline across all models at temperature 0. The open-weight competitors are Mistral-7B-Instruct-v0.3, Mistral-Nemo-12B, Llama-3.1-8B and Gemma-2-9B:

BenchmarkPhi-3.5-miniMistral-7BMistral-Nemo-12BLlama-3.1-8BGemma-2-9B
MMLU
General knowledge
6960.367.268.171.3
BigBench Hard CoT
Hard reasoning
6933.460.263.463.5
GPQA
Expert science
30.415.628.626.329.2
GSM8K
School math
86.254.484.282.484.9
MATH
Competition math
48.51931.247.650.9
HumanEval
Python coding
62.835.463.466.561
MBPP
Basic Python
69.650.468.169.469.3

A 3.8B model takes four of the seven rows against models two to three times its size, with the clearest leads on reasoning and math. Gemma-2-9B keeps the knowledge-heavy rows and Llama-3.1-8B leads HumanEval; Microsoft's full table also includes Gemini 1.5 Flash and GPT-4o-mini, which sit ahead on most benchmarks.

Two more vendor tables matter locally. On RULER, a retrieval benchmark for long context, it averages 84.1 from 4K to 128K and holds 63.6 at the full 128K: behind Llama-3.1-8B-Instruct at 88.3, far ahead of Mistral-Nemo-12B, which averages 66.2 and falls to 19.0. On RepoQA, long-context code understanding across five languages, it averages 77 against 71 for Llama-3.1-8B and 62 for Mistral-7B. Multilingual is the weaker column: a 55.2 average against 47.9 for Mistral-7B but 59.6 for Gemma-2-9B, though on the Korean set in Appendix B it averages 35.62 against 29.29 for Llama-3.1-8B-Instruct.

Phi-3.5-mini-instruct hardware requirements

The system requirement to check is memory, and here almost anything qualifies. The builds below come from the community repo bartowski/Phi-3.5-mini-instruct-GGUF.

MemoryBuild to pickFile size
2 GBIQ2_M1.32 GB
3 GBQ3_K_M1.96 GB
4 GBQ4_K_M2.39 GB
5 GBQ5_K_M2.82 GB
6 GBQ6_K3.14 GB
8 GB and upQ8_04.06 GB

Neighbouring files differ by a few hundred megabytes, so when two builds both fit, take the larger one. That matters most at the IQ2 end of the ladder, where quality falls fastest. Leave headroom beyond the file size if you plan to push the 128K context. If the format is new to you, start with what GGUF is, and see the best local LLMs for a 16 GB Mac for what else fits alongside it.

If you run the original weights instead, Microsoft's reference stack is transformers 4.43.0 with torch 2.3.1, accelerate 0.31.0 and flash_attn 2.5.8, and that flash attention path was tested on A100, A6000 and H100. The vendor's own sample generates greedily, temperature 0 with sampling off, the same setting behind the numbers above.

How to run Phi-3.5-mini-instruct in Atomic Chat

Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.

  1. Download Atomic Chat for your platform and open it.
  2. Search for Phi-3.5-mini-instruct in the model browser and open Download Options.
  3. Pick the build that fits the memory you have, then start a chat.

The rest of the family lives on our Phi models page, and the successor in the same size class is Phi-4-mini-instruct.

Phi-3.5-mini-instruct license

Phi-3.5-mini-instruct is released under the MIT license, the most permissive of the common open licenses. It allows commercial use, modification and redistribution with no royalties; the only separate ground rule in the repo is Microsoft's trademark and brand guidelines.

Get the weights from Hugging Face

pip install -U transformers accelerate
huggingface-cli download microsoft/Phi-3.5-mini-instruct
# or via Ollama:
ollama run phi3.5
curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "microsoft/Phi-3.5-mini-instruct",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("microsoft/Phi-3.5-mini-instruct", trust_remote_code=True)
tokenizer = AutoTokenizer.from_pretrained("microsoft/Phi-3.5-mini-instruct")
msgs = [{"role": "user", "content": "Explain quantum computing simply."}]
ids = tokenizer.apply_chat_template(msgs, return_tensors="pt")
out = model.generate(ids, max_new_tokens=256)
print(tokenizer.decode(out[0]))
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "http://localhost:8000/v1", apiKey: "local" });
const res = await client.chat.completions.create({
  model: "microsoft/Phi-3.5-mini-instruct",
  messages: [{ role: "user", content: "Hello!" }],
});
console.log(res.choices[0].message.content);
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

Phi-3.5-mini-instruct has 3.8 billion parameters. It is a dense decoder-only Transformer, so all 3.8B parameters are active during inference. It was trained on 3.4 trillion tokens.

Phi-3.5-mini-instruct supports a 128K token context length. This makes it suitable for long-context tasks such as long document and meeting summarization, long document QA, and information retrieval over large inputs.

Yes. Microsoft released Phi-3.5-mini-instruct under the MIT license, which permits free commercial and research use, modification, and redistribution. The weights are available on Hugging Face and the model is intended for commercial and research use across multiple languages.

Phi-3.5-mini-instruct supports 23 languages, including Arabic, Chinese, Czech, Danish, Dutch, English, Finnish, French, German, Hebrew, Hungarian, Italian, Japanese, Korean, Norwegian, Polish, Portuguese, Russian, Spanish, Swedish, Thai, Turkish, and Ukrainian. It was trained primarily on English, so non-English performance can be lower.

Because it has only 3.8B parameters, Phi-3.5-mini-instruct runs on modest hardware. A 4-bit quantized build fits in roughly 3-4 GB of VRAM and runs on most consumer GPUs, and it can also run on CPU through llama.cpp or Ollama. Full FP16 precision needs about 8 GB of VRAM.