Phi-4

Updated
24.08.2026
Reasoning
Code

A 14B open model from Microsoft Research tuned for math, reasoning, and code, competitive with much larger LLMs.

At a glance

  • License: MIT
  • Parameters: 14B dense decoder-only
  • Context length: 16K tokens
  • Modalities: Text input and output, English first
  • Minimum hardware: 8 GB memory with the IQ3_M GGUF (6.91 GB)

What is Phi-4?

Phi-4 is a 14B dense decoder-only Transformer from Microsoft Research, released on December 12, 2024 under the MIT license. Microsoft trained it on 9.8T tokens, including synthetic textbook-style data written to teach math, coding, common sense reasoning and world knowledge, plus rigorously filtered public documents, acquired academic books and Q&A datasets. The stated primary use cases are memory and compute constrained environments, latency bound scenarios, and reasoning and logic: in other words, the profile of a model built to run on your own machine.

SpecificationPhi-4
DeveloperMicrosoft Research
Parameters14B dense (14.66B exact)
ArchitectureDense decoder-only Transformer
Context window16K tokens
Inputs and outputsText in, text out, best suited to chat format prompts
LanguagesEnglish
Chat template<|im_start|>role<|im_sep|> ... <|im_end|>
Training data9.8T tokens
Training compute1,920 H100-80G GPUs, 21 days
Training windowOctober 2024 to November 2024
Post-trainingSupervised fine-tuning and direct preference optimization
Knowledge cutoffJune 2024
Release dateDecember 12, 2024
LicenseMIT

The data recipe is the distinctive part. Instead of scaling the parameter count, Microsoft extends the mix used for Phi-3: public documents filtered to contain the correct level of knowledge, synthetic textbook-like data for math, code, common-sense reasoning and world knowledge, and chat format supervised data covering instruction following, truthfulness and helpfulness. Multilingual text is only about 8% of the mix, and Microsoft is explicit that Phi-4 is not intended for multilingual use, so treat it as an English model.

Phi-4 benchmarks

Microsoft published these numbers on the model card, measured with OpenAI's SimpleEval, against the predecessor Phi-3 14B, the same-size Qwen 2.5 14B Instruct, GPT-4o, and the much larger Llama 3.3 70B Instruct:

BenchmarkPhi-4Phi-3 14BQwen 2.5 14BLlama 3.3 70BGPT-4o
MMLU
Academic knowledge
84.877.979.986.388.1
GPQA
Expert science
56.131.242.949.150.6
MGSM
Multilingual math
80.653.579.689.190.4
MATH
Competition math
80.444.675.666.374.6
HumanEval
Code generation
82.667.872.178.990.6
SimpleQA
Factual knowledge
3.07.65.420.939.4
DROP
Complex reasoning
75.568.385.590.280.9

Phi-4 wins GPQA and MATH outright, ahead of GPT-4o and the five times larger Llama 3.3 70B, while the bigger models keep MMLU, MGSM, HumanEval and DROP. The clear weakness is factual recall: at 3.0 on SimpleQA the model should look facts up, not remember them, so pair it with retrieval for knowledge work.

Two things qualify that table. Microsoft ran the comparison on simple-evals because it is reproducible, and flags that its strict formatting gives Llama models trouble: Meta itself reports 77 on MATH and 88 on HumanEval for Llama 3.3 70B, against the 66.3 and 78.9 measured here. The second caveat is code scope: Microsoft states that the majority of Phi-4's code data is Python built on common packages such as typing, math, random, collections, datetime and itertools, and recommends verifying API use by hand when the model writes another language or reaches outside that set.

Safety got its own post-training stage on data covering helpfulness, harmlessness and specific safety categories, and before release Microsoft's independent AI Red Team probed the model with jailbreaks, encoding-based attacks, multi-turn attacks and adversarial suffix attacks.

Phi-4 hardware requirements

The system requirement to check is memory: the builds below are Microsoft's own GGUF files, published as microsoft/phi-4-gguf.

MemoryBuild to pickFile size
5 GBTQ2_04.23 GB
6 GBQ2_K5.55 GB
8 GBIQ3_M6.91 GB
10 GBIQ4_XS8.01 GB
12 GBQ4_K9.05 GB
16 GBQ6_K12.03 GB
24 GBQ8_015.58 GB
32 GB and upbf1629.32 GB

Neighbouring files differ by a gigabyte or two, so when two builds both fit, take the larger one. Quality falls fastest at the bottom of that table: the ternary builds, TQ1_0 at 3.59 GB and TQ2_0 at 4.23 GB, are the only way to fit a 14B model under 5 GB, and they pay for it, so go there only when nothing else fits.

On sampling, Microsoft sets temperature 0 in the inference parameters on the model card, and the reference stack it ships for the original weights is a transformers text-generation pipeline with torch_dtype and device_map both on auto. If the GGUF format is new to you, start with what GGUF is, and see the best local LLMs for a 16 GB Mac for what else runs well in that memory class.

How to run Phi-4 in Atomic Chat

Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.

  1. Download Atomic Chat for your platform and open it.
  2. Search for Phi-4 in the model browser and open Download Options.
  3. Pick the build that fits the memory you have, then start a chat.

For the rest of the family, see every Phi model you can run locally, or the smaller sibling Phi-4-mini-instruct.

Phi-4 license

Phi-4 is released under the MIT license, one of the most permissive licenses an open model ships with. It permits commercial use, modification and redistribution with no royalties, so you can build products on top of the model and run it on your own hardware without a usage fee.

Get the weights from Hugging Face

pip install -U transformers
huggingface-cli download microsoft/phi-4
# or with Ollama:
ollama run phi4
curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "microsoft/phi-4",
    "messages": [{"role": "user", "content": "Solve: 24*17"}]
  }'
import transformers
pipeline = transformers.pipeline(
    "text-generation",
    model="microsoft/phi-4",
    model_kwargs={"torch_dtype": "auto"},
    device_map="auto",
)
out = pipeline([{"role": "user", "content": "How should I explain the Internet?"}], max_new_tokens=256)
print(out[0]["generated_text"])
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "http://localhost:8000/v1", apiKey: "local" });
const res = await client.chat.completions.create({
  model: "microsoft/phi-4",
  messages: [{ role: "user", content: "Write a Python function for fibonacci." }],
});
console.log(res.choices[0].message.content);
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

Phi-4 is a 14-billion-parameter dense decoder-only language model from Microsoft Research, released on December 12, 2024. It was trained on a blend of synthetic "textbook-like" data, filtered web documents, and academic Q&A sets, with a focus on reasoning quality over raw scale. Phi-4 is tuned for math, coding, and logical reasoning, and is best used with chat-formatted prompts.

At 4-bit quantization (Q4_K_M) Phi-4 needs roughly 9 GB of VRAM, so a 12 GB GPU runs it comfortably. On an 8 GB card it slightly overflows and Ollama or llama.cpp will offload the extra layers to system RAM, which works but slows generation. Apple Silicon with unified memory and CPU offloading are also supported.

Yes. Microsoft released Phi-4 under the permissive MIT license, which allows commercial use, modification, and redistribution. The weights are available on Hugging Face and through Ollama, Azure AI Foundry, and GitHub Models at no cost.

Phi-4 has a 16K-token context window. That is enough for long chat sessions, multi-step reasoning problems, and moderately sized documents, though it is shorter than the 128K windows offered by some larger competing models.

Despite being only 14B parameters, Phi-4 is competitive with much larger models on reasoning and math. It scores 84.8 on MMLU, 80.4 on MATH, and 56.1 on GPQA, beating Qwen 2.5 14B and outperforming GPT-4o on the GPQA science benchmark. Its weak spot is factual recall, where SimpleQA scores are low, so it is better at reasoning than at memorized world knowledge.