What is Phi-4?
Phi-4 is a 14B dense decoder-only Transformer from Microsoft Research, released on December 12, 2024 under the MIT license. Microsoft trained it on 9.8T tokens, including synthetic textbook-style data written to teach math, coding, common sense reasoning and world knowledge, plus rigorously filtered public documents, acquired academic books and Q&A datasets. The stated primary use cases are memory and compute constrained environments, latency bound scenarios, and reasoning and logic: in other words, the profile of a model built to run on your own machine.
| Specification | Phi-4 |
|---|---|
| Developer | Microsoft Research |
| Parameters | 14B dense (14.66B exact) |
| Architecture | Dense decoder-only Transformer |
| Context window | 16K tokens |
| Inputs and outputs | Text in, text out, best suited to chat format prompts |
| Languages | English |
| Chat template | <|im_start|>role<|im_sep|> ... <|im_end|> |
| Training data | 9.8T tokens |
| Training compute | 1,920 H100-80G GPUs, 21 days |
| Training window | October 2024 to November 2024 |
| Post-training | Supervised fine-tuning and direct preference optimization |
| Knowledge cutoff | June 2024 |
| Release date | December 12, 2024 |
| License | MIT |
The data recipe is the distinctive part. Instead of scaling the parameter count, Microsoft extends the mix used for Phi-3: public documents filtered to contain the correct level of knowledge, synthetic textbook-like data for math, code, common-sense reasoning and world knowledge, and chat format supervised data covering instruction following, truthfulness and helpfulness. Multilingual text is only about 8% of the mix, and Microsoft is explicit that Phi-4 is not intended for multilingual use, so treat it as an English model.
Phi-4 benchmarks
Microsoft published these numbers on the model card, measured with OpenAI's SimpleEval, against the predecessor Phi-3 14B, the same-size Qwen 2.5 14B Instruct, GPT-4o, and the much larger Llama 3.3 70B Instruct:
| Benchmark | Phi-4 | Phi-3 14B | Qwen 2.5 14B | Llama 3.3 70B | GPT-4o |
|---|---|---|---|---|---|
MMLU Academic knowledge | 84.8 | 77.9 | 79.9 | 86.3 | 88.1 |
GPQA Expert science | 56.1 | 31.2 | 42.9 | 49.1 | 50.6 |
MGSM Multilingual math | 80.6 | 53.5 | 79.6 | 89.1 | 90.4 |
MATH Competition math | 80.4 | 44.6 | 75.6 | 66.3 | 74.6 |
HumanEval Code generation | 82.6 | 67.8 | 72.1 | 78.9 | 90.6 |
SimpleQA Factual knowledge | 3.0 | 7.6 | 5.4 | 20.9 | 39.4 |
DROP Complex reasoning | 75.5 | 68.3 | 85.5 | 90.2 | 80.9 |
Phi-4 wins GPQA and MATH outright, ahead of GPT-4o and the five times larger Llama 3.3 70B, while the bigger models keep MMLU, MGSM, HumanEval and DROP. The clear weakness is factual recall: at 3.0 on SimpleQA the model should look facts up, not remember them, so pair it with retrieval for knowledge work.
Two things qualify that table. Microsoft ran the comparison on simple-evals because it is reproducible, and flags that its strict formatting gives Llama models trouble: Meta itself reports 77 on MATH and 88 on HumanEval for Llama 3.3 70B, against the 66.3 and 78.9 measured here. The second caveat is code scope: Microsoft states that the majority of Phi-4's code data is Python built on common packages such as typing, math, random, collections, datetime and itertools, and recommends verifying API use by hand when the model writes another language or reaches outside that set.
Safety got its own post-training stage on data covering helpfulness, harmlessness and specific safety categories, and before release Microsoft's independent AI Red Team probed the model with jailbreaks, encoding-based attacks, multi-turn attacks and adversarial suffix attacks.
Phi-4 hardware requirements
The system requirement to check is memory: the builds below are Microsoft's own GGUF files, published as microsoft/phi-4-gguf.
| Memory | Build to pick | File size |
|---|---|---|
| 5 GB | TQ2_0 | 4.23 GB |
| 6 GB | Q2_K | 5.55 GB |
| 8 GB | IQ3_M | 6.91 GB |
| 10 GB | IQ4_XS | 8.01 GB |
| 12 GB | Q4_K | 9.05 GB |
| 16 GB | Q6_K | 12.03 GB |
| 24 GB | Q8_0 | 15.58 GB |
| 32 GB and up | bf16 | 29.32 GB |
Neighbouring files differ by a gigabyte or two, so when two builds both fit, take the larger one. Quality falls fastest at the bottom of that table: the ternary builds, TQ1_0 at 3.59 GB and TQ2_0 at 4.23 GB, are the only way to fit a 14B model under 5 GB, and they pay for it, so go there only when nothing else fits.
On sampling, Microsoft sets temperature 0 in the inference parameters on the model card, and the reference stack it ships for the original weights is a transformers text-generation pipeline with torch_dtype and device_map both on auto. If the GGUF format is new to you, start with what GGUF is, and see the best local LLMs for a 16 GB Mac for what else runs well in that memory class.
How to run Phi-4 in Atomic Chat
Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.
- Download Atomic Chat for your platform and open it.
- Search for Phi-4 in the model browser and open Download Options.
- Pick the build that fits the memory you have, then start a chat.
For the rest of the family, see every Phi model you can run locally, or the smaller sibling Phi-4-mini-instruct.
Phi-4 license
Phi-4 is released under the MIT license, one of the most permissive licenses an open model ships with. It permits commercial use, modification and redistribution with no royalties, so you can build products on top of the model and run it on your own hardware without a usage fee.
