Microsoft ships Phi-4-Reasoning at 14B parameters with MIT license
Microsoft released Phi-4-Reasoning, a 14B dense model trained with chain-of-thought methods. It scores 80.6 on AIME 2025 and 75.3 on GPQA Diamond, beating many models twice its size on reasoning benchmarks. The MIT license makes it one of the most permissive model releases of the month. Runs on a single consumer GPU with about 8-10GB VRAM at 4-bit quantization.
Microsoft has released Phi-4-Reasoning, a 14 billion parameter dense model aimed at reasoning-heavy tasks, and published it under the MIT license. The model is available through Microsoft Research on Hugging Face, which means developers can download it, inspect it, and use it under one of the most permissive open-source licenses available.
Phi-4-Reasoning is part of Microsoft’s Phi family, a line of relatively small models that tries to get more performance out of fewer parameters. That matters because many of the best-known large language models are much bigger, which usually means more memory, more compute, and more cost to run.
The company says the model was trained with chain-of-thought methods, a training approach that teaches a model to work through problems step by step rather than jump straight to an answer. In practice, that often improves performance on math, science, and logic tasks, especially when benchmarked against problems that require multi-step reasoning.
Microsoft is highlighting the model’s results on two benchmark suites that are commonly used to test reasoning. On AIME 2025, a benchmark based on challenging math problems, Phi-4-Reasoning scored 80.6. On GPQA Diamond, a difficult question-answering benchmark focused on graduate-level science questions, it scored 75.3.
Those numbers place it ahead of many models that are roughly twice its size, according to Microsoft’s release notes and benchmark framing. That comparison is part of the appeal of the Phi line: smaller models that are easier to run, but still competitive on targeted tasks where reasoning quality matters more than general breadth.
The size also matters on the deployment side. According to the model card and release summary, Phi-4-Reasoning can run on a single consumer GPU with roughly 8 to 10GB of VRAM when quantized to 4-bit precision. Quantization reduces the memory footprint by storing model weights in fewer bits, which is one of the standard ways to make larger models usable on smaller hardware.
That puts the model within reach of a wider range of teams, including developers working on local tools, private deployments, and AI systems that do not need a giant model for every step. A 14B dense model is still large by everyday application standards, but it is much more practical than frontier models that require datacenter-scale infrastructure.
The MIT license is another notable part of the release. Compared with more restrictive model licenses, MIT is simple and permissive, and it is often favored by people building commercial software, internal tools, and self-hosted systems because it places few restrictions on reuse.
Microsoft has been shipping the Phi series as a way to show that small models can be surprisingly capable when trained carefully on high-quality data and targeted tasks. Phi-4-Reasoning fits that pattern, with an emphasis on benchmark performance, low enough resource demands for local deployment, and a license that makes the model easier to adopt in real products and internal systems.