Microsoft releases Phi-4-Reasoning - 14B model under MIT license beats models twice its size
Microsoft released Phi-4-Reasoning, a 14B parameter dense model under the MIT license. It scores 80.6 on AIME 2025 and 75.3 on GPQA Diamond, beating many models twice its size on reasoning tasks through chain-of-thought training.
Microsoft released Phi-4-Reasoning on April 10, a 14B parameter dense model that punches well above its weight on reasoning benchmarks. The model scores 80.6 on AIME 2025 and 75.3 on GPQA Diamond, numbers that beat many models twice its size on mathematical and scientific reasoning tasks.
The approach is chain-of-thought training. Rather than scaling up parameters, Microsoft trained Phi-4 on carefully filtered synthetic data that emphasizes step-by-step reasoning. The training process teaches the model to break problems into intermediate steps, show its work, and self-correct along the way. The result is a model that thinks through problems methodically rather than jumping to conclusions.
The MIT license is the most notable aspect for builders. MIT is the most permissive open-source license available. There are no usage restrictions, no revenue thresholds, no registration requirements, and no obligation to share modifications. You can use it in commercial products, modify it, redistribute it, and never owe Microsoft anything. This removes all legal ambiguity for anyone building products on open models.
At 14B parameters with a dense architecture, the model runs on a single consumer GPU. An RTX 3090 with 24GB handles it comfortably, even without aggressive quantization. Inference is fast enough for interactive use, which matters for coding tools and chat interfaces where latency affects the experience.
The tradeoff is that Phi-4-Reasoning excels specifically at reasoning tasks. It is not a general-purpose chatbot, and it is not optimized for creative writing, open-ended conversation, or long-form content generation. Use it where structured thinking matters: planning, analysis, code logic, mathematical proof, multi-step problem solving, and decision trees.
For agent operators, this is a useful specialist model. You could route reasoning-heavy tasks to Phi-4-Reasoning while using a general-purpose model for everything else. The low resource requirements mean running it alongside another model is practical even on a single machine. A multi-model setup where simple queries go to a fast 7B model and hard queries go to Phi-4-Reasoning gives you the best of both worlds.
Microsoft's approach with Phi-4 challenges the assumption that bigger is always better. When training data quality and methodology are strong enough, a 14B model can compete with 70B alternatives on the tasks it was designed for.