Microsoft releases Phi-4-Reasoning - 14B model under MIT license beats models twice its size

Microsoft released Phi-4-Reasoning, a 14B parameter dense model under the MIT license. It scores 80.6 on AIME 2025 and 75.3 on GPQA Diamond, beating many models twice its size on reasoning tasks through chain-of-thought training.

Microsoft releases Phi-4-Reasoning - 14B model under MIT license beats models twice its size

Microsoft released Phi-4-Reasoning on April 10, a 14B parameter dense model that punches well above its weight on reasoning benchmarks. The model scores 80.6 on AIME 2025 and 75.3 on GPQA Diamond, numbers that beat many models twice its size on mathematical and scientific reasoning tasks.

The approach is chain-of-thought training. Rather than scaling up parameters, Microsoft trained Phi-4 on carefully filtered synthetic data that emphasizes step-by-step reasoning. The training process teaches the model to break problems into intermediate steps, show its work, and self-correct along the way. The result is a model that thinks through problems methodically rather than jumping to conclusions.

⚡ New to this?

Phi-4-Reasoning proves that a smaller model trained on better data can outperform much larger ones on reasoning tasks. The MIT license means anyone can use it for anything, including commercial products, with no restrictions.

🦞 OpenClaw angle

At 14B parameters, Phi-4-Reasoning runs comfortably on a single RTX 3090. Try it via Ollama (ollama pull phi4-reasoning) for tasks where reasoning quality matters more than speed. It is a good fit for agent tasks that involve multi-step logic, planning, or math.

The MIT license is the most notable aspect for builders. MIT is the most permissive open-source license available. There are no usage restrictions, no revenue thresholds, no registration requirements, and no obligation to share modifications. You can use it in commercial products, modify it, redistribute it, and never owe Microsoft anything. This removes all legal ambiguity for anyone building products on open models.

At 14B parameters with a dense architecture, the model runs on a single consumer GPU. An RTX 3090 with 24GB handles it comfortably, even without aggressive quantization. Inference is fast enough for interactive use, which matters for coding tools and chat interfaces where latency affects the experience.

The tradeoff is that Phi-4-Reasoning excels specifically at reasoning tasks. It is not a general-purpose chatbot, and it is not optimized for creative writing, open-ended conversation, or long-form content generation. Use it where structured thinking matters: planning, analysis, code logic, mathematical proof, multi-step problem solving, and decision trees.

For agent operators, this is a useful specialist model. You could route reasoning-heavy tasks to Phi-4-Reasoning while using a general-purpose model for everything else. The low resource requirements mean running it alongside another model is practical even on a single machine. A multi-model setup where simple queries go to a fast 7B model and hard queries go to Phi-4-Reasoning gives you the best of both worlds.

Microsoft's approach with Phi-4 challenges the assumption that bigger is always better. When training data quality and methodology are strong enough, a 14B model can compete with 70B alternatives on the tasks it was designed for.

Source: Microsoft ↗

More from AI News