Microsoft unveils MAI-Thinking-1 and MAI-Code-1-Flash
Microsoft announced two new text models on June 2: MAI-Thinking-1, a reasoning model, and MAI-Code-1-Flash, a code-focused model for GitHub Copilot and Visual Studio Code. Microsoft says both were trained on clean, commercially licensed data, though the technical paper for MAI-Thinking-1 also describes large-scale web crawling and filtering.
Microsoft announced two new text large language models on June 2: MAI-Thinking-1 and MAI-Code-1-Flash. The company said MAI-Thinking-1 is a reasoning model with 1 trillion parameters and 35 billion active parameters, while MAI-Code-1-Flash has 137 billion parameters with 5 billion active parameters. Microsoft said MAI-Thinking-1 is available to “select early partners,” and MAI-Code-1-Flash is rolling out to GitHub Copilot individual users in Visual Studio Code.
The announcement drew attention because Microsoft is positioning both models as relatively efficient. The source notes that the author was initially struck by the fact that MAI-Thinking-1 is only 35B active parameters, which is smaller than many widely used models, even though Microsoft says it performs well in its internal tests.
⚡ New to this?
This matters because MAI-Thinking-1 and MAI-Code-1-Flash are Microsoft’s own AI models, which can affect products like Copilot and Visual Studio Code. A large language model, or LLM, is the system that predicts and generates text; “active parameters” is the part of the model that actually gets used for a given request, which can be smaller than the model’s total size.
The data claims matter too. “Licensed data” means data the company says it has permission to use, while web crawling means collecting large amounts of public web pages for training.
🦞 OpenClaw angle
If you run self-hosted coding or reasoning agents, watch the split between total size and active parameters when comparing models. A smaller active-parameter model can still be a strong choice if your hardware budget is tight, so test throughput and latency on your own workload instead of assuming bigger is better.
Also review your data policy if you train or fine-tune internal models. Microsoft’s disclosure shows that licensing and web-crawl filtering are now part of model selection, not just a legal footnote, so keep provenance records for any corpus you build or ingest.
Microsoft claims MAI-Thinking-1 was “trained from the ground up on enterprise grade, clean and commercially licensed data, without distillation from third-party models.” For MAI-Code-1-Flash, Microsoft said it was “built end-to-end” using “clean and appropriately licensed data.” Those claims matter because model makers are under pressure to show where training data comes from and whether it was licensed properly.
The initial writeup was later corrected after the model sizes were misread. The MAI-Code-1-Flash model card lists it as a 137B model with 5B active parameters, and the MAI-Thinking-1 technical paper describes it as a 1T model with 35B active parameters.
Microsoft also said MAI-Thinking-1 was preferred to Anthropic’s Sonnet 4.6 in blind human side-by-side evaluations. That is Microsoft’s own benchmark claim, and the source does not include independent verification of the result.
The bigger question raised by the technical paper is what Microsoft meant by “clean” data. According to the paper, MAI-Thinking-1 was trained on a large crawl of the public web. Microsoft said the majority of its web HTML corpus comes from a proprietary crawl, and that it initially gathered about 1.2 trillion pages before filtering.
Microsoft said it used policy filters to remove adult content and piracy-related domains, then applied an AI-content detection model and manual review to exclude domains with extensive AI-generated content. After filtering and deduplication, the corpus was reduced to 794 billion pages.
The paper also says Microsoft processed Common Crawl through the same pipeline. After filtering, deduplication, and merging with the proprietary web corpus, the Common Crawl portion contained 24.2 billion pages.
The source text says this means Microsoft’s models are not free from the same web-crawling and licensing issues that affect other major LLMs. The author also said the technical paper was more detailed than the initial announcement suggested, and that the first published notes had underestimated the model sizes because of confusion over active versus total parameters.