Microsoft chief executive Satya Nadella has issued a stark warning to companies that rely on artificial intelligence models from major providers. In a blog post published Sunday, he argues that organizations are unknowingly paying for AI twice: once through token fees and again by handing over proprietary business knowledge that can be used against them.
The debate is not new. Venture capitalists and tech executives have worried that startups and enterprises using AI models from labs like OpenAI and Anthropic are exposing sensitive data to companies that may eventually compete with them. Palantir CEO Alex Karp and investor Jason Calacanis are among those who have voiced similar concerns. But Nadella's intervention carries particular weight because Microsoft is a major investor in OpenAI and has partnered with Anthropic.
Key Facts From Nadella's Warning
- Nadella says AI buyers pay for intelligence twice: with money and with proprietary knowledge they must reveal to the model.
- Enterprises teach models through prompts, agent tool usage, and corrections, creating institutional know-how that could be used by competitors.
- Nadella argues that if AI labs can freely train on public internet data, companies should be allowed to study and distill those models in return.
- He urges businesses to retain ownership of their prompts, feedback, and interaction data, and to build “orchestration layers” to switch between AI providers.
- Industry data suggests a shift toward open source models running on-premises, as companies seek lower costs and more control.
The Trojan Horse Fear in Silicon Valley
The concern that AI labs act like Trojan horses has circulated in Silicon Valley for years. The logic is straightforward: when a startup integrates a proprietary model into its workflow, it sends prompts containing business secrets, financial metrics, product plans, and customer data to the model provider. The provider can then use that knowledge to improve its own models or even build competing products. The risks grow as AI agents become more sophisticated and gain access to internal tools and databases.
Nadella's blog post crystallizes this fear. He uses the term “buyers” to describe AI customers and warns that they are paying twice. The first payment is obvious: the usage fees charged for model calls. The second is more insidious: the data that customers must disclose in order to make the model useful. The more context, examples, and corrections a company provides, the better the model performs. But that same input becomes part of the model's knowledge base.
What Nadella Means by “Exhaust”
In his post, Nadella explains that models learn from what he calls exhaust. This includes the prompts people write, the actions AI agents take, and especially the corrections humans make when the model produces an error. Every correction is distilled into institutional know-how. Over time, a model becomes familiar with a company's internal vocabulary, decision-making patterns, and operational quirks. That kind of knowledge is remarkably difficult for a competitor to acquire through other means, yet enterprises give it away as a byproduct of using AI.
This is especially dangerous for companies in competitive industries. If a model provider trains on customer data from multiple players in the same market, it could inadvertently encode the strategies of one customer into the responses it gives another. Even if the provider never intentionally uses the data, the legal and reputational risks are significant.
Distillation and the Fairness Argument
The Microsoft CEO also addresses distillation, the practice of using a model's outputs to understand how it works and to train a new, often cheaper model. In February, Anthropic accused Chinese open source developers of sending millions of prompts to Claude to improve their own models. Anthropic urged the U.S. government to address this through export controls. Nadella argues that this position is hypocritical.
AI model makers have long benefited from fair use rights to train on publicly available internet data. If it is acceptable for them to scrape the world's information, Nadella suggests, it should be acceptable for others to learn from their models. He writes, “While the great innovation that comes from model providers having fair use rights to train models on public data is needed, I find it ironic that the status quo is to then turn around and impose restrictive terms on distillation.”
Nadella's Proposed Solution
Nadella's answer reflects his role as the head of a major cloud provider. He wants companies to retain ownership of their data, including prompts, feedback, and interaction records. He recommends building what he calls “proprietary learning environments” on the cloud, where their data is likely stored anyway. He also urges companies to create “orchestration layers” that allow them to switch easily between AI models from different providers. AI gateways, which sit between users and multiple models, have become increasingly popular for this purpose.
Skeptics may see self-interest in this advice. Microsoft owns Azure, one of the largest cloud platforms, and would benefit if more enterprises built their learning environments there. But the broader point about data ownership has struck a chord with IT leaders, especially those in highly regulated industries such as finance, healthcare, and government.
The Push Toward Open Source and On-Premises AI
Nadella never explicitly says “open source,” but his argument strongly implies it. Open source models can be downloaded and run on a company's own servers, giving the organization full control over both the model and its data. No prompts or corrections leave the building, and no proprietary knowledge is shared with a third-party lab.
Idit Levine, founder and CEO of Solo.io, a company that makes networking and security software for enterprise AI systems, says she sees this shift happening with her own customers. After experimenting with proprietary model makers, they start asking whether an open source model can do almost 90 percent of what the big model does at a much lower cost. “They understand that, and they can control it,” she says. Solo.io's technology was selected last year to power the Linux Foundation's Agentgateway project, and its customers include T-Mobile, ADP, and SAP.
Growing Adoption of Open Models
Other companies are seeing the same movement. Vercel, a platform for building and hosting websites, recently added AI model-switching tools. OpenRouter helps developers route requests across different AI models. Both report surging traffic to open source models. Vercel says open models accounted for 29 percent of all traffic routed through its gateway last month.
The appeal is straightforward. Open models can be downloaded, fine-tuned with internal data, and run on infrastructure the company already owns. This avoids per-token fees and ensures that proprietary knowledge stays in-house. The tradeoff is that open models may not match the largest proprietary models on every benchmark, but for many enterprise use cases the difference is acceptable.
What Enterprises Should Do Now
Enterprises that use AI need to weigh the benefits of cutting-edge models against the long-term risk of data leakage and vendor lock-in. Nadella's warning suggests that even leaders in the AI industry recognize the problem. Companies should audit how their data is being used by model providers, review contract terms for clauses that allow training on customer usage, and consider whether they have an exit strategy.
Some organizations will decide that the performance of the biggest models is worth the risk. Others will move to open source alternatives that provide “good enough” performance with far more control. The balance may shift as the gap between proprietary and open models narrows.
Model providers that reserve the right to learn from customer usage and interaction data create a particular concern. Nadella singles out these practices as especially troubling. For companies that operate in sensitive sectors, the safest approach may be to avoid sharing such data altogether.
Nadella's message is clear: “In consuming intelligence, you are creating intelligence. And what you create should belong to you.” That idea could become a rallying cry for data sovereignty and open source AI, even if it comes from the leader of a company with deep ties to the very labs he is warning against.
Source: TechCrunch News