How do the AI labs start making money?
My response to Dwarkesh's second question
Below is my submission to question #2 in Dwarkesh’s essay contest. Along with how my opinion has changed slightly since then.
What’s the most plausible story where foundation model companies actually start making money? If you consider each individual model as a company, then its profits may be able to pay back the training cost. But of course, if you don’t train a bigger, more expensive model immediately, then you stop making money after 3 months. So when does the profit start? Maybe at some point scaling will plateau, but if progress at the frontier has slowed down, then the combination of distillation and low switching costs (cloud margins result from high switching costs) makes it really easy for open source to catch up to the labs, eating into their margins. So how do the labs actually start making money?
How do the foundation model labs actually start making money? If we isolate the model layer in Jensen's 5-layer cake, the answer isn't obvious. The model layer sits in a structurally hard position, with high upfront costs and easy replication. History doesn't have many examples of companies making money with economics like this.
The closest historical analogies are drug discovery and operating systems. The former relies on patents to make money, which doesn’t work in AI because abstract algorithms are largely unpatentable, and the actual product (model weights) can be functionally replicated through distillation. The latter (Windows specifically) relied on exclusive OEM licensing which led to embedded network effects. Neither condition exists in AI: there’s no OEM equivalent, and even if there were, you can switch models almost instantly.
Since monetization at the model layer isn’t possible, any foundation model company must go up a layer, to applications, in order to make money.
This shouldn’t be controversial. All frontier lab revenue thus far has been made at the application layer through chat bots and APIs. But the question is whether enough money can be made at the application layer to justify the costs incurred at the model layer. The two variables that matter are pricing power and future model development costs. Let’s look at three scenarios:
Jagged model progression - A singular frontier lab maintains a meaningful lead in perpetuity. (High pricing power, high future model development costs)
Broad model progression - Business as usual. Frontier labs trade off having the best model at any one time. (Moderate pricing power, high future model development costs)
Commoditization - Model progress slows to the point that all models, including open source ones, are effectively the same. (Low pricing power, moderate future model development costs)
Scenario 1, while unlikely, would certainly allow long-term profitability, despite maintaining high R&D spend. The winning lab could use their pricing power to ensure profitability since people are willing to pay for the best model. Scenario 2 and 3, on the other hand, see their pricing power diminish while still incurring model training costs. They also seem to be the most likely outcomes.
It doesn’t matter if commoditization happens, scaling plateaus, or even if the current landscape persists, the frontier model labs must get application pricing power in order to be profitable. And achieving pricing power is about finding sustainable moats.
Building a deep and wide B2C moat is inherently difficult. The largest consumer apps maintain their users through deeply entrenched network effects (Facebook, WhatsApp) or switching costs (Gmail). Building moats like these will require the foundation model labs to build generationally great consumer products that are at best tangential to their core competency. Saying this will be difficult is a massive understatement. This is a little like conquering the world to give your people more living space.
B2B software moats are less about network effects and more about switching costs. However, businesses differ from consumers in that they have the propensity and incentive to build. With total cost of ownership dropping for custom software, businesses will opt to build more of their own applications themselves. Given this, the modularity of the different coding models is stark - not only does each one largely work the same, but you can switch one out for another with virtually no disruption at all.
However, there is a potential switching cost moat for both B2B and B2C products: memory. If the current AI platforms leverage their current advantage to gain enough information about you to dramatically improve your or your firm’s experience, then they could use this as a long-term defensible moat. But this has yet to become meaningful so far.
The possible memory moat raises a question - do you even need to own the model itself to take advantage of memory lock-in? Or is owning the application data enough? You can broaden the question even further to get to the crux of this essay: does owning a model improve the frontier lab’s ability to deploy applications? We know using one does, but it’s unclear if owning one does too.
This is basically a vertical integration question. We have plenty of examples where owning commoditized lower levels of a stack improves the viability of higher levels. In 2020, Apple ditched Intel and started designing their own chips. Doing this both reduced upstream supplier issues and improved the performance of their operating systems and applications. Tesla is famous for doing the same thing with many of their components.
I don’t see synergies like this with the AI model and application layers. The application-model interface is just text in, text out, the same API boundary whether you own the model or not. Apple’s value from owning silicon comes from the chip-OS interface being co-designed at the hardware level, where every layer can be tuned against every other. There’s no analogous integration surface in AI. There’s no manufacturing cost advantage, and the inherent modularity means there won’t be integrated co-design benefits either.
The answer to whether foundation model labs make money as foundation model labs is no. Owning a foundation model doesn’t structurally advantage you in capturing the application moats that generate pricing power, and pricing power is what justifies the cost of the next model.
In the more likely cases, the companies positioned to profit from foundation models are the ones that already have application businesses big enough to absorb the cost. Google is an example of a foundation model lab that is profitable. Therefore, Anthropic and OpenAI simply need to become Google: either build generational application businesses themselves or accept that the natural endpoint is either acquisition or nationalization. Neither outcome is failure, exactly, but neither matches the foundation model lab as a standalone business model that current valuations imply.
Post-Submission Thoughts
My thoughts have expanded a little since I originally wrote that essay a few months ago. For one, I read Michael Li’s winning essay on this question (and suggest you do too) which unsurprisingly proposed something that I missed. The conclusion in my submission was pretty pessimistic, but Li’s essay offers another path: try to own some of the value you are creating. The example he uses is the Hong Kong Mass Transit Railway, which is one of the only transit systems that is self-sustaining, pays dividends, and doesn’t rely on government support.
The economics of mass transit make paying back the development costs virtually impossible with rider fares alone. So how does MTR do it? They purchase and develop land that is made more valuable from being connected to their line, and capture that value.
Li says the AI version of this is forward-deployed integration. This means they need to use their models to build solutions that can then capture the value they’re creating, rather than just selling low margin tokens. If your model can aid drug discovery, you should try to capture a share of the drug’s value rather than selling low-margin tokens to whoever does that instead.
I think this is the answer. And apparently, so do the labs. The path forward is to get closer to where the value is created, instead of selling cheap subway tickets when you’re competing with the bus.

