I went down a rabbit hole following the money behind Frontier Models. I started with trying to better understand what seemed fairly intuitive - why are companies making increasingly capable models open weight... and I came out with a broader question on how we are measuring competition, who can carry the cost of Frontier intelligence, where else in the value chain is the bill moving, and whether the economic support structures we see now are the path forward.
Over the past few days, there have been some pretty revealing developments:
Nvidia is spending $7B on a deal with Poolside.ai ($6B for a non-exclusive license + $1B equity investment + over 100 engineers to accelerate its Nemotron open-weight models). Poolside remains independent. Its people and technology are being pulled into Nvidia’s Nemotron efforts.
Alibaba announced a $10.2B share sale. It’s planning to use all of the proceeds on AI - chips, infra and models. It’s doing this even though heavy AI spending brought down its Q2 quarterly profit by 75%.
Stripe agreed to buy OpenRouter for $8B. OpenRouter doesn’t build frontier models. It routes requests between them and calls itself the “Stripe for LLMs”. It was already handling more than 10T tokens a day across 400+ models, for over 10M developers and companies.
Now these look like three very different stories.
Chip company spends billions around open-model efforts.
Tech conglomerate raises equity to finance just about everything from silicon to models.
Payments company pays billions for the layer that helps customers choose between all those models.
But put together, they’re pointing to an open question - with all the spending, where exactly does the return on intelligence need to turn up?
Company A and Company B can produce very similar capability models, and still have completely different prices. One needs the model itself to generate enough cash to fund its next generation. the other benefits if the model is widely used, because it sells the chips, cloud, advertising, devices or services around it.
So who can afford to make intelligence the cheapest may not be the company that produces it most cost effectively. It may be the company with the strongest economic system around it.
The visible AI race we are watching is model against model. Underneath there is a second competition forming between the economic systems that support these models.
I started here - why are increasingly capable models turning up as open weights - the first obvious answer was Complements.
Nvidia sells GPUs. Meta has advertising and distribution. Alibaba has cloud and commerce. Google has TPUs, Cloud, Android and devices.
The concept of giving something away, or making it cheaper, to increase the value of something else isn’t new. It makes sense. But let’s look at the capital structure behind these model companies.
I looked at about 20 open-weight model families (because my weekends are just so interesting) to see if this is the answer. Roughly, this is what turned up:
Alibaba owning Qwen is very different from ASML investing in Mistral, which is again very different from a government supporting domestic AI capability. And different from a normal venture investor who expects the AI company itself to eventually generate returns.
There’s no one common funding model, but what shows up is how few of these open models actually sit in a simple enough loop where model profits fund repeated frontier development.
(X Axis - relative strength of non model pathway to revenue, Y Axis - relative economic capacity to support repeated development, Z Axis - AI Index Model scores - This is purely a subjective analysis for demonstrating the points in the article - use it directionally)
So is this an open-weights strategy - where the goal is to reduce scarcity around the model and have something else pay for it. i.e. complements?
Turns out closed models also need large economic systems
Now, OpenAI can charge directly for its intelligence. It has one of the strongest consumer AI brands in the world, an enormous API business, and growing enterprise revenue. It has still ended up raising $110B this year (Amazon, Nvidia, SoftBank)
Anthropic’s annual revenue run rate is over $65B showing that frontier intelligence can generate very substantial direct revenue. But Anthropic is also expanding its revolving credit facility by over $10B on top of its existing capital and infrastructure relationships.
xAI raised $20 billion in its Series E funding round before it got acquired by SpaceX, and DeepMind has always had Alphabet behind it.
At what point does building frontier AI start looking less like software and more like infrastructure?
The frontier has researchers, data, inference, massive training runs, post-training, failed experiments, networking, power and physical infrastructure. Some of those costs are falling very quickly.
But the ambition of companies building at this frontier must also keep rising.
Yesterday’s frontier is getting cheaper while staying at the frontier remains extremely expensive - it makes it harder to assume that the winner will simply be the one who has the best model - but who has the economic system that can support the next frontier, and the next one, and the one after that.
Mistral is genuinely independent and with a real business (annual revenue run rate crossed $400M in 2026 with 100+ enterprise customers). But the perimeter around Mistral is now much wider than Mistral itself. ASML is its largest shareholder, Microsoft supports infrastructure and distribution, and Mistral is building its own compute. It is also co-developing the first Nemotron coalition base model with NVIDIA, while NVIDIA is also bringing Poolside technology and engineers into the Nemotron effort.
Alibaba doesn’t need Qwen to carry its overall AI Strategy by itself. It can spread costs across cloud, chips and applications. Sarvam (India) saw HCLTech investing around $150M for a 10.5% stake in it. The two are engaged in a $1.48B AI data-centre project with the Odisha government. The stack includes the data centre, GPUs, models and applications. The returns can be gotten across several layers.
I think while Model Leader boards are useful, they don’t tell us the full story of what sits behind the models that look similar on the board. They can be products of completely different economic systems.
One company needs to make money only from intelligence..Another can make money somewhere else because intelligence gets used…A third justifies the cost through national capability…Yet another uses public equity…And another can spread costs across infrastructure, distribution and customers...
Capability of course matters. But once the models are capable enough to be substitutes for a workload, pricing, distribution, capital, and the ability to keep funding the next generation starts to matter as well,
Their models may compete directly on capability. Their businesses could be competing on very different terms.
Second Order: Open weights can reduce concentration at the model layer but increase the value of routing, compute, distribution or proprietary data elsewhere.
OpenRouter doesn’t build models at all. The more models that exist, the harder it becomes for an app developer to answer which model should handle this request? Different models have different prices, speeds, context windows, capabilities and provider availability. It has built an intermediation layer supporting 100s of models from multiple providers through one interface, processing over 10 Trillion tokens / Day. Now Stripe is paying $8B to own it.
More competition and commoditization at the model layer is creating a valuable set of aggregation points on top of it. And you’ll see it not just in training but inference -
Fireworks Series D funding was $1.5B at $17.5B. It reports $1B annualiazed revenue and 40T+ tokens/day.
Baseten has raised $1.5B at $13B, reporting 20x revenue growth for deployment and serving.
Cloud providers capture compute.
Hugging Face captures discovery and developer distribution.
Device companies capture value when good local models make their hardware more useful.
The bill isn’t disappearing, it’s changing home.
The ability to afford making intelligence cheaper is in itself a competitive advantage
If you own something that will be useful when intelligence gets cheaper, you’ve got your reasons to help the intelligence get cheaper.
Not to say that all roads lead to concentration eventually - Open weights will enable self hosting, clouds will compete, accelerators are growing, inference pricing is falling fast, infrastructure is getting more competitive - customers can move between models much quicker than if every application depends only on proprietary APIs.
But there is an asymmetry - some competitors can tolerate weaker economics at the model layer than others can, and that’s a significant competitive strength.
There is a counter argument though:
Arcee -
Its Trinity large model has 400B+ total parameters with 13B active per token. It’s trained on 2000+ NVIDIA B300s, and final pre-training took under 33 days. The full 6 month development lifecycle across 4 models has taken it less than $20M all inclusive (compute, salaries, data, storage, operations)!
That is a completely different capital structure from the large frontier programmes. But it’s not just the bill size but the invoiced items - sparser architecture, synthetic data, RL environments - breaking the pieces up.
Also, DeepSeek -
Its V3 docs say it needs only 2.8M H800 GPU hours for its full training. It says training costs were $5.6 million, but that’s not the total development cost (SemiAnalysis has independently estiamted $500M for historic GPU investment and $1.6B of server Capex) - but everyone is agreeing that the achievements in efficiency of these companies is definitely looking real.
Models will get cheaper. But will the Frontier keep moving faster than the speed at which yesterday’s frontier gets cheaper?
Architectural efficiency of models like Arcee also come at the time that raw pre-training scaling laws are starting to face diminishing returns (Check out the Cameron R. Wolfe, Ph.D. post on frontier labs actively working towards post-training and inference efficiency - it’s a very dense read, but I learned a lot)
Follow the return
So in this little bit of a meandering journey, I originally wanted to learn why increasingly capable models were being made open. The bigger question is how companies will keep financing Frontier AI in the first place. Giant balance sheets, infra partners, public markets, sovereign capital , and yes, complements.
These differences can change how long a company keep investing, what the model itself needs to fund, and how aggressively it can price and distribute that intelligence.
If the Frontier stays expensive, this model is the path forward to stay in the game. But if the cost structure of building frontier models is evolving as we can see from companies like Arcee - Better data with post training, RL environments, synthetic data, are some of the approaches making things more cost efficient and modular - which path accelerates/sustains, whether this migrates the cost, or lowers the capital bar remains to be seen.
Pallavi Chari , Moving Parts






