This website uses cookies

Read our Privacy policy and Terms of use for more information.

In partnership with

Heterogeneous orchestration is here.

And it has arrived with a hundred million dollars. Some of it the UK government's. And every launch partner is a rival of the one chipmaker everyone else builds around.

If successful, this could be the most important idea in AI infrastructure.

Or, a very expensive exercise in orchestrating complexity.

I'm Ben Baldieri. Every week, I break down what's moving in GPU compute, AI infrastructure, and the data centres that power it all.

Here's what's inside this week:

Let's get into it.

Today's issue is brought to you by Rafay

The GPU stays independent because sponsors cover the bills, not because they shape the copy. Rafay works the orchestration and GPU-management layer beneath all of this, so if you build, buy or fund the physical stack of AI, they are worth a look.

AI Infrastructure Leadership Summit

Hosted by Rafay, join fellow tech executives for this two-day summit bringing together leaders from the world's most innovative neoclouds, telcos, NVIDIA Cloud Partners, sovereign AI initiatives, and enterprise AI platforms to share real-world strategies for commercializing AI infrastructure, expanding margins, and building the next generation of AI businesses. *Registrations subject to approval

Callosum Raises $100M to Route AI Workloads Across Rival Chips

The same week a payments giant paid $7 billion for the layer that routes between models, a British startup raised $100 million to route between the chips beneath them.

Callosum has raised $100 million in what it calls a seed, led by Atomico with Plural, DCVC and the UK Sovereign AI Fund, the state fund's first equity cheque. It builds a heterogeneous integrator, a layer that splits a workload and routes each piece to the model and chip that runs it best on cost, speed and power, with Cerebras and Korea's Rebellions as launch partners. It is the OpenRouter thesis one layer down, routing between the chips rather than the models above them.

Why this matters:

  • The value is migrating to whoever orchestrates. Owning the chips matters less than owning the layer that decides which chip runs what, and Callosum is betting that layer becomes a business of its own.

  • Heterogeneity is the anti-lock-in bet dressed as sovereignty. None of Callosum's launch partners is NVIDIA, and the Sovereign AI Fund spending its first equity cheque here, not on the chips the state's separate £1.1 billion plan targets, is Britain backing the orchestration layer over the silicon.

  • The caveat is the stage. Orchestrating live workloads across different chips at production latency is what everyone who has tried heterogeneous compute has found brutal; the thesis reads right, but the partner logos are a launch rather than a track record.

Nebius Prices an Upsized $5 Billion Convertible Note to Fund Its GPU Buildout

Nebius asked the debt market for $4.5 billion to build AI data centres and walked away with $5 billion.

Nebius has priced an upsized $5 billion offering of convertible senior notes, up from the $4.5 billion it first proposed, in tranches due 2030 and 2034. The money funds data centres and GPUs. This is the same play as CoreWeave and Lambda in Issue #119: borrow against future demand to fund GPUs it has not installed, and bet the capacity is contracted before the interest compounds.

Why this matters:

  • The debt market is still voting for AI infrastructure: the same appetite carried CoreWeave to $35 billion in Issue #119, and this is Nebius's turn at the window.

  • Convertibles push the risk out and down: a low coupon now, dilution to shareholders later if the stock rises, and if it does not, plain debt due in 2030 and 2034 against GPUs that age faster than that.

  • The money to build is coming from everywhere except operating profit: the UK state stands behind DataVita below, AWS pays cash, Nebius taps convertibles; different routes, same buildout, little cash flow under any of it.

Groq Raises $350M With NVIDIA in the Round, Turning From Chip Challenger to NVIDIA Neocloud

Groq made its name on a chip built to beat NVIDIA at inference, and just raised $350 million, NVIDIA included, to deploy NVIDIA instead.

Groq has closed a $350 million Series A led by Disruptive at a $3.5 billion valuation, with NVIDIA's participation planned but not yet closed. It follows the $650 million Groq raised in June, so a company that made its name trying to beat NVIDIA has pulled in $1 billion in three months. The money goes to NVIDIA clusters for training and inference, scaling Groq from 54MW to more than 200MW in 2027, and it now sits inside NVIDIA's channel as a certified NVIDIA Cloud Partner.

Why this matters:

  • Vendor financing from Issue #119 in miniature: NVIDIA joining a round whose proceeds buy NVIDIA compute is the same circular structure, at $350 million rather than $500 billion, and direct rather than through a platform.

  • The most credible non-NVIDIA inference bet just became an NVIDIA reseller. Its LPU was the alt-silicon story others were measured against; buying NVIDIA clusters says the money now sits in running NVIDIA at scale rather than out-engineering it.

  • For NVIDIA, it is cheap insurance. A few hundred million turns its sharpest inference-chip rival into a paying customer and shrinks the field still trying to beat it.

Stripe Buys OpenRouter for More Than $7 Billion, and With It the Meter on AI Inference

OpenRouter billed itself as the Stripe of AI, and Stripe just agreed to pay more than $7 billion to make it literal.

Stripe has agreed to buy OpenRouter for more than $7 billion, both companies now confirming it after Bloomberg first reported it; Stripe has not disclosed the price, which has been put north of $7 billion. OpenRouter, founded in 2023, routes requests across more than 400 models for a claimed 8 million users, and was worth $1.3 billion at a Series B in May, a fivefold markup in ninety days. Founder Alex Atallah co-founded OpenSea, the NFT marketplace whose usage cratered through 2022.

Why this matters:

  • The margin is moving to the meter rather than the model. Seven billion dollars for a router that trains nothing, owns no GPUs and holds no weights says the value sits at the toll booth between demand and compute.

  • OpenRouter is commoditisation turned into a business. It only works because models are now interchangeable enough to route to the cheapest that clears the bar, the slide we tracked through DeepSeek and Qwen in Issue #118 and Meta's give-away in Issue #119, now priced at $7 billion.

  • Whoever owns the routing layer sees the demand signal first. For anyone selling GPUs, clouds or models, a payments giant now sits between them and 8 million users, holding the record of which models, and which silicon, win the work.

Cerebras Launches the CS-4, Claiming Up to 30x Faster Inference Than GPUs

The same week Groq turned to reselling NVIDIA, Cerebras launched a wafer-scale system it says runs inference up to thirty times faster than GPUs.

Cerebras has launched the CS-4, its fourth-generation wafer-scale system, three WSE-3 Turbo processors in a redesigned rack it calls Nexus. It claims up to 30x faster inference than GPUs and more than 1,000 tokens a second on models above 10 trillion parameters. The WSE-3 Turbo is the existing wafer pushed harder rather than new; the gain comes from the rack, which moves power conversion about 100 times closer to the silicon to roughly double the power reaching the wafer. First shipments begin this quarter.

Why this matters:

  • Two non-NVIDIA inference bets went opposite ways in one week. Groq turned its capital toward NVIDIA racks; Cerebras answered by overclocking its own wafer and claiming 30x. The alt-silicon field is narrowing to the few still building against NVIDIA rather than around it.

  • Speed is becoming the product: OpenRouter routes to whichever model is cheapest and quickest, and Cerebras sells the hardware that wins the quickest half of that call. We watched OpenAI's Codex-Spark hit 1,000 tokens a second on Cerebras silicon in Issue #89; the CS-4 now claims that on models beyond 10 trillion parameters.

  • The number to watch is independent. Until someone measures 30x in production it stays a claim rather than a result. What is real is the option: buyers who care about latency now have a second non-NVIDIA answer.

AWS Pledges $6 Billion for a Louisiana Campus, Taking Its State Bet to $18 Billion

The neoclouds are borrowing to build; AWS just committed another $6 billion to a Louisiana campus and offered to pay for the power grid itself.

AWS has pledged $6 billion for a data-centre campus at the 313-acre Resilient Technology Park in Shreveport, Louisiana, developed by Stack Infrastructure, taking its Louisiana spend to $18 billion across three sites. It will fund the water and wastewater work, and with Southwestern Electric Power fully fund the grid upgrades the campus needs. It has not said how much capacity the campus will hold.

Why this matters:

  • Hyperscaler capex is the counterweight to the debt stories. Paying for a $6 billion campus and its grid from operating cash flow is what the neoclouds cannot do, the same buildout without the interest bill Nebius and CoreWeave carry.

  • Power is the real commitment. Funding the utility's grid upgrades points at the true long pole in an AI campus, new generation and transmission rather than chips; the binding constraint is increasingly the megawatt.

  • The missing number is the point. Amazon guides to roughly $220 billion of capex this year yet will not say how many megawatts Shreveport buys, the recurring gap between hyperscaler spending and the units to size it.

DataVita Closes £300M for Its AI Growth Zone, Every Megawatt Pre-Leased to CoreWeave

The UK just put a £202 million state guarantee behind an AI data centre for the first time, and CoreWeave has already pre-leased the lot.

DataVita has closed £300 million of financing for its North Lanarkshire campus, anchored by a £202 million guarantee from the National Wealth Fund, the state lender's first backing of a compute project. The guarantee covers about 80% of a £252.5 million senior debt tranche led by ING, ABN AMRO and Santander. The money funds a DV1 expansion at Chapelhall and a new DV3 build, every megawatt pre-leased to CoreWeave on a 15-year term, so the whole buildout is contracted before it opens.

We have tracked this site since Issue #86, when it became Scotland's first AI Growth Zone, an £8 billion project with 3,400 jobs; by Issue #89 CoreWeave was the named tenant and DataVita had won a £45 million Glasgow City Council contract.

Why this matters:

  • A national government just put its balance sheet behind an AI data centre. Guaranteeing 80% of the senior debt does for a British build what NVIDIA is institutionalising privately in Issue #119: making a depreciating, single-tenant asset financeable by backstopping the downside.

  • The sole tenant is carrying about $35 billion of debt. CoreWeave's 15-year lease is what makes DV3 bankable, and a UK state guarantee now sits behind capacity whose revenue line depends entirely on that one balance sheet holding.

  • AI Growth Zones are Britain's answer to Fermi's Texas and China's superclusters. Fermi needed a neocloud to sign before its campus made sense and China builds the floor itself; the UK uses a state guarantee to pull private debt into the ground faster than the market would alone.

Everything Else

p.s. The GPU Daily now goes out every morning. Opt in below if you want the physical layer of AI in your inbox daily.

Want The GPU Daily in your inbox?

Monday to Friday, a fast morning brief on the physical layer of AI. Opt-in only, pick below.

Login or Subscribe to participate

Reply

Avatar

or to participate