This website uses cookies

Read our Privacy policy and Terms of use for more information.

NVIDIA spent the week buying the parts of AI it doesn't already sell.

A reported $12.9 billion bid for Hugging Face, the home of open-source models. The same week, word that its own systems are getting more than 15% dearer.

Buy the commons.

Charge everyone more for the shovels.

I'm Ben Baldieri. Every week, I break down what's moving in GPU compute, AI infrastructure, and the data centres that power it all.

Here's what's inside this week:

Let's get into it.

NVIDIA Closes In on Hugging Face for a Reported $12.9 Billion

The chipmaker every open model runs on is moving to buy the place they all live.

NVIDIA has reportedly agreed to acquire Hugging Face for about $12.9 billion, though neither side has confirmed it and nothing is signed. Hugging Face is where the open-source AI world lives: the default host for model weights, datasets and the libraries most labs build on, with millions of models mirrored on its servers. It is the same hub that turned down NVIDIA's $500 million investment in 2023, when a single dominant backer was the one thing it did not want; it last raised at a $4.5 billion valuation and says it is near profitability on roughly $150 million of revenue. With OpenAI, Google, Amazon and Anthropic all building their own silicon, the place every open model is downloaded from is the one layer of the stack NVIDIA does not yet control.

Why this matters:

  • It is vertical integration dressed as a bolt-on: NVIDIA already sells the silicon every open model trains on, and owning where they are distributed lets it set the defaults and steer workloads before a rival's chip gets a look-in.

  • It doubles as a quiet return to the cloud business NVIDIA exited, with Hugging Face's inference endpoints a place to resell compute its customers have paid for and left idle, the same keep-it-in-house instinct that ran through #119.

  • The commons stops being neutral. The hub the whole field treats as shared ground would sit inside the company that already supplies the silicon beneath it, handing one vendor a say over the layer everyone else builds on.

Anthropic Rents $45 Billion of Compute From Nscale, Days Before Its IPO

The offtake that makes the float landed a fortnight before the float.

Anthropic has signed a reported $45 billion, six-year deal for about 460MW at Nscale's West Virginia campus, one of the largest commitments a frontier lab has ever made to a neocloud. It landed days before Nscale was said to be targeting a US IPO of up to $3 billion. This is the same Nscale we have tracked through a lost Dutch lawsuit (#72), FT-reported loan defaults (#99) and questions over its UK Stargate role (#96); the cheque re-papers all of it into an anchor tenant just in time to list. For Anthropic it is the latest ten-figure compute grab in months, after $10 billion to Volta (#118), $5 billion to AMD, and expanded capacity from Amazon and SpaceX.

Why this matters:

  • The offtake is the IPO. A $45 billion contracted revenue line is what lets bankers price a neocloud with a difficult file, and it is now the template every neocloud needs before it can list.

  • Anthropic is doing to neoclouds what NVIDIA does to labs: seed the demand, then let the counterparty raise the debt and pour the concrete (#115). The lab keeps its balance sheet clean while someone else carries the depreciation.

  • The concentration risk just moved to West Virginia. 460MW of single-tenant capacity is only as good as Anthropic's revenue in 2029, so Nscale's whole float now rides on one customer staying solvent for six years.

CoreWeave Puts Hudson River Trading on Vera Rubin Ahead of the Hyperscalers

The newest NVIDIA silicon is going to a trading firm before it goes to a hyperscaler.

Hudson River Trading has picked CoreWeave to build a research platform on NVIDIA's Vera Rubin NVL72, a multi-year deal CoreWeave calls multi-billion-dollar but leaves unpriced. HRT is one of the world's largest algorithmic trading firms, and it wants the newest silicon to shave microseconds off the research that sets its trades. It joins the first wave of tenants on the Vera Rubin racks CoreWeave was first to bring up (#109), alongside existing quant customers Jane Street and IMC. Vera Rubin only reached volume production this quarter, so HRT is running frontier hardware that AWS and Azure are still queuing for.

Why this matters:

  • Neoclouds are moving up-market, not just up-scale: a quant fund chasing latency is a different, stickier customer than a lab renting training capacity by the quarter.

  • First-silicon access is the whole pitch. Getting HRT onto Vera Rubin ahead of AWS and Azure is worth more than the headline price, because it proves CoreWeave can beat the hyperscalers to the newest hardware.

  • A financial-services book is the diversification a single-thesis neocloud needs. As the labs consolidate their compute into a few mega-deals, trading firms give CoreWeave revenue that does not ride on any one model winning.

A Stealth Model Called “Ox Alpha” Turns Out to Be China's Best Open Coding Model

The anonymous model topping the leaderboards was a Chinese lab all along.

Ox Alpha appeared on OpenRouter as a free, unattributed model and topped multiple coding leaderboards for a fortnight before Z.ai confirmed it was GLM-5.3-Flash, per TechCrunch and TechNode. Dropping a model anonymously, letting the boards rank it on merit, then claiming it is becoming a recognised launch tactic for the Chinese labs. It is a 320-billion-parameter model with 18 billion active, the first natively multimodal release in the GLM-5 line, with open weights on Hugging Face. The number that stings, via Artificial Analysis: it scores 57 on the intelligence index at roughly 9 cents a task, against GPT-5.6 Sol at 59 for 67 cents. That is a 7-to-10x cost gap for a point or two of capability.

Why this matters:

  • The open-weight crown is Chinese now, issue after issue: Kimi K3 (#115), Qwen3.8 (#118), now GLM-5.3. The frontier of free, self-hostable models is being set in China while the US labs keep their best work closed.

  • Cheap-and-good pulls inference off the metered clouds. At 9 cents a task against 67, enterprises drowning in coding-agent bills gain a reason to self-host, which chips at the per-token economics OpenAI and Anthropic are built on.

  • It is a chip story too: GLM-5.3 is trained and served entirely on Chinese silicon, so every workload that routes to it is one that never touches an NVIDIA GPU, export controls or not.

The Price of an NVIDIA AI System Just Jumped More Than 15%

The memory crunch just reached the one line item buyers can't route around.

Server builders have notified customers, per Bloomberg, that prices for systems built around NVIDIA's chips are rising more than 15%, with soaring memory the driver rather than the GPUs themselves. The shortage in HBM and DRAM has run through supply all year and is only now reaching the finished rack. The notices hit systems shipping early next year, on Grace Blackwell and the forthcoming Vera Rubin; they are private customer communications, not a public price list, and NVIDIA declined to comment. At current demand for Blackwell and Rubin, buyers have close to zero leverage, so the memory bill flows straight through to them.

Why this matters:

  • Pricing power has moved to memory, not NVIDIA. HBM and DRAM are set by three suppliers, and the shortage means whoever controls the stacks now sets the ceiling on how fast the whole industry can build.

  • The cost lands on the neoclouds, already carrying tens of billions against depreciating GPUs (CoreWeave's $35 billion, #119). A 15% jump on the system is 15% more debt to service against assets worth less each quarter.

  • It is a tailwind for the alt-silicon field. Every 15% on an NVIDIA system narrows the gap to Meta's MTIA and Groq, and shortages, not benchmarks, are usually what push buyers to try the alternative.

Alibaba Raises $10.2 Billion for AI, and the Stock Falls

China's biggest cloud tapped the market for AI capex and got marked down for it.

Alibaba raised HK$80 billion (about $10.2 billion) through a share placement, with all proceeds earmarked for full-stack AI, and the stock slid on the news. It is the same investor reflex that met the Western hyperscalers' capex guidance through Q2 (Alphabet at $45 billion a quarter, #116): approval of the ambition, unease at the bill. The money funds Alibaba Cloud's build-out against ByteDance and the US giants, and feeds both its Qwen models and the home-grown accelerators it is leaning on as export controls tighten.

Why this matters:

  • The capex-fatigue trade has reached China. Two years of “more spend, higher stock” has flipped to “show me the return” on both sides of the Pacific, and the build-out no longer gets a free pass.

  • Equity, not debt, is the China route: where Amazon sold $25 billion of bonds (#114), Alibaba dilutes, because a firm cut off from cheap dollar borrowing funds its build the only way it can.

  • “Full-stack” is code for self-sufficiency. Feeding Qwen and home-grown accelerators is how Alibaba hedges export controls, and every model and chip it builds in-house is one less thing Washington can switch off.

OpenAI Gets the Green Light for 3.2GW of Georgia Power

OpenAI just locked up an industrial region's worth of power before breaking ground.

OpenAI has secured approval for a 3.2GW power deal to supply a planned data centre in Effingham County, Georgia, with Georgia Power supplying the electricity and OpenAI funding the full cost. It is an enormous block of power, committed years ahead of the build. It fits OpenAI's roughly $750 billion infrastructure plan (#116): underwrite the supply directly because the interconnection queue cannot move fast enough on its own. Georgia Power serves the load and recovers the cost, OpenAI takes the power, and the Georgia Public Service Commission has signed off.

Why this matters:

  • Power, not chips, is the binding constraint now. The labs are buying electricity wholesale a decade out, before a single GPU is installed, because the megawatt is what the build-out is rationed by.

  • Funding the utility's build is how you jump the interconnection queue, the same move behind the recent behind-the-meter gas and nuclear deals, and it is becoming the price of entry for a frontier data centre.

  • The Southeast is the new frontier: Georgia is absorbing load that Virginia, Texas and the Northeast increasingly cannot, quietly redrawing the map of where American AI actually gets built.

Everything Else

p.s. The GPU Daily now goes out every morning. Opt in below if you want the physical layer of AI in your inbox daily.

Want The GPU Daily in your inbox?

Monday to Friday, a fast morning brief on the physical layer of AI. Opt-in only, pick below.

Login or Subscribe to participate

Reply

Avatar

or to participate