General Compute, a US startup running an inference-focused AI cloud, announced on July 17, 2026 that it has secured a committed debt facility of up to $400 million from investment firm Upper90 Capital Management1. The facility starts with an initial commitment of $100 million and scales up to $400 million as customer demand grows1.
What stands out is not the amount but the collateral. The loan is backed not by NVIDIA GPUs but by SambaNova inference-specific chips, and it appears to be the first major AI loan with inference-specific silicon as the primary collateral2 — a sign that the structure of AI infrastructure financing is beginning to shift.
The Deal: Staged Debt for an “Inference Neocloud”
General Compute is a young startup founded by CEO Finn Puklowski and CTO Jason Goodison, which raised a $15 million seed round as recently as May 20262. It is building an “inference neocloud” — a cloud provider dedicated to inference — around SambaNova silicon, offering managed versions of open-source language models via API3. According to its press release, the company runs models from OpenAI, DeepSeek, and MiniMax1.
On performance, the company claims inference 16x faster than standard GPU clouds, 7x faster time-to-first-token, and 8.5x higher output throughput at 1,000 tokens per second1. The SambaNova chips the company operates (SN40 and SN50) run air-cooled at 20kW per rack with no liquid cooling required1, which makes them easier to deploy in ordinary existing data centers. The company reportedly plans 15 megawatts’ worth of air-cooled rack capacity by the fourth quarter3. These figures are the company’s own claims, however, and have not been verified by third parties.
The Context: GPU-Backed Lending Goes Beyond NVIDIA
Borrowing against chips has become a common playbook for AI infrastructure operators since GPU cloud giant CoreWeave proved the business model2. Upper90’s CEO Billy Libby, a former Goldman Sachs quantitative trader, is one of the pioneers of this market, having financed Crusoe Energy’s GPU purchases back in 20212. Now that same GPU-lending pioneer has accepted non-NVIDIA inference silicon as collateral. “We think open source models are going to be important, and we went and looked for a player last year that was in inference,” Libby said2.
The backdrop is the industry’s shift in center of gravity from training to inference. Training is a one-off capital outlay, while inference happens every time a user calls an AI service and generates recurring revenue — an asset class whose repayment stream lenders can underwrite. On the semiconductor side, capital keeps pouring into AI infrastructure, with TSMC posting record earnings on AI demand and SK Hynix completing a record IPO on AI memory demand. This deal extends that flow of capital to non-NVIDIA silicon.
Puklowski calls the loan “the first signal of capital organizing itself and the fragmenting of Nvidia’s monopolistic dominance”2. That is the founder’s own framing, and nothing at this scale threatens NVIDIA’s position today. Still, the fact that lenders have started treating non-NVIDIA chips as assets with collateral value marks a new phase in AI infrastructure finance.
Inference Price Competition Is Coming
For readers who use AI in their business, this trend could put downward pressure on inference costs over the medium term. As capital flows into inference-specific chips and non-GPU clouds come online, competition on API pricing and hosting costs becomes easier to sustain. As with Databricks’ fundraise at a $188 billion valuation, financing activity in the infrastructure layer eventually feeds through to the prices and choices available to users.
For companies that want to self-host open-weight models, a broader inference-dedicated infrastructure market also means more alternatives to owning GPU servers outright (see our local LLM hardware guide for the requirements of running models yourself). Whether inference-chip-backed lending works as a financial structure will be a test case that shapes how fast this segment can grow.
Sources
- General Compute Secures Up to $400 Million to Scale the World’s Fastest Inference Neocloud - General Compute press release (July 17, 2026)
- Why the first GPU financiers are turning to inference chips in a $400 million deal - TechCrunch
- Inference cloud operator General Compute raises $400M in debt financing - SiliconANGLE