Gimlet Labs just closed a $300 million Series B led by Andreessen Horowitz, and the round values the AI infrastructure startup at $3 billion. Six months ago, in March 2026, the San Francisco company raised an $80 million Series A. Going from $80 million to $300 million in half a year is the kind of jump that signals investors think this problem is urgent, not just interesting. Gimlet's total funding now stands at $392 million.
What does Gimlet actually build? It calls itself the industry's first "multi-silicon inference cloud" for agentic AI. Rather than running an AI model's inference workload on one kind of chip from start to finish, Gimlet's software splits that workload into phases and sends each phase to whatever hardware handles it best: Nvidia, AMD or Intel GPUs, Arm-based processors, or specialty accelerators built by Cerebras and d-Matrix. The company is chip-agnostic on purpose. It isn't betting that any single vendor wins the AI hardware race. It's betting that no one does.
RelatedWonderful's $550M Series C Doubles Valuation to $5B
Arm and M12, Microsoft's venture fund, are new to this round. Sapphire Ventures, Menlo Ventures and Factory all returned from the Series A. That's a notable lineup: a chip designer and a hyperscaler-adjacent fund both putting money behind a company whose entire pitch is that customers shouldn't need to standardize on one type of silicon.
Gimlet says the traction backs up the valuation. Since the Series A, the company has added billions of dollars in contracted revenue, built a data-center pipeline it describes as gigawatt-scale, and is moving toward managing hundreds of megawatts of compute capacity.
What does a "multi-silicon inference cloud" actually mean?
Break an AI inference request into its actual steps and the idea gets concrete fast. There's a prefill phase, where the model processes everything in the prompt at once. That's compute-heavy and tends to reward raw GPU throughput. Then there's decode, where the model generates output tokens one at a time. That phase is bound more by memory bandwidth than raw compute, and it can often run cheaper on different hardware entirely. There's also scheduling and batching overhead, which doesn't need a GPU at all and can sit on a CPU.
Cerebras' wafer-scale chips carry enormous on-chip memory, which suits phases that need a huge amount of model state available without shuttling data on and off the chip. d-Matrix built its in-memory compute architecture specifically to cut the latency and power cost of decode. Gimlet's software watches a workload move through these phases and, in real time, decides where each piece runs based on cost and speed at that moment, instead of locking the entire job onto whatever chip happens to be installed in the data center.
That's what "disaggregating phases of inference across different chips" means in practice. It isn't about avoiding Nvidia. It's about not needing to commit to Nvidia, or AMD, or a specific accelerator, as a permanent architectural decision made once and lived with for years.
Why does chip-agnostic infrastructure matter right now?
Most of the money in AI infrastructure has gone toward two bets: buy as many Nvidia GPUs as possible, or build a rival chip and hope enough workloads eventually move over. Both bets assume the winning move is picking the right silicon. Gimlet's bet is that the winning move is the software layer that never has to pick, because it can shift a workload to whichever chip is cheapest or fastest for that specific job at that specific moment.
That distinction matters more now than it did two years ago because inference spend is starting to dwarf training spend as agentic AI products multiply. Training happens periodically and can be planned around whatever chips a company already has queued up. Running inference for millions of agent calls a day happens continuously, at a volume where a single percentage point of cost savings compounds fast. That's exactly the kind of spend where flexibility to route around an expensive or capacity-constrained chip matters, rather than being stuck with a GPU fleet committed to eighteen months earlier.
It's also a direct hedge against Nvidia's CUDA lock-in: the software ecosystem that makes switching away from Nvidia hardware expensive even when a competitor's chip is cheaper or faster for a given job. If Gimlet's orchestration layer can genuinely abstract that away, chip choice stops being a multi-year commitment and becomes a live optimization problem instead.
Who is Gimlet Labs betting against?
Not any one company. Gimlet is betting against the idea that infrastructure value concentrates in owning or exclusively buying one type of chip at all. That puts it in an odd position relative to its own investors: Arm makes chips, Microsoft (through M12) is both a huge Nvidia customer and a chip developer in its own right with Maia, and Cerebras and d-Matrix are themselves hardware vendors whose chips Gimlet routes work to. Everyone in this round benefits if the market decides orchestration, not any single chip, is where flexibility and margin live.
RelatedEtched hits $10.3B valuation on its Sohu inference chip bet
| Gimlet Labs (multi-silicon) | Single-vendor GPU cloud (e.g. CoreWeave-style) | Custom silicon bet (e.g. a hyperscaler building its own chip) | |
|---|---|---|---|
| Chip strategy | Routes each inference phase to whichever chip fits: Nvidia, AMD, Intel, Arm, Cerebras, d-Matrix | Standardizes on one vendor's GPUs for the whole workload | Builds proprietary silicon to cut reliance on outside vendors |
| Lock-in risk | Designed to avoid it by staying chip-agnostic | High: tied to that vendor's pricing, supply and software stack | Trades one lock-in for another, now dependent on your own chip roadmap |
| Cost flexibility | Can shift workload phases to whatever is cheapest at that moment | Fixed to one vendor's pricing and availability | Fixed to internal chip supply, hard to flex against outside market pricing |
| Bet | The orchestration software becomes the valuable layer, not any single chip | The vendor's hardware and ecosystem keep winning | Owning silicon end to end beats buying it from anyone |
What it means for the market
The signal here is less about Gimlet specifically and more about where infrastructure money thinks the next bottleneck sits. A16z, Arm and Microsoft's M12 aren't betting that one chip architecture wins outright. They're betting that the orchestration layer above the chips becomes the valuable, durable piece of the stack, regardless of which vendor has the best chip in any given quarter. For Arm, that's a way to make sure its architecture shows up in inference workloads without owning the whole stack itself. For Microsoft's venture arm, it's a hedge that doesn't require Microsoft's own infrastructure bets, its Maia chips, its enormous Nvidia purchases, to pick one winner internally.
It's worth tracking who else raises money making a similar argument in the coming months. GenZTech keeps a running list of rounds like this one on the Funding Tracker, and this raise slots into the broader picture on the biggest AI funding rounds page.
Whether the billions in contracted revenue Gimlet cites hold up once agentic AI adoption moves past pilot projects and into steady, large-scale usage. Whether Nvidia responds by making it cheaper to leave CUDA rather than deepening the lock-in further. And whether other multi-silicon orchestration startups show up to compete directly with Gimlet, rather than leaving the whole category to one company at a $3 billion valuation.
Our take
Whether chip-agnostic orchestration turns into a durable moat or gets absorbed by the hyperscalers is the real question, and it isn't obvious yet which way it goes. The case for durability: neutral orchestration works best for customers who explicitly don't want to be locked into one cloud's default silicon stack, the same reason Databricks and Snowflake became valuable sitting above AWS, Azure and Google Cloud instead of being folded into any of them. If Gimlet keeps adding new chip integrations faster than any single hyperscaler wants to build a matching internal tool, that neutrality becomes the product.
The case against it is just as straightforward. Every hyperscaler already runs mixed fleets internally and already has a direct financial incentive to build this kind of phase-aware scheduler for its own infrastructure, since it saves them money rather than paying a third party for it. If AWS, Google or Microsoft ship a good-enough internal version of what Gimlet does, a chunk of the market that might have paid Gimlet simply won't need to. The gigawatt-scale pipeline and the contracted revenue are the strongest evidence in Gimlet's favor right now, real customers choosing to pay for this instead of waiting on a hyperscaler to build it in-house. Whether that holds once the hyperscalers' own tools catch up is the thing worth watching closest.
- FundingPulse 2.0, "Gimlet Labs Raises $300 Million Series B Led By Andreessen Horowitz At $3 Billion Valuation" : funding figures and investor list
- OfficialGlobeNewswire (via Manila Times), "Now Valued at $3 Billion, Gimlet Labs Raises $300 Million in Series B" : company announcement and traction figures
- AnalysisQuasa, "Gimlet Raises $300M for Multi-Silicon AI Inference" : context on orchestration versus single-chip strategies
Original analysis by GenZTech Team. Source: Pulse 2.0.
