AMD is acquiring Toronto startup Taalas, which builds chips with AI model weights etched directly into the silicon instead of streamed from HBM. AMD plans to pair them with its Helios racks so GPUs handle prompt processing while Taalas parts generate tokens, with the deal expected to close in Q4 2026.

Read the full story: AMD buys Taalas, betting on AI models etched into silicon →

Transcript

AMD just bought a company whose whole idea is that you should stop storing AI models in memory. Taalas etches the weights straight into the chip. Physically, in the metal layers. There's no high bandwidth memory in the design at all, and that matters, because HBM is the expensive, power hungry, supply constrained part of every AI chip shipping right now. For generating tokens it mostly exists to feed weights to the compute. Delete the fetch, delete the bottleneck. They claim forty eight times an Nvidia GPU at a tenth of the power, though that's their number and nobody independent has run it. The catch is real though. That chip runs one model. Forever. A new model means a new trip to the fab. AMD's plan is to split the job: GPUs read your prompt, these things write the answer. It's a cheap bet with a very large payoff if model churn ever slows down.