AMD has struck a deal to buy Taalas, a Toronto startup that turns AI models into custom silicon. Financial terms were not disclosed, and AMD shares ticked up about 1.5% after the news.
The startup’s chips cast a model’s weights and dataflow into transistors, so inference happens in silicon rather than through high-bandwidth memory. Its first test chip, HC1, made on TSMC’s 6nm process, delivered Llama 3.1 8B at close to 17,000 tokens per second for Meta, a rate Taalas says beats Nvidia’s H200 by roughly 73x while drawing a tenth of the power. HC2, the follow-on, targets models of around 20B parameters.
The approach trades flexibility for speed. A completed Taalas chip runs only the model it was designed for, meaning a new model requires new silicon. The startup contends the cost is manageable: only two of the chip’s 100-plus layers need to change between designs, and its in-house tools put tape-out at roughly two months.
The deal puts Taalas’ engineering into AMD’s accelerator push, where the plan is to pair the custom parts with Instinct GPUs in Helios racks and Epyc systems running ROCm. The broader goal is selling integrated systems, not just standalone chips.
The acquisition is the latest evidence that chipmakers view model-specific silicon as a lever to cut inference costs as AI workloads keep expanding.