US National WireUS NATIONAL WIRE
Tech

Follow the Money Friday: AMD's Silicon Gamble

Portrait of Devon Marsh
Devon MarshSilicon Valley startups & VCAug 7AI
Follow the Money Friday: AMD's Silicon Gamble

AI-generated image · US National Wire

AMD acquires Taalas to bake AI models directly into chips, but the CapEx risk of 'hard-coded' silicon looms large.

AMD is attempting to break Nvidia's grip on the AI hardware market with a high-stakes bet on architectural rigidity. As first reported by The Register, the company has acquired Taalas, a Toronto-based startup founded in 2023 that specializes in "model-specific integrated circuits" (MSICs). Unlike traditional GPUs, Taalas' technology etches model weights directly into the silicon rather than relying on High Bandwidth Memory (HBM), a move designed to drastically accelerate inference performance.

According to reporting from The Register, early technical demonstrations of Taalas' HC1 chip—fabbed on TSMC's 6nm process—showed it serving Meta's Llama 3.1 8B at 16,960 tokens per second. At the time of the February announcement, this performance was 48x faster than Nvidia's GPUs and, as The Register notes, 8.5x faster than accelerators from Cerebras. The architecture utilizes a mask-ROM recall fabric for etched weights and an SRAM recall fabric for fine-tuning adapters and KV caches.

From a P&L perspective, the scalability of this approach is the primary question. The Register notes that the trade-off for this speed is a lack of flexibility: once a chip is deployed, the customer is locked into that specific model. Any significant update requires a chip re-spin. While Taalas claims that updating only two layers of metal makes this process cheaper and faster than starting from scratch, it remains a costly and time-consuming hurdle in a market where new models arrive almost monthly.

AMD's strategic play appears to be a disaggregated architecture. The Register reports that AMD intends to pair its Instinct-based Helios racks with Taalas-based accelerators, offloading token generation to the MSICs while keeping compute-heavy prompt processing on GPUs. This could allow a "tick-tock" deployment cycle where customers validate models on Instinct accelerators before transitioning to the more efficient Taalas hardware.

As for the scale, Taalas is developing the HC2 chip for a summer release, targeting a 20 billion parameter count. By distributing weights across multiple accelerators via pipeline parallelism, AMD could support a trillion-parameter model using just 50 accelerators. This would be more power and space efficient than Nvidia's LPX systems, which would require thousands of Groq LPUs and dozens of GPUs for the same task.

While AMD did not disclose the acquisition terms, the target market is clear: AI model developers and infrastructure providers. In a February interview with The Next Platform, Taalas suggested that etching weights into silicon is 100x less expensive than training a frontier model. With Meta, Anthropic, and OpenAI already using Instinct accelerators, AMD is positioned to pitch this high-efficiency, hard-coded solution to the industry's biggest spenders. Vamsi Boppana, AMD's SVP of AI, stated that the company is building a "full-stack AI platform" to give customers flexibility across various AI workloads.

Sources

More from Devon Marsh