AMD is buying Canadian startup Taalas, which hard-codes model weights directly into inference chips. That makes them extremely fast but locks each chip to a single model. A demo chip hit over 16,000 tokens per second per user running Llama 3.1-8B. Google is reportedly working on a similar approach for Gemini.<br /> The article AMD acquires Taalas, a startup that bakes AI models directly into silicon appeared first on The Decoder. [...]
AMD has reached a definitive agreement to acquire Taalas, a Toronto-based startup whose chips are custom-built around individual AI models, the company announced on August 6, 2026. The deal folds spec [...]
Google is developing "Frozen v2," a server chip that bakes the Gemini architecture directly into hardware. According to internal sources, it could be 6 to 10 times more efficient than curren [...]
Google has integrated "Computer Use" directly into Gemini 3.5 Flash, letting the model operate computers, browsers, and mobile devices on its own. On the OSWorld benchmark, it scores 78.4, p [...]
AMD's decision to start off with mid-range RDNA 4 GPUs now seems prescient. NVIDIA's high-end RTX 5090 and 5080 are already selling well beyond their absurdly high prices, if you can find an [...]
You might know the story by now: Framework makes repairable, modular laptops where you can sub in new components for old or broken ones. It’s been two years since the company debuted an AMD mainboar [...]