
Tenstorrent#73
About


Tenstorrent is an AI and RISC-V computing company that designs high-performance AI processors, chiplets and open-source software as an alternative to GPU-based AI hardware. It is led by veteran chip architect Jim Keller.
From X
New Tenstorrent cluster hot from the kitchen > 1TB of VRAM > 3TB DDR5 RAM > 32TB SSD Storage New product, will share more later P.S. Can you find the cat in the picture? https://t.co/JFYT0tsWd2
72 is actually a small number @tenstorrent is building big computers. 600 AI processors now, 1000-2000 soon. We program it as one computer. @satyanadella https://t.co/Jl708sIWqX
I’ve been thinking about the usual order of problem solving in computers 1) Compute 2) Memory 3) IO Scaling with MPI or similar is often an exercise for the reader @tenstorrent solved it in this order 1) Data placement and movement for Scaling 2) IO 3) Memory 4) Compute Compute is the easiest and best understood so making it last makes sense This was hard, actually, but puts the hard part first. Results are surprisingly good
Everyone can now run BoltzGen, the state-of-the-art system for drug design created by the genius @HannesStaerk , on @tenstorrent hardware with incredible performance, and the same accuracy as GPUs. While Boltz-2 can predict how strong a potential drug binds to a target, BoltzGen generates potential drugs (binders) from scratch given a target. Think: Boltz-2 is the evaluator, BoltzGen is the generator. BoltzGen is integrated into TT-Boltz now and runs on Wormhole and Blackhole at any scale: single card, QuietBox, Galaxy servers, Galaxy clusters, anything. The video below shows BoltzGen designing binders, fully parallelized across the 4 cards of Tenstorrent QuietBox. Tenstorrent is the perfect fit for the @boltz_bio models and infinitely scalable. I want to run those models at unprecedented scales. A big step towards a drug designed on Tenstorrent hardware.
Tenstorrent just dropped serious benchmarks on DeepSeek 671B, and they are worth watching. On decode, they hit 350 tokens per second. That is more than double Fireworks and Google Vertex at 144 tsu, and over 12x faster than Novita at 28 tsu. On prefill (100k sequence length), they clock 4.0 seconds — right in the top tier, behind only Google Vertex at 1.4 seconds while beating everyone else between 6.0 and 8.5 seconds. The cost story is even bigger. At high throughput, Tenstorrent delivers at $6 per million tokens while NVIDIA GPU setups jump to $30 and keep climbing. They compare Galaxy Blackhole with the GB300 NVL72 rack from NVIDIA, quoting SemiAnalysis benchmarks. The advantage comes from Tenstorrent’s purpose-built inference architecture. It is engineered from the ground up to keep cost per token low even as throughput scales, unlike general-purpose GPUs that were originally designed for training workloads. For DeepSeek 671B specifically, this translates into dramatically better efficiency on the metrics that matter most to AI companies: speed + real dollar cost at high volume. This is the structural edge they are betting the entire inference market on, and the real question is whether Tenstorrent could serve inference at scale, provided that they gain traction within the infrastructure industry.
Yay. I can finally share my thesis: Porting Boltz-2 to Tenstorrent Accelerators. - Thesis: https://t.co/vckiS5E8Qu - Presentation: https://t.co/PhLfJV8SG3 Huge thanks to all my friends at Tenstorrent and Boltz, to my supervisor Isaac, and to Prof. Gerndt. To the best of our knowledge, this work represents the first accurate implementation of a state-of-theart biomolecular structure and binding affinity prediction model on accelerators beyond GPUs and TPUs. Moreover, it achieves competitive performance per dollar compared to current GPU implementations, demonstrating the viability of alternative accelerators for computational biology.
News
Polls
Will Tenstorrent's RISC-V AI chips reach $1B in annual revenue by 2028?
Founders



