Taalas HC1: 17,000 tokens/sec on Llama 3.1 8B vs Nvidia H200's 233 tokens/sec. 73x faster at one-tenth the power. Each chip runs ONE model, hardwired into the transistors.

you are viewing a single comment's thread
view the rest of the comments
[–] 4 points 5 months ago

Would be great, but feels unlikely, most of the gains they're making rely on the lack of versatility.

  • source
  • parent