It hasn't been quantized, then. I've run 70B models on my consumer graphics card at a reasonably good tokens-per-second rate.
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
It hasn't been quantized, then. I've run 70B models on my consumer graphics card at a reasonably good tokens-per-second rate.