I believe the full size DeepSeek-R1 require about 1200 GB of VRAM. But there are many configurations that require much less. Quantization, MoE and other hacks. I don't have much experience with MoE, however I find that quantization tend to decrease performance significantly. At least with models from Mistral.
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
replies: