you are viewing a single comment's thread
view the rest of the comments
[–] 3 points 20 hours ago (1 child)

I run it using LM Studio, which defaults to Q4 quantization, I think. I was able to put about 10 layers on the GPU with 64k token context. That put me at about 9.1 GB VRAM usage, leaving some room for Video playback xD

  • source
  • parent
  • hideshow 2 child comments