Since implementation of the --fit parameter and its relatives, and --fit on becoming the default, llama.cpp intelligently decides what to offload. For me, it made --n-cpu-moe obsolete.
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
replies: