They invented a hybrid attention design that drastically reduces the amount of memory needed for the KV cache at inference time. Like, dividing it by 10. And memory is a large part of the cost of inference.
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
replies: