I mean, it’s not shit at everything; it can be quite useful in the right context (GitHub Copilot is a prime example). Still, it doesn’t surprise me that these first-party LLM benchmarks are full of smoke and mirrors.
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
replies: