I see that you don’t know what an algorithm is. Because every search engine before Google used algorithms, just not Pagerank.
Tell me what you are seeing. This page you linked shows all top models for AI research at above 80% and the top ones at 90%.
On a single benchmark, GPQA, that was first released in 2022 and is now widely considered to be on its way out. This is extremely common in LLM benchmarking: every couple years, old methodologies are weeded out because most recent LLMs score 90% or higher. This could be for many reasons, but more than likely LLMs are trained on the specific corpus these benchmarks test against, which is akin to taking an exam knowing the answers beforehand.
The aggregates tell a very different story.
On the NYT article, you missed where it says that on the latest Gemini version, more than half of the summaries are ungrounded, meaning that users couldn’t possibly verify the accuracy of the response given the linked site. They also often provide additional information that’s not true.
I’m not even going to bother with the “I’m fine paying with my data” bit, because think that’s just a morally broken take.