There is feedback that makes it back to their servers: the next prompt.
Though from my understanding, a lot of the agentic tools do a lot of prompt generation, so that would be a weakness of just logging everything and using subsequent prompts to evaluate earlier ones. They'd have no idea whether each prompt is coming from an actual user or another LLM.