Whilst true about anyone can scrape data off Reddit, I think it's more of a pain since before the API updates the rate limit was 2 API calls per second. You also have to find or create a scraper. With Lemmy, you follow the instructions (copy and paste) on join-lemmy.org to create your instance and you're done. Both methods you have to configure it to subscribe to communities, so they're about the same.
In the EU at least there is a right to be forgotten, so yeah, Reddit and other platforms are forced to delete the data on request. I'm not sure how the same can be applied to a distributed network like Lemmy.
There were publicly available archives of Reddit. The last time I checked, you couldn't find the latest submissions and comments. Maybe things have changed, maybe newer alternatives have appeared.