As of a few minutes ago old.reddit.com now seems to requre logging in to access it.

Shame as it was still the best place to browse for info about some niche interests.

Edit: Seems like they might be testing the change as it is hit or miss at the moment.

Edit: Adding this to My flilters in uBlock Origin seems to fix the issue for the time being:

||reddit.com^$header=referer:

you are viewing a single comment's thread
view the rest of the comments
[–] 67 points 4 days ago (5 children)

Fuck reddit.

If you need the old info download the archive. 99% of anything newer than 3 years is going to be bot slop. There's zero reason to go to reddit intentionally.

  • source
  • hideshow 10 child comments
  • [–] 22 points 4 days ago (2 children)

    Reddit archive? Whereabouts would I acquire that?

  • source
  • parent
  • hideshow 4 child comments
  • [–] 35 points 4 days ago (1 child)

    Steve is a jealous bitch and has made it harder to find archives of that content which he does not nor has ever owned. He's like Smaug but with brain damage and an incredibly tiny penis.

    Here's one, you'll have to hunt to find the big ones at this point.

    https://academictorrents.com/details/85a5bd50e4c365f8df70240ffd4ecc7dec59912b

  • source
  • parent
  • hideshow 2 child comments
  • [–] 14 points 4 days ago (1 child)

    I'd wager that older dumps are higher quality, before the mass redactions and deletions of the various exoduses. If you're looking for up to date information on current topics, it's time to look for (or build) other places.

  • source
  • parent
  • hideshow 2 child comments
  • [–] 22 points 4 days ago

    These are Pushshift torrents from 2005 to 2025: https://sciop.net/datasets/reddit. There are two data sets, one by month and the other by subreddit. Each is about 3.5 TB compressed and based on the compression ratio I've seen, about 30 TB uncompressed. This is without any sort of media, just text.

    Reddit sent AcademicTorrents a request to take them down so AI models couldn't use them as a way to avoid paying Reddit hundreds of millions of dollars for training data.

  • source
  • parent
  • [–] [S] 16 points 4 days ago (1 child)

    Agree.

    Most of it sure, but there are still some active subreddits with human involvement I wish would switch to Lemmy.

  • source
  • parent
  • hideshow 2 child comments
  • [–] 9 points 4 days ago

    Same. I have 1 or 2 subreddits I check for info/advice that I wish would just migrate over here. It just isn't worth opening the reddit app and dealing with the worst of humanity being pushed on me anymore.

  • source
  • parent
  • [–] 2 points 3 days ago (1 child)

    There's zero reason to go to reddit intentionally

    Know of a viable replacement for r/bapcsalescanada?

  • source
  • parent
  • hideshow 2 child comments
  • [–] 0 points 3 days ago (1 child)

    Run your own discussion board and literally tell people about it.

  • source
  • parent
  • hideshow 2 child comments
  • [–] 2 points 3 days ago

    I don't really have an interest in running a public board, and the reason I use(d) that subreddit was because I am also not interested in looking for the deals myself.

    However, from another comment, redlib looks like a decent solution for the time being.

  • source
  • parent
  • [–] 4 points 4 days ago (3 children)

    How feasible is it (legally, technically) to host a federated, read-only copy of Reddit data from before The Enshittening?

  • source
  • parent
  • hideshow 6 child comments
  • [–] 3 points 3 days ago (1 child)

    Why bother?

    Its all long since been injested by every gen AI model around, just ask one.

  • source
  • parent
  • hideshow 2 child comments
  • [–] 5 points 3 days ago* (1 child)

    The problem with AI is that while 80% of the time it will provide correct information, but 20% of time it will generate very convincingly sounded bullshit.

    This is the danger of it.. You test it few times and then you start trusting it. Kind of like that farmer that destroyed his whole crop.

  • source
  • parent
  • hideshow 2 child comments
  • [–] 2 points 3 days ago (1 child)

    I agree. However I don't think answers sourced from reddit are any more accurate.

  • source
  • parent
  • hideshow 2 child comments
  • [–] 3 points 3 days ago (1 child)

    Yeeeesss..., but typically it is easier to determine how reliable the person is from the test of the context.

    Since Gemini is fed by the reason directly or might be fun exercise to ask it about yourself and while in some questions it will answer correctly or just loves to invent new things.

    This works especially well if the person owning the account did not worry as much about privacy. So you hear many true things and also a lot of bullshit.

  • source
  • parent
  • hideshow 2 child comments
  • [–] 1 point 3 days ago

    typically it is easier to determine how reliable the person is from the test of the context.

    sure. that's an advantage of reading off reddit.

    an advantage of asking an LLM is that it can consider multiple reddit posts and stack overflow and whatever else.

    you're right that there's an eternal problem of determining whether information is accurate or just confidently incorrect, but I don't think that applies exclusively to chat bots.

  • source
  • parent