cross-posted from: https://kbin.earth/m/[email protected]/t/3120649

As AI tech companies increasingly buy and destroy books to feed to their AI models, Anna’s Archive is calling for volunteers to help preserve them for the public record.

  • Tollana1234567@lemmy.today
    link
    fedilink
    English
    arrow-up
    0
    ·
    18 hours ago

    the AI ran out of Reddit slop material to scan, because reddit has been purging alot of content/accounts even suspected of spamming, so they go for physical books now.

        • mojofrododojo@lemmy.world
          link
          fedilink
          English
          arrow-up
          0
          ·
          17 hours ago

          that’s what I’m saying. I’ve never understood the hoopla of reddit selling it’s corpus, aside from being fucking gross (spot on for spez) it’s a shitton of garbage. there’s gold in there, but it’s absolutely buried in shite

          • Tollana1234567@lemmy.today
            link
            fedilink
            English
            arrow-up
            0
            ·
            16 hours ago

            im guessing reddit has been banning so aggressively lately, that AI scraping isnt getting as much “useful data” as much anymore (chatgtp/google dropped "references for reddit in thier slop summaries). so they are trying to force logins now to keep up with the user generated content so they can squeeze every last penny they can from reddit, i dont think ads were ever a big presence in reddit? its the data. people try to post content, like posts, subreddits, those can instantly get removed by redit filters. i think lost alot of traffic recently to, which was likely AI scraping.

          • 3abas@lemmy.world
            link
            fedilink
            English
            arrow-up
            0
            ·
            16 hours ago

            For an AI to troubleshoot problems with your computer (something it’s really good at today), it needs content like reddit’s. Books don’t and won’t detail weird software issues and their workarounds. One the current training set becomes irrelevant to whatever software people are running, this is a specific area where Reddit’s data is way more valuable than books.

            Of course Reddit locking down access to that information leads to fewer people contributing. Reddit is dead, it’s just struggling to go down.

            • mojofrododojo@lemmy.world
              link
              fedilink
              English
              arrow-up
              0
              ·
              16 hours ago

              there’s gold in there, but it’s absolutely buried in shite

              yep, there is a lot of nuanced depth, but measured against the overall sea of chuds lol… and yep, the moment they started putting up walls and controlling the community it was doomed. it’s gonna thrash for a while then, when the only ones left are advertisers and bots, it’ll shit itself and stop twitching.

  • WakeUpSmashPots@lemmy.today
    link
    fedilink
    English
    arrow-up
    0
    ·
    18 hours ago

    I hate this timeline… It’s things like this that make me hope we actually do live in a simulation, and the devs are about to release a major bug fix…

  • dil@lemmy.zip
    link
    fedilink
    English
    arrow-up
    0
    ·
    1 day ago

    Yall keep overreacting to this bs, what’s even a rare book with no other copies, every article I saw says the books were scheduled to be destroyed or taken to a dump, and those are the books they are swooping up. I have a feeling rare means like 1st edition or one with typos or some sht. These are books ppl already weren’t buying that they are grabving in mass. All knowledge and all books aren’t important, plenty of below average intelligence ppl dropping slop long before ai became a thing. Any book that you value obviously already has tons of physical copies and digital backups.

    This isn’t them looting some ancient libraries/museums taking some hidden/secret knowledge from humanity like yall are making up in your heads.

      • dil@lemmy.zip
        link
        fedilink
        English
        arrow-up
        0
        ·
        6 hours ago

        Okay hero, defend us all with your words on a social media site with maybe 40,000 ppl max that do nothing, keyboard activists to the rescue

    • KairuByte@lemmy.dbzer0.com
      link
      fedilink
      English
      arrow-up
      0
      ·
      1 day ago

      every article I saw says the books were scheduled to be destroyed or taken to a dump

      Do you have a source for this? Because I’ve seen literally nothing about this being the way books are being sourced. And no offense but this reads way too close to weasel words. For example: Every article I’ve read on the subject says my pants are sentient and enjoys my legs being in them. Also, I’ve read no articles on the subject so that isn’t a lie.

      • partofthevoice@lemmy.zip
        link
        fedilink
        English
        arrow-up
        0
        ·
        20 hours ago

        Fair point. Beside the point, is it technically correct to make a positive assertion about a negative thing — only because any positive assertion about that negative thing is inherently part of an empty set?

        Meaning like… our logic technically rules based on “no available contradictions, even if there is no evidence” as opposed to “at least one evidence is required?”

        • KairuByte@lemmy.dbzer0.com
          link
          fedilink
          English
          arrow-up
          0
          ·
          20 hours ago

          In certain situations I’d agree that the absence of contradiction is evidence enough, but those situations are pretty specific. For example, we don’t have absolute proof that orange juice doesn’t cause cancer, our “proof” comes from the lack of reliable evidence showing that it does, along with relevant research. But that doesn’t mean every claim can be treated the same way. But that doesn’t mean every claim can be treated the same way, especially a claim based only on “every article I saw,” without a source being provided.

    • Knock_Knock_Lemmy_In@lemmy.world
      link
      fedilink
      English
      arrow-up
      0
      ·
      1 day ago

      every article I saw says the books were scheduled to be destroyed or taken to a dump, and those are the books they are swooping up.

      Sounds like propaganda to me.

  • Smoogs@lemmy.world
    link
    fedilink
    English
    arrow-up
    0
    ·
    1 day ago

    destroying (A) physical copy. in order to copy the (1) copy of a book(which most already exist as pdf). kinda like how people do it when they scan the book.

    far cry from “destroying all books” which these rage bait titles.

    bigger tragedy is how humans seem shittier at communication as AI learns. probably a more accurate title.

    • KairuByte@lemmy.dbzer0.com
      link
      fedilink
      English
      arrow-up
      0
      ·
      1 day ago

      There are many, many non destructive scanning methods. My understanding is that most archival projects are using the non destructive options.

      The easy option is to just slice off the binding and run it through a scanner.

      The easy option is what is being talked about here, and those books are then fed into AI as training material. The digital copies aren’t “making knowledge permanent” as they are tucked away into a digital corner virtually no one has access to. They are as permanent as the source code for Fortnite.

      If a PDF already exists, as you claim, then why go to the trouble of obtaining, destroying and scanning the books again?

      • Smoogs@lemmy.world
        link
        fedilink
        English
        arrow-up
        0
        ·
        10 hours ago

        There’s also Kobo… Kindle… epubs…

        Not just pdfs.

        Heck you can get audio books. Most books are now sold with recording.

        This has been happening for decades.

        Are all you people complaining here just not book readers and just here for pure out rage?

        • KairuByte@lemmy.dbzer0.com
          link
          fedilink
          English
          arrow-up
          0
          ·
          9 hours ago

          Ok? That doesn’t answer my question, if anything it makes it more obvious how you are incorrect:

          If a [insert digital format name] already exists, as you claim, then why go to the trouble of obtaining, destroying, and scanning the books again?

          Your response reads like there are more options than just PDF, so it’s fine to destroy the books.

  • melsaskca@lemmy.ca
    link
    fedilink
    English
    arrow-up
    0
    ·
    1 day ago

    I guess the next step is picking specific people to memorize specific books and pass that knowledge down to an assistant/acolyte who will keep the memories alive. Rinse and repeat. Fahrenheit 451 anybody?

  • pseudo@jlai.lu
    link
    fedilink
    English
    arrow-up
    0
    ·
    2 days ago

    AI training is distroying books

    Every time I hear that claim it sounds weird to me. You can’t say that without an explanation.

    Sure, scanning a print copy can mean distroying it but a book as a work still exist. We are very far from having scanned every books which physical copy are numerous. If AI trains on every romance novel published in the 80’s, it is far from distroying books buy distroying physical copy.

    Now if AI company are searching for rare books to scan and distroy, it is much more worrying but finding rare editions or print copies of books is a trade and the normal citizen can’t jumps in to save the day.

    • Test_Tickles@lemmy.world
      link
      fedilink
      English
      arrow-up
      0
      ·
      1 day ago

      They are literally searching for rare books, paying premium costs for them, then cutting off their spines because it is faster to scan them that way.
      Not only do they not need to destroy the books to scan them, but Google already did this 2 decades ago, and they did it non-destructively.

      • dil@lemmy.zip
        link
        fedilink
        English
        arrow-up
        0
        ·
        1 day ago

        they arent paying extra, they are buying books planned for destruction anyways?

      • pseudo@jlai.lu
        link
        fedilink
        English
        arrow-up
        0
        ·
        1 day ago

        I just went on and read more article of the topics people shared in the comment. That’s sounds worrying indeed but my point is more that Annans Archive appeal is very abstract in a way that it présent something that doesn’t have to be a problem in a very sensationnalising manner instead of being clear about what is whappening and what we can do.

    • Toga77@lemmy.world
      link
      fedilink
      English
      arrow-up
      0
      ·
      1 day ago

      They’re destroying things that don’t have to be destroyed and compiling the information into privately owned megastructures they’ll then charge us to access while we own nothing and they can then alter the books.

      Nah I’m good. Everything these AI companies stand for is fucking evil and malicious.

      For every AI that researches a legit medical issue, they start literally burning fucking books, replacing jobs, screwing basically everyone, and making the world overall much worse.

      If you can’t see why these companies destroying any amount of books is bad, idk how to help you. Maybe look up a list of who has ever mass destroyed books. There’s not a lot of good people there.

      • Smoogs@lemmy.world
        link
        fedilink
        English
        arrow-up
        0
        ·
        1 day ago

        Humans have been scanning books for decades. All books created today exist electronically before getting physically printed.

        Type writers are a museum piece. not a utility anymore. Calm down.

        • KairuByte@lemmy.dbzer0.com
          link
          fedilink
          English
          arrow-up
          0
          ·
          1 day ago

          So you have the digital source of literally any book? Literally any single book you yourself didn’t write or publish.

          Because the source code for GTA6 exists, but you sure as shit don’t have it.

          • Smoogs@lemmy.world
            link
            fedilink
            English
            arrow-up
            0
            ·
            22 hours ago

            And how about you? Do you have the hard originals of every book? Why do you suddenly care only now that books are being copied?? They’ve been getting scanned for half a century now. You only care now to be outraged about it something for today

            • KairuByte@lemmy.dbzer0.com
              link
              fedilink
              English
              arrow-up
              0
              ·
              21 hours ago

              I don’t think that’s the argument you think it is.

              The outrage is over the books being destructively scanned, en mass, and held in private after the fact. Very, very, very different than the majority of scanning which is non destructive, and usually shared.

              • Smoogs@lemmy.world
                link
                fedilink
                English
                arrow-up
                0
                ·
                10 hours ago

                ‘En mass’

                No. That’s not what that means.

                They destroyed (a) copy they bought. That is not en mass.

                you didn’t buy it. It’s their copy.

                They didn’t destroy all copies. Go buy yourself one of you wanted it so bad. You can support the author too.

                • KairuByte@lemmy.dbzer0.com
                  link
                  fedilink
                  English
                  arrow-up
                  0
                  ·
                  9 hours ago

                  If they buy 100k books and destroy them all, even if every one is unique, that is en mass. I don’t know why I have to explain that to you, but there you go.

        • pseudo@jlai.lu
          link
          fedilink
          English
          arrow-up
          0
          ·
          1 day ago

          All books created today exist electronically before (a copy) gets physically printed of a book that remains existing.

          Not really. There is a lot of full manuscrit books made to this day. By scholar in region where electricity is not as much available. By student compile hand-written notes from oral teaching. By people composing small things for family and friends or just for themselves. By activist and locally involed people making just on or a few copy to be share in physicial Space around them.

          That’s a huge part of what exist outside the publishing industry.

          • Smoogs@lemmy.world
            link
            fedilink
            English
            arrow-up
            0
            ·
            1 day ago

            And we used copiers and voice recorders for that since before computers. That has been used since the 70s. Even earlier.

            scanning that and adding it to the common knowledge and making it accessible is something we’ve been doing since computers were a thing decades before today.

            Students now still use recordings that can be parsed into text.

            Suddenly it’s an issue now because someone needs so badly to be outraged by literally everything that can be distorted into a rage bait title like this one.

          • Smoogs@lemmy.world
            link
            fedilink
            English
            arrow-up
            0
            ·
            1 day ago

            “gullible”

            hurry, you better take up reading books to learn the meaning of these words you’re using before the AI ‘destroys’ the useless copy you never wanted nor thought about until now.

    • melsaskca@lemmy.ca
      link
      fedilink
      English
      arrow-up
      0
      ·
      1 day ago

      All those books seem to end up in goodwill eventually. What better place to gather them all for destruction.

      • pseudo@jlai.lu
        link
        fedilink
        English
        arrow-up
        0
        ·
        1 day ago

        Many should probably be distroy without training a LLM or we will get models that are aggressivly mysoginist and who verbally abuse female users, on top of every other problem they cause.

  • baconsunday@lemmy.zip
    link
    fedilink
    English
    arrow-up
    0
    ·
    2 days ago

    With modern phones having a document scan option with our phone cameras, is it possible we can scan and combine into a PDF file?

    I am going to test it myself, see how easy it is, and find a way to ensure anyone can follow with low cognitive load, but hopefully it’s as easy as I think to do.

    The sad part, is to do so would require a slow tedious task for books with hundreds of pages. It would be nice to see a large group tackle chunks at a time to the point we eventually have groups dedicated to genres and scouring to find them to preserve them.

  • TrackinDaKraken@lemmy.world
    link
    fedilink
    English
    arrow-up
    0
    ·
    2 days ago

    If you read the blog post, they don’t provide help or guidance for regular bookworms on how to scan. No information about what they expect.

    Also, flatbed scanning a whole book takes many hours. It’s one page at a time, line up each page, press the book flat and scan, for hundreds of pages.

    You’d likely need to buy, or have access to a book scanner. Which is fine, there are cheap models, but who knows if that would work for them? They don’t say. It doesn’t seem like they’re expecting help from you or me.

    • EastofEdson@lemmy.ca
      link
      fedilink
      English
      arrow-up
      0
      ·
      1 day ago

      I was wondering how feasible this is for an average Joe. My assumption was not very, thank you for confirming.

    • heartSagan5@lemmy.zip
      link
      fedilink
      English
      arrow-up
      0
      ·
      2 days ago

      Allegedly, it’s the rare, “undesirable” prints - the ones that just didn’t sell? Again, allegedly. But where’s a title disclosure?

    • matlag@sh.itjust.works
      link
      fedilink
      English
      arrow-up
      0
      ·
      2 days ago

      It’s worse than that. I’m sure they’re getting wet dreams about eradicating as many books as possible so that they can finally sell AI slop books at a premium price, claiming that will be “the last way to get something in [author] style!”.

      • BlindPenguin@lemmy.world
        link
        fedilink
        English
        arrow-up
        0
        ·
        2 days ago

        I’m rather thinking that hey want to make their slop machines the single source of truth. Much easier to get control over information, when all information is located in your own house.

    • Catoblepas@lemmy.blahaj.zone
      link
      fedilink
      English
      arrow-up
      0
      ·
      1 day ago

      If your local library doesn’t own a book scanner, any colleges near you may have one. You usually don’t have to be a student to walk into the library, and the book scanners I’ve used don’t require a login. YMMV on whether this is how it’s set up where you are.

      • vala@lemmy.dbzer0.com
        link
        fedilink
        English
        arrow-up
        0
        ·
        1 day ago

        I might try this. Im concerned that some of the books may still be under copyright. Not all of my rare books are that old. I collect indie poetry books and comics as well as old religious texts/pamphlets.

        Also concerned that some might be literally some of the only remaining copies and I really don’t want to risk damaging them.

    • tb_@lemmy.world
      link
      fedilink
      English
      arrow-up
      0
      ·
      2 days ago

      Those who want to conduct large-scale scans and upload many titles can reach out to the archive for support, as the archive says that it “can help pay for the scanning fees and other rewards.”

      • vala@lemmy.dbzer0.com
        link
        fedilink
        English
        arrow-up
        0
        ·
        2 days ago

        Hmm, it’s just a few books but I do think they’re pretty rare. Not something I’m willing to send away. I wonder if a library would be willing to help?

        • njordomir@lemmy.world
          link
          fedilink
          English
          arrow-up
          0
          ·
          2 days ago

          My local library has some large format scanners. I’ve considered using them to do my photo albums. I have the individual photos scanned and backed up, but the layout and margin notes are worth preserving also.