• ricecake@sh.itjust.works
    link
    fedilink
    arrow-up
    0
    ·
    2 days ago

    I’m not sure that’s been extensively tested in courts. The document you referenced below appears to be as-yet not officially published, so I don’t believe it actually qualifies as an official position yet, but the bigger issue is that it’s untested in court.

    This thread is a response to an AI court case where the ruling was that training on copy written works is fair use.

    https://www.reuters.com/sustainability/boards-policy-regulation/us-judge-approves-15-billion-anthropic-copyright-settlement-with-authors-2025-09-25/

    Alsup ruled in June that Anthropic made fair use of the authors’ work to train Claude, but found that the company violated their rights by saving more than 7 million pirated books to a “central library” that would not necessarily be used for that purpose

    Regardless, you do make good points and I think we agree that the end state is “they shouldn’t be able to do that”. I have concerns that using existing standards that take copying too literally results in some unintended ambiguity, and situations where AI training is incidentally blocked, but so is stuff like “opening a news article on a computer”, which does the same things the copyright office report highlights as infringement.
    I think we’d be in a much more agreeable place if we just legally state that a commercial AI tools training isn’t fair use. That lets you have nuance like “search engine? It’s a statistical model, but not generative: allowed. AI agent? Statistical model that’s generating content as opposed to classification or ranking: not allowed”.