AI training only makes a copy incidentally to what they’re doing and then it’s deleted. It’s the same standard that makes viewing a photo on an artists website legal.
That’s not what the U.S. Copyright office says about training. They hold that it does implicate the copyright of reproduction. Meaning: If you train on a protected work without a license you are violating copyright, and if that’s not a fair use then you are breaking the law.
Training ~ viewing might be an analogy used by “AI” brands, but it is not legal reality.
I’m not sure that’s been extensively tested in courts. The document you referenced below appears to be as-yet not officially published, so I don’t believe it actually qualifies as an official position yet, but the bigger issue is that it’s untested in court.
This thread is a response to an AI court case where the ruling was that training on copy written works is fair use.
Alsup ruled in June that Anthropic made fair use of the authors’ work to train Claude, but found that the company violated their rights by saving more than 7 million pirated books to a “central library” that would not necessarily be used for that purpose
Regardless, you do make good points and I think we agree that the end state is “they shouldn’t be able to do that”. I have concerns that using existing standards that take copying too literally results in some unintended ambiguity, and situations where AI training is incidentally blocked, but so is stuff like “opening a news article on a computer”, which does the same things the copyright office report highlights as infringement.
I think we’d be in a much more agreeable place if we just legally state that a commercial AI tools training isn’t fair use. That lets you have nuance like “search engine? It’s a statistical model, but not generative: allowed. AI agent? Statistical model that’s generating content as opposed to classification or ranking: not allowed”.
Its a slopper who wants to project these spreadsheets as ‘conscious’ when what’s happening is they’re essentially being transcoded into statistical models.
Yup. Give a Markov chain multi-billion parameters and you can get some surprisingly cogent results.
I will freely admit that current LLM architectures include several innovations that make them not actually Markov chains, but it’s still statistics and linear algebra. I don’t know what thought is, but I’m quite unconvinced that LLMs (or any current generative AI architecture) is doing it.
I don’t know what thought is, but I’m quite unconvinced that LLMs (or any current generative AI architecture) is doing it.
Okay, why not? I also don’t know what thought is, so I don’t think it’s possible to say if an LLM is or is not doing it. And giving wrong or incoherent answers doesn’t invalidate it as thought or your local stoner buddy would be considered brain dead.
I’ve not seen evidence of it in any of my interactions with generative AI, which have pretty universally been bad. I feel like it has something to do with autonomous spontaneity. I recognize it in animals I can’t communicate well with, but I found it lacking in the LLM that I tried to play a TTRPG with. It would be easier for me to be convinced, if I really had a better understanding of what thought is. It’s hard for me to be convinced because while I understand LLMs and diffusion networks better than most people*, I don’t think I understand thought so I recognize the gap.
Also I’m not sure I agree with your final assertion, the stoner buddy is plenty wrong, but there is a coherency there. When coherency disappears entirely from human thought that’s usually a seizure or stroke. Even as confusing as they are dreams and acid trips often have a coherency while you are in them, if not one that’s easily described when recalling the experience.
*: My formal AI training ended before big data met ML, so it’s woefully out of date. I am quite the computer geek tho, it’s just my passion tends toward languages, type systems, and proof assistants. So, better than most, but not an expert by any means.
Nah. Shut the fuck up slopper. If you’re going to insult us and then ask chatgpt to win the argument, which you always do and it always misses the point, I’m not going to answer your question.
The fact is these systems can do what a lot of humans do. That’s not because the matrix multiplication is identical to thinking, but because most of these humans have never thought in their lives, do not have interiority, and are not people in any way that matters. Prove you’re conscious if you want me to address you as such, fucking slopper.
I could’ve sworn a court case decided otherwise. Literally EVERY AI model in existence right now is commiting copyright theft on a massive scale if that’s the interpretation the courts took. Which is why I have a hard time buying it. I fear it’s reached the idea of normalcy in people’s minds and we’ll never see it illegal.
There’s been a couple court cases (that I know of / at least), and one judge was accepting on the argument that model training was a “fair use” while the other was not. I think both of those rulings came down prior to the publication of the U.S. Copyright Office guidelines.
Also, I’m not 100% sure that the U.S. Copyright Office is an authority here. The DOJ and/or Federal judiciary would have the authority to interpret the copyright laws: The DOJ to decide to prosecute, and the judiciary to make binding rulings and/or advise juries. I’m sure both the DOJ and the judiciary will give a lot of weight to the guidelines, but the guidelines aren’t actually the law.
In any case, you can read the guidelines and make your own decisions: https://www.copyright.gov/ai/ Part 3 is about training, and I think the damning bits are III, B and D. Part 2 is about outputs, and I think the damning bits are II, B and D.2. (My summaries: 1. Training infringes 2. Outputs that are substantially similar infringe 3. models get no copyright 4. prompts are NOT ‘human creative effort’ and thus are insufficient to establish copyright 5. human creative effort still gets copyright protections, even when generative AI is used as a tool in the creative process.)
It is likely that commercial generative AI is in violation of a lot of copyrights, yes. Research projects are fair use, but only as long as they stay research projects.
That’s not what the U.S. Copyright office says about training. They hold that it does implicate the copyright of reproduction. Meaning: If you train on a protected work without a license you are violating copyright, and if that’s not a fair use then you are breaking the law.
Training ~ viewing might be an analogy used by “AI” brands, but it is not legal reality.
I’m not sure that’s been extensively tested in courts. The document you referenced below appears to be as-yet not officially published, so I don’t believe it actually qualifies as an official position yet, but the bigger issue is that it’s untested in court.
This thread is a response to an AI court case where the ruling was that training on copy written works is fair use.
https://www.reuters.com/sustainability/boards-policy-regulation/us-judge-approves-15-billion-anthropic-copyright-settlement-with-authors-2025-09-25/
Regardless, you do make good points and I think we agree that the end state is “they shouldn’t be able to do that”. I have concerns that using existing standards that take copying too literally results in some unintended ambiguity, and situations where AI training is incidentally blocked, but so is stuff like “opening a news article on a computer”, which does the same things the copyright office report highlights as infringement.
I think we’d be in a much more agreeable place if we just legally state that a commercial AI tools training isn’t fair use. That lets you have nuance like “search engine? It’s a statistical model, but not generative: allowed. AI agent? Statistical model that’s generating content as opposed to classification or ranking: not allowed”.
Its a slopper who wants to project these spreadsheets as ‘conscious’ when what’s happening is they’re essentially being transcoded into statistical models.
Yup. Give a Markov chain multi-billion parameters and you can get some surprisingly cogent results.
I will freely admit that current LLM architectures include several innovations that make them not actually Markov chains, but it’s still statistics and linear algebra. I don’t know what thought is, but I’m quite unconvinced that LLMs (or any current generative AI architecture) is doing it.
Okay, why not? I also don’t know what thought is, so I don’t think it’s possible to say if an LLM is or is not doing it. And giving wrong or incoherent answers doesn’t invalidate it as thought or your local stoner buddy would be considered brain dead.
I’ve not seen evidence of it in any of my interactions with generative AI, which have pretty universally been bad. I feel like it has something to do with autonomous spontaneity. I recognize it in animals I can’t communicate well with, but I found it lacking in the LLM that I tried to play a TTRPG with. It would be easier for me to be convinced, if I really had a better understanding of what thought is. It’s hard for me to be convinced because while I understand LLMs and diffusion networks better than most people*, I don’t think I understand thought so I recognize the gap.
Also I’m not sure I agree with your final assertion, the stoner buddy is plenty wrong, but there is a coherency there. When coherency disappears entirely from human thought that’s usually a seizure or stroke. Even as confusing as they are dreams and acid trips often have a coherency while you are in them, if not one that’s easily described when recalling the experience.
*: My formal AI training ended before big data met ML, so it’s woefully out of date. I am quite the computer geek tho, it’s just my passion tends toward languages, type systems, and proof assistants. So, better than most, but not an expert by any means.
Nah. Shut the fuck up slopper. If you’re going to insult us and then ask chatgpt to win the argument, which you always do and it always misses the point, I’m not going to answer your question.
The fact is these systems can do what a lot of humans do. That’s not because the matrix multiplication is identical to thinking, but because most of these humans have never thought in their lives, do not have interiority, and are not people in any way that matters. Prove you’re conscious if you want me to address you as such, fucking slopper.
I could’ve sworn a court case decided otherwise. Literally EVERY AI model in existence right now is commiting copyright theft on a massive scale if that’s the interpretation the courts took. Which is why I have a hard time buying it. I fear it’s reached the idea of normalcy in people’s minds and we’ll never see it illegal.
There’s been a couple court cases (that I know of / at least), and one judge was accepting on the argument that model training was a “fair use” while the other was not. I think both of those rulings came down prior to the publication of the U.S. Copyright Office guidelines.
Also, I’m not 100% sure that the U.S. Copyright Office is an authority here. The DOJ and/or Federal judiciary would have the authority to interpret the copyright laws: The DOJ to decide to prosecute, and the judiciary to make binding rulings and/or advise juries. I’m sure both the DOJ and the judiciary will give a lot of weight to the guidelines, but the guidelines aren’t actually the law.
In any case, you can read the guidelines and make your own decisions: https://www.copyright.gov/ai/ Part 3 is about training, and I think the damning bits are III, B and D. Part 2 is about outputs, and I think the damning bits are II, B and D.2. (My summaries: 1. Training infringes 2. Outputs that are substantially similar infringe 3. models get no copyright 4. prompts are NOT ‘human creative effort’ and thus are insufficient to establish copyright 5. human creative effort still gets copyright protections, even when generative AI is used as a tool in the creative process.)
It is likely that commercial generative AI is in violation of a lot of copyrights, yes. Research projects are fair use, but only as long as they stay research projects.