But first, my personal problem with Ai doesn’t matter to this disussion, as I was talking about what people could have an issue with, not about my personal problems.
Okay. Then the problem that people, not you particularly, have with AI isn’t with the technology itself nor its use, but the methods by which the companies “grow and harvest it”, so to speak.
Which I don’t even think is true. People have just lost their minds over this. If a company unethically steals data and publishes a book with that data, we don’t say that there are legitimate concerns using books.
We would be having a problem with it if every book was regurgitated stolen data. That’s kind of the core of how all of these generative AI models work to the point that their creators have argued in court that they wouldn’t exist at all if they weren’t allowed to steal all the training data.
It can be trained on the exact same free data that we can be trained on. Just like us, it doesn’t have to bypass paywalls. I think they’re arguing that if I can look at the Wikipedia and tell you that an avalanche in Pakistan killed ten climbers, then AI should be allowed to as well.
If a company unethically steals data and publishes a book with that data, we don’t say that there are legitimate concerns using books.
Off course we say its a legitimate concern if someone unethically steals data and publishes a book with that data… The point of scraping and stealing data without respecting its license, training their software with it and using it commercially (or even non commercially) without giving even credit and not respecting its original license?
the problem that people … have with AI isn’t with the technology itself nor its use, but the methods by which the companies “grow and harvest it”, so to speak.
No, its not just the methods. I am talking about deep problems with the technology, the methods how companies go with it, and then its usage from company and end user. Ai trained with those data and used makes it impossible to respect the original license, for the user. So in this case, the company using the Ai AND the user of the model violate licenses. Maybe. Maybe not, if the data allows it. But we don’t know and CANNOT check and know for sure, that’s the problem. Just because you don’t understand the issue or don’t care does not make it less of a problem.
As an example with software licenses like GPL requires that any derivative works using this source code (such as Linux or whatever my personal program is licensed under GPL) has to be GPL licenses (Open Source) too. That is not debatable. That is how this license work. Now the Ai can be trained on this data, using the source code, and then? See where this goes? That is only one generalized example.
Okay. Then the problem that people, not you particularly, have with AI isn’t with the technology itself nor its use, but the methods by which the companies “grow and harvest it”, so to speak.
Which I don’t even think is true. People have just lost their minds over this. If a company unethically steals data and publishes a book with that data, we don’t say that there are legitimate concerns using books.
We would be having a problem with it if every book was regurgitated stolen data. That’s kind of the core of how all of these generative AI models work to the point that their creators have argued in court that they wouldn’t exist at all if they weren’t allowed to steal all the training data.
It can be trained on the exact same free data that we can be trained on. Just like us, it doesn’t have to bypass paywalls. I think they’re arguing that if I can look at the Wikipedia and tell you that an avalanche in Pakistan killed ten climbers, then AI should be allowed to as well.
Off course we say its a legitimate concern if someone unethically steals data and publishes a book with that data… The point of scraping and stealing data without respecting its license, training their software with it and using it commercially (or even non commercially) without giving even credit and not respecting its original license?
No, its not just the methods. I am talking about deep problems with the technology, the methods how companies go with it, and then its usage from company and end user. Ai trained with those data and used makes it impossible to respect the original license, for the user. So in this case, the company using the Ai AND the user of the model violate licenses. Maybe. Maybe not, if the data allows it. But we don’t know and CANNOT check and know for sure, that’s the problem. Just because you don’t understand the issue or don’t care does not make it less of a problem.
As an example with software licenses like GPL requires that any derivative works using this source code (such as Linux or whatever my personal program is licensed under GPL) has to be GPL licenses (Open Source) too. That is not debatable. That is how this license work. Now the Ai can be trained on this data, using the source code, and then? See where this goes? That is only one generalized example.
We have a legitimate concern with that particular book, not all books.