Comment on 🤔 Interesting

<- View Parent
balsoft@lemmy.ml ⁨2⁊ ⁨weeks⁊ ago

If a company unlawfully … reproduced protected works in a way that violates copyright law, then it has broken the law and should be prosecuted.

Merely training an LLM on copyrighted material without holder’s permission (and then distributing the weights or selling access to inference on those weights) is a violation of copyright law if it were to be applied consistently (ignoring the fair use argument, which I’ll get back to). That is, if you applied any other computational process in this way, the result would be a derivative work. The reason it’s “different” this time is that the people violating the law are richer than those who wrote the law in the first place, not because of any legal argument.

As of right now, there are multiple lawsuits against major LLM developers. In some of those cases, the courts have ruled that training on publicly available data can qualify as fair use.

If a court ruled that it’s “fair use”, that actually lends more credence to the idea that LLM weights are a derivative work - “fair use” is a defense for copyright infringement that only makes sense in this case if the new work is an unauthorized derivative of the original.

Whether it’s actually fair use or not is another question (I can see the fair use argument for open-weight non-commercial models, not so much for commercial offerings).

BTW, I’m not even necessarily anti-AI (at least the open-weight, local models). I use a local model in my job almost daily, and also I think it’s mostly good that the entirety of FOSS corpus is available for download in a compressed and easily remixable form. I’m just pointing out the hypocrisy of the legal system which applies its already unjust copyright law (and most other laws) only against poor people.

source
Sort:hotnewtop