Sharing is caring!

The rapid advancement of artificial intelligence (AI) has spurred an array of innovations, including in the realm of generative AI. This technology, exemplified by models like OpenAI’s ChatGPT, Google’s proprietary AI, and the lesser-known Claude, has shown the capability to create content spanning text, images, and even music.

However, this capability has brought about a host of legal challenges concerning copyright infringement. The crux of the issue lies in the AI’s ability to reproduce existing content, either verbatim or with substantial similarity in meaning, and the usage of copyrighted material for training these AI models.

In the following points I want to go through the existing or upcoming AI legal case prototypes:

I. Reproduction of Exact Content Parts

Generative AI can reproduce the exact content from books, or other copyrighted materials. Hence the core nature of generative AI would prevent this behavior’s. The problem lies when there are a lot of references to an exact quote in a training dataset. If something is popular enough it may repeat in the training dataset, which may cause the appearance of the exact quote also in the answers.

This reproduction can happen either by directly outputting verbatim excerpts from these materials or by generating highly similar or derivative works. This raises a significant concern for copyright holders, who fear a loss of control over their intellectual property.

While I’m almost 100% sure ChatGPT couldn’t reproduce a whole book, not even 10 pages from a famous written text, the reproduction concern is real. Another field where this question emerged is the world of coding. There was a case against Microsoft, Github and OpenAI where the problem lied, if you give a similar coding problem to GitHub Copi­lot the result could be almost the same every time. Of course when you code Javascript programmers are looking for solution and if anything works properly it could be applied to another website. But in this case we are speaking about identical or almost identical code lines. And nothing is more frustrating when you as a programmer see your own copyrighted code coming out from Copilot as it happened with Tim Davis.

II. Style or Essence Reproduction

What if I’m Picasso and my style could be copied by a generative AI tool? This is exactly the situation today. Stable Diffusion, Midjourney are capable to reproduce certain style easily. Of course the new image won’t be a Picasso, but will have the exact

Artists are already started several class action lawsuits on this ground with less success. It seems that similarity and style is hardly definable part of an art.

This recent case in the US highlighted the complex nature of this issue. A judge ruled against a group of artists in a copyright suit concerning AI-generated art. The artists argued that the AI had infringed upon their copyrights by creating derivative works without permission. However, the judge’s decision indicated that the legal framework surrounding AI-generated content still has many grey areas that require further clarification.

III. Traffic Cannibalization

Another realm of concern is the ability of these AI to reproduce the meaning of original content, especially from news outlets. This ability can lead to traffic cannibalization for these outlets, as platforms like Google could potentially display AI-generated summaries or rephrasing, diverting traffic away from the original sources. This threatens the revenue streams of news publishers and other content creators.

In case of Google a long time existing topic is with info snippets. Google extracts information from website in order to keep the visitor on Google’s platform, which will result lower ad income at the destination site, where the info comes from.

With the appearance of Claude.ai and ChatGPT with web access, the situation will get worse.

IV. Use of Copyrighted Data for Training

One famous early case was when Getty Images claimed that their stock photo collection were used as training dataset for Stable Diffusion.

Another case was started by several famous authors including George R.R. Martin and Jodi Picoult which books were involved in an AI training dataset called “Books3”. This case is against OpenAI even when the “Books3” dataset was available free on the web.

The training of generative AI models often requires vast datasets, which sometimes include copyrighted material. The inclusion of such material in the training set without proper licensing or permission is a major point of contention, as it may contribute to the AI’s ability to generate similar or derivative works. But here comes the billion dollar question: Does training generative AI with copyrighted data fall under the purview of “fair use”?

V. Impersonation and Deep Fakes

I personally feel this is one of the dark side of generative AI. The emergence of sophisticated generative AI has given rise to a new form of impersonation: deep fakes. These are hyper-realistic digital manipulations that can depict individuals saying or doing things they never actually did. This technology has significant implications for personal privacy, consent, and the integrity of information.

One notable case, reported by Variety highlighting the potential misuse of this technology involved Scarlett Johansson, a prominent Hollywood actress. Johansson took legal action against an AI app, Lisa AI, which used her image and name in an online advertisement without her consent. The advertisement, which was 22 seconds long, featured Johansson’s likeness to promote the app’s services in creating digital avatars reminiscent of the 1990s yearbook photos.

The debate around deep fakes and impersonation is complex, entwining questions of copyright, consent, and the authenticity of digital content. It pushes us to consider not just the legality but also the morality of using someone’s likeness without explicit permission, particularly when it involves public figures who rely on their image as part of their profession. As AI continues to evolve, so too must our approaches to governance and the protection of individual rights in the digital realm.

Potential Solutions

Various solutions have been proposed to mitigate these issues. One such proposal is the establishment of revenue-sharing agreements or payment structures for citations. These arrangements could provide a means for compensating original content creators when their work is used or reproduced by AI. This approach could foster a collaborative environment between AI developers and content creators, ensuring fair compensation and adherence to copyright laws.

Final Thoughts on Legal Cases against Generative AI

The evolving legal landscape surrounding generative AI and copyright infringement is a testament to the profound impact AI has on our society. As lawmakers, tech developers, and content creators grapple with these issues, the legal battles underscore the urgent need for a well-defined regulatory framework to ensure a harmonious coexistence between AI innovation and intellectual property rights.

Sharing is caring!