A painter finishes a new work and shares it with the world. An illustrator posts a sketch that took hours to create. A photographer uploads a single image capturing a moment they wanted to preserve. Within seconds, these works can travel farther than any artist from previous generations could have imagined, crossing borders, reaching new audiences, and becoming part of the shared creative landscape that exists online. But today, artworks shared online may reach more than just human audiences. Alongside collectors, admirers, and fellow creators, a new kind of observer has emerged: artificial intelligence (AI) systems that scan, collect, and learn from the images that make up our digital world.
Hidden within the enormous datasets used to train these systems are billions of images gathered from across the internet, including countless copyrighted works created by artists who never knew their work had become part of an AI model’s education. The conversation surrounding AI and art extends beyond whether machines can create. It asks a more fundamental question – what did these systems learn from, and who decided they could learn it?
Copyright law has always existed in the space between protection and progress. It gives artists and authors control over how their creative works are reproduced, distributed, and adapted. At the same time, it recognizes that creativity does not happen in isolation. Artists have always learned from one another, studying techniques, responding to movements, borrowing ideas, and transforming what came before into something new. The challenge for copyright law is identifying where inspiration ends and unlawful copying begins.
Generative AI has complicated this balance in unprecedented ways. A human artist may encounter influences gradually over a lifetime, filtering them through personal experience, interpretation, and creative judgment. AI systems, by contrast, can analyze millions of images at a scale and speed impossible for an individual creator, identifying patterns across vast collections of human expression without experiencing the cultural, emotional, or personal contexts from which those works emerged. This difference in scale raises a new challenge for copyright law. It asks if a machine’s ability to learn from existing creative works should be treated as a natural extension of artistic influence or as a form of copying that requires permission.
Before an AI system can generate a new image, it must first be trained on existing ones. Like all forms of creative development, AI models are shaped by what came before, but the scale and speed of this process are unlike anything that has existed. Generative image models such as Stable Diffusion, Midjourney, and DALL·E are trained on enormous datasets of image-text pairs collected from publicly available web sources including artist’s websites, Instagram, and other social media platforms. These datasets are known to include copyrighted artworks alongside photographs, illustrations, and other visual media. During training, images are converted into numerical representations that allow the model to learn statistical relationships between visual features and language. Once training is complete, the model does not retain image files or function as a searchable archive of artworks. This technical reality has led to public confusion about why copyright infringement is alleged at all. That confusion arises from a misunderstanding of copyright law.
Copyright infringement does not require that a final product contain or reproduce a copyrighted work. Instead, infringement analysis focuses on whether unauthorized copying occurred at any stage of the process. Under U.S. law, the exclusive right of reproduction is triggered when a copyrighted work is copied in a fixed form for more than a transitory duration, even if the copy is later transformed or discarded. Training a generative AI model necessarily involves copying. Images must be downloaded, stored on servers, and processed in order to extract visual features. Although these images are later reduced to statistical parameters and removed from the deployed system, the act of reproduction has already occurred. From a legal perspective, the absence of stored images in the trained model does not negate the copying that took place during dataset creation and training. This distinction explains why infringement claims persist despite accurate descriptions of how AI models function. AI developers generally do not deny that copying occurs during training. Rather, they argue that such copying is legally permissible under the doctrine of fair use.
AI companies rely heavily on the fair use defense which is codified in 17 U.S.C. § 107. The doctrine recognizes that certain unauthorized uses of copyrighted works may be permitted when they serve broader purposes that outweigh the copyright owner’s exclusive rights. Courts evaluate fair use using four factors:
- Purpose and character of the use: Is the use transformative, educational, commercial, or noncommercial?
- Nature of the copyrighted work: Is the work primarily creative or factual? Published or unpublished?
- Amount and substantiality of the work used: How much of the original work was copied, and was the copied material central to the work?
- Effect on the potential market: Does the use harm the value or potential market for the original work?
AI companies and developers argue on the first factor, that training is transformative because it teaches a machine to recognize patterns rather than directly replicating the original works for human consumption. They contend that these models do not simply reproduce individual artworks. Instead, they analyze relationships among millions of images, extracting patterns and associations that allow the system to generate new content. From this perspective, copying during training is not the ultimate purpose of the use, but rather an intermediate step toward creating a new technological tool. AI developers point to Google Books as support for the argument that copying copyrighted material as part of a larger technological process may qualify as fair use. One frequently discussed example is Authors Guild v. Google, Inc., where Google scanned millions of books to create a searchable database. Although Google copied entire works, the Second Circuit held that the use was transformative because Google did not provide the public with replacements for the original books. The copies merely enabled a new function of allowing users to search, locate, and discover information contained within those works. Like Google’s search system, generative AI models do not simply reproduce the original images used during training, but analyze patterns within those works to create new outputs. However, the analogy is imperfect. Unlike Google Books, which did not produce outputs that could directly substitute for the original works, generative AI can produce images that closely resemble copyrighted art, compete in the same commercial markets, and even imitate individual artists’ styles.
This distinction makes the fourth fair use factor, market impact, particularly significant. If AI-generated works reduce demand for commissions, licensing opportunities, or other forms of creative labor, courts may view AI training differently from earlier examples of transformative copying. For now, courts have not provided a definitive answer on whether AI training qualifies as fair use. Future decisions will require balancing two important goals: encouraging technological innovation while protecting the economic and creative interests of the artists whose work contributes to these systems.
The fair use analysis becomes even more complicated when the discussion moves beyond individual artworks and into something less tangible: style. For many artists, style is what makes their work recognizable. It emerges gradually through years of experimentation. Through It becomes a visual language shaped by repeated acts of creation, refinement, and response to the world around them. Yet under current U.S. copyright law, style itself is not protected.
Copyright law draws a distinction between protected expression and unprotected ideas, concepts, methods, and techniques. The law protects the specific creative choices embodied in an individual artwork such as the particular arrangement of elements, the details of a composition, and the expression captured in a finished work. It does not protect broader artistic approaches, aesthetic characteristics, or a general visual style. This principle is rooted in Section 102(b) of the Copyright Act, which provides that copyright protection does not extend to “ideas, procedures, processes, systems, methods of operation, concepts, principles, or discoveries,” regardless of the form in which they are expressed.
This limitation reflects one of copyright law’s foundational principles. Creativity depends on influence, exchange, and the ability to build upon what came before. If artists could claim ownership over entire styles, artistic development would become impossible. A painter could not create work inspired by Impressionism. A filmmaker could not adopt the visual language of film noir. A musician could not build upon the traditions of jazz, rock, or classical composition. This approach embodies the constitutional purpose of copyright itself. The Copyright Clause grants authors exclusive rights to their works not as an end in itself, but as a means of promoting the progress of science and the arts. Copyright seeks to reward individual creativity while preserving the ability of creators to learn from, respond to, and transform existing works. The law therefore protects the expression embodied in individual works while leaving room for artistic influence and experimentation. Historically, this distinction has functioned because human imitation is limited. Artists may be inspired by another creator’s style, but that influence is filtered through individual judgment, experience, and interpretation.
A human artist may spend years studying another artist’s work before developing a related aesthetic. An AI system can analyze thousands of examples of an artist’s work and generate outputs that reproduce recognizable stylistic characteristics almost instantly. What was once a gradual process of influence becomes a scalable technological capability. The concern is not that AI systems can produce images that resemble existing artworks. The bigger issue is that they can absorb and reproduce the creative patterns that make an artist’s work distinctive, including the visual decisions, techniques, and aesthetic choices developed through years of human practice.
The limitations of copyright law and rapid adoption of generative AI has forced artists into a revolutionary defensive posture. Traditional copyright enforcement mechanisms including litigation, takedown notices, and DMCA claims, operate slowly, require substantial resources, and are poorly suited to addressing large-scale automated data scraping. By the time an artist becomes aware that their work has been incorporated into a training dataset, the harm has often already occurred. As a result, artists have increasingly turned to a combination of technical countermeasures, strategic self-help practices, and policy advocacy to mitigate risk in real time rather than relying solely on ex post legal remedies.
Glaze, developed by researchers at the University of Chicago, is a defensive tool designed to prevent AI systems from reliably learning an artist’s stylistic patterns. Rather than blocking access to an image, Glaze subtly alters the image at the pixel level in ways that are imperceptible to human viewers but disruptive to machine learning models. Technically, Glaze works by introducing adversarial perturbations– small, carefully calculated changes that exploit known weaknesses in neural networks. When a cloaked image is included in a training dataset, the AI model misinterprets stylistic cues such as brushstroke texture, edge transitions, or shading gradients. As a result, the model may associate those cues with incorrect or inconsistent representations. For example, an illustrator might upload a Glaze-processed digital painting to an online portfolio. To human viewers, the image appears unchanged. To an AI model, however, the stylistic signals are distorted. If the model later attempts to generate an image “in the style” of that artist, the output may appear muddled, inconsistent, or stylistically inaccurate. If it is an impressionist style painting, it could confuse the system by making it think it is an abstract painting. In this way, Glaze does not prevent copying outright but degrades the model’s ability to imitate a specific artistic identity.
Nightshade extends Glaze’s defensive logic but adopts a more aggressive posture. Instead of merely obscuring stylistic features, Nightshade actively corrupts the semantic associations that AI models learn during training. Nightshade works by embedding adversarial signals that cause a model to associate an image with the wrong concept. For instance, an image that clearly depicts a dog may be altered so that an AI model interprets it as a cat. Individually, such errors may appear minor. But when many Nightshade-altered images enter a training dataset, the cumulative effect can degrade model accuracy and reliability, and the actual subject matter is never copied. From the artist’s perspective, this represents a form of data deterrence. Because AI companies rely on massive volumes of scraped data, the presence of poisoned inputs introduces risk: degraded outputs, misclassifications, and reputational harm to the model itself. In theory, widespread adoption of such tools could increase the cost of indiscriminate scraping. However, researchers caution that this strategy is not a permanent solution. AI developers can attempt to identify adversarial patterns, filter poisoned data, or retrain models using curated datasets. As a result, Nightshade exemplifies an emerging arms race between creators seeking to protect their work and developers seeking to neutralize defensive measures.
Not all responses rely on altering the artwork itself. Some artists have also experimented with limiting the amount of information available for AI systems to analyze. For example, some creators intentionally upload lower-resolution images, crop works, or reduce visual detail when sharing their work online. Because generative AI systems rely on large amounts of high-quality visual data to identify patterns, lower-fidelity images may be less useful for training purposes. However, these approaches require artists to make difficult compromises. Reducing image quality may limit unauthorized AI use, but it can also affect an artist’s ability to showcase their work, attract collectors, secure commissions, or participate in licensing opportunities. In attempting to protect their creative output, artists may inadvertently restrict the very online visibility that helps sustain their careers.
Some platforms have also introduced tools that allow creators to communicate their preferences regarding AI training. For example, DeviantArt introduced an opt-out preference that allows artists to indicate that their work should not be used for AI dataset training. Similar “do not train” signals have also emerged across other digital platforms. However, these systems largely function as declarations of preference rather than guaranteed protections. Their effectiveness depends on whether AI developers and data collectors recognize and respect those signals, and they cannot necessarily prevent copies collected from other sources.
Other tools focus not on disrupting AI systems, but on increasing transparency. Have I Been Trained?, developed by Spawning, allows artists and creators to search whether their work appears in certain AI training datasets. The tool does not prevent the use of images or provide legal protection, but it addresses one of the important frustrations expressed by creators which is the difficulty of knowing whether their work has been included in datasets used to train AI models.
A different approach emphasizes establishing provenance rather than preventing use. The Coalition for Content Provenance and Authenticity (C2PA) and related Content Credentials systems aim to create a record of where digital content originated, how it was created, and whether it has been modified. Rather than blocking AI training, these systems seek to create greater transparency around the lifecycle of digital works and provide creators and audiences with more information about the origin of content.
Together, these technologies represent a change in the relationship between artists and digital platforms. For decades, the internet allowed creators to share their work more widely and connect with audiences around the world. Now, some artists are turning to technology defensively by creating barriers, monitoring how their work is used, and documenting ownership in response to systems that may incorporate their creations into future AI models. At the same time, tools like Glaze, Nightshade, and provenance systems reveal a broader limitation. Individual artists are being asked to address a large-scale technological transformation through individual solutions. A painter, illustrator, or photographer should not necessarily need expertise in cybersecurity, data tracking, or digital provenance simply to maintain control over how their work is used.
Some creators and policymakers have argued that AI developers should disclose the materials used to train their systems, allowing artists to understand whether their work has been included. Others have advocated for licensing frameworks that would permit AI companies to access copyrighted works while ensuring that creators receive recognition and compensation.
Licensing could represent one possible path forward, although its long-term viability remains uncertain. Rather than relying solely on disputes after AI systems have already been trained, licensing frameworks could allow creators and rights holders to negotiate how their works are used, while providing AI developers with clearer access to creative material. These models are already beginning to emerge. Adobe Firefly provides one example of an approach built around licensed training data. Adobe has stated that Firefly’s models were trained on licensed content, including Adobe Stock images and public-domain content. Adobe has also stated that Adobe Stock contributors whose works are used in Firefly’s training datasets may receive compensation through its contributor programs. This model attempts to create a clearer relationship between creators and AI developers by establishing permission and compensation mechanisms before creative works are incorporated into AI systems.
Licensing arrangements are also being explored at the level of major entertainment companies. In 2025, Disney and OpenAI announced a partnership that would allow OpenAI’s Sora platform to generate videos featuring more than 200 characters from Disney, Marvel, Pixar, and Star Wars. The agreement represented a significant example of a major rights holder choosing negotiated access over restriction or litigation, suggesting one possible path for how AI companies and copyright owners might collaborate. At the same time, the later uncertainty surrounding Sora’s future demonstrated the challenges of relying on licensing alone as a long-term solution. Business priorities, technological development, and market conditions may all affect whether these arrangements remain viable. These examples also reveal an important limitation of licensing as a complete solution. Companies like Disney possess centralized intellectual property portfolios and significant negotiating power, while individual artists often lack the same ability to negotiate terms for the use of their work. Even if licensing becomes more common, difficult questions remain. How should artists be compensated when their work represents only one contribution among millions of images used to train a model? Should compensation be based on individual works, dataset participation, or the commercial success of the resulting system? Who should be responsible for negotiating these agreements– individual creators, collective licensing organizations, platforms that host creative work, or the companies developing AI models?
These questions also demonstrate why the AI copyright discussion cannot be reduced to a simple conflict between artists and technology companies. Both art and technology have historically developed through experimentation, adaptation, and exchange. The challenge is creating a framework where innovation can continue without treating human creativity as an unlimited resource.
The future of generative AI will depend not only on what these systems can produce, but also on how they are built. If artificial intelligence is shaped by decades of human expression, then the artists behind that expression must remain visible and valued. Transparency, consent, compensation, and meaningful protections will determine whether generative AI becomes a tool that expands creative possibility or a system that benefits from human work without acknowledging its source.
Copyright law will continue to evolve, and courts will determine how existing principles apply to emerging technologies. However, the larger issue extends beyond legal doctrine. It asks what creators are owed in a world where machines can learn from vast amounts of human expression. The next generation of artists may create alongside artificial intelligence, just as previous generations created alongside photography, computers, and digital tools. For that partnership to succeed, the relationship between human creators and intelligent systems must be built not only on innovation, but also on respect.
Before a machine can create something new, it must first learn from something human.