如您希望下载本文的PDF版本,请点击文末“阅读原文”获取。
Whether training AI models with copyrighted materials constitutes copyright infringement is a heavily debated and litigated topic in China and around the world. In this article, we examine the matter with a step-by-step breakdown of the technical process for training AI models and reveal that copyrighted works may be stored only briefly in the memory of computing devices. Additionally, we discuss how AI model training temporarily uses stored copyrighted works for "understanding" and "extracting" concepts and ideas, rather than retaining particular expressions for "independent economic value," and what this means under copyright laws.
The article primarily focuses on Chinese law, while also briefly mentioning U.S. and EU laws.
01
Large Model Training Uses Copyrighted Materials
In June 2017, Google's machine translation team published a groundbreaking paper, "Attention Is All You Need," which completely abandoned recurrent neural network (RNN) and convolutional neural network (CNN) structures in favor of an attention-based Transformer architecture for machine translation tasks.[1] This ignited a new wave of AI development, leading numerous developers to create powerful large models (Large Models) based on the Transformer architecture, such as OpenAI's GPT Family, Google's PaLM Family, Meta's Llama Family, and xAI's Grok.
Compared to earlier deep neural networks (DNN), Large Models possess more neurons, deeper hidden layers, and more complex neural structures. To endow these Large Models with comprehensive intelligence, developers train them using significantly larger datasets. As a result, these datasets inevitably include numerous copyrighted works, such as novels, paintings, videos, and more.
This raises copyright concerns as there is a view that Large Models "remember" their training data. In other words, these models may reproduce and save copyrighted works, allowing them to generate expressions identical or similar to the training data during task execution.
02
Large Model Training May Only Briefly Retain Training Data
While legal professionals may not fully understand the technical processes behind how data is used in model training, we surveyed various technical materials in the AI field for a better understanding. According to these materials, the large model training process can be divided into three stages, during which training data is only briefly retained within the model—typically for far less than one second.[2]
Below are the three stages for large model training:
(1)
First, most large model training uses a deep "End-to-End Learning" learning approach. A unit of training data is imported into the large model and paired with another unit of training data to form an "input-output" pair for End-to-End Learning. Both are temporarily stored in the device's random-access memory (RAM), such as the high bandwidth memory (HBM) stack in Nvidia's H100 chip.
(2)
Then, the computing device performs a series of transformations and computations on the training data and modifies the corresponding neurons' parameters in the large model based on the computation results. This process is extremely brief, considering the H100 chip's INT8 Tensor Core can reach computing speeds of 4000 TOPS,[3] meaning it can perform forty trillion calculations per second. Processing an "input-output" pair may take far less than one second.[4] Once the computation of one pair is completed, the next "input-output" pair is automatically imported into the large model, overwriting the previous training data in the RAM.
(3)
Finally, after all computations based on the "input-output" pairs are completed (usually repeated numerous times), the End-to-End Learning concludes, and the parameters of the neurons in the large model are fixed. Developers can then input specific content (including text, images, videos, etc.) as prompts into the large model, which will generate corresponding content based on the prompts, such as translations, continuations, classifications, summaries, etc.
The above analysis demonstrates that training data does not create a stable copy in the training process. Instead, it is stored in the device's memory, forming a temporary copy, and each temporary copy is automatically overwritten by subsequent training data in a very short time.
03
Such Temporary Storage May Not Qualify as Reproduction Under Chinese Copyright Law
Under Chinese copyright law, generating temporary copies of works in computer memory may be regarded as temporary reproduction rather than reproduction in some cases, which does not constitute copyright infringement.
While Chinese Copyright Law does not directly address whether temporarily storing data constitutes reproduction,[5] Chinese courts and scholars have established a legal standard of "temporary reproduction" to examine this issue. Specifically, temporary storage should be deemed as temporary reproduction and not qualified as reproduction if:
(1)
The duration of the storage of involved works is transient, and
(2)
These temporarily stored works have no independent economic value.[6]
The case of Beijing Zhongqingwen Cultural Media Co., Ltd. v. a leading company in the AI industry is illustrative.[7] The defendant provided a web search service that transcodes web pages to adapt to smartphone screens. In this process, the defendant generated temporary storage of the original web page content comprising the plaintiff’s copyrighted books. In this case, the Beijing High Court set out the legal standard for temporary reproduction by stressing that temporary copies (i.e., cached works) were not permanently stored and should have no independent economic value. The court found that WAP search services do not constitute reproduction.
The Beijing High Court noted that although the defendant temporarily stored webpages containing the asserted works, it did not retain and "permanently store" these pages after users viewed them.[8] Consequently, the temporarily stored pages lacked "independent economic value" and did not constitute reproduction under copyright law:[9]
In the context of WAP search services involving the format conversion of original web pages, temporary storage of the original web page content is typically generated. This means that the search service provider does not retain the temporarily stored web pages after the network user has finished browsing. In such cases, because the temporarily stored web pages do not have independent economic value, they do not constitute reproduction or public provision of the works under copyright law. The services provided by WAP search providers are thus considered to be search and link services. However, if WAP search service providers do not temporarily store web pages but instead permanently store the relevant web pages on their servers, this may still constitute directinfringement of the used works. According to the evidence, since the complete URL addresses of the transcoded pages displayed when different mobile terminals log into defendant's Mobile Search to browse the same involved book were not the same and transcoding failures occurred, this court ruled that defendant did not store the relevant web pages in the WAP search service based on the preponderance of the evidence rule in civil litigation.[10]
Other courts have similarly held that temporary storage of works should not be regarded as reproduction. The Tianjin No. 3 Intermediate Court issued a decision in 2024, finding that transcoding copies of other's works online is not reproduction that constitutes infringement as no permanent copies of asserted works are retained on servers or disks.[11] Recent court rulings reflect those of earlier decisions, such as the 2010 decision issued by the Shenzhen Futian District Court, which held that temporary reproduction in WAP technology is not reproduction as long as the network service provider deletes the temporarily copied content in a timely manner after the transcoding process is completed.[12]
In contrast to cases in which the defendants promptly deleted temporarily copied content is the criminal copyright case against Beijing Yi Cha Wu Xian Information Technology Co., Ltd., and Yu. In this case, the defendants failed to delete cached works promptly and continued to provide the cached works to other users. The Shanghai Pudong New Area People's Court ruled that doing so gave the temporarily stored works independent economic value and constituted reproduction under copyright law:
"Yicha.com," after transmitting its so-called "temporary copies" to users triggering the "transcoding," did not immediately delete the corresponding content from the server's hard drive. The copied novel content could still be reused by other users. During the provision of novel reading services, Yicha.com not only performed webpage format conversion but also stored the converted webpage content on its server, allowing later users to directly obtain it from the server. This behavior clearly exceeded the necessary process for transcoding technology, and the so-called "temporary copies" had independent economic value. Therefore, Yicha.com's novel service model constituted direct provision of work content, and the defenses by the defendant and its counsel were not upheld.[13]
The concept of "temporary reproduction" has also long attracted academic attention in the digital age. Leading scholars tend to similarly agree that creating a fleeting temporary copy in computer memory should be regarded as temporary reproduction and not as reproduction, for two reasons: (1) creating temporary copies in memory is merely a technical phenomenon accompanying non-reproductive actions like online reading, viewing, and use of works; and (2) such temporary copies themselves do not have independent economic value or circulation potential.[14]
Under this guidance, temporary storage of data in model training may not rise to the level of reproduction under copyright law in China. As analyzed above, when works are used for training, they are temporarily stored, but this process does not produce a work copy available for dissemination or further reproduction. The temporary copies only exist in the device's memory and are automatically overwritten and deleted after a very short duration (far less than a second). No one can further reproduce or disseminate such temporary copies, meaning the copies themselves have no independent economic value. In this sense, developers' training activities in China—generating temporary copies of works in the memory—should be regarded as temporary reproduction and does not constitute copyright infringement under Chinese copyright law.
04
Temporary Reproduction in Other Jurisdictions
International copyright conventions have not established uniform legal rules regarding temporary reproduction, meaning each country (or region) may design its systems based on its copyright laws.[15] The United States and the European Union (EU) share a view similar to China on this issue: Some forms of temporary storage of works should be exempted from the reproduction right. In distinguishing temporary reproduction from reproduction, the U.S. and the EU consistently focus on whether the temporary storage is promptly deleted within a short duration and whether it has independent economic value; this aligns with Chinese judicial practice.
The EU tends to require that both factors be satisfied simultaneously for a process to qualify as "temporary reproduction," while the U.S. courts lean toward an intrinsic relationship between the two factors: If the temporary storage exists for a very short time and is automatically deleted, it can be reasonably inferred that it cannot be used for further dissemination or reproduction.
1. Rules and Practices in the United States
The U.S. Copyright Law does not provide a clear definition of temporary reproduction. However, Section 101 of the U.S. Copyright Act[16] defines concepts such as copies and fixation:
"Copies" are material objects, other than phonorecords, in which a work is fixed by any method now known or later developed, and from which the work can be perceived, reproduced, or otherwise communicated, either directly or with the aid of a machine or device. The term "copies" includes the material object, other than a phonorecord, in which the work is first fixed.
...
A work is "fixed" in a tangible medium of expression when its embodiment in a copy or phonorecord, by or under the authority of the author, is sufficiently permanent or stable to permit it to be perceived, reproduced, or otherwise communicated for a period of more than transitory duration. A work consisting of sounds, images, or both, that are being transmitted, is "fixed" for purposes of this title if a fixation of the work is being made simultaneously with its transmission.
Accordingly, a copy must be fixed in a medium for a duration longer than transitory, allowing it to be perceived, reproduced, or otherwise communicated in a permanent or stable manner. If the object produced by an action does not meet the duration requirement, the action naturally does not fall under the control of the reproduction right.
In Cartoon Network LP, LLLP v. CSC Holdings, Inc.,[17] the United States Court of Appeals for the Second Circuit clarified how long the "transitory duration" in the context of temporary reproduction might be. In this case, the defendant provided a service called Remote Storage Digital Video Recorder (RS-DVR), which allowed users to remotely store cable television programs on its central hard drive and access and replay recorded programs through their home televisions. To offer RS-DVR services, the defendant set up a Broadband Media Router (BMR) and a primary ingest buffer. The defendant would automatically load the plaintiff's programs into these devices and store the data in the BMR for no more than 1.2 seconds and in the primary ingest buffer for no more than 0.1 seconds at any time. Therefore, every 1.2 seconds, the information in the BMR would be automatically cleared and replaced, and every 0.1 seconds, the information in the primary ingest buffer would be automatically cleared and replaced.[18]
The Second Circuit ruled that the defendant's temporary storage of video clips did not infringe the reproduction right, stating that no byte of data was stored in the buffer for more than 1.2 seconds and each byte was automatically and rapidly overwritten by other data during the system's operation. These facts demonstrated that the works involved existed in the buffer for such a short duration that they did not meet the legal duration requirement for reproduction. Therefore, the temporary storage of works in the RS-DVR service did not produce a copy in the sense of copyright law and did not infringe the copyright holder's reproduction right.[19]
2. Rules and Practices in the European Union
The EU defines temporary reproduction through legislation.[20] It explicitly acknowledges that "temporary reproduction" is subject to reproduction rights[21] and provides exceptions[22] only if such acts of "temporary reproduction" are
(1)
Transient or incidental;
(2)
An integral and essential part of a technological process and whose sole purpose is to enable: (a) a transmission in a network between third parties by an intermediary, or (b) a lawful use of a work or other subject-matter to be made; and
(3)
Have no independent economic significance.[23]
The European Court of Justice, in its ruling on Infopaq International A/S v. Danske Dagblades Forening,[24] further elaborated on the above requirements. The Court held that an act of reproduction may be exempt from the reproduction right only if it fulfils five conditions:
(1)
The act is temporary;
(2)
It is transient or incidental;
(3)
It is an integral and essential part of a technological process;
(4)
The sole purpose of that process is to enable a transmission in a network between third parties by an intermediary of a lawful use of a work or protected subject matter; and
(5)
The act has no independent economic significance.[25]
Infopaq, a Danish company specializing in tracking and analyzing commercial information, scanned various newspapers and periodicals into a database and then searched articles in the database for client-submitted keywords. During scanning, the computer program temporarily generated a TIFF (Tagged Image File Format) document and automatically transmitted it to an OCR (Optical Character Recognition) server. The OCR server converted the TIFF document into text format in real-time for editing by any word processing program. At the end of this process, the TIFF document was deleted.[26]
The European Court of Justice further found that the temporary copies (i.e., the TIFF document) generated from these two temporary storages were automatically deleted from the computer memory. They did not exceed what was necessary for the proper completion of that technological process, and there was no human intervention. Thus, the storage process could be considered "transient," fulfilling all the five conditions for exemption from reproduction rights.[27]
Comment
A detailed investigation of the Large Model training process reveals that data is only transiently stored in computer memory when building the models. Since copyrighted works are only stored in the device's memory temporarily and then are quickly, automatically, and non-manually deleted during the training process, such use would likely be considered "temporary reproduction" rather than "reproduction" under Chinese Copyright Law. Similarly, copyright systems in the U.S. and the EU tend to adopt comparable standards regarding temporary copies. Large Model training produces temporary copies of copyrighted works to understand and extract ideas and concepts therein, rather than retaining specific "expressions" protected by copyright. This process should not be deemed as copyright infringement.
*This article has been published in PLI Current: The Journal of PLI Press, https://plus.pli.edu.
Slide down to view
Authors
Gordon Gao
International Partner
Intellectual Property Group
gordon.gao@cn.kwm.com
Areas of Practice:intellectual property litigation involving patents, trade secrets, trademarks, copyrights, patent right abuse, end-of-patent-life litigation for pharmaceutical products, and appeals to the PRC Supreme Court for the above types of cases
Dr. Gao advises multinational technology companies on intellectual property protection and enforcement strategies, and has handled many famous cases, including more than 10 groups of cases (a total of over 80 cases) on appeal to the PRC Supreme Court.
Xiaoyi (Sherry) Yao
International Partner
Intellectual Property Group
sherry.yao@cn.kwm.com
Areas of Practice:intellectual property litigation, with a focus on patent invalidation and patent litigation in relation to pharmaceutical and medical devices.
Dr. Yao is proficient in advising on IP-related transactions. Her clients include well-known domestic and foreign companies across fields such as pharmaceutical, medical devices, and electronic communications. Dr. Yao has years of experience in scientific research with a solid technical background in pharmaceutical related fields. In the past 14 years of practice, she has been focusing on IP matters in the pharmaceutical industry. Especially when practicing in the United States, she has handled a number of patent litigation cases under the drug patent linkage system.
转载声明:好文共赏,如需转载,请直接在公众号后台或下方留言区留言获取授权。
封面来源:问题少年002·杜飞辰

