AI LAW RADAR · Daily Last verified 22 Aug 2026

Topic dossier

Training-data sourcing & disclosure rules

Rules on where model training data may come from and what has to be said about it — text-and-data-mining permissions, scraping limits and public training-content summaries. 7 obligations across 5 jurisdictions — 5 in force. Next dated deadline: 1 Mar 2027.

Training data has become a regulated input in its own right. Three distinct duties are emerging: permissions that define when protected material may be mined or copied for model development; sourcing limits on scraped and third-party data; and disclosure duties that force a provider to describe publicly what a model was trained on. The EU AI Act's public training-content summary and the statutory mining and development exceptions appearing in national copyright law are the leading examples. The instruments below are the ones AI Law Radar tracks under this theme, each dated to its last check against the primary source.

The Register

7 obligations · 5 jurisdictions

European Union 3

EU Comprehensive

EU — General TDM Exception with Rightholder Opt-Out (DSM Directive Art. 4)

Binds Anyone carrying out text and data mining on lawfully accessible works in the EU, including commercial AI model training (grants a permission that lapses for any work whose use the rightholder has expressly reserved under Art. 4(3)). Art. 4 requires Member States to allow reproductions and extractions of lawfully accessible works for text and data mining by anyone, for any purpose including commercial AI training, and lets copies be retained as long as the mining needs them. The permission applies only where rightholders have not expressly reserved the use in an appropriate manner — machine-readable means for content made publicly available online. This opt-out is the reservation that AI Act Art. 53(1)(c) then obliges general-purpose AI model providers to identify and respect.

LEGAL PERMISSION with an opt-out — not a compliance obligation in itself. Directive (EU) 2019/790 entered into force on 7 June 2019 (Art. 31: twentieth day after publication in OJ L 130 of 17 May 2019) and Art. 29(1) set the Member State transposition deadline at 7 June 2021, which is the date carried here; as a directive it takes effect through national implementing law, so the precise wording and any national nuance vary by Member State. Art. 4(1) verbatim: 'Member States shall provide for an exception or limitation to the rights provided for in Article 5(a) and Article 7(1) of Directive 96/9/EC, Article 2 of Directive 2001/29/EC, Article 4(1)(a) and (b) of Directive 2009/24/EC and Article 15(1) of this Directive for reproductions and extractions of lawfully accessible works and other subject matter for the purposes of text and data mining.' Art. 4(3) verbatim: 'The exception or limitation provided for in paragraph 1 shall apply on condition that the use of works and other subject matter referred to in that paragraph has not been expressly reserved by their rightholders in an appropriate manner, such as machine-readable means in the case of content made publicly available online.' Art. 2(2) defines text and data mining as 'any automated analytical technique aimed at analysing text and data in digital form in order to generate information which includes but is not limited to patterns, trends and correlations'. Art. 4(4) preserves the separate scientific-research exception in Art. 3, which carries no opt-out.

Stated maximum penalty — N/A — permissive exception (no penalty attaches to mining within Art. 4; mining a work whose use has been reserved under Art. 4(3) falls outside the exception and is dealt with as ordinary copyright infringement under national law)

In force · 7 Jun 2021 checked 17 Aug 2026 EU DSM Copyright Directive (EU) 2019/790 Art. 4 ↗ high confidence
EU Comprehensive

GPAI model provider obligations (Art. 53) — docs, copyright policy, training summary

Binds Providers of general-purpose AI models. Technical documentation (Annex XI), downstream-provider information (Annex XII), a copyright-and-related-rights policy identifying Art. 4(3) DSM rights reservations, and a public summary of training content on the AI Office template.

Stated maximum penalty — Up to 3% turnover or €15M

In force · 2 Aug 2025 checked 17 Aug 2026 EU AI Act ↗ high confidence
EU Comprehensive

GPAI models placed on the market before 2 Aug 2025 — legacy compliance deadline (Art. 111(3))

Binds Providers of general-purpose AI models placed on the EU market before 2 August 2025. The AI Act's GPAI duties bite on legacy models on 2 August 2027: general-purpose AI models placed on the EU market before 2 August 2025 have until that date to be brought into line with the Regulation. Until then, legacy models sit outside the Chapter V obligations that have bound newly placed models since 2 August 2025.

Article 111(3) of Regulation (EU) 2024/1689, unchanged by the Digital Omnibus on AI: Regulation (EU) 2026/1744 Article 1(39) amends only Article 111(2) (replaced) and adds Article 111(4) (synthetic-content marking retrofit, 2 Dec 2026) — paragraph 3 and its 2 August 2027 date are untouched (OJ L, 24.7.2026). This is the counterpart to the GPAI duties that bound new models from 2 August 2025 (Art. 53, Annex XI/XII, copyright policy, training-data summary) and to Commission enforcement powers live since 2 August 2026 (Art. 101). Scope is the model, not the system: a legacy model that is substantially modified is treated as newly placed on the market and loses the grace period.

Stated maximum penalty — Up to 3% turnover or €15M

Applies 2 Aug 2027 checked 17 Aug 2026 EU AI Act Art. 111(3) ↗ high confidence

Japan 1

Japan Comprehensive

Japan — AI Training / Information Analysis Exception (Copyright Act Art.30-4)

Binds Anyone in Japan reproducing or otherwise exploiting copyright works for information analysis, including AI model training (grants a statutory permission, subject to the Art.30-4 proviso and to Art.47-5 limits on downstream enjoyment use). Art.30-4 permits exploitation of a published or unpublished work, by any means and to the extent deemed necessary, where the purpose is not to enjoy the ideas or sentiments expressed in it — item (ii) names information analysis expressly, which covers machine-learning training. Subject to a proviso: the exception falls away where, in light of the type and use of the work and the manner of exploitation, it would unreasonably prejudice the copyright owner's interests. Legal permission, not a compliance obligation.

LEGAL PERMISSION — not a compliance obligation. Art.30-4 (Act No. 48 of 1970) in its current form was inserted by the 2018 amendment (Act No. 30 of 2018), whose supplementary provisions set entry into force at 1 January 2019 (Heisei 31). The chapeau allows exploitation 'to the extent deemed necessary' where the purpose is not self- or third-party enjoyment of the expressed ideas or sentiments; item (ii) covers information analysis, defined in the statute as extracting and comparing, classifying or otherwise analysing language, sound, image or other elements from a large number of works or a large volume of information. The proviso is the operative limit: no exception where the exploitation would unreasonably prejudice the copyright owner's interests in light of the type and use of the work and the manner of exploitation. Art.47-5(2) and Art.113(9) then withdraw the shelter from anyone who later uses an Art.30-4 copy for enjoyment purposes. Re-checked against the current consolidated e-Gov text on 2026-08-09: no amendment since 2024 touches Art.30-4 — the 2024-2026 amending Acts (Reiwa 6 No. 55, Reiwa 7 No. 27, Reiwa 8 Nos. 37 and 48) leave the article unchanged.

Stated maximum penalty — N/A — permissive exception (no penalty attaches to exploitation within Art.30-4; the general infringement ceiling is Art.119(1), up to 10 years' imprisonment and/or a JPY 10,000,000 fine, and Art.124(1)(i), up to JPY 300,000,000 for a corporate body)

In force · 1 Jan 2019 checked 12 Aug 2026 JP Copyright Act (Act No. 48 of 1970) Art.30-4 ↗ high confidence

Russia 1

Russia Binding

243-FZ art. 10(1) — whoever lets you use a large foundational model has to tell you who owns the output

Binds Any person that provides the ability to use a large foundational AI model as defined in art. 3(2) — not fewer than 1 billion parameters, general-purpose across a large number of tasks, and serving as the basis for creating and refining other software. The duty is expressed without a nationality, size or turnover limb, in contrast to arts. 6 to 8, which apply only to Russian legal persons developing models granted sovereign or national status. Art. 1(3) reserves to other federal laws and presidential acts the setting of special rules for defence, state security, operational-search activity, public order and property protection, public and road safety including counter-terrorism, anti-money-laundering and counter-terrorist-financing, emergency prevention, diplomatic and consular service and state administration, so those uses may be governed differently.. Federal Law No. 243-FZ of 26 July 2026 «О поддержке развития технологий искусственного интеллекта в Российской Федерации» is Russia's first AI statute, and this is its broadest genuine duty. Art. 10(1) requires «лица, предоставляющие возможность применения больших фундаментальных моделей искусственного интеллекта» — the persons who make a large foundational model available for use — to notify the user of two things unless another federal rule provides otherwise: to whom the rights belong in the results of intellectual activity obtained with the help of the model, and on what conditions the user is given access to, use of, and retention of those results, retention being qualified by technical possibility. Unlike arts. 6 to 8, the duty is not confined to sovereign or national models or to Russian developers, so it reaches any provider offering such a model to users in Russia. Its scope is set entirely by the art. 3(2) definition: a large foundational model is a computer program intended to perform intellectual tasks at a level comparable to or exceeding human intellectual activity, using algorithms and trained on data sets to infer patterns, supply information, take decisions or forecast results against human-set goals, simultaneously serving as the basis for creating and refining various kinds of software, containing not fewer than 1 billion parameters and applied to a large number of different tasks. Every cumulative limb has to be met, so smaller and narrow-purpose models fall outside the Law altogether. Art. 10(2) sits alongside as a permission rather than a duty: accessing information contained in copyright and neighbouring-rights objects for the practical application of what they contain, including machine extraction, comparison, classification and analysis of patterns, trends and correlations, and short-term reproduction in machine memory, is declared not to infringe — but only where it is done exclusively to train a sovereign and (or) national large foundational model, and only where the developer uses a lawfully obtained copy or the work had been communicated to the public and was available for analysis without technical restriction. A text-and-data-mining exception that is available only to models holding a state-conferred status is an unusual shape and worth noting when comparing it with the EU and Singapore exceptions.

The date on which this duty starts is 1 March 2027, not the 1 September 2026 date reported as the commencement of the Law. Art. 13(1) does put the Law in force on 1 September 2026, but art. 13(2) then defers arts. 8, 9 and 10 in full, together with art. 5(2) points 3 to 5 and art. 6 parts 2 to 5, to 1 March 2027. What actually commences on 1 September 2026 is the subject matter, aims, definitions and principles in arts. 1 to 4, the coordination and support-measure powers in art. 5(1) and art. 5(2) points 1 and 2, the statement of purpose in art. 6(1), the art. 7 list of what a status-holding developer may do, the bare liability referral in art. 11, and arts. 12 and 13 — none of which places a compliance duty on anyone. The official record confirms the position: the pravo.gov.ru register carries the Law as «Не вступил в силу» with a single original redaction marked «вступает в силу 01.09.2026». Adopted by the State Duma on 8 July 2026, approved by the Federation Council on 17 July 2026, officially published on the legal-information portal on 26 July 2026 under number 0001202607260003, and reproduced at Собрание законодательства РФ 2026 No. 30 item 4089 and in «Российская газета» of 31 July 2026.

Stated maximum penalty — None is stated in the Law. Art. 11 is a bare referral — participants in relations in the field of development, deployment and application of large foundational models bear responsibility «в соответствии с законодательством Российской Федерации» for breaches of the Law and of the acts adopted under it — and as at 21 August 2026 the Code of Administrative Offences carries no article addressed to large foundational AI models, so no monetary band attaches to art. 10(1). Where the failure to notify also amounts to a consumer-information failure or a personal-data breach, the existing KoAP articles apply on their own terms. This entry deliberately states no figure rather than importing one from an adjacent regime.

Applies 1 Mar 2027 checked 21 Aug 2026 243-FZ art. 10(1) ↗ high confidence

Saudi Arabia 1

Saudi Arabia Binding

Saudi Arabia — AI Training Exemption (Copyright Law Art.26)

Binds Developers of AI products and algorithms reproducing copyrighted works in Saudi Arabia. The statutory permission is conditioned on lawful publication of the work, lawful acquisition of the original copy, and copying limited to the purpose (Law Art.26(4)), and on the six further controls in Art.30 of the Implementing Regulation, of which Art.30(3) binds the developing entity to keep records of the type, source, purpose and date of use of each work used and to produce them on request to a competent body examining a dispute over that use.. Art.26(4) permits reproduction of an original work for developing AI products and algorithms without author authorization or compensation, subject to three statutory conditions: the work was lawfully published, the original copy was lawfully obtained, and copying stays within what the purpose requires. The Implementing Regulation published on 31 Jul 2026 adds Art.30, which subjects that exception to six cumulative controls — among them a bar on relying on it for purely commercial use, a bar on unnecessary inclusion of the work in the final products, and an affirmative duty on the developing entity to keep records of every work used. A conditioned permission carrying one standing compliance duty, rather than an unconditioned freedom.

LEGAL PERMISSION — not a compliance obligation. Royal Decree No. M/169 (Copyright Law) Art.26(4) permits reproduction of an original work for AI product and algorithm development without author consent or compensation, on three statutory conditions (lawful publication; lawful acquisition of the original copy; copying limited to the purpose). IN FORCE since 12 Aug 2026. Art.61: the Law enters into force 180 days after publication in the Official Gazette (Umm Al-Qura issue 5144, 13 Feb 2026) = 12 Aug 2026. The Implementing Regulation (اللائحة التنفيذية لنظام حقوق المؤلف, 98 articles in 13 chapters) was published in Umm Al-Qura on 17/02/1448, corresponding to 31 Jul 2026 (https://www.uqn.gov.sa/decisions-and-regulations/4001498). Art.60 of the Law governs when it bites: the Council issues the Regulation within 180 days of the Law's issuance «ويُعمل بها من تاريخ نفاذه» — it applies from the date the Law itself enters into force. The Regulation carries no commencement article of its own; Chapter Thirteen (final provisions) runs Arts.95-98 and closes with the Authority's chief executive issuing implementing decisions. So Art.30 binds from 12 Aug 2026, the date on this row, not from the date of the Regulation's publication. Chapter Seven is headed «برامج الحاسب الآلي واستخدامات الذكاء الاصطناعي» (computer programs and artificial-intelligence uses) and its Art.30 subjects the Art.26(4) exception to six cumulative controls, read one by one in the gazette text: (1) copying and analysis are confined to the extent necessary for the purpose of developing AI algorithms or products, and do not extend to republication, distribution or direct commercial exploitation of the work; (2) the work may not be used «في إطار تجاري بحت» — within a purely commercial frame — unless that use is insubstantial in relation to the work or does not affect its normal exploitation; (3) the developing entity is obliged to keep records showing the type of work used, its source, the purpose of use and the date of use, and to produce them on request to any competent body examining a dispute relating to that use; (4) the use must not cause unjustified harm to the author's legitimate interests and must not affect the opportunity to exploit the work or obtain material return from it, which restates the three-step test the Law applies to Arts.26-36 through Art.37(1); (5) adaptation, republication, making the work available to the public, and unnecessary inclusion of the work in the final products are prohibited without the rightholder's permission, unless the work has passed into the public domain; and (6) where the work contains elements under separate protection or independent rights, the provisions governing those elements continue to apply. Condition (3) is the operative compliance item on this row — an affirmative, standing record-keeping duty on the developer. Conditions (2) and (5) are the material limits on the exception's scope, and neither was captured when the row was first written. Divergence carried on the entry rather than escalated: Art.26(4) of the Law states three conditions and imposes no commercial limitation, while Art.30(2) of the Regulation bars reliance on the exception for purely commercial use. These are two instruments of different rank, the junior one supplying the detail Art.60 of the Law directs it to supply, not two sources contradicting each other on a fact.

Stated maximum penalty — N/A — permissive exemption (no penalty attaches to a user acting within Art.26(4); the Law's general infringement ceiling is SAR 1,000,000 and/or 1 year, doubled on repeat offence)

In force · 12 Aug 2026 checked 22 Aug 2026 SA Copyright Law (Royal Decree M/169) ↗ high confidence

Singapore 1

Singapore Binding

Singapore — Computational Data Analysis Exception (Copyright Act 2021 s.244)

Binds Anyone in Singapore copying or communicating works or recordings of protected performances for computational data analysis, including training a computer program (grants a permitted use conditioned on single-purpose use, no onward supply, lawful access, and a non-infringing source copy). s.244 makes it a permitted use to copy a work or a recording of a protected performance for computational data analysis, which s.243 defines to include using the material as an example to improve how a computer program functions — i.e. model training. Four conditions: the copy serves only that purpose, it is not supplied onward except for result verification or collaborative research, the person has lawful access to the source copy, and that source copy is not a knowingly infringing one. No rightholder opt-out, and contract terms purporting to exclude the exception are void under s.187.

LEGAL PERMISSION — not a compliance obligation. Copyright Act 2021 (Act 22 of 2021) commenced 21 November 2021; s.2 of the Act is framed by reference to that date. s.243 defines computational data analysis to include (a) using a computer program to identify, extract and analyse information or data from the work or recording, and (b) using the work or recording as an example of a type of information or data to improve the functioning of a computer program in relation to that type of data — the statutory illustration is training a program to recognise images. s.244(2) conditions the permitted use on: the copy being made only for that analysis or for preparing the material for it; no other use of the copy; no onward supply except to verify results or for collaborative research or study; lawful access to the source copy; and the source copy being non-infringing (or the user neither knowing nor reasonably able to know otherwise). The statutory illustrations name circumventing paywalls and breaching database terms of use as defeating lawful access. s.244(3) confirms storage and retention count as copying; s.244(4) extends the permission to communication to the public of a copy made under s.244(1).

Stated maximum penalty — N/A — permissive exception (no penalty attaches to a permitted use under s.244; the Act's general criminal ceiling for wilful commercial-scale infringement is a fine and imprisonment under Part 9 Division 6, and civil remedies including statutory damages remain available for use falling outside the conditions)

In force · 21 Nov 2021 checked 20 Aug 2026 SG Copyright Act 2021 (Act 22 of 2021) s.244 ↗ high confidence

Questions & answers

From the data

Do AI providers have to disclose what data a model was trained on?

In the EU, yes for general-purpose AI models: Article 53(1)(d) of the AI Act requires a sufficiently detailed public summary of the content used for training, following a template provided by the AI Office. Under Article 113(b) that duty has applied since 2 August 2025. Most other jurisdictions currently have no equivalent public-summary requirement.

What is a text-and-data-mining (TDM) exception?

A statutory carve-out permitting the automated analysis of lawfully accessible works without separate authorisation. The EU version sits in Articles 3 and 4 of Directive (EU) 2019/790, with Article 4(3) letting rightsholders reserve the commercial case in machine-readable form. Separate statutory permissions aimed specifically at AI development — such as Saudi Arabia's Copyright Law Article 26(4) — are tracked here as well.

Is scraping public web data for AI training regulated?

Copyright, database rights and data-protection law can each apply independently, and lawful accessibility of a page does not by itself settle the copying question. The rows above list the instruments that speak directly to training-data sourcing; where a jurisdiction has enacted no specific rule, general law applies and no row appears.

Which jurisdictions does AI Law Radar track for training data & tdm?

We currently track training data & tdm obligations across 5 jurisdictions: European Union, Japan, Russia, Saudi Arabia and Singapore. Each is dated and linked to its primary source on this page.