Japan Adopts Voluntary AI Training Data Disclosure Code

Japan’s government has adopted a voluntary Principles Code urging generative AI providers to disclose their training processes, the types of data used and collection methods. Companies that accept all or part of the code must notify the government and make specified information available online, extending the framework to overseas providers serving Japan.
Japan's government adopted voluntary guiding principles for generative AI operators on Aug. 25, calling on such businesses to disclose outlines of the data and methods used to train AI tools. The Principles Code applies to domestic businesses and overseas operators that provide AI systems or services in Japan, according to Jiji Press reporting published by The Japan Times and Nippon.com.
The code is not legally binding. Jiji Press reported that its stated purpose is to increase AI transparency while supporting intellectual-property protection and innovation as generative AI use expands.
What participating operators would disclose
Businesses that accept all or part of the code will notify the government and disclose information on their websites about model learning processes, the types of training data used, and the methods used to collect that data, according to Jiji Press.
The framework also calls for participating operators, subject to specified conditions, to disclose whether training data include material that could create copyright-infringement concerns when users or rights holders request that information. This is particularly relevant to generative systems trained on large web-scale corpora, where model developers may not publicly enumerate every underlying work.
Draft details addressed pirate sources and creator requests
Earlier coverage of the draft rules by Japan Today and Anadolu Agency described additional operating expectations. Operators were asked to publish an overview of automated online data-collection programs and avoid collecting from pirate sites.
Japan Today reported that the draft contemplated disclosure requests from rights holders in areas such as manga, music, and film, as well as from users who generate AI content. One cited scenario involved a rights holder preparing potential litigation after finding online AI-generated material that resembles its work.
The draft also asked operators to establish mechanisms for informing users whether training data contained copyrighted works similar to generated content, Japan Today reported. At the same time, the guidelines stated that compulsory disclosure would not be sought for trade secrets or security-related information. The draft carried no penalties for noncompliance, but companies would be asked to explain their reasons, according to Japan Today and Anadolu Agency.
Practical relevance for AI providers
The scope covering foreign providers makes the code relevant to model developers, API platforms, and application vendors that serve Japanese users, even if they are headquartered elsewhere. The reported framework is voluntary, but it creates a government-backed transparency baseline around dataset provenance and copyright-related requests.
For ML teams, comparable transparency regimes commonly increase the value of maintaining records on source categories, web-crawling controls, licensing status, dataset versions, and data-removal workflows. Those records can support externally published summaries without requiring disclosure of model weights, proprietary collection systems, or sensitive security information.
The code also illustrates a broader policy tension for generative AI: rights holders seek enough visibility to investigate possible copying, while providers seek to protect trade secrets and preserve practical access to large-scale training data. Japan's adopted principles address that tension through voluntary disclosure and conditional requests rather than direct statutory penalties.
Key Points
- 1Japan adopted a voluntary AI disclosure code covering domestic and foreign providers, extending training-data transparency expectations to services offered in Japan.
- 2Participants are asked to publish training-data types, collection methods, and learning-process outlines, creating a governance requirement short of full dataset disclosure.
- 3Comparable transparency frameworks make provenance records and copyright-request workflows more operationally important for teams deploying web-trained generative models.
Scoring Rationale
The voluntary code establishes a notable transparency framework for generative AI providers serving Japan, including overseas operators. It matters to ML and platform teams because training-data provenance, collection controls, and rights-holder response processes are increasingly operational policy concerns, though the framework is not legally binding.
Sources
Public references used for this report.
Practice interview problems based on real data
1,625 SQL & Python problems across 15 industry datasets — the exact type of data you work with.
Try 250 free problems

