🔍 Medical histories. Criminal records. In Japan, if the purpose is training an AI model, you may no longer need to ask the person first. Read as a headline, it sounds like a scandal, and that is roughly how it played in the Diet. Read as statute, the design turns out to be more deliberate than the headline, and unlike anything else in force anywhere.

What actually passed

The amended Act on the Protection of Personal Information cleared Japan's upper house on July 10, 2026. The chamber's published tally records 240 votes cast, 146 in favor and 94 against. By party, the Liberal Democrats, the Democratic Party for the People, Ishin and Team Mirai voted yes. The Constitutional Democrats, Komeito, Sanseito, the Communists, Reiwa, Okinawa no Kaze and the Social Democrats voted no, while the Conservative Party of Japan and the unaffiliated members split. The line ran neither along government versus opposition nor along left versus right.

Its centerpiece is what practitioners call the statistical special provision. Where the purpose is compiling statistics or developing AI models, businesses may acquire personal data and pass it to third parties without the consent of the people it describes. Sensitive categories such as medical history, criminal record, race and creed sit inside that exception rather than outside it. A policy document published by the commission's secretariat in April 2026 had already stated that statistics creation covers AI development that can be organized as such. That does not mean every AI project qualifies. The statutory definition imposes conditions, among them that the end product must not be information about individuals.

The law takes effect within two years of promulgation, by July 2028. Most of the substance has yet to be written. At a press conference on July 31, 2026, the Personal Information Protection Commission said it would publish in August the full picture of the issues its rules and guidelines will cover, and open a public comment period. Secretary-General Kiyoshi Sawaki said the commission would spell out anticipated cases as clearly as possible, and that it also wanted to publish criteria for how the new surcharge orders will operate.

What the exception really unlocks: crawling and record linkage

"AI developers can use anything" is a long way from what the text says. Yoichiro Itakura, an attorney at Hikari Sogoh Law Offices who works on privacy law, breaks the provision into two functions. He responded twice to pre-legislative hearings on the bill and says plainly that this makes him neither an outsider nor entirely neutral.

The first function is web crawling. Under current law, acquiring sensitive personal information without consent is unlawful, with an exception where the person themselves or a news organization has published it. But when a colleague posts that someone is off sick, picking up that post is already an acquisition without the subject's consent. Crawling the open web to train a generative model, you cannot reliably filter every post of that kind out. The first half of the exception exists to give that act a lawful route.

The second function is record linkage across companies. This is the part most often misread. According to Itakura, much of AI development was already lawful. A company working alone can build statistics or train a model on its own data without even notifying the people involved, and the finished statistics or trained model generally fall outside the privacy law altogether so long as they are not information about identifiable individuals, which means they can already be handed to other firms.

What was blocked was two companies matching their records to establish that A's customer and B's customer are the same person, then building statistics or training data from the combined view. The restriction on third-party provision stood in the way. The second half of the exception opens that door.

What makes Japan different, part one: privacy law brought into step with copyright law

Article 30-4 of Japan's Copyright Act, introduced by the 2018 amendment and in force since January 1, 2019, permits exploiting a work for purposes such as information analysis where the point is not to personally enjoy, or let others enjoy, the thoughts or sentiments expressed in it. It draws no line between commercial and non-commercial use, and it is widely described as among the most permissive limitations on copyright anywhere with respect to machine learning.

Itakura notes that crawling raises the same problem under copyright law, that Article 30-4 already handles it, and that the privacy law's answer to crawling is in step with the copyright reasoning.

That single observation identifies the most Japanese thing about this amendment. One act, scraping the open web indiscriminately, has now been cleared by statute on both the copyright side and the personal data side, using the same underlying logic: the machine is not consuming this the way a human consumes it. Copyright law calls it non-enjoyment use. Privacy law calls it statistics creation. No other major jurisdiction has closed both halves of that circle by legislation.

Europe is standing somewhere else. In Opinion 28/2024, adopted on December 17, 2024, the European Data Protection Board confirmed that legitimate interest can serve as a legal basis for developing and deploying AI models, but left the judgment to a case-by-case assessment rather than granting it as a category. South Korea's Personal Information Protection Act allows pseudonymized data to be processed without consent for statistics and scientific research, but pseudonymization is the technical step that opens the door in the first place. Japan draws the line by purpose, Korea by technique, Europe case by case.

What makes Japan different, part two: disclosure and contracts instead of technology

The second distinctive feature is how the brake was built.

A business using the provision may not use the data for anything other than the statistical purpose. It must publish its own name and the substance of what it is doing. Where data changes hands, both provider and recipient publish those items, the provider also publishes the name of the recipient, and the two sides must sign a written agreement limiting the purpose. Security obligations extend beyond database-organized personal data to personal information generally. Breaching any of this exposes a company to the newly created surcharge order.

The constraint, in other words, is disclosure and contract. It is not technology. Neither technical measures that would let statistics be produced without a human ever seeing raw records, nor the involvement of a provider of privacy-enhancing technologies such as secure computation, were made mandatory.

Itakura has spent years arguing that such technologies deserve proper incentives, and during the hearings he floated an alternative modeled on the certified intermediaries used under Japan's next-generation medical infrastructure law, limiting who may perform the processing. What emerged instead routes personal information from company A and company B into company X, which builds the statistics, with disclosure and contract standing in as the guarantee. Itakura himself writes that the design may have loosened things further than it needed to.

The Japan Association for Medical Informatics, in an opinion dated June 19, 2026, argued that nothing in the scheme assesses in advance whether a receiving third party can actually handle the data safely. Itakura's counter is that no such advance assessment exists for outsourced data processing either, nor for the academic research exception widely relied on in medicine.

Is the fight about profiling, or about enforcement?

The loudest objection is that models built under this provision will end up sorting people. Ryoji Mori, an attorney who sat on the commission's review panel, points to the volume of statistical inferences the approach will generate: fragments like late-night shopping, browsing history and how often a person drinks, bundled into a conclusion that this person is likely to fall behind on payments, and generated by the thousand. The attributes feeding such inferences include ones that are not obvious and ones the person is unaware of. Mori's own proposal is that, apart from uses clearly benefiting the individual, applying these inferences to actual people should be prohibited for the time being. Tightening the rules on profiling was a live issue while the amendment was being drafted, but Masaharu Koda, professor emeritus at Kanagawa University, says that despite many experts urging its inclusion it vanished without ever being debated at the commission.

Itakura considers the criticism built on several misunderstandings. Once you apply a finished model or statistic to actual personal data, the exception no longer applies and the ordinary rules return: disclosure of purpose, the ban on use beyond that purpose, the restriction on third-party provision. Rules on profiling were introduced in the 2020 amendment, including the prohibition on improper use, so saying there is no regulation is simply wrong. And Sawaki told a lower house special committee on May 12, 2026 that building and releasing an AI model, using personal information, while foreseeing that it risks inducing unlawful discrimination would itself be illegal, since only uses posing little risk of harm to individual rights can qualify as statistics creation at all.

The two men converge in the end. Whether any of it works depends on enforcement.

The surcharge is new to Japanese privacy law. It requires that more than 1,000 people be affected and that the company obtained money or other financial benefit in exchange for the violation. The amount equals the improper gain, rises by half again for repeat offenders, and drops by fifty percent for a company that reports itself. Data breaches caused by security failures were deliberately left out. The instrument targets deliberate profit-taking, not operational mistakes.

The awkward part is that the duties the provision creates are precisely the kind the commission has hardly ever enforced. In the figures Itakura cites, of 88 administrative guidance actions in the fourth quarter of fiscal 2025, 84, or 95 percent, involved security-management failures or failures to supervise contractors. Three concerned unlawful third-party provision and one improper acquisition. Formal orders are rarer still: one in the first half of fiscal 2025, none in fiscal 2024 or 2023. A promise to police purpose limitation and disclosure duties has thin history behind it.

Clause 14 of the lower house supplementary resolution called for further strengthening the commission's staffing and budget in light of the workload the amendment creates. Itakura's own conclusion is that whether the scheme earns trust comes down to whether a supervisory capacity that actually catches violations gets built.

What August will reveal

The statutory text is settled. What remains movable is the set of issues the commission publishes in August and the rules and guidelines that grow out of it.

Three things are worth watching. First, what gets designated as statistics creation, and whether the statutory limit to uses posing little risk of harm turns out to be a real boundary. Second, the granularity of the disclosure duty, since publishing a company name and a description means very little or quite a lot depending on how much detail is required. Third, whether the surcharge criteria arrive with enough specificity to compensate for how seldom the commission has used its existing powers.

For companies developing AI in Japan, and for organizations outside Japan handling data about people in Japan, the August document is where a compliance plan running to 2028 starts.

Japan has become one of the very few countries to answer the question of what it means for a machine to read, by legislation, on both the copyright side and the privacy side. Whether that answer was right will be decided not by the text but by two years of rulemaking. How does your country handle it: consent, mandatory anonymization, or a judgment call every time?

References