← All dispatches
Dispatches · #intelligence · ScrapeOps

AI Training Data Disclosure: What EU Enforcement Means

August 4, 2026 · Abhishek Gupta
EU AI Act training data disclosure enforcement infographic showing GPAI fines and the August 2026 deadline

Two days ago, the European AI Office got the legal power to fine a general-purpose AI model provider €15 million for not publishing what it trained on. Two days before that deadline hit, OpenAI published a compliance statement that covered safety frameworks and watermarking in detail — and left out the one chapter that requires a public training-data summary.

That's not an oversight you make by accident this close to enforcement. It's a bet that "we said something" beats "we said the specific thing regulators asked for," and it tells you how unresolved AI training data disclosure still is even at the companies with the most lawyers.

The short version

  • The EU AI Act's GPAI enforcement powers activated on August 2, 2026 — the European AI Office can now request documentation, run evaluations, and impose fines.
  • Maximum fines are €15M or 3% of global revenue for GPAI obligation breaches, and €35M or 7% for prohibited AI uses.
  • OpenAI's July 2026 statement, "Advancing Responsible AI Across Europe," addressed the Code's Safety and Watermarking chapters but not its Copyright chapter — the one requiring a public training-data summary.
  • Providers of GPAI models released before August 2025 have until August 2, 2027 to comply; every model shipped after that date is already on the clock.
  • The compliance mechanism runs through machine-readable signals — robots.txt and ai.txt — meaning provenance now has to be traceable at the crawl level, not just declared in a PDF.

What Actually Changed on August 2, 2026?

Nothing about the law changed — the obligations have applied to GPAI providers since August 2025. What changed is enforcement: the European AI Office can now investigate, demand documents, and fine non-compliant providers, turning a paper requirement into an operational risk.

For a year, "publish a training-data summary" was a box GPAI providers could leave unchecked with no consequence. As of this week, it's a box regulators can ask about, and can penalize you for leaving unchecked. That's the actual news — not a new rule, but a new cost to ignoring an old one.

OpenAI's Filing Skipped the Chapter That Mattered

OpenAI's statement, published in late July 2026, walked through two of the GPAI Code of Practice's three chapters in real detail. The third — Copyright, which requires a publicly available training-data summary and a documented copyright compliance policy — wasn't addressed (Tech Times).

The timing does the talking. A company with OpenAI's compliance resources doesn't miss a chapter of a law it's been tracking for a year — it chooses which chapter to lead with days before the enforcement window opens. GPT-5 and every model released after August 2025 falls squarely inside the obligation this omission avoids.

What Are the Actual Penalties?

Obligation categoryMaximum fineApplies to
GPAI obligations (incl. training-data summary)€15M or 3% of global revenueModels released after Aug 2, 2025
Prohibited AI practices€35M or 7% of global revenueAll providers, regardless of model age
Legacy GPAI models (pre-Aug 2025)Same tiers, deferredGrace period until Aug 2, 2027

The fine structure (EU AI Act enforcement tracker) is steep enough that "we'll fix the disclosure later" stops being a rational default for any provider with meaningful EU revenue. It's cheaper to build the provenance trail than to defend its absence.

Why This Is a Structural Problem, Not a Paperwork One

The reason a company like OpenAI can plausibly skip a training-data summary isn't secrecy for its own sake — it's that most large-scale scraping pipelines were never built to produce a clean, attributable provenance record in the first place. Data gets pulled from thousands of sources, deduplicated inconsistently, and merged into training sets where the original source, license, and crawl date are the first things to get lost.

That's the same failure mode we built ScrapeOps around, just from the acquisition side instead of the compliance side: if your pipeline can't tell you where a piece of data came from and whether it's a duplicate of something you already have, it also can't tell a regulator that on request. Deduplicated, source-traceable acquisition isn't just cleaner data — as of this week, in the EU, it's closer to a legal requirement than a nice-to-have.

We've tracked the broader shift from open scraping to licensed, permissioned access in more detail at dekryptlabs.com/dispatches, and the research behind how we think about data provenance lives at dekryptlabs.com/research.

Where This Goes Next

Expect more GPAI providers to publish partial compliance statements over the next few months, and expect the European AI Office's first information requests to target exactly the gaps those statements leave open. The providers that treat provenance as infrastructure — not as a document to write after the fact — are the ones who won't be scrambling when the request lands.

That's the quiet shift underneath this story: training data stopped being a private implementation detail sometime in the last year, and started being something you have to be able to show your work on. The companies that already built for that will barely notice the deadline. The ones that didn't just found out it has teeth.

Frequently Asked Questions

What is the EU AI Act's training-data disclosure requirement? GPAI model providers must publish a public training-data summary using the European Commission's template and maintain a documented copyright compliance policy, including respecting opt-out signals like robots.txt and ai.txt.

When did EU AI Act enforcement for training data actually start? The obligation existed since August 2, 2025, but the European AI Office's enforcement powers — investigations, information requests, and fines — only activated on August 2, 2026, one year later.

How much can a company be fined for not disclosing training data? Up to €15 million or 3% of global annual revenue for GPAI obligation breaches, including missing or incomplete training-data summaries. Prohibited AI practices carry fines up to €35 million or 7% of global revenue.

Does this apply to AI models built before August 2025? Yes, but on a delay. Providers of GPAI models released before August 2, 2025 have until August 2, 2027 to reach full compliance before enforcement applies to them.

Abhishek Gupta is Co-Founder at Dekrypt Labs, building ScrapeOps — the data acquisition engine that turns any question into clean, deduplicated, comprehension-ready sources. dekryptlabs.com