GDPR Compliance for AI — Build One File That Can Testify

A living AI system record links personal data, model lineage, human review and remedies, creating proof that…

GDPR Compliance for AI — Build One File That Can Testify

The short version

  • GDPR compliance for AI requires a living record connecting purpose, personal data, models, decisions, evidence and remedies.
  • EU hosting alone does not ensure data sovereignty; access, reuse, copies, lineage and deletion paths also matter.
  • Meaningful human review needs context, authority, challenge routes and logs proving people can interrupt automated decisions.

Uber’s software didn’t just score drivers. It cut off their income without human review.

On August 21, 2026, the Dutch Data Protection Authority announced an €825 million fine, against a GDPR ceiling of 4% of worldwide annual turnover. Uber appealed, but anyone shipping AI that scores workers, candidates, customers or creators should check the wiring now.

Delayed AI Act deadlines won’t save sloppy GDPR compliance. Neither will an EU AI Act compliance checker: no questionnaire can reconstruct what happened after someone pasted a CV into a model at midnight.

I want one record connecting purpose, personal data, model, decision and remedy. I call it the living AI system record. Mine follows Arianna, a fictional recruitment assistant that summarizes CVs and ranks candidates while HR says “efficiency” with suspicious enthusiasm.

Build one AI system record and make it prove everything

Arianna starts as one inventory row, before anyone uploads a CV. I record its purpose, affected people, internal owner, vendor, inputs, outputs, retention periods, processing locations and influence over hiring. A tool that tidies interview notes carries different risk from a ranking engine whose lowest-scoring candidates disappear before a recruiter knows they applied. That difference shapes lawful basis, access rights, deletion routes, security measures and the need for a data protection impact assessment. When the vendor changes its model or adds a subprocessor, I update the row instead of excavating Slack. Without it, every compliance meeting starts with six people debating what the product does.

Then I run two legal tests. GDPR applies when Arianna processes personal data, even if it’s glorified autocomplete in a blazer. The AI Act asks whether Arianna is an AI system and how its intended use is classified. A spreadsheet of candidate names can trigger GDPR duties while falling outside the AI Act. Industrial forecasting based entirely on machine telemetry may do the reverse. Recruitment software commonly triggers both, so the record needs two legal conclusions—not one beige blob marked “EU compliance.”

The roles also split. Under GDPR, the employer is usually the controller and a vendor following its instructions may be a processor. Under the AI Act, the vendor may be the provider and the employer the deployer. A company that substantially modifies Arianna or sells it under its own name can gain provider obligations while remaining the GDPR controller.

One “compliance owner” column will collapse here. Writing “vendor” in a green cell won’t help. Conditional formatting has yet to win a regulatory appeal.

An EU AI Act compliance checker remains useful for initial scoping. It can ask about intended use, flag possible classifications and list documents to investigate. It cannot see what employees entered, establish the controller’s lawful basis or verify that a recruiter challenged Arianna’s recommendation. I use it to start the interview, then attach evidence for every answer to the living record.

Some evidence can support multiple assessments, but each has a different job. A GDPR impact assessment examines risks from personal-data processing. The AI Act’s fundamental-rights impact assessment applies to specified deployers using covered high-risk systems. Renaming one PDF after lunch is not legal alchemy.

Data sovereignty starts beyond the server location

I used to give EU hosting too much credit. Frankfurt sounds comforting, especially when American vendors treat “Europe” as an availability-zone dropdown. But sovereignty also depends on who can access data, how it’s reused and whether deletion reaches every copy.

Geography answers one question. Not the whole exam.

Follow one sentence from a candidate’s CV. Arianna puts it in a prompt, which the application may log before sending it to the model provider. The provider may retain separate operational or safety logs under its terms, while a retrieval store keeps another copy for future queries. If it enters a fine-tuning dataset, deleting the original CV may leave those versions behind. An open-weight model may later be modified, combined with another model or released as a descendant. Every step creates another place where the information may survive and another party that must respond when the candidate exercises a GDPR right. My system record therefore tracks the model version, contract terms, known copies and a deletion path far beyond the shiny Frankfurt endpoint.

Open-weight AI makes this especially spicy. France’s CNIL explains that models develop complicated genealogies through fine-tuning and combination. Its Genmod demonstrator maps ancestor and descendant models, helping investigators find related releases that may also contain memorized personal data. In its August 26, 2026 update, the CNIL said an unlimited-depth genealogy search averages about 20 seconds, though it gave no previous duration for comparison. Fast graph exploration shows investigators where to look. It doesn’t prove a particular descendant contains one person’s address.

CNIL investigators examine an AI model genealogy network, tracing descendant paths and an isolated branch for GDPR compliance.

Lineage belongs in Arianna’s file. I record its base model, adapters, fine-tunes and known derivatives, connecting them to the relevant datasets and processing purposes. When a deletion request arrives, I know where to investigate instead of emailing the vendor the technical equivalent of “boh, maybe?”

Genmod also reveals a limit. A genealogical connection shows a route through which memorized data may have persisted; further investigation must determine whether a specific item did. Public model metadata may be incomplete. Nobody currently knows which derivative open-weight models, if any, retain a particular person’s memorized data unless someone tests those models against it.

I record that uncertainty. Compliance evidence gets dangerous when someone quietly upgrades “unknown” to “cleared.”

Human review has to interrupt the machine

Arianna’s output eventually reaches a recruiter. The system record says whether it summarizes a CV, assigns a rank, recommends rejection or automatically removes the applicant. Those verbs matter more than a hundred pages of vendor marketing. Even a “recommendation” becomes a decision when recruiters accept every score while speed-running the queue before lunch.

Here’s the mechanism regulators care about. Software processes information about someone and produces a significantly affecting outcome. If that outcome is applied without human assessment, the person faces a solely automated decision. Uber’s systems monitored driver behaviour and customer ratings; fraud signals or persistently low scores then triggered temporary or permanent account deactivation. Deactivation stopped drivers earning through the platform, so the effect wasn’t theoretical or buried in a privacy policy. The Dutch authority found inadequate information for drivers about automated decision-making and no human intervention. Those findings produced the announced fine. A human-review button hidden in an admin panel solves nothing unless someone with authority meaningfully uses it.

Monique Verdier, deputy chair of the Autoriteit Persoonsgegevens, described the impact in the authority’s August announcement:

Uber has committed serious infringements. Drivers were deactivated without pardon. From one moment to the next, they no longer had any income through Uber. That's forbidden. A computer should not make decisions on its own that have major consequences for you. These decisions should have been looked at first by a human being.

Meaningful review leaves evidence. Reviewers need the inputs and enough context to spot inconsistencies, time to consider information beyond the model’s output, and authority to reverse the result without asking the algorithm to grade its own homework. The affected person needs understandable information and a usable challenge route. I log reversals, escalations, written reasons and intervention times. A workflow diagram proves only that someone can draw rectangles.

The penalty needs context. The Dutch authority put Uber’s 2025 global turnover at about €45 billion, while GDPR fines can reach 4% of worldwide annual turnover. The examined conduct occurred between 2018 and 2022. Uber says it discontinued those policies and its current process includes human reviews, safeguards and an appeal opportunity.

Uber’s objection deserves a fair hearing, especially because the full Dutch decision remained unpublished when specialist coverage examined the announcement. We still lack the authority’s detailed Article 22 analysis and fine calculation. Nobody knows whether Uber’s appeal will alter, annul or uphold the penalty. The public record also can’t show whether its current human-review process works meaningfully in practice or merely exists on paper.

The investigation began after 171 French drivers reported their experiences to the Ligue des droits de l’Homme, which complained to France’s CNIL. Because Uber’s European headquarters are in the Netherlands, the Dutch authority investigated through the GDPR one-stop-shop mechanism. I’m unapologetically pro-European, and this is useful EU coordination: people can report a problem in one member state and trigger Union-wide enforcement.

Twenty-seven disconnected digital markets wouldn’t make Europeans safer or our companies more competitive. Shared rights need institutions that can carry evidence across borders.

European AI needs evidence that travels

A European provider can improve contractual control and jurisdictional clarity. Nationality alone can’t make Arianna GDPR-compliant. “Made in Europe” should mean I can inspect the terms, understand model changes and move my data when the relationship ends. Otherwise, it’s artisanal compliance prosciutto: lovely packaging, questionable nutritional value.

My vendor review starts with the records Arianna will later need. I request retention and training terms, subprocessor details, security documentation and rights-request support. I check who can access data outside the EU and whether the provider can export it in a usable format. Then I ask what happens when a model is withdrawn, the contract ends or a regulatory restriction blocks service. Every answer enters the system record beside dated evidence, surviving even after the lawyer whose inbox held everything changes jobs. Switching providers becomes planned work, not a founder emergency over Slack at 2 a.m. A European AI champion should win this review on substance, not because procurement waved a little blue flag over the paperwork.

Mistral shows why brand, model lineage and corporate status need separate entries. The CNIL includes Mistral Medium in its open-weight genealogy demonstrator, showing that derivative relationships can be mapped. Inclusion gives investigators a trail. It isn’t a regulatory blessing.

People searching for “Mistral AI stock” deserve a straight answer: available primary material doesn’t establish that its equity is publicly traded. Official corporate disclosures and exchange listings can settle that. Model downloads, regulatory participation and similarly named financial products cannot.

I apply the same restraint to “Mistral AI valuation.” No verified current valuation appears in the available primary material, and private-company value can’t be inferred from strategic importance or model popularity. A defensible figure needs dated financing terms or an official disclosure explaining what it measures. Finance Twitter has enough imaginary cap tables.

Europe should build a shared compliance layer around records like Arianna’s: common evidence formats, cross-border regulatory access and rights that work wherever a provider is based. By 2028, I expect serious European AI procurement to require portable system records alongside security documentation.

Companies unable to produce one will discover that “trust us” is Europe’s most expensive model architecture.

Frequently asked questions

What does GDPR compliance for AI require?

GDPR compliance for AI requires a living record connecting the system’s purpose, personal data, legal roles, model version and lineage, decision influence, human review, retention, processing locations, deletion routes and remedies. Each conclusion should link to dated evidence demonstrating that the stated controls operate in practice.

Is EU hosting enough for data sovereignty?

EU hosting alone is not enough for data sovereignty. Organizations must also know who can access personal data, whether vendors reuse it, where operational and safety logs are retained, which copies exist in retrieval or training systems, and whether deletion requests reach models and descendants.

Can an EU AI Act compliance checker ensure GDPR compliance?

An EU AI Act compliance checker can support initial scoping by identifying intended uses, possible classifications and documents to investigate. It cannot establish a controller’s lawful basis, discover what employees entered, verify meaningful human review or reconstruct processing events without system records and supporting evidence.

Sources

Related reading

Luca

Luca

Luca by the way is the personal blog of Los Angeles based entrepreneur Luca Capula. A true Italian who lives between Torino and LA.

More posts →