EU AI Act compliance checker — Can you trust the PASS?

A green score can hide stale law and self-reported claims. Learn what a checker must observe, preserve and…

EU AI Act compliance checker — Can you trust the PASS?

The short version

  • An EU AI Act compliance checker proves only what its evidence actually observed and preserved.
  • Evidence packs need legal citations, timestamps, rule versions, deployment artifacts, reviewer decisions, and unresolved questions.
  • Portable, versioned evidence lets counsel, auditors, and regulators verify claims instead of trusting a green score.

PASS lit up the dashboard in Duolingo green, as if somebody had remembered that il gufo means “the owl.” The meeting immediately decided Brussels could no longer hurt us.

Then I asked whether the EU AI Act compliance checker had inspected runtime logs, the live interface or vendor documentation. Nope. A founder had completed a form between meetings using answers from the product’s sales team. That PASS had the nutritional value of airport focaccia.

A checker can scope a system, map likely obligations and organize evidence. Its confidence should never exceed what it observed. Questionnaire checkers convert self-reported answers into a risk tier. Evidence-bound checkers connect conclusions to legal text and deployed artifacts while preserving dates and software versions. Both have a job. Only the second produces a file I would hand to counsel or a regulator without developing a mysterious cough.

A green score has a tiny field of view

A useful checker identifies my role for a specific deployment, determines the likely risk route and records evidence for every result. Typing “Claude” or “Mistral” into a form tells me almost nothing: either could draft internal emails or screen job applicants. The deployment controls the analysis—its intended purpose, affected people, accessible data, connected tools and possible actions. I must also establish whether I supplied it, deployed somebody else’s product or substantially changed its purpose. Regulation AI’s guidance on third-party systems explains why repurposing can move a deployer toward provider obligations. A company-wide label flattens this into compliance soup.

Bar chart comparing current figures against their baselines: Organisations with significant exposure… 86 % versus 14 %, Organisations reporting end-to-end… 14 % versus 86 %, reported order accuracy for the IBM… 80 % versus 95 %, Increase in available GPU KV-cache-token… 37 % versus 8 %.

The machinery should start with an inventory of deployed interfaces and capabilities, then determine my role for each system. It tests the product route in Annex I and use cases in Annex III, including employment, education, essential services and justice. Intended purpose and the exception under Article 6 still need case-specific analysis. The checker then maps duties attached to my role and transparency requirements currently in force. Every transition must retain supporting evidence because one wrong system fact can poison everything downstream. The output should separate applicable from likely inapplicable duties. If a missing fact blocks the answer, show it in red—preferably with less confetti than the PASS screen.

Dates quietly wreck competent assessments. The Commission’s AI Pact timeline says the AI Office and national authorities assumed enforcement powers on 2 August 2026, while some high-risk requirements remain scheduled for December 2027 after earlier AI Act provisions took effect. Results become stale when laws, official guidance or product tools change. Kosmoy’s market analysis recommends asking which legal version supported each historical assessment. I would preserve the consolidated regulation, relevant guidance and rule-engine version beside the evaluation date. Plenty of dashboards update their rules but leave last month’s cheerful PASS untouched, which is molto comodo for the vendor.

An honest checker also abstains. Vigilia’s published methodology limits its strongest finding to what an automated visitor observed on one page at one moment and explicitly disclaims certification:

The strongest sentence the record can support is this one: on 9 September 2026, the chat surface at aivigilia.com stated that it was an AI system before any interaction.

I trust “not observable” more than guesses.

The useful checker brings receipts

I would test competing tools on an EU-facing support agent that identifies users, reads account records, drafts replies and issues refunds after human approval. It feels ordinary. That is how software acquires extraordinary permissions while everyone debates the button radius.

With a questionnaire checker, I select “chatbot,” confirm personal data is involved, attest that users receive a disclosure, upload a policy and answer governance questions. Then I get a score. With an evidence-bound checker, I provide an interface capture and owner record, document where data goes and map every tool operation to its approval condition. The output connects claims to evidence and assigns unresolved questions to actual people. That matters when somebody asks who verified the disclosure. “Gianni clicked yes” is a weak audit trail, even if Gianni is delightful.

The action surface determines what I collect. LangGuard’s operational guidance notes that agents can read data, call tools and modify records before triggering wider workflows. An accounts-payable integration might separately read an invoice, create and approve a payment request, or change bank details. Grouping these powers under “finance integration” hides wildly different consequences. Evidence that a human approves payment requests says nothing about who can first change the bank account. The checker must trace each action from permission through approval to the resulting log. Otherwise, it documents the policy while missing the mechanism that empties the account.

The stronger checker starts with my description but treats it as an assertion needing support. It inspects the customer-facing disclosure, records vendor terms and links each permitted tool call to its approval condition. It separates what the provider supplied from what I implemented as deployer. Under the AI Act’s transparency article, an upstream capability does not prove my interface delivered the required notice. Each conclusion then links to its source, with gaps exposed. If the checker sees one screenshot and two policy files, its verdict must remain inside that tiny box. Software does not gain omniscience because the CSS looks expensive.

Provenance makes artifacts checkable. I want the collector and origin recorded with the collection method, amendment history and date. A company-uploaded PDF remains a company assertion. A timestamped capture tied to a rule version and correction record gives an auditor something reproducible.

Vigilia hashes pages, documents and screenshots, retaining superseded captures after remediation. It concedes that a hash only proves the captured file did not later change; it cannot prove the original capture was honest or complete.

A record that only ever showed the pass would be an advertisement.

NovaFabric goes further with signed, timestamped and replayable evidence bundles for agent runs. Its preprint reported complete declared-stream capture in roughly two-thirds of tested scenarios, versus the complete capture its design aims to provide. Mocked replay completed only two of ten tool-using workloads; the authors attributed failures to missing tool-response substitution. Tamper evidence preserves a receipt. It cannot inspect the kitchen or establish events beyond its trusted computing base.

EU AI Act regulation under document scanner as reviewers replace pages beside archival boxes in digitisation lab.

My compliance-checker torture test

I would give every vendor the same messy production case: third-party APIs, personal data, several interfaces and a human approval step, wrapped around an ownership boundary nobody can explain without searching Slack. Actual companies are ambiguity wearing a Notion account.

I start with inventory and role determination. I ask which exact facts block a stronger classification, then check whether each came from self-reporting or an observed event. Next, I request a sample evidence pack with legal citations, timestamps, engine versions, reviewer decisions and unresolved questions. Then I change one material detail—perhaps adding a refund tool or replacing the model vendor. A strong checker identifies stale conclusions and reruns affected rules. Counsel reviews legal judgment; engineers verify operational evidence. Accountability remains with people who understand the deployment and can sign their names.

The Commission’s voluntary AI Pact offers a sensible preparation baseline: establish AI governance, map systems likely to be high-risk and improve staff literacy. The Commission also says its pledges carry no legal obligations. An AI Pact logo demonstrates preparation, not that tonight’s controls worked—much as a framed hygiene certificate tells me little about tonight’s oysters.

AICONFORM is the European project I am watching most closely. Its planned 24-month run from the 1 June 2026 launch is a development schedule; completed compliance automation lies ahead. Regulatory duties span EU and sectoral legislation, use legal language and change over time, so AICONFORM proposes using AI and natural-language processing to extract them. It would formalize duties as machine-readable, executable rules connected through a compliance knowledge graph. A rule engine could map system facts to obligations, generate structured documentation and preserve conformity-work audit trails. Continuous monitoring would detect legal or operational changes that invalidate earlier claims. Human reviewers would remain responsible for role determination, classification and accountable decisions. Where a system poses risk or fails to comply, the Commission says market-surveillance authorities can remotely monitor it, access documentation and datasets, examine source code, request corrective action and impose penalties.

Nobody knows whether AICONFORM will accurately extract legal duties, classify messy systems or survive real-world pilots. The source describes work in progress, not results. I support the ambition because Europe needs shared compliance infrastructure. I would rather fund transparent experimentation than buy another American dashboard with an EU flag pasted into the footer.

Former European Commissioner Thierry Breton captured Europe’s digital mood in an address reported in Thierry Breton aux DSIN de l’année:

La confiance s’est effondrée.

The repair belongs at European scale. Twenty-seven incompatible evidence formats would be a painfully on-brand own goal.

GDPR, Mistral and the limits of a checker

An AI Act result cannot establish GDPR compliant AI by itself. GDPR separately requires analysis of lawful basis, purpose limitation, individual rights, retention, processor relationships, security and international transfers. Some artifacts can support both reviews: a data inventory or processing agreement may answer overlapping questions, but the legal tests remain distinct. AICONFORM reflects this by treating the AI Act, GDPR, NIS2 and eIDAS as separate sources within interoperable workflows. A checker should mark reusable evidence and route unanswered privacy questions to the right reviewer.

The boundary becomes comic with investment searches. An EU AI Act compliance checker cannot tell me whether Mistral AI stock is publicly tradable or establish the current Mistral AI valuation. Those answers require current company records, financing disclosures and reliable market reporting. The supplied evidence establishes neither, so I refuse to manufacture an answer to seduce Google. A compliance score belongs nowhere near an investment decision.

A Mistral AI vs Claude comparison can still alter the assessment because a model swap may change processing terms, logging capabilities, update policies or intended-purpose limits. The brand gets no compliance halo. I would collect relevant evidence from either vendor and reassess whenever those deployment facts changed.

I want Mistral and other European AI champions to win. I also accept the strongest sovereignty objection: keeping data in Europe offers limited protection when providers remain exposed to foreign laws such as the CLOUD Act or FISA, as Ifri has argued. Complete technological independence is unrealistic for most organizations. Selective control over critical systems, strategic partnerships and credible exit options offer Europe a practical path. European models and cloud capacity still matter because dependency without leverage is terrible industrial policy. Shared compliance infrastructure belongs in that strategy; it lowers the cost of building once and selling across the Union.

When the remaining high-risk requirements arrive, serious EU procurement teams should reject standalone green scores and demand portable evidence packs tied to legal versions. The startup in Palermo, auditor in Paris and regulator in Helsinki should inspect the same record. That requires European standards, European infrastructure and, yes, more integration—not twenty-seven governments reinventing the spreadsheet.

If Europe builds that shared language together, today’s compliance dashboard will join the airport focaccia: glossy, green and left untouched.

Frequently asked questions

What does an EU AI Act compliance checker actually prove?

An EU AI Act compliance checker proves only conclusions supported by the facts and artifacts it observed. Questionnaire results primarily map self-reported answers to likely duties, while evidence-bound results can connect legal rules to deployed interfaces, logs, vendor documents, dates, software versions, reviewer decisions and unresolved gaps.

Can an AI Act compliance checker prove AI is GDPR compliant?

An AI Act compliance result does not establish GDPR-compliant AI. GDPR requires separate analysis of lawful basis, purpose limitation, individual rights, retention, processor relationships, security and international transfers. Data inventories and processing agreements may support both reviews, but each legal framework applies distinct tests and requires its own unanswered questions to be resolved.

What evidence should an EU AI Act compliance assessment include?

An EU AI Act compliance assessment should include deployed interface captures, system ownership records, vendor terms, data flows, tool permissions, approval conditions, legal citations, collection dates, rule-engine versions, reviewer decisions and unresolved questions. Superseded evidence should remain available so auditors can reproduce the assessment and identify stale conclusions.

Sources

Related reading

Luca

Luca

Luca by the way is the personal blog of Los Angeles based entrepreneur Luca Capula. A true Italian who lives between Torino and LA.

More posts →