Commentary

Interpretability as compliance evidence

Bruno Coelho Reliability & Assurance

What happened. A pre-registered study ran 15,840 defensible analytic choices across one model (GPT-2) and one task. The EU AI Act Annex IV style statement you would write from the result flipped across 73% of them.

Why this is relevant. Standardising the single most influential choice, the evaluation metric, left 59% still flipping. The circuits two competent analysts found overlapped by 4%. The evidence fails the authors’ own “filability” criterion at every tolerance a conformity assessment body would plausibly accept.

The question to ask. Have two people (not AI) run the same interpretability analysis independently and compare the statements they produce. If those differ, you have a reproducibility problem even before a compliance problem.

What is next. The study was done with a “small” model; further studies with “bigger” models are needed in addition to the interpretability mechanisms.