Ir para a Avaliação PREreview

Avalilação PREreview de Who judges the judges? Governance from metrics: a runtime framework for continuous LLM compliance monitoring

Publicado
DOI
10.5281/zenodo.23129468
Licença
CC0 1.0

Summary

This paper attacks what it calls the compliance fiction: the industry practice of treating regulatory conformity as a binary verdict declared at deployment time, when agentic and generative systems are dynamic — their behavior evolves with use, context, and model updates. The author introduces governance from metrics, the principle that compliance should be derived as a continuous signal from runtime observability, and presents govllm, an open-source framework implementing governance-driven routing in which model selection is determined by accumulated compliance scores rather than latency or cost alone. Central to the design is a panel of regulatory judges — small language models specialized per criterion (EU AI Act, GDPR, ANSSI, accessibility) — whose inter-judge disagreement is reframed not as noise but as a regulatory uncertainty signal warranting human arbitration. The framework is validated on a ground-truth corpus of 49 annotated prompt/response pairs across five regulatory criteria, evaluated by four small language models (1.7B–7B parameters) running fully on-premise. Agreement rates range from 51.5% (mistral:7b) to 69.1% (phi4-mini), with no single model dominating across all criteria — empirically motivating the Profile-as-jury design. The paper further documents three structural failure modes in small regulatory judges and a judge-specific position bias degrading agreement by up to 25 percentage points across question-order conditions.

Strengths

The "compliance fiction" framing is the paper's core contribution, and it is exactly right. I audit ML and AI systems for production readiness, and the gap this paper names is one I see constantly: a conformity assessment performed at time t0 is treated as a durable property of the system, while the system's behavior keeps moving. Naming it a fiction — and showing it is structurally incompatible with the EU AI Act's demand for continuous human oversight (art. 14) and lifecycle risk management (art. 9) — gives practitioners language for an argument they currently lose to schedule pressure. This framing alone justifies the paper. On-premise evaluation as regulatory necessity, not design preference. The observation that routing compliance evaluation through external APIs may itself violate the obligations being assessed (GDPR art. 44 on international transfers) is sharp and under-appreciated. Most governance tooling assumes cloud evaluation is available; this paper correctly identifies that for regulated sectors it often is not, and designs around the harder constraint. That is the kind of architectural honesty production work requires. Inter-judge disagreement reframed as signal rather than noise. The existing literature (PoLL, cascaded evaluation) treats variance across judges as something to suppress through aggregation. This paper's move — disagreement among specialized judges marks the compliance grey zone and should trigger human arbitration — is genuinely novel and operationally useful. Finding 5 (bimodal disagreement on adversarial prompts identifying the grey zone more reliably than hard prompts) is the most interesting empirical result in the paper, and it falls directly out of this reframing. The limitations discussion is admirably honest. Section 6.3 quantifies its own weakness with unusual candor: ±30pp confidence intervals at the case level, single-annotator ground truth, 3 of 24 orderings tested, lifecycle drift detection not validated on production data. A paper that tells you exactly where its numbers are soft earns trust for the numbers it stands behind. The open-source release (framework + 49-case corpus) makes the honesty checkable, which is the right combination. Checklist-based validity anchored to jurisprudential sources. Binary checklists tied to CNIL, ANSSI, and AI Act provisions are a credible attempt at making "validity" (does the judge measure compliance?) separable from "reliability" (does the judge agree with itself?). The null self-preference finding — plausibly explained by checklist framing suppressing fluency-preference mechanisms — is a worthwhile, if preliminary, contribution to the judge-bias literature.

Major comments/issues

1. The corpus is too small for the comparative claims built on it. 49 cases (10 or fewer per criterion) with the author's own ±30pp case-level confidence intervals means the headline ranking — phi4-mini 69.1% vs. mistral:7b 51.5% — sits inside overlapping uncertainty. The paper is candid about this in section 6.3, but the main text still presents judge rankings, criterion difficulty orderings (Tables 7–8), and the Profile-as-jury optimal assignment as findings rather than hypotheses. Either the corpus needs substantial expansion before these comparisons are made, or every comparative claim needs explicit hedging that survives into the abstract. As written, a reader who skips section 6.3 will come away with rankings the data cannot support.

2. Single-annotator ground truth is the deeper validity problem. The author constructed and annotated all 49 cases alone; inter-annotator agreement (Cohen's kappa) with domain experts is unmeasured. For criteria like human_oversight and non_manipulation — where the paper itself notes the compliant/violating boundary is "inherently contextual" (section 7.4) — one researcher's mapping from legal obligation to binary question is not ground truth, it is one interpretation. The checklist is anchored to regulatory texts, but anchoring is not validation. Expert annotation with reported kappa is, as the author notes, "a necessary step toward a publishable reference benchmark" — I would go further: it is a necessary step before the agreement rates (as opposed to the agreement methodology) can be taken seriously.

3. Position bias is measured on 3 of 24 possible orderings. The "conditional robustness" narrative — phi4-mini immune to reversal but collapsing under permutation — rests on a single permuted ordering (q2→q4→q1→q3). With the full permutation space unexplored, we cannot know whether the observed pattern reflects a general property of the judge or an idiosyncrasy of that one permutation interacting with those specific cases. The finding is intriguing but should be labeled exploratory; the three-ordering design is a pilot, not a characterization.

4. Two of the six claimed contributions lack empirical validation. The governed qualification lifecycle (test → human gate → production → quarantine) and trajectory-based routing are described as implemented and operational but "not empirically validated against production data in this study." A framework paper can certainly propose unvalidated components, but they should then be framed as architectural proposals, not contributions on equal footing with the validated judge panel. The conclusion's claim list should distinguish validated from proposed.

5. The compliance gate needs an error-rate analysis. The per-use-case minimum score threshold automatically excludes underperforming models from routing — "without human intervention on every routing decision." Given judge agreement rates of 50–80%, the gate will produce false exclusions (compliant models routed away, with real latency/cost/availability consequences) at some rate the paper never quantifies. A policy-as-code mechanism that acts on a statistical signal needs its own false-positive analysis. What is the gate's precision, and who reviews its exclusions?

6. No cost or latency analysis of panel judging at production scale. The architecture runs a panel of judges over production interactions to produce continuous compliance scores. Four small language models per interaction, on-premise, is not free — and the paper's own routing proposal would multiply this across use cases. The evaluation literature the paper cites (PoLL) explicitly trades accuracy against a 7x cost reduction; this paper makes no analogous accounting. For a framework whose selling point is operational governance, the operational cost of the governance itself is first-order information. Even rough per-interaction inference costs on the reported hardware would help.

7. The "parameter count is a poor proxy" claim is n=4. Pearson r=-0.39 across four models, with the author's own "indicative only; no statistical significance is implied" caveat, does not support a general claim about scaling and governance. It is consistent with the judge-side finding (phi4-mini beating mistral:7b), and the direction is plausible, but the paper should present this as a suggestion motivating larger-panel studies — which, to be fair, section 7.3 partly does — rather than as an established result.

Minor comments/issues

- The checklist authorship problem (section 7.4) deserves a concrete next step: the paper could release the annotation guidelines alongside the corpus so independent experts can replicate or contest the mappings, turning a limitation into a community process.

- The Incoherence-B false-positive inflation from negation constructions is disclosed honestly; the 29-marker filter reducing phi4-mini's transparency false-positive rate from 21.2% to 13.8% is good practice. But it also suggests the incoherence metric is measuring prompt-phrasing artifacts as much as judge incoherence — worth a sentence on construct validity.

- The French-language cases (n=4 of 49) are insufficient to support any bilingual claim; the paper is careful not to overclaim here, but the "intentionally bilingual" framing in section 6.3 slightly oversells n=4.

- mistral:7b's order-insensitivity at 51.5% global agreement is characterized as "structural rather than positional" — elegant, but at that agreement level it may simply be a floor effect. A chance-agreement baseline per criterion would clarify whether mistral is judging badly or barely judging at all.

- The geographic self-preference discussion (section 7.3) is thoughtful future work; given the panel spans US/FR/CN model families, even a descriptive (non-inferential) look at the existing judge-by-generator matrix for geographic patterns would have been a nice addition, though I accept the statistical power is not there.

Overall assessment

A valuable and timely paper on a real problem — the compliance fiction it names is something I have watched organizations live inside, declaring conformity at deployment while the system drifts out from under the declaration. The governance-from-metrics principle, the disagreement-as-signal reframing, and the on-premise-as-necessity argument are genuine contributions, and the limitations section is a model of candor. My major comments ask for hedging or expansion commensurate with a 49-case single-annotator corpus, empirical grounding for the lifecycle and routing components, and the operational analyses (gate error rates, panel inference cost) a deployment-oriented framework needs. I would be glad to see this published once those are addressed.

Competing interests

The author declares that they have no competing interests.

Use of Artificial Intelligence (AI)

The author declares that they used generative AI to come up with new ideas for their review.