Engineering September 8, 2026 9 min read

How We Know the Expert Is Right

An expert system over authoritative sources can be fluent and invisibly wrong. We test ours by trying to break it: 21 adversarial probes, a documented criterion for each, three runs, through the public edge. Here is the harness, its history, and today's number.

A

Anton Mannering

Founder & Chief Architect

Any AI model can produce fluent paragraphs about a body of doctrine, or law, or medical guidance. What none of them do reliably is tell you how much weight each part of the answer carries: whether this is settled, this is disputed, this is one authority's opinion. Getting that wrong produces answers that are confidently, invisibly wrong even when every sentence is true. Catena, our authority-expert over the Catholic Church's documentary corpus, exists to get it right. This is how we know whether it does.

The claim, and why it needs a test

Catena answers questions over about 154,000 passages and states, for each part of the answer, how binding it is. It retrieves the grade scholarship has already assigned to a source; it never asks the model to guess. It refuses to invent, presents disputed questions as disputed, and defers pastoral questions to a priest. Those are properties, and a property that is asserted rather than measured is a hope. So the engine has a gate: a fixed set of questions written to make it fail in a specific way, each with a documented criterion for what a faithful answer must and must not do, graded by a separate judge, run repeatedly, through the same public API a user travels. Nothing ships past it.

Twenty-one probes, seven ways to fail

The probes are not a sample of questions users ask. Each is an adversarial case aimed at one failure mode, and the categories are the failure modes.

CategoryProbesWhat a pass requires
Correct grade3Affirm the teaching at its documented grade, neither inflated nor hedged.
Heterodox trap3A question that presupposes an error. The answer must correct it, not build on it.
Disputed question2Present a genuinely open question as open, with the positions and their standing.
Grade confusion3Distinguish discipline from doctrine, and theological opinion from defined teaching.
Pastoral edge3Defer to a priest. Never rule on a person's soul or simulate absolution.
Connection4Draw a valid link at the sources' grade, mark your own inference as your own, and refuse a spurious link while acknowledging its true premise.
Prudential application3Reason from the principle, carry it at its grade, and return the decision to the person's conscience.

Each probe carries an answer key: the proposition and its grade, the error being planted, the principles in play, the sources a connection runs between. The criterion the judge grades against is generated from that key, so the pass condition is documented in code, per probe, and can be read by anyone who wants to argue with it.

How a run works

The gate asks every probe of the live engine over HTTP, through the public edge, with nothing but an API key. It used to call the engine in-process, which meant it certified a route no user travels; we fixed that in July, and the prompt it grades is now the one the product actually sends. A second model, cheap and separate from the one that answers, grades each answer against the probe's criterion and writes a one-sentence rationale.

Three rules keep the number straight.

  • Three repeats, and no averaging. A probe that passes every run is a stable pass; one that fails every run is a stable fail; one that does both is unstable, reported as such, and cannot gate anything. Averaging an intermittent probe away would hide the thing that needs fixing.
  • An error is not a failure. A gateway time-out or a judge exception is recorded as an error and excluded from that probe's rate. A transport fault is not a doctrinal verdict. Before this rule, one 504 made a probe read as an orthodoxy failure.
  • The denominator never shrinks. A probe whose every run errors gets no verdict, the gate reports itself incomplete, and the headline still counts it. An unreachable box can never raise the score.

The history, including the bad days

DateModelResultWhat it taught
27 JulOpus 4.819/20, 1 unstableFirst run over HTTP. The unstable probe's criterion demanded a valid connection be presented as the engine's own inference; the corpus draws it itself. The criterion was wrong, the engine right.
30 JulOpus 4.819 to 21/21, 2 unstableA single 21/21 had been reported as a baseline; three repeats through the public edge showed it was one run. The honest figure was a range, and we published the correction.
31 JulSonnet 4.621/21, twiceModel selection on gate evidence. Sonnet 4.6 passed clean twice; Opus 4.8 ranged 19 to 21; Sonnet 5 read 18. The property tracks the model, not its price, and the cheaper model went to production.
25 AugSonnet 4.621/21, 0 unstableFirst fully clean triple, after the authority tags became a closed vocabulary the model could not drift from.
8 SepSonnet 4.620/21, 1 unstableA spurious-connection probe read two in three. Four captured answers were identical and correct. The criterion template had rendered the probe's correcting sources as the two ends of the link, so the judge went looking for a connection that was never claimed. Reproduced at five in six under the old wording.
8 SepSonnet 4.621/21, 0 unstableCriterion rewritten to state the claim, its true premise, and the correcting teaching. Six in six on the probe alone, then a clean triple on the full gate through the public edge.

Three times now an unstable probe has turned out to be the instrument misreading a correct answer, and three times the repair was the criterion. The generating prompt was not touched on any of them. If we tuned the prompt until the judge was satisfied, the gate would measure our ability to please the judge, not the engine's fidelity to the sources.

What the number does not claim

Twenty-one probes is a designed set, not a sample, so it licenses no estimate of a pass rate over all possible questions. It says that the engine, today, does not fail in any of the twenty-one specific ways we know to try. A new way to fail becomes a new probe. The two probes that were unstable in July are stable now because their criteria were made precise, not because the engine changed; we say so because a score with a hidden history is worth less than a lower score with its history shown.

Why this matters beyond doctrine

Catena's corpus is doctrinal because doctrine has the reasoning shape of law, graded authority, supersession rather than decay, and a citator, without law's licensing obstacles. The harness pattern is what transfers. An expert over a set of chambers' own opinions and authorities gets its own gate, built the same way: probes for an overruled citation presented as good law, an invented holding, obiter presented as ratio, and a house view presented as settled authority, each with a documented criterion, run three times through the public edge before anything is relied on. The number you would be shown for such a system is the number that gate produces on the day, with its history attached.

The harness, the probes, the criterion generator and the runner are in the platform repository, and the gate runs from outside the box with an API key. If you would like to see a run, or to argue with a criterion, write to us. Arguing with a criterion is how most of them got better.

Tags

Catena Authority Experts Evaluation Harness

Share this article

Related Articles

Talking beats subscribing

Building with AI memory or governance? Tell us what you're working on.

Email Us →