# AI verification bench — 3.4 Six blind-labelled candidates cover access evidence, capacity arithmetic and data interpretation. They were authored during the 2026-10-03 build, including a false authorisation claim and an unsupported capacity improvement. They are recorded examples, not live calls. Imported outputs are displayed as escaped text with separately declared claim/value fields. The deterministic rules inspect those declared fields and citation IDs; they do not infer meaning from prose. **Run current checks** records citation structure, a capacity reference check where applicable, and a declared red flag. Display/UI, source and independent-review judgements remain separate. **Record review levels** requires a reviewer and rationale; a candidate author cannot self-record independent review. Reviewer identities remain self-declared. The app deliberately says “Complete local review record”, not universally verified, and produces no cross-task ranking. Capacity reference: alpha = min(1, budgetA / 30, 32 / 16), passenger = 4 alpha, freight = alpha, in movements per representative hour. At budgetA 24 the reference is 3.2 and 0.8; raising the already nonbinding B resource alone cannot improve it. The test is a teaching arithmetic check, not a train-path model. **Save new rubric and test version** invalidates old current comparisons while preserving their exact candidate/rubric/test snapshots and results. Imports cannot forge a deterministic result; validation recomputes it from its original snapshot. Compared with a star score, the bench shows missing kinds of evidence. Imported-output evaluation is functional; live batch calls, costs, representative evaluation and authenticated independent reviewers remain outside this first version. Tests include the hand reference, severe counterexamples, separated levels, stale comparisons, forged results, inert HTML and reopen.