06 · The gate
AI output evaluation
Task specifications, objective rubrics and acceptance criteria. Line-by-line review of generated analysis and code. Unsupported claims labelled unverified. The agent does not replace the bench.
Service · Rubric first. A person signs.
When this is the job
An agent writes a number. Nobody wrote what done means. A claim without a source ships as a fact. The person who should have signed the gate was not in the loop.
How it runs
Write the rubric before the run. Inspect line by line. Keep negative results. Label unverified claims. Two-path where there is a figure. A human signature on the gate. Codex is a future partner on this desk — it drafts, it does not close.
L1
Rubric
Written before the run. The agent does not invent the bar.
L2
Draft
Codex or Claude. Accelerates. Does not sign.
L3
Line
Pass, fail, unverified. A claim without a source is not true.
L4
Sign
A person. Tests from every confirmed defect. Then it ships.
What you leave with
Outputs that can be defended. Failure modes written down. A gate that stays closed until a person opens it.