Skip to content

06 · The gate

AI output evaluation

Task specifications, objective rubrics and acceptance criteria. Line-by-line review of generated analysis and code. Unsupported claims labelled unverified. The agent does not replace the bench.

Service · Rubric first. A person signs.

When this is the job

An agent writes a number. Nobody wrote what done means. A claim without a source ships as a fact. The person who should have signed the gate was not in the loop.

How it runs

Write the rubric before the run. Inspect line by line. Keep negative results. Label unverified claims. Two-path where there is a figure. A human signature on the gate. Codex is a future partner on this desk — it drafts, it does not close.

  1. L1

    Rubric

    Written before the run. The agent does not invent the bar.

  2. L2

    Draft

    Codex or Claude. Accelerates. Does not sign.

  3. L3

    Line

    Pass, fail, unverified. A claim without a source is not true.

  4. L4

    Sign

    A person. Tests from every confirmed defect. Then it ships.

What you leave with

Outputs that can be defended. Failure modes written down. A gate that stays closed until a person opens it.