This is a recording of a real run.
Every event below was emitted by the agent, every number was computed by it,
and the raw JSON further down is exactly what HumanStandard returned. Step
through it with the controls at the bottom — nothing moves on its own
— or read the memo it produced.
Plan
waiting to start
Credits—Runtime—
Live
The agent states its plan, then works
through it. Every escalation carries its reason.
What HumanStandard returned
The stream above says “16/16 scored”. It does not
say what any of the sixteen were. These are the verdicts themselves, worst
first — one row per real API response.
Track
Tier
Verdict
Confidence
AI score
Attribution
The response, in full
Click any row above to read exactly what came back for
that track. Nothing is reordered or trimmed.