Capture · Review · Replay

A passing check deserves an explanation.

Turn a bad LLM response into a reviewed regression case. Inspect the answer, the judge's evidence, and the rule that produced the verdict.

SYNTHETIC CASE EXPLORER Compare a literal fixture with a recorded OpenAI evaluation. No provider calls or API key required when using this explorer.

Pick a response

The same policy, with different ways to get the answer right or wrong.

Loading cases…

Inspect a case

APPROVED TEACHING POLICY

Runner verdict
Policy label

Judge evidence

01 / Keep the failure

Capture inputs, output and context in your own SQLite workspace. Flag a response worth reviewing.

02 / Review the expectation

Write criteria a colleague can inspect. Approve the YAML before it becomes a check.

03 / Challenge the judge

Rerun your application. Inspect false passes and false failures, not only an aggregate score.

What this demo proves, and what it does not

Twelve selected fixtures demonstrate the local runner, structured evidence parser and aggregation rule. The literal judge rejects valid paraphrases and can accept contradictions. The recorded OpenAI mode shows one completed twelve-request evaluation from 23 September 2026, not a live service or a repeated benchmark. Policy labels were authored for these examples; this is not a population benchmark or a general estimate of semantic-judge accuracy.

The local product can run a live OpenAI judge with your own environment key. This public explorer has no key and does not accept customer prompts or uploads. It is not the private local pilot dashboard.

Privacy and deployment

No accounts, cookies, analytics scripts or visitor prompt storage are used by this application. The hosting provider may process ordinary request logs. Render's free service may sleep when idle; its filesystem is temporary. Demo state is reconstructed from bundled synthetic fixtures.