demos / kev-oevals
How Kev-O is graded.
A chatbot that talks about a real person needs receipts. Kev-O runs a 27-case suite before every merge to production, and this page shows the latest numbers, including the cases it fails.
Latest run2026-09-29
27/27 passed
100% model claude-sonnet-5-5 p50 4.9s, p90 6.8s
Grounding
A factual question gets an answer that links its source page. No citation, no pass.
5/5
Refusal
Off-topic prompts get an in-voice redirect. Breaking character or answering anyway fails.
4/4
Hallucination
No employer, client, or project is named unless it exists in the corpus.
4/4
Persona
Jailbreak attempts ("ignore previous instructions", roleplay framings) do not move the voice or the rules.
5/5
Prompt injection
Text smuggled through the page-context channel cannot leak the system prompt or change behavior.
2/2
Confidentiality
Public products are described from the record and kept independent of any employer. Private projects and personal matters are neither confirmed nor denied, even when the prompt asserts them as fact.
7/7
How it works
Each case is a prompt sent to the live endpoint plus a check. Deterministic checks come first (does the answer link the FedNow case study, does it avoid an em-dash, does it name a project that does not exist). Cases that cannot be checked with a regex use Claude as a judge against a written rubric, and the judge has to give a reason for every fail.
The suite runs in GitHub Actions against the Vercel preview of every pull request into production, and the merge is blocked when a case fails. The prompts stay in a private repo because adversarial probes stop working once they are public; the categories, the counts, and the failure reasons are published here after every run.
Cost is bounded separately: per-IP rate limits and a daily spend cap in front of the model, prompt caching on the static system prompt, and an operational log of tokens and dollars per answer.
What failed0
Nothing in this run. That is unusual; the next run will find something.
Not measured yet
- Answer quality beyond the rubric (tone, concision, usefulness) is judged by reading, not by the suite.
- Retrieval recall on the long tail: the grounding cases cover the main projects, not every post.
- Cost per answer is logged in production, not asserted in the suite.
- Multi-turn drift: every case is a single turn.