Field Guide · Craft skills
Evaluate without theater
Lightweight eval habits without fake benchmark theater.
The trap
Teams either invent a fake leaderboard or skip evaluation entirely. Both fail. You need a small, honest loop matched to risk.
A lightweight eval
- Golden prompt - one saved input that represents a real job
- Expected checks - three things that must be true in a good output
- Fail examples - one known bad pattern (invented citation, leaked secret tone, wrong owner)
- Owner - who says ship or wait
- Cadence - when you re-run (model update day, weekly, before launch)
What theater looks like
- Benchmarks you cannot map to a customer job
- “The model scored 92%” with no task definition
- Eval owned by nobody
- Changing the rubric until the demo passes
On model-update day
Freeze critical workflows long enough to re-run golden prompts. Note diffs. Assign an owner. Silence for a morning beats a broken Friday send. See What should teams do on model-update day?.
Practice
Create one golden prompt for your highest-risk AI workflow. Run it today. Store the output and the pass/fail notes.
See also
- Pulse: What is human-in-the-loop without theater?
- Pulse: What should teams do on model-update day?
- Field Guide: Team norms
- Field Guide: Verify