Field Guide · Judgment and ops
Name the failure modes
Common ways AI help goes wrong at work - and how to notice them early.
Failure is specific
Vague fear (“AI is unreliable”) does not help a team. Name the mode. Named modes get checks. Unnamed fear gets theater.
This chapter is an ops vocabulary for postmortems and standups. Use the labels in tickets and retros so the lesson travels.
The modes
Confident wrong
Fluent prose, broken fact. Numbers, names, dates, and legal-sounding citations are the usual casualties.
Early signs: Nobody can point to a source. The answer “sounds right.” Urgency to paste is high.
Counter: Require a citation or primary check for every claim that can hurt someone. Prefer “I don’t know” over a polished guess.
Silent omission
Looks complete. Skipped a constraint you stated (or should have stated). Policies, edge cases, and “do not” lists vanish quietly.
Early signs: The draft is shorter than the brief implied. A stakeholder says “but we also need X” after you thought you were done.
Counter: Convert constraints into a checklist. Diff the brief against the output line by line before review.
Tool theater
The workflow called a tool (search, calculator, ticket system) and then ignored the result. The write-up still sounds confident.
Early signs: Tool output and final answer disagree. Logs show a call but no one can explain what changed.
Counter: Paste the tool result into the review note. If the narrative conflicts, stop and reconcile.
Scope creep
You asked for a draft. You shipped a policy, a promise, or a customer commitment.
Early signs: Language shifted from “proposal” to “we will.” Recipients treat the draft as binding.
Counter: Label artifacts: draft / recommendation / decision. Only decision owners send decision language.
Fluency debt
The text is smooth. The team no longer understands the system it describes. Next week nobody can maintain it.
Early signs: “The model wrote it” is the only explanation. Onboarding docs are model dumps.
Counter: Require a human outline before a long generation. Keep a short ops log of what actually changed.
Citation cosplay
Looks sourced. Links are wrong, behind a paywall nobody checked, or invented. Worse when the topic is contested.
Early signs: URLs 404. Titles and authors do not match. Preprints presented as settled fact.
Counter: Open every citation. Label preprints. Prefer primary agencies and peer-reviewed work for Evidence Challenges. See site citation hygiene in daily ops.
Early warning signs (any mode)
- Nobody can explain the last edit
- No one can find the source for a number
- The workflow only works for the person who invented it
- Review is “looks good” with no checklist
- Speed pressure replaces verification
Response playbook
When a mode fires:
- Slow down - freeze outbound until the risky claim is checked
- Add a check - one human, one primary source, or one tool reconciliation
- Shrink the blast radius - unsend, correct, or narrow who received the draft
- Document once - one paragraph in the team log: mode, miss, new rule
- Inherit the lesson - update the team charter or Guide link, not just chat memory
Practice
Pick a recent AI-assisted miss (yours or a public one). Name the mode. Write the five-step response as if it were your ticket. Share the label with your team so the vocabulary sticks.