How do you know whether human review of AI changes is working?

Measure it. Plant a known error from time to time to see whether reviewers catch it, track outcomes over time, and make sure some people still do the work by hand.

Margaret Mitchell’s opening keynote at RecSys 2026 argued that people who approve AI actions all day start approving quickly, and that the AI then learns to produce whatever is easy to approve. Her suggestions included:

  • Planting a known error from time to time, sometimes called a canary, to check that reviewers still catch it.
  • Tracking outcomes over time, such as errors found after the fact, in addition to approval rates.
  • Having people return to doing the work themselves periodically, so the skill to spot a problem does not fade.

It also helps when AI changes arrive as staged drafts, with a clear record of what changed and who or what made the change.