Evals and Metrics

BrightOS, 2026. Built it, owned it.

How we know whether the AI is actually helping guests and staff: one set of trusted metrics, and regular reviews of real conversations.

What I did

  • Replayed real conversations through a new model before switching. It dropped replies it should have sent, so we didn’t ship it.
  • Caught a quality score that read 42% when the real number was 67%, so the team fixed the right problem instead of a false one.
  • Set a rule: no AI-judge numbers until humans have labeled enough to check them.

What happened

Every model change is tested on real guests’ conversations before it reaches real guests.

See BrightOS

By Kyle Blanchard. Kyle Blanchard is a Los Angeles product leader and AI consultant who builds and evaluates AI agents for hospitality, healthcare, and startups. Work with me.