Evals and Metrics
BrightOS, 2026. Built it, owned it.
How we know whether the AI is actually helping guests and staff: one set of trusted metrics, and regular reviews of real conversations.
What I did
- Replayed real conversations through a new model before switching. It dropped replies it should have sent, so we didn’t ship it.
- Caught a quality score that read 42% when the real number was 67%, so the team fixed the right problem instead of a false one.
- Set a rule: no AI-judge numbers until humans have labeled enough to check them.
What happened
Every model change is tested on real guests’ conversations before it reaches real guests.