Charts comparing model performance across segments

2 April 2026

Offline metrics that survive a sceptical review

We often see teams present a single F1 score and expect clearance for a pilot. Reviewers then ask which customer segment fails, how the holdout was stratified, and what happens when the class balance shifts next month.

A review-ready pack names the decision the model supports, lists the cost asymmetry of errors, and shows performance on at least three operationally meaningful slices. If a slice is thin, say so — hiding thin cells is worse than admitting uncertainty.

None of this requires new platforms. Spreadsheets and carefully named notebooks are enough if the narrative is honest.

Back to field notes