Grade by hand

Ten tickets, one confusion matrix

A classifier tagged each support ticket "urgent" or not. The gold label is the human's call. Flip a prediction and watch every score move.

Click a ticket to flip the model's prediction between urgent and not urgent. Its cell (TP / FP / FN / TN) and all four scores update live.

Confusion matrix

Scores

Accuracy0.00
Precision0.00
Recall0.00
F10.00

Why accuracy lies

The lazy baseline that never says "urgent"

A model of 20 tickets that always predicts "not urgent." Slide how many are truly urgent and watch accuracy stay high while recall sits at zero.

Accuracy
0.00
looks great
Recall
0.00
catches nothing

Precision here is 0 / 0 (the baseline never predicts "urgent," so there is nothing to be right about), and it is guarded to 0.00. When zero tickets are urgent, recall is 0 / 0 too, also guarded to 0.00. Same rule as the code: an empty denominator scores 0.00, never a crash.