Model disagreement helps spot tricky cases by measuring how confused models...
https://deansinspiringperspective.hexaforgey.com/posts/objective-mismatch-examples-between-sensitivity-and-balanced-accuracy
Model disagreement helps spot tricky cases by measuring how confused models are, using metrics like entropy or ensemble variance. When the disagreement is high, say the top 1-2% of inputs, those cases get flagged for expert review or targeted labeling