When models disagree on predictions, it signals uncertain or risky inputs worth...
https://online-wiki.win/index.php/Objective_Mismatch_Examples_Between_Sensitivity_and_Balanced_Accuracy
When models disagree on predictions, it signals uncertain or risky inputs worth flagging. Measuring ensemble variance helps spot these cases. By routing the top 1-2% of high-variance inputs for human review, you catch potential errors early