Model disagreement highlights inputs where predictions differ most, flagging...
https://wiki-velo.win/index.php/How_to_Measure_Reviewer_Agreement_on_Escalated_Cases
Model disagreement highlights inputs where predictions differ most, flagging potential risk. By tracking ensemble variance, we spot cases with high uncertainty. Routing the top 1-2% of these disputed cases for human review helps catch errors early