When models disagree a lot—measured by high ensemble variance or low...
https://wiki-room.win/index.php/What_Metrics_Should_I_Track_Besides_Accuracy_for_Risky_AI%3F
When models disagree a lot—measured by high ensemble variance or low margin—it’s a red flag that the input might be risky or unusual. We can catch those tricky cases early by monitoring disagreement scores