Evaluator disagreement is information
Google Research has published work asking how many human raters an artificial intelligence benchmark needs. The question matters because many model evaluations rely on people to judge qualities that cannot be reduced to exact matching: helpfulness, factuality, relevance, safety, style, or the quality of an explanation.
Adding raters can improve reliability. But disagreement is not always noise that disappears when the sample grows. Sometimes it is the finding.