Skip to content

Decision Support

Evaluator disagreement is information

Google Research has published work asking how many human raters an artificial intelligence benchmark needs. The question matters because many model evaluations rely on people to judge qualities that cannot be reduced to exact matching: helpfulness, factuality, relevance, safety, style, or the quality of an explanation.

Adding raters can improve reliability. But disagreement is not always noise that disappears when the sample grows. Sometimes it is the finding.

A benchmark score needs an uncertainty model

The National Institute of Standards and Technology (NIST) has published a report on expanding artificial intelligence evaluation with statistical models. Its central implication is easy to state and surprisingly easy to neglect: an evaluation result is an estimate.

Organizations often present model scores to one decimal place while leaving the population, sampling assumptions, and uncertainty largely invisible. That precision can exceed what the evidence supports.

Human-like perception is not human judgment

Google DeepMind's November 11 research shows that visual artificial intelligence (AI) models can learn to organize images more like people do. The work uses human “odd-one-out” judgments to reshape the conceptual relationships inside vision models and reports gains in human alignment, few-shot learning, and robustness to distribution shift.

That is meaningful progress. It is also a useful occasion to distinguish human-like perception from human judgment.

Expert judgment belongs in the system design

A new study from International Business Machines (IBM) researchers examines how machine-learning predictions might be adjusted when a domain expert's judgment conflicts with the model, particularly when a case is poorly represented in the training data. The work addresses a practical reality: experts and models often disagree for reasons that neither an accuracy score nor an appeal to experience can settle alone.

The disagreement should be treated as information.

Trustworthy decision support must be tested in the moment

The Defense Advanced Research Projects Agency (DARPA) has selected teams for its In the Moment program, which is exploring how machines might support difficult decisions when established rules are incomplete. Initial research focuses include mass-casualty triage and other settings where time, uncertainty, and competing values make judgment unusually demanding.

This is a serious test of human-centered artificial intelligence: not whether a model can produce an answer, but whether a human–machine team can make a better decision under pressure without obscuring who remains responsible.

Search has become a knowledge-verification problem

Microsoft has introduced a new version of Bing that combines search with a conversational artificial intelligence system. Instead of returning only a ranked list of links, it can synthesize an answer, respond to follow-up questions, and show sources alongside the conversation.

This interface is convenient because it compresses the distance between a question and a usable explanation. It is risky for exactly the same reason.

READER-NEUTRAL SUBSCRIPTION

Follow Field Notes via RSS.

Copy this address into the RSS reader you already use. New notes will appear there automatically—no account, email address, or tracking required.