Skip to content
Article reader Listen + reading controls
LISTEN + READ YOUR WAY

Article reader

Preparing the reader…

0:00 0:00
Reading settings
Text size
100%

Expert judgment belongs in the system design

A new study from International Business Machines (IBM) researchers examines how machine-learning predictions might be adjusted when a domain expert's judgment conflicts with the model, particularly when a case is poorly represented in the training data. The work addresses a practical reality: experts and models often disagree for reasons that neither an accuracy score nor an appeal to experience can settle alone.

The disagreement should be treated as information.

Models and experts know different things

A machine-learning (ML) model can detect patterns across more examples than a person can review. A domain expert can recognize context that was never encoded in the dataset: a changed procedure, unusual environment, missing variable, or rare combination that looks ordinary statistically and wrong operationally.

Both can also fail. A model may generalize poorly beyond its training distribution. An expert may anchor on a memorable case, apply an outdated rule, or express confidence without evidence.

Human-centered artificial intelligence (AI) should not assume one source always outranks the other. It should help the organization understand when each is likely to be informative.

Representation is a useful trigger

The IBM study considers how well a new case is represented in training data when deciding how much weight to give expert judgment. The idea is compelling because it connects disagreement to a reason. If a case closely resembles the data on which the model performed well, the model's prediction carries stronger evidence. If the case is unfamiliar, expert context may deserve more influence.

Representation is not the only factor. The expert's relevant experience, the consequence of error, data quality, and time available also matter. But the general pattern is valuable: reliance should be conditional, not global.

Lee and See's research on trust in automation calls this appropriate reliance. The goal is not to make users trust the system more. It is to help them trust it when the evidence warrants and challenge it when conditions change.

Design disagreement into the workflow

Many interfaces force a binary choice: accept the model or override it. A better workflow captures the structure of disagreement:

  • the model prediction and relevant confidence or familiarity signal;
  • the expert's alternative and reason;
  • missing or conflicting evidence;
  • the decision taken and accountable person; and
  • the eventual outcome, when observable.

This record supports learning. If experts repeatedly override correctly in one class of cases, the model, features, or operating boundary may need revision. If overrides are consistently harmful, training or interface design may need attention.

The override should not disappear as an unexplained exception.

Preserve expertise without turning it into folklore

Experienced practitioners often hold tacit knowledge that is difficult to formalize. Nonaka's theory of organizational knowledge creation describes movement between tacit and explicit knowledge. An AI-supported disagreement process can help: ask experts to name the cue, condition, or analogy behind an override, then convert recurring patterns into scenarios, features, rules, or training.

Not every judgment can be reduced to a rule. The purpose is to make enough of the reasoning visible that others can examine and learn from it.

Give experts an evidence obligation too

Human authority should remain clear in consequential decisions, but expertise should not be immune to challenge. Require a brief rationale for material overrides, especially when they contradict strong model evidence. Review patterns at the team level rather than punishing individuals for appropriate dissent.

Leaders should protect the ability to say, “This case is different,” while asking, “Different in what way, and what should we learn?”

The IBM research offers a useful technical method for blending prediction and judgment. The broader organizational lesson is more durable. Expert knowledge should not be bolted onto the process as a final approval. It should shape data, evaluation, interface, exceptions, and learning from the beginning.

A trustworthy decision-support system does not eliminate disagreement. It makes disagreement productive.

Sources and research trail

READER-NEUTRAL SUBSCRIPTION

Follow Field Notes via RSS.

Copy this address into the RSS reader you already use. New notes will appear there automatically—no account, email address, or tracking required.