Skip to content

2026

Evaluator disagreement is information

Google Research has published work asking how many human raters an artificial intelligence benchmark needs. The question matters because many model evaluations rely on people to judge qualities that cannot be reduced to exact matching: helpfulness, factuality, relevance, safety, style, or the quality of an explanation.

Adding raters can improve reliability. But disagreement is not always noise that disappears when the sample grows. Sometimes it is the finding.

Independent evaluation needs a secure place to work

The Center for Artificial Intelligence Standards and Innovation (CAISI) at the National Institute of Standards and Technology (NIST) has entered a cooperative research and development agreement with OpenMined. The collaboration is intended to advance secure methods for evaluating artificial intelligence systems.

The agreement points at a recurring barrier to credible assurance: evaluators need access to meaningful systems and evidence, while model developers, customers, and government organizations need to protect intellectual property, personal information, security-sensitive data, and operational methods.

Technology transition begins when the prototype changes owners

The Defense Advanced Research Projects Agency (DARPA) has transferred an experimental H-60Mx Black Hawk equipped with Sikorsky autonomy technology to the U.S. Army for advanced operational testing. The transition marks the culmination of the Aircrew Labor In-Cockpit Automation System program.

It is a substantial technical milestone. It is also the beginning of a different kind of work: transferring enough knowledge, authority, and learning capacity for a receiving organization to make the capability its own.

Monitoring deployed AI is a knowledge practice

The National Institute of Standards and Technology (NIST) has published a report on the challenges of monitoring deployed artificial intelligence systems. It addresses a growing operational reality: predeployment testing cannot anticipate every combination of user, data, environment, and system change.

Monitoring is the bridge between what a team expected and what the deployed system is actually doing. Building that bridge requires more than a dashboard.

Stateful agents make memory a governance problem

OpenAI and Amazon have announced a strategic partnership that includes plans to co-develop a stateful runtime environment for artificial intelligence agents. The phrase “stateful” deserves attention. An agent that can preserve context across steps and sessions may be more useful than one that repeatedly starts from zero.

It may also accumulate assumptions, permissions, and mistakes that no one intended to become durable.

A benchmark score needs an uncertainty model

The National Institute of Standards and Technology (NIST) has published a report on expanding artificial intelligence evaluation with statistical models. Its central implication is easy to state and surprisingly easy to neglect: an evaluation result is an estimate.

Organizations often present model scores to one decimal place while leaving the population, sampling assumptions, and uncertainty largely invisible. That precision can exceed what the evidence supports.

A scientific companion should strengthen the evidence chain

Google DeepMind has described new results using Gemini Deep Think for mathematical and scientific discovery. The work is another indication that advanced models can contribute more than polished explanations: they can explore candidate approaches, connect ideas, and help experts work through difficult problems.

The most useful interpretation is not that the scientist is leaving the loop. It is that the loop itself can become richer—if the system preserves the evidence needed for expert challenge.

READER-NEUTRAL SUBSCRIPTION

Follow Field Notes via RSS.

Copy this address into the RSS reader you already use. New notes will appear there automatically—no account, email address, or tracking required.