Skip to content
Article reader Listen + reading controls
LISTEN + READ YOUR WAY

Article reader

Preparing the reader…

0:00 0:00
Reading settings
Text size
100%

AI Safety Needs Public Measurement Infrastructure

The early debate over funding the U.S. Artificial Intelligence Safety Institute could be read as an ordinary appropriations story: lawmakers sought initial resources so NIST could recruit specialists, convene a consortium, and begin work on AI evaluation and safety standards.

The more consequential question is what kind of institution the country expects to build.

An AI Safety Institute should not be a policy office that comments on company evaluations, nor a consortium whose consensus is defined by its largest members. It should function as public measurement infrastructure: an institution capable of developing independent methods, testing important claims, making results reproducible, and creating a shared technical basis for decisions that markets cannot supply on their own.

NIST formally launched the AI Safety Institute Consortium in February 2024 with more than 200 participating organizations from industry, academia, government, and civil society. Its initial priorities included red teaming, capability evaluation, risk management, synthetic-content provenance, safety, and security. NIST later articulated a strategic vision centered on advancing the science of AI safety, disseminating safety practices, and supporting institutions and coordination.1

Those are ambitious objectives. They require durable scientific capacity, not episodic consultation.

Measurement Is a Public Good

AI developers have strong incentives to evaluate their systems. They also face incentives to choose favorable metrics, withhold sensitive failure information, define the comparison, or prioritize risks that align with product plans.

Independent public measurement creates value that no single firm can fully capture:

  • common definitions and test methods;
  • reference datasets and environments;
  • interlaboratory comparison;
  • protocols for uncertainty and reproducibility;
  • neutral analysis of competing claims;
  • and public confidence that evidence is not solely produced by the seller.

This is familiar territory for NIST. Measurement science underpins commerce and safety because organizations need shared ways to determine whether a claim means the same thing across products and laboratories.

AI complicates the task. The object being measured may be nondeterministic, rapidly updated, sensitive to prompting and context, and deployed as part of an adaptive socio-technical system. The correct response is not to abandon measurement. It is to build methods that state their scope and uncertainty honestly.

The Institute Must Distinguish Four Kinds of Evidence

AI-safety debates often collapse different questions into a single “evaluation” score. A mature institute should distinguish:

  1. Capability evidence: What can the model or system do under specified conditions?
  2. Propensity evidence: How likely is it to exhibit a behavior across realistic interactions, users, and environments?
  3. Control evidence: Which technical and organizational measures prevent, detect, constrain, or recover from harmful behavior?
  4. Consequence evidence: What harm could occur in the deployed system, given actual access, incentives, dependencies, and human behavior?

A model may possess a hazardous capability but rarely express it under normal controls. A seemingly modest capability may create severe consequences when connected to tools, sensitive data, or high-scale automation. Safety decisions require the chain, not one benchmark.

The NIST AI Risk Management Framework already emphasizes contextual mapping, measurement, governance, and management. The institute can deepen the measurement layer while preserving the framework’s socio-technical boundary.

Consortium Participation Is Valuable—and a Governance Risk

Industry participation is indispensable. Frontier developers possess systems, compute, data, and operational knowledge that public researchers cannot easily reproduce. Civil society and independent researchers bring different threat models and affected-population perspectives. Government agencies understand public missions and authorities.

The consortium’s strength is access to all of those forms of knowledge. Its risk is capture.

To preserve legitimacy, the institute should make clear:

  • how research priorities are set;
  • which members contribute evidence and which influence decisions;
  • how conflicts of interest are disclosed and managed;
  • when methods and results will be public;
  • how smaller organizations participate without being out-resourced;
  • how dissent is documented;
  • and which evaluations NIST performs independently.

Consensus is not always the right output. If methods produce disagreement, the institute should expose the source of that disagreement rather than negotiate it into a vague standard.

Reproducibility Is an Architecture Problem

Reproducing an AI evaluation requires more than publishing prompts. Results may depend on model version, system prompt, sampling parameters, tool permissions, retrieval sources, content filters, hardware, orchestration code, and provider-side changes.

Public measurement infrastructure needs an evaluation provenance model that records:

  • model and endpoint identity;
  • date and version;
  • configuration and policy layers;
  • input construction and transformations;
  • tool and data access;
  • evaluator model or human protocol;
  • scoring method and uncertainty;
  • repeated-trial distribution;
  • and material environmental dependencies.

Where providers cannot expose proprietary internals, they should still support cryptographically and procedurally credible identification of the evaluated artifact. Otherwise, neither regulators nor customers can know whether later behavior belongs to the system that was tested.

The Institute Should Build Evaluation Supply Chains

No central institution can test every model, domain, language, population, and application. NIST should develop a distributed evaluation ecosystem.

That ecosystem needs:

  • accredited or otherwise qualified independent laboratories;
  • common protocols and proficiency testing;
  • secure mechanisms for evaluating sensitive models and risks;
  • shared tooling and reference scenarios;
  • funding for academic and civil-society evaluators;
  • pathways for domain regulators to add sector-specific requirements;
  • and reporting formats that allow evidence to travel into procurement, authorization, and oversight.

The institute’s role is partly to evaluate and partly to make trustworthy evaluation possible elsewhere.

This is especially important for government agencies. They need to distinguish general model evidence from application evidence. A foundation model may pass a national evaluation and still be unsafe for a particular benefits decision, intelligence workflow, clinical context, or critical-infrastructure control system.

Funding Determines Independence

An institution dependent on voluntary industry contributions for access, compute, or personnel may hesitate to pursue questions that impose costs on contributors. Stable appropriations are not administrative overhead; they are part of the institute’s independence.

Funding must support:

  • permanent scientific and engineering staff;
  • secure compute and evaluation environments;
  • access to frontier and open models;
  • grants and contracts for independent researchers;
  • protected handling of vulnerability and incident information;
  • rapid technical response to new capabilities;
  • and sustained participation in international measurement work.

The scale should reflect the role. AI systems attract investments measured in billions of dollars. A public institute charged with assessing consequential claims cannot be expected to operate as a small coordination office and remain technically authoritative.

Standards Should Preserve Learning

Prematurely rigid standards can freeze immature metrics and encourage superficial compliance. The institute should separate:

  • exploratory research methods;
  • recommended evaluation practices;
  • stable technical standards;
  • and regulatory or procurement requirements.

Methods should mature through repeated application, interlaboratory comparison, adversarial challenge, and evidence that they predict real-world behavior. Standards should define minimum comparability without pretending a passing result eliminates risk.

This staged model allows the field to learn while giving decision-makers progressively stronger evidence.

The Strategic Inference

AI safety is often framed as a race between technological development and regulation. Measurement infrastructure changes the terms. It creates a shared layer through which developers, customers, researchers, regulators, and the public can test claims and learn from failures.

That layer is pro-innovation when it distinguishes real capability from marketing, makes procurement evidence more portable, and reduces duplication. It is pro-safety when it exposes uncertainty, funds independent challenge, and gives institutions a basis for constraining uses that lack adequate evidence.

The essential resource is not a consortium membership list. It is an independent public institution with the expertise, compute, access, governance, and long-term funding to make consequential AI claims contestable.

This essay was substantially revised in July 2026 to replace the original appropriations summary with an evidence-based institutional analysis. It incorporates the consortium and strategy announced after the original January 2024 post.

For related work on AI evaluation, governance, and Responsible AI Development Operations, see my research and portfolio or connect with me on LinkedIn.

References


  1. National Institute of Standards and Technology, “U.S. Secretary of Commerce Gina Raimondo Releases Strategic Vision on AI Safety,” May 21, 2024. 

READER-NEUTRAL SUBSCRIPTION

Follow Field Notes via RSS.

Copy this address into the RSS reader you already use. New notes will appear there automatically—no account, email address, or tracking required.