Article reader Listen + reading controls
Article reader
Preparing the reader…
Reading settings
AI safety needs better questions before better rules¶
The National Institute of Standards and Technology (NIST) has issued a request for information (RFI) on the safe, secure, and trustworthy development and use of artificial intelligence. Responses will inform future guidance on evaluation, red teaming, risk management, and related measurement challenges.
The request arrives after a year of rapid capability releases and equally rapid calls for guardrails. Before guidance becomes more specific, the field needs to ask more precise questions about systems, evidence, and use.
“Is it safe?” is too broad to measure¶
Artificial intelligence (AI) safety depends on the capability, application, users, environment, and consequence. A model can be acceptable for brainstorming and unacceptable for an autonomous action. A control can reduce one risk while increasing another.
Useful questions sound narrower:
- Does the system recognize when source evidence is insufficient for this decision?
- Can a user distinguish retrieved fact from generated synthesis?
- Does a tool-connected assistant remain within delegated authority under adversarial input?
- Which changes require renewed evaluation?
- How quickly can the organization detect, contain, and learn from a failure?
Those questions lead to testable claims and actionable guidance.
Measurement needs a theory of use¶
Benchmarks and red-team exercises can reveal important behavior. Their meaning depends on how the system will be used. A dangerous capability may matter even when it appears rarely. A high average accuracy may be inadequate when the errors cluster around one population or operating condition.
The NIST Artificial Intelligence Risk Management Framework establishes a useful sequence: Map the context, Measure what matters there, and Manage the resulting risk under a governing structure. Future guidance should deepen that connection rather than produce a universal score.
Every recommended metric should state the claim it supports, conditions under which it was validated, and known blind spots.
Red teaming must connect to control improvement¶
Red teaming has become a common answer to AI risk. The term can cover very different activities: testing harmful content, probing security boundaries, eliciting dangerous capability, manipulating tools, or studying user behavior.
Guidance should distinguish those objectives and require a route from finding to response. For each result, organizations should record the exploited condition, consequence, reproducibility, mitigation, residual risk, and owner. Repeated findings should alter requirements, evaluation sets, and incident planning.
A red-team report without a product and governance feedback loop is an interesting document, not a safety system.
Submit operational evidence, including failure¶
Organizations responding to the RFI can contribute more than policy preferences. The most useful input will include structured cases:
- intended use and system boundary;
- failure or threat observed;
- control applied;
- evidence used to judge the control;
- effect on users and workflow;
- new failure or burden introduced; and
- unresolved measurement question.
Negative results matter. If a popular control created warning fatigue, blocked legitimate domain language, or failed under a realistic attack, NIST needs that evidence before the control is repeated as a best practice.
Build guidance that can evolve¶
The field will learn faster than static guidance can be rewritten. Standards should identify stable principles and support modular profiles, playbooks, test methods, and public repositories that can evolve with evidence.
Wenger's work on communities of practice is relevant. Mature practice grows through a community that shares language, artifacts, experience, and judgment. NIST can provide infrastructure for that community, but practitioners must contribute the hard-won knowledge of implementation.
Use the RFI internally¶
Even organizations that do not submit a public response can use the questions as an internal audit. Bring product, security, data, human-factors, legal, and domain teams together. Identify where the organization has evidence, where it relies on assertion, and where no owner exists.
The exercise will likely reveal that the missing element is not another rule. It is a decision, measurement, or feedback loop that has not yet been designed.
NIST's request is an opportunity to slow the conversation just enough to improve it. Safe and trustworthy AI will not come from rules written at the level of the slogan. It will come from better questions, testable claims, shared evidence, and organizations prepared to revise what they believe when the system meets reality.
Sources and research trail¶
- National Institute of Standards and Technology, “NIST Calls for Information to Support Safe, Secure and Trustworthy Development and Use of Artificial Intelligence” (December 19, 2023).
- National Institute of Standards and Technology, Artificial Intelligence Risk Management Framework (AI RMF 1.0) (2023).
- Wenger, Communities of Practice: Learning, Meaning, and Identity (1998).
- Leveson, Engineering a Safer World (2011).