Article reader Listen + reading controls
Article reader
Preparing the reader…
Reading settings
A Safety Report Is Not an Assurance Case¶
Revised and substantially expanded July 17, 2026, with the later policy record made explicit.
In January 2024, the U.S. government began implementing a novel requirement from Executive Order 14110: companies developing certain powerful dual-use foundation models were directed, under the Defense Production Act, to report defined development activities and provide information about training, ownership and protection of model weights, and red-team testing.
The move addressed a real information asymmetry. Frontier-model developers could observe capabilities, incidents, infrastructure, and internal test results that the government could not readily see. Voluntary disclosure alone was unlikely to produce consistent coverage, especially when safety findings might affect competitive positioning or invite scrutiny.
But the title of the original version of this post—“AI Companies to Begin Sharing Safety Test Reports”—made the policy sound more complete than it was. Receiving a report is not equivalent to understanding a system. A red-team result is not a safety certificate. A compliance submission does not establish that a model is acceptably safe for every downstream use.
Reporting is best understood as one sensor in a larger assurance system. Its value depends on the quality of the questions, comparability of the evidence, expertise of the reviewer, protections around sensitive information, and—most importantly—the government's ability to act on what it learns.
What the 2023 Executive Order Required¶
The authoritative historical source is Executive Order 14110, issued October 30, 2023. Section 4.2 invoked the Defense Production Act and directed the Secretary of Commerce to require companies developing or demonstrating an intent to develop potential dual-use foundation models above defined technical conditions to provide information on:
- ongoing or planned development and the physical and cybersecurity protections applied to training;
- ownership and possession of model weights and measures used to secure them;
- results from relevant red-team testing, including efforts related to dangerous capabilities, software vulnerabilities, and the model's ability to evade control;
- certain large-scale computing clusters and their security measures.
The reporting concept was not a general approval regime for every AI model. It targeted a narrow set of models and infrastructure believed to warrant visibility because of scale and potential dual-use capability. The contemporary C4ISRNET coverage that prompted the original post placed the requirement within a broader discussion of domestic and international AI governance.
The policy context later changed. President Trump revoked Executive Order 14110 in January 2025 through Executive Order 14179 and directed review of actions taken under it. This essay does not treat the 2024 reporting mandate as current law. It examines the enduring design problem that the mandate attempted to solve.
The Government's Problem Was Observability¶
Frontier AI creates a governance asymmetry familiar from other high-consequence technical domains. The producer controls the system and sees rich internal evidence. The public bears some portion of the external risk but sees marketing, selected benchmark results, research publications, and occasional incident disclosures.
Markets do not reliably eliminate that asymmetry. Customers may lack negotiating power or the expertise to demand comparable evidence. Researchers may not have access to weights, training data, infrastructure, or realistic deployment telemetry. Providers may reasonably withhold details that would expose intellectual property, personal information, or security vulnerabilities. The result is not simply secrecy; it is fragmented, noncomparable knowledge.
Mandatory reporting can improve observability by establishing common triggers and information rights. Yet it introduces its own risks:
- thresholds may become obsolete as algorithms and hardware improve;
- developers may optimize against the formal trigger rather than underlying risk;
- heterogeneous tests may be summarized as though they are comparable;
- highly sensitive submissions can become valuable intelligence targets;
- reviewers may be overwhelmed by evidence they lack tools or personnel to interpret;
- a submission can create the appearance of oversight without a decision process behind it.
The policy should therefore be judged not by the number of reports collected but by whether it changes risk identification, mitigation, and accountability.
Red Teaming Answers a Bounded Question¶
Red teaming is essential, but its meaning is frequently inflated. A red team explores how a system might fail under a particular threat model, level of access, time budget, evaluator skill, tool environment, and set of prohibited or hazardous outcomes. It can discover important vulnerabilities. It cannot demonstrate the absence of unknown failure modes.
The statement “the model was red-teamed” is almost content-free unless the recipient knows:
- the model and system version tested;
- the access provided to evaluators;
- the threat actors and capabilities represented;
- the covered harm categories and operational contexts;
- tools, retrieval, memory, and permissions available to the system;
- duration, sampling strategy, and stopping conditions;
- baseline and acceptance criteria;
- discovered failures, severity judgments, and unresolved disagreement;
- mitigations applied and the evidence from retesting;
- limitations and scenarios excluded from the exercise.
The same base model can present radically different risk when embedded in a consumer chat interface, a coding assistant, a biological design workflow, or a cyber agent with network tools. A report about model-level behavior should not be silently generalized into a claim about every composed system.
NIST's Secure Software Development Practices for Generative AI and Dual-Use Foundation Models reinforces this lifecycle view. Secure development encompasses provenance, protection, verification, vulnerability response, and release practices—not a single predeployment exercise.
Replace “Safety Report” With a Portfolio of Assurance Claims¶
Government reviewers need evidence organized around decisions. A useful submission should make bounded claims and connect each claim to evidence, limitations, ownership, and required action.
For example:
Under the tested configuration and access conditions, the model did not materially increase the success rate of a defined class of novice cyber misuse beyond specified baselines.
That claim is still incomplete, but it is testable. It identifies a population, capability, configuration, outcome, and comparison. Reviewers can ask whether the test set was representative, whether experts would show a different effect, how the result changes with tools, and what model updates invalidate the evidence.
A mature assurance portfolio would cover several evidence types:
- Development-process evidence: data provenance, secure development practices, access to model weights, insider-risk controls, dependency management, and reproducibility.
- Capability evidence: structured evaluations of covered hazardous and beneficial capabilities, including uncertainty and methodological limits.
- Adversarial evidence: red-team findings, abuse testing, prompt-injection and tool-use testing, and mitigation retests.
- Deployment evidence: actual system architecture, user population, permissions, monitoring, human decision rights, and failure containment.
- Operational evidence: incidents, near misses, anomalous behavior, abuse patterns, performance drift, and the effectiveness of response.
- Change evidence: model, policy, tool, and architecture updates that may invalidate prior conclusions.
No individual document establishes “safety.” Together, the evidence can support or refute specific decisions: whether to continue training, release weights, expand access, authorize a deployment, require mitigation, or restrict a use.
Thresholds Should Trigger Scrutiny, Not Define Risk¶
Compute-based thresholds are administratively attractive because they can be measured. They are also imperfect proxies. Training efficiency improves. Fine-tuning, scaffolding, retrieval, tool access, and test-time computation can change effective capability. Smaller specialized systems may create substantial risk in narrow domains, while a larger model may be used within a tightly constrained environment.
A reporting regime should combine bright-line thresholds with evidence-based escalation. Compute or cluster size can trigger an initial duty. Additional triggers might include demonstrated dangerous capability, deployment scale, access to high-impact tools, major security incidents, or material changes in release strategy.
This layered design avoids asking a single number to carry the entire policy. Thresholds create predictable coverage; capability and context determine the depth of review.
The reporting architecture should also be revised on a known cadence. A threshold that cannot adapt becomes either a loophole or an indiscriminate burden.
The State Needs Capacity to Interpret What It Collects¶
Information authority without analytical capacity produces a compliance warehouse.
Reviewing frontier-model evidence requires machine learning, cybersecurity, statistics, domain science, intelligence, software assurance, law, and human-factors expertise. The government must be able to reproduce selected findings, commission independent tests, protect sensitive submissions, compare evidence across providers, and connect technical conclusions to policy authorities.
That suggests a federated review model:
- a central function maintains schemas, intake, secure handling, common methods, and cross-provider analysis;
- mission and domain agencies supply contextual expertise;
- national laboratories and qualified independent evaluators reproduce or extend selected testing;
- intelligence and security organizations assess threat-relevant implications under appropriate controls;
- policy owners document decisions and follow-up requirements;
- providers receive clear channels for questions, corrections, incident notification, and protected vulnerability disclosure.
The system also needs auditability. Future reviewers should be able to reconstruct which evidence was available, which experts assessed it, what uncertainty was recorded, why a decision was made, and what later information changed that decision.
Confidentiality and Public Legitimacy Must Coexist¶
Detailed frontier-model reporting may contain trade secrets, security controls, vulnerabilities, or information that could enable misuse. Full publication would be irresponsible. Total opacity would also undermine legitimacy and independent learning.
A tiered disclosure model can preserve both values:
- sensitive technical submissions remain protected and access-controlled;
- government publishes reporting schemas, coverage rules, review methods, and aggregate compliance statistics;
- providers publish standardized high-level assurance summaries with bounded claims and limitations;
- significant incidents receive timely public disclosure where lawful and safe;
- independent researchers obtain structured access to systems and evidence through governed mechanisms;
- oversight bodies receive enough detail to evaluate the effectiveness and consistency of the regime.
Public trust should not require exposing exploit instructions. It does require evidence that the government is doing more than receiving documents.
The Durable Lesson Survives the Policy Change¶
The 2024 DPA requirement was one attempt to solve a persistent governance problem: the organizations best positioned to observe frontier AI risk are also parties with commercial, strategic, and reputational interests in how that risk is described.
Reasonable policymakers can disagree about statutory authority, coverage, burden, thresholds, confidentiality, and the appropriate balance between mandatory and voluntary measures. Those disagreements do not remove the need for credible information.
Whatever policy mechanism replaces or supersedes the 2024 approach should meet a practical test:
- Does the government learn about material capabilities and incidents early enough to act?
- Is evidence comparable enough to support decisions across providers?
- Can independent experts challenge important claims?
- Are downstream system and deployment risks visible, not only base-model tests?
- Do submissions trigger documented mitigation, monitoring, or escalation?
- Can the framework adapt as technical scaling relationships change?
- Is sensitive information protected without making oversight performative?
If the answer is no, the policy has created reporting activity without assurance.
The broader principle applies well beyond frontier-model regulation. In enterprise and public-sector AI, leaders routinely receive dashboards, model cards, risk assessments, and test reports that appear authoritative but do not support a defined decision. The remedy is to begin with the claim the organization needs to make, then design the evidence and accountability necessary to defend it.
That claim-centered approach is part of my work in responsible AI development operations, technical governance, and mission-focused AI strategy. If your organization needs to turn AI documentation into a living assurance system, you can connect with me on LinkedIn or explore the engineering work at Xendev Labs.