Skip to content

Trustworthy AI

World models need worlds worth trusting

Robots and autonomous vehicles need experience. The hard question is where that experience should come from when real-world data is expensive, dangerous, rare, or incomplete.

NVIDIA's January 6 announcement of Cosmos world foundation models offers one answer: generate and manipulate simulated physical environments at scale. The company presents the platform as infrastructure for training and evaluating physical artificial intelligence (AI) systems, including robotics and autonomous vehicles.

AI disclosures in political ads are necessary—and insufficient

A label on a political advertisement can tell us that artificial intelligence helped make it. It cannot tell us whether the message is true, who authorized the representation, how materially the content was altered, or whether millions of people saw it before the label appeared.

That is the problem the Artificial Intelligence (AI) Transparency in Elections Act tried to address in 2024. Senate Bill 3875 would have directed the Federal Election Commission (FEC) to require disclosures when covered political communications contained content “substantially generated” by AI. The bipartisan proposal recognized a real gap: voters could encounter a synthetic voice, image, or video without knowing that part of the apparent evidence had never occurred.

The bill advanced out of committee and reached the Senate calendar, but it did not become law before the 118th Congress ended. The later record makes the underlying design question more useful, not less. What would an effective disclosure regime need to accomplish—and what should no one expect a label to solve?

Federal AI governance needs a memory

Federal artificial intelligence (AI) policy changed substantially between 2024 and 2025. The need to know which systems government uses, who owns them, how they affect people, and what evidence supports them did not.

That is the enduring idea behind the Federal AI Governance and Transparency Act. Introduced as House bill H.R. 7532 in March 2024, the bipartisan proposal would have consolidated several federal AI governance requirements in statute. It directed agencies to create governance charters for certain systems, strengthened the Office of Management and Budget's government-wide role, expanded public visibility, and required contractors to provide information agencies would need for oversight.

The bill advanced out of committee by a 36–3 vote and was reported to the House in December 2024. It did not become law before the Congress ended.

Its most useful contribution was not a particular form or office. It was the recognition that accountable AI requires an institutional memory: a durable connection between the system, its public purpose, the decisions made about it, and the evidence available to challenge those decisions.

Federal AI standards should standardize evidence, not freeze design

Federal agencies need a common way to show that an artificial intelligence (AI) system is understood, controlled, and worthy of use. They do not need Washington to prescribe one architecture, model class, or development method for every mission.

That tension sat inside the Federal AI Governance and Transparency Act introduced in 2024. House bill H.R. 7532 proposed a government-wide structure for AI governance, including agency charters, inventories, risk practices, workforce training, oversight, and updates to federal acquisition rules.

The bill advanced through the House Oversight Committee and was formally reported late in 2024, but it did not become law. The question it raised remains unresolved: what should federal AI standards make uniform, and where should they preserve variation?

The right answer is to standardize the interfaces of accountability—the evidence agencies retain, the decisions they document, and the signals they exchange—without freezing the technical design beneath them.

Bringing AI into Mission-Critical Systems: Unifying Mission, Systems, and Data

Originally published in 2024; substantially revised in 2026 to deepen the analysis and incorporate additional sources.

Organizations often describe AI integration as if a model were a component that could be inserted into an existing system: connect an API, provide data, expose a prediction, and declare the capability operational. That mental model is useful for demonstrations and dangerously incomplete for mission-critical environments.

Operational AI is not a model-integration problem. It is a systems-engineering problem spanning mission outcomes, software architecture, data stewardship, human judgment, security, assurance, and the mechanisms through which the system learns after deployment.

The central challenge is not getting a model to run. It is making the larger mission system trustworthy and adaptable when one of its components behaves probabilistically.

Building Trustworthy AI Systems: From Aspiration to Evidence

Originally published in 2024; substantially revised in 2026 to deepen the analysis and incorporate additional sources.

DARPA’s pursuit of AI systems that can be trusted raises a deceptively difficult question: what, precisely, are we claiming when we call an AI system trustworthy?

Trustworthiness is often presented as a list of desirable attributes—reliability, robustness, explainability, fairness, security, safety, accountability. Those attributes are important, but a list is not an assurance argument. A system can perform well on an aggregate benchmark and fail under a mission-relevant distribution shift. It can produce an explanation that sounds coherent without helping a user detect error. It can satisfy a formal control while leaving responsibility fragmented across organizations.

Trustworthy AI is not a permanent label attached to a model. It is a bounded, evidence-backed claim about how a socio-technical system behaves under specified conditions.

Scaling Trustworthy AI in Government Requires an Operating System

Originally published in 2024; substantially revised in 2026 to deepen the analysis and incorporate additional sources.

Government agencies do not lack AI ideas. They lack repeatable mechanisms for turning a promising use case into a capability that can be evaluated, authorized, adopted, monitored, and improved.

The usual response is to scale the technology: add compute, models, data pipelines, or platform capacity. Those investments matter. But when every program defines its own risk process, evidence package, human-oversight model, security interpretation, and approval path, the organization scales experimentation while preserving the bottlenecks that prevent adoption.

Trustworthy AI scales when the enterprise standardizes the work around the model—not merely access to the model.

READER-NEUTRAL SUBSCRIPTION

Follow Field Notes via RSS.

Copy this address into the RSS reader you already use. New notes will appear there automatically—no account, email address, or tracking required.