Microsoft's February 19 introduction of Muse presents an artificial intelligence (AI) model capable of generating video-game visuals and controller actions. It is trained as a World and Human Action Model (WHAM), learning both how an environment changes and how people act within it.
Gaming is the immediate application. The larger signal is that generative simulation is moving closer to an interactive design material.
Software has spent decades waiting for people to click the buttons. OpenAI's January 23 preview of Operator presents an artificial intelligence agent that can see web pages and act through a browser on a user's behalf.
The interface is new. The organizational question is old: What is this actor allowed to do?
The first consequential defense artificial intelligence story of 2025 does not arrive as a new model or a weapons demonstration. It arrives as an invitation to test.
The Department of Defense's (DoD) Chief Digital and Artificial Intelligence Office (CDAO) begins January with a crowdsourced assurance pilot in military medicine. The setting matters. A medical system can perform impressively on average and still fail a clinician or patient at exactly the wrong moment. Its quality cannot be separated from the people, workflow, uncertainty, and consequences around it.
Originally published in 2024; substantially revised in 2026 to deepen the analysis and incorporate additional sources.
DARPA’s pursuit of AI systems that can be trusted raises a deceptively difficult question: what, precisely, are we claiming when we call an AI system trustworthy?
Trustworthiness is often presented as a list of desirable attributes—reliability, robustness, explainability, fairness, security, safety, accountability. Those attributes are important, but a list is not an assurance argument. A system can perform well on an aggregate benchmark and fail under a mission-relevant distribution shift. It can produce an explanation that sounds coherent without helping a user detect error. It can satisfy a formal control while leaving responsibility fragmented across organizations.
Trustworthy AI is not a permanent label attached to a model. It is a bounded, evidence-backed claim about how a socio-technical system behaves under specified conditions.
A new study from International Business Machines (IBM) researchers examines how machine-learning predictions might be adjusted when a domain expert's judgment conflicts with the model, particularly when a case is poorly represented in the training data. The work addresses a practical reality: experts and models often disagree for reasons that neither an accuracy score nor an appeal to experience can settle alone.
The disagreement should be treated as information.
The Defense Advanced Research Projects Agency (DARPA) is seeking technology for its Rapid Experimental Missionized Autonomy (REMA) program. The objective is to add adaptable autonomy to commercial drones so they can continue a predefined mission when communication with the operator is lost.
Loss of connection is often described as a communications problem. For an autonomous system, it is also an authority problem: what may the machine continue to do when the person can no longer supervise it?
Microsoft Research has published an overview of responsible artificial intelligence work on multimodal systems—models that analyze or generate across text, images, audio, and other forms of data. The research highlights a practical problem: risks can appear in the combination even when each input looks acceptable on its own.
Evaluation must follow the system across modalities, interactions, and real-world effects.
Anthropic has released Claude 2 with a context window that can accept roughly 100,000 tokens—enough for hundreds of pages of material in one prompt. The immediate attraction is obvious: a user can bring a long report, technical documentation, or even a book into a conversation without dividing it into tiny fragments.
More context changes what a language model can see. It does not guarantee that the model will attend to the right thing.