Skip to content
Article reader Listen + reading controls
LISTEN + READ YOUR WAY

Article reader

Preparing the reader…

0:00 0:00
Reading settings
Text size
100%

Charting the Pentagon’s Course: Scale AI Leading Testing and Evaluation of Large Language Models Scale AI, in collaboration with the Pentagon’s Chief Digital and Artificial Intelligence Office, is spearheading the development of a comprehensive framework for testing and evaluating large language models. The innovative one-year contract aims to deploy AI safely, measure model performance, offer real-time feedback for warfighters, and create specialized evaluation sets for military support applications. This cutting-edge initiative addresses the potential of generative AI to revolutionize military planning and decision-making. By working on developing holdout datasets, engaging DOD insiders, and automating model assessments, the partnership endeavors to enhance the robustness and resilience of AI systems in classified environments. Scale AI’s strategic approach towards testing and evaluating generative AI models will enable the DoD to harness the technology responsibly and support military applications effectively.

Scale AI: Charting the Pentagon’s Course in Testing and Evaluating Large Language Models

San Francisco-based Scale AI has received a one-year contract from the Pentagon’s Chief Digital and Artificial Intelligence Office (CDAO) to develop a comprehensive testing and evaluation (T&E) framework for generative AI.

Generative AI and Its Potential Challenges

Generative AI, including large language models that can generate text, software code, images and other media based on human prompts, is an emerging technology. Despite its potential, it poses significant challenges for the Department of Defense due to the complexities involved and a lack of universal AI safety standards. The T&E processes currently in use aim to ensure that systems, platforms, and technologies perform safely and reliably before being fully deployed. However, these processes struggle to comprehensively test and evaluate AI-based systems, such as generative models.

Aiming for a Comprehensive Framework

Under the new contract, Scale AI will work towards providing CDAO with a framework that allows the safe deployment of AI, measures model performance, and provides real-time feedback for warfighters. Specialists from Scale AI will also create specialised public sector evaluation sets to test AI models for military support applications.

Building Trustworthy AI

The contract further aims to develop “holdout datasets”, including input from DoD insiders, as part of the process of testing and evaluating large language models (LLMs). Involving DoD insiders for reviewing the responses generated by these LLMs will help ensure that the output is as reliable as humanly possible. This iterative improvement process eventually will help create models that automatically alert officials when they deviate from established standards.

Conclusion

The partnership between Scale AI and The Pentagon is a significant stride towards ensuring that AI models deployed within the defense sector perform efficiently, safely, and in a trustworthy manner. By integrating AI safely, we can improve decision-making and planning within the military, thereby enhancing our overall defense capabilities. Refer to the original DefenseScoop article for more details.

READER-NEUTRAL SUBSCRIPTION

Follow Field Notes via RSS.

Copy this address into the RSS reader you already use. New notes will appear there automatically—no account, email address, or tracking required.