Charting the Pentagon’s Course: Scale AI Leading Testing and Evaluation of Large Language Models Scale AI, in collaboration with the Pentagon’s Chief Digital and Artificial Intelligence Office, is spearheading the development of a comprehensive framework for testing and evaluating large language models. The innovative one-year contract aims to deploy AI safely, measure model performance, offer real-time feedback for warfighters, and create specialized evaluation sets for military support applications. This cutting-edge initiative addresses the potential of generative AI to revolutionize military planning and decision-making. By working on developing holdout datasets, engaging DOD insiders, and automating model assessments, the partnership endeavors to enhance the robustness and resilience of AI systems in classified environments. Scale AI’s strategic approach towards testing and evaluating generative AI models will enable the DoD to harness the technology responsibly and support military applications effectively.
Scale AI: Charting the Pentagon’s Course in Testing and Evaluating Large Language Models
San Francisco-based Scale AI has received a one-year contract from the Pentagon’s Chief Digital and Artificial Intelligence Office (CDAO) to develop a comprehensive testing and evaluation (T&E) framework for generative AI.