Model evaluation needs an early-warning function
Researchers from Google DeepMind and several partner organizations have proposed a framework for evaluating general-purpose artificial intelligence models for dangerous capabilities and misalignment. Their central idea is to test for emerging risks early enough that developers can change training, security, or deployment decisions before a capability becomes difficult to contain.
That makes evaluation more than a scorekeeping function. It becomes an early-warning system.