The research found two machine-learning controllers with the same architecture and the same 99.8% test score could differ by 406-fold in operating cost when tested in simulated ‘real-world’ conditions.
The finding points to a simple lesson as AI moves into critical infrastructure: don't just test whether an AI gets the answer right, test what happens when it acts on that answer.
White received the international IEEE CASE Best Student Paper Award for the research, which examined machine-learning systems designed to rapidly reconfigure simulated microgrids – small electricity networks that must constantly respond to changes in electricity demand, generation and equipment.
Unlike a conventional AI system that makes a prediction and stops, these controllers make decisions that change the system they are controlling. Those changes then become the basis for the controller's next decision.
White's research found that conventional AI performance measures can mask major differences in how controllers perform when operating in this dynamic environment.
Across 480 simulations, near-perfect test scores concealed differences of several orders of magnitude in the cost of operating the system.