21 September 2026

AI can score almost perfectly and still potentially cost 400 times more – here's why

New research from an Elevate scholar has found that an AI system can score almost 100% in testing and still perform dramatically differently when put in charge of an electricity system designed to simulate the real world.

New research by ATSE Elevate: Boosting diversity in STEM scholarship recipient and University of Queensland researcher Meaghan White has found that an AI system can score almost 100% in testing and still perform dramatically differently when put in charge of an electricity system designed to simulate the real world.

The research found two machine-learning controllers with the same architecture and the same 99.8% test score could differ by 406-fold in operating cost when tested in simulated ‘real-world’ conditions.

The finding points to a simple lesson as AI moves into critical infrastructure: don't just test whether an AI gets the answer right, test what happens when it acts on that answer.

White received the international IEEE CASE Best Student Paper Award for the research, which examined machine-learning systems designed to rapidly reconfigure simulated microgrids – small electricity networks that must constantly respond to changes in electricity demand, generation and equipment.

Unlike a conventional AI system that makes a prediction and stops, these controllers make decisions that change the system they are controlling. Those changes then become the basis for the controller's next decision.

White's research found that conventional AI performance measures can mask major differences in how controllers perform when operating in this dynamic environment.

Across 480 simulations, near-perfect test scores concealed differences of several orders of magnitude in the cost of operating the system.

Meaghan White Elevate Scholar Photo
Elevate: Boosting diversity in STEM scholarship recipient and University of Queensland researcher Meaghan White

The research suggests AI controllers should be assessed in the environment in which they will operate, alongside conventional measures.

Testing should consider factors including cost, feasibility, system behaviour and the ability to recover when conditions change.
Importantly, the research also found that some approaches to improving the robustness of AI controllers worked extremely well when matched to the right system. The best-performing controller met the researchers' requirements in 28 of 30 scenarios.

White said the research highlighted the importance of testing AI beyond conventional performance measures.

“A 99.8% score can sound almost perfect, but for AI systems designed to make decisions in the real world, that number doesn't tell the whole story. We need to understand what happens when those decisions change the system itself.”

White said the Elevate scholarship program had also supported her development beyond her technical research.

Elevate has connected my research with a wider community of women in STEM and encouraged me to think about how engineering leadership includes communicating risk clearly and translating evidence into better practice.

Meaghan White

 

Jemima Kang Elevate Scholar
17
SEP
2026
Discover another Elevate scholar
From psychology to AI: Elevate scholar using big data to understand how the mental health landscape is changing

An Elevate: Boosting diversity in STEM scholarship is helping an emerging STEM researcher uncover how our understanding of mental health is shifting and what that means for how we respond. 

Elevate
Diversity & inclusion
Women in STEM