Researchers have developed a technology capable of recognizing when artificial intelligence (AI) has misinterpreted human intentions by analyzing brain signals. The goal of this development is to enable AI systems to recognize the state of 'that was not what I meant' without the need for verbal clarification from the user.
The work was conducted by scientists from Korea Advanced Institute of Science and Technology (KAIST) in collaboration with Microsoft Research Asia. This technology, named Neural Value Alignment (NVA), uses a brain-computer interface to analyze data obtained through electroencephalography (EEG), allowing AI to correct its behavior in real time.
The core idea arose from a common problem in human-machine interaction: the same human action can correspond to different goals, and the same goal can be achieved in multiple ways. For example, a person picking up a cup does not necessarily mean the intention to drink from it; the person might want to pass the object, wash it, or simply place it elsewhere.
Similarly, if the goal is to quench thirst, a person might pick up a cup, find a bottle of water, or ask someone to bring a drink. This ambiguity between action and goal leads to AI systems potentially interpreting the action correctly but being wrong about the true desire of the human, or understanding the goal but choosing a different method of achieving it than expected by the user.
During experiments, researchers discovered different patterns in brain waves when RPE, SPE, or both states occurred simultaneously. They then applied deep learning methods to the EEG signals captured. This allowed the system to classify the human's reaction to AI actions.
In practice, this technology can distinguish between two scenarios: 'you misunderstood my goal' and 'you understood the goal, but did it in the wrong way.' The decoded signals are then used as feedback for the AI's decision-making process. Upon detecting SPE, the system interprets that the goal is correct, but the strategy needs modification. Meanwhile, RPE indicates that the goal itself was misunderstood, and the AI must try to ascertain the human's intention again.
According to KAIST, simulations showed that the method adapts faster than existing approaches, even under uncertain conditions such as a sudden change in the human's goal or the absence of some neural feedback. The study was published in August 2026 in the journal IEEE Transactions on Cybernetics under the title Neural Value Alignment: Human–AI Collaboration Under Goal–Action Ambiguity.
