University calendar

Inner Speech Decoding from EEG: A Comparative Study of Deep Learning Architectures for Brain–Computer Interfaces

Wednesday, August 19, 2026 at 1:00pm to 2:00pm

Zoom (please contact: pnadipalli@umassd.edu or ashok.patel@umassd.edu for Zoom information)
Dr. Ashok Patel
apatel38@umassd.edu

Advisor: Dr. Ashokkumar R. Patel - Department of Computer & Information Science
 
Committee Members:
Dr. Yuchou Chang – Department of Computer & Information Science
Dr. Debarun Das – Department of Computer & Information Science
 
Abstract:
Inner speech - the silent production of words in the mind, without any movement or sound — is an appealing control signal for a brain–computer interface (BCI), because the command is the thought. For someone who has lost the ability to speak or move, a decoder that reads intended words directly would be far more natural than the indirect mental tasks most BCIs rely on. Reading inner speech from scalp electroencephalography (EEG) is, however, extremely hard: the signals are weak and non-stationary, the neural traces of covert language are faint, and reported four-class accuracies in the literature rarely climb much past the low thirties. In this regime, the meaningful question is not whether a model reaches high accuracy; none reliably do but whether a design choice yields a real, statistically reliable signal, assessed honestly.
 
This thesis compares three standard architectures - a compact convolutional network (EEGNet), an LSTM recurrent network, and a self-attention Transformer - against a hybrid model that feeds a shared convolutional front end into parallel recurrent and self-attention branches and fuses them before classification. All four are evaluated identically on the public "Thinking Out Loud" inner-speech dataset (Nieto et al., 2022), under subject-dependent five-fold cross-validation on the four-class directional-word task (Up, Down, Right, Left), with on-the-fly augmentation and a fixed seed. The unit of statistical analysis is the subject (n = 10). The hybrid model attains the highest mean accuracy, 28.5% (SD 3.1%), and is the only model whose accuracy is statistically significantly above the 25% chance level (Wilcoxon signed-rank p = 0.014; one-sample t-test p = 0.006). The three baselines do not reach significance against chance. In direct paired comparisons the hybrid is not significantly better than any individual baseline, and an ablation shows that each single branch performs at roughly the level of its corresponding baseline, with the two-branch fusion adding a small, non-significant improvement. A subject-independent analysis falls to chance, consistent with the well-documented failure of cross-subject generalization for inner speech. These results are consistent with prior decoding studies on this dataset. The contribution is therefore a controlled, like-for-like benchmark of four architectures under one protocol, with rigorous statistical assessment: it shows that on this difficult task the hybrid is the only architecture to clear chance, while the architectures are otherwise statistically indistinguishable from one another.

For further questions, please contact Professor Ashokkumar R. Patel at ashok.patel@umassd.edu

Back to top of page