Researchers reported in Nature Neuroscience on Sept. 14 that a single high-density electrocorticography implant can support speech and gesture decoding from the same sensorimotor interface in people with severe paralysis. Isolated movement decoding was evaluated across three participants; Bravo-3 withdrew before the full-body avatar experiments, leaving Bravo-1r and Bravo-6 for the simultaneous speech-and-gesture work.
The real-time results make the technical achievement easier to judge. In the proof-of-concept conversational paradigm, Bravo-6 averaged 85% gesture accuracy and 75% speech accuracy across five blocks, against 9.1% chance for each decoder. Bravo-1r recorded median 100% accuracy for both gesture and speech across three blocks, with chance levels of 20% and 16.7%, respectively.
Those are small-sample results, and the broader task was harder. In a 12-block simultaneous copy task for Bravo-6, real-time accuracy was 66% for gestures and 70% for speech across ten gestures and ten phrases. Bravo-1r’s simultaneous performance was evaluated offline, where mean accuracy reached 88.3% for gestures and 84.0% for speech. The participants also used different attempt strategies — Bravo-1r silently attempted speech and minimally attempted gestures, while Bravo-6 overtly attempted speech and imagined gestures — so their percentages are not clean head-to-head comparators. The paper offers far more than a qualitative avatar demonstration while still falling short of a population-performance estimate.
The engineering result sits in the training regime. Models trained only on isolated speech or gesture attempts generalized poorly when the participant attempted both at once. Context-inclusive training — using isolated and simultaneous examples — sharply reduced false negatives across behavioural contexts, while cross-modality training reduced false activations when the participant attempted the other modality.
That matters because natural communication is multi-effector. A useful assistive interface cannot assume that speech, facial movement and gesture occur in clean laboratory channels one at a time. The product has to distinguish concurrent intent without firing one decoder because the user engaged another.
One cortical interface has now supported multiple expressive outputs in two people at accuracies materially above chance, including closed-loop conversational use. Reproducibility across more participants, longer home use and changing behavioural contexts is now the boundary between an impressive feasibility result and a clinically useful reliability curve.
