Leakage-Audited Riemannian Classification and Guardrail-Constrained AI Narration for Low-Channel EEG
- Publicado
- Servidor
- Zenodo
- DOI
- 10.5281/zenodo.21729488
Closed-loop Brain-Computer Interfaces (BCIs) mapping low-channel frontopolar scalp potentials face severe translational bottlenecks due to biological non-stationarities, high-amplitude ocular and electromyographic (EMG) artifacts, and systemic calibration drift. During dry-run optimization, developers routinely evaluate processing pipelines on oversimplified, deterministicsinusoidal simulations, creating an “algorithmic tautology” where decoders perform flawlessly on synthetic inputs but fail when exposed to the chaotic, non-linear dynamics of biological scalp potentials. Furthermore, the integration of generative AI layers to narrate neural index trends introduces a high-risk vector for the hallucination of speculative clinical diagnostics.
To break these validation bottlenecks, we present a unified, hardware-agnostic, low-latency (200 ms) closed-loop BCI software platform and a rigorous verification methodology. We establish a generic parametric specification framework for an actively shielded, dry-electrode acquisition front-end, proving that its physical constraints are tolerated by our pipeline. We evaluateBlind Source Separation (BSS) denoising layers using SOBI and FastICA under extreme noise conditions (SNR = −18.68 dB), achieving a +19.99 dB net gain in signal-to-noise ratio; a 60 trial paired analysis finds no statistically significant performance difference between the two algorithms (Wilcoxon p = 0.729), indicating algorithm selection should be driven by deployment constraints rather than expected accuracy differences. To audit the pipeline against data leakage, we execute a Leave-One-Subject-Out (LOSO) machine learning tournament across 50 human subjects from the PhysioNet BCI2000 dataset [7], demonstrating that a Riemannian Tangent Space Alignment (TSA) Random Forest pipeline achieves 74.72% classification accuracy (+26.84 percentage points over a shuffled-label negative control) after removal of a hardcoded per-subject noise injection identified and corrected during verification (Section 10). We expose a critical TSA temporal leakage bug that yielded an artificial 76.20% accuracy on pure Gaussian noise, providing an algorithmic case study on the necessity of negative-control gates. We extend this pipeline to real-time closed-loop control using a biophysically authentic, stochastic second-order autoregressive AR(2) simulator that models thalamocortical networks.
We correct two widespread errors in real-time manifold tracking, replacing flat Euclidean exponential averages with true geodesic step updates on the curved manifold of Symmetric Positive Definite (SPD) covariance matrices M(4) (α = 0.01). Real-time evaluation over 2,153 sliding windows (7.2 min) verifies a Signal Quality Index (SQI) of 99.74% (computed over the fullunfiltered session) and clear state separability (Cohen’s d = 4.967; a Mann–Whitney p-value is also computed but is not independently interpretable given the ∼90% sample overlap between consecutive 200 ms-step windows, and is reported only for completeness — see Section 6.3). To break the simulation circularity loop, we cross-validate the entire pipeline on raw biological potentials from the 48-subject human STEW dataset [8], globally centered to remove DC bias (∼4,297 µV). By conducting a systematic sensitivity sweep, we demonstrate that legacy Euclidean arithmetic averaging suffers from a volatile, cumulative matrix volume swelling artifact — ranging from 360× to 92,898× determinant-based inflation across window counts of 100to the full pooled cohort (Table II median: ∼ 16,000×) — whereas our iterative Riemannian Fr´echet Mean optimizer stably converges (19–42 iterations across all window counts) and strictly conserves the underlying biological volume (1.000× volume preservation, consistent across all tested conditions).
Finally, we introduce a dual-layer AI guardrail (logit-biasing combined with grammar constrained decoding) that restricts generative feedback to safe, deterministic template structures, preventing the generation of clinical or diagnostic terminology by mathematical construction. We characterise the guardrail’s structural coverage against 20 adversarial prompt patterns spanning diagnostic, treatment, and pathological-state categories, demonstrating by construction that the constrained output grammar cannot express any of these categories regardless of the underlying model’s alignment level. This unified framework establishes a mathematically rigorous, biologically validated platform and verification standard, with physical hardware-in-the-loop and human-subject pilot studies explicitly deferred to planned follow-up investigations.