A Hybrid CNN-Transformer Framework for Multimodal Biomedical Signal Classification

Authors

  • Hedi Ammar Gusemi Department of Electrical Engineering, College of Engineering, Qassim University, Saudi Arabia Author

DOI:

https://doi.org/10.65606/41/2026/83

Abstract

Classifying biomedical signals from multiple physiological modalities can improve diagnostic accuracy, yet fusing heterogeneous signals such as electroencephalography (EEG), electrocardiography (ECG), and electromyography (EMG) remains difficult. Differences in sampling rates, waveform morphology, and noise levels make straightforward feature concatenation unreliable. This paper presents HybridBioNet, a framework that pairs convolutional neural networks with transformer encoders for multimodal biomedical signal classification. Each modality passes through a dedicated branch: one-dimensional convolutional layers extract local morphological features while transformer encoders model long-range temporal dependencies. A cross-modal attention fusion module then learns to weight and combine representations across modalities. We evaluate HybridBioNet on three multimodal and two single-modality public benchmarks spanning emotion recognition, stress detection, arrhythmia classification, seizure detection, and hand gesture recognition. On multimodal tasks, HybridBioNet compares favorably with existing methods, improving accuracy by 0.9 to 5.1 percentage points over unimodal baselines. On single-modality benchmarks, the CNN-Transformer backbone alone achieves results competitive with recent task-specific architectures. Ablation experiments show that both the convolutional and transformer components contribute to performance, with cross-modal attention fusion providing the largest marginal gain on multimodal datasets. These findings indicate that hybrid architectures with attention-based fusion offer a generalizable approach to multimodal biomedical signal analysis.

Downloads

Download data is not yet available.

Downloads

Published

2026-06-30