Sensor Geometry as a Flow-Matching Prior for Multi-Channel Brain Signals

Massachusetts Institute of Technology

Building the graph-Matérn source prior from sensor positions. (i) Connect each sensor to its \(k\) nearest neighbors, with edge weights that decay with distance. (ii) Form the normalized graph Laplacian, whose eigenvectors are spatial patterns over the sensors ordered from smooth to rough. (iii) Weight each pattern by the Matérn density \(\phi(\lambda_m)=(1+\tau\lambda_m)^{-\alpha}\), which gives smooth patterns more variance and rough ones less. (iv) Assemble the source covariance \(\boldsymbol{\Sigma}_0\) from which the flow draws its starting samples. Every step uses only the sensor coordinates and never the recorded signals.

Abstract

Flow-matching models start from an isotropic Gaussian source, the standard choice when the correlation structure of the data is unknown in advance. For multi-channel brain recordings, however, part of this structure is known in advance. Electrodes sit at fixed positions on the head, and volume conduction through the skull and scalp makes nearby electrodes co-vary in a way that is shared across subjects. Existing EEG generative models nonetheless leave the network to learn this from scratch.

We put this structure into the source instead. From the sensor coordinates alone, we build a \(k\)-nearest-neighbor graph and take a graph-Matérn function of its Laplacian as the source covariance, so the flow starts from spatially coherent patterns rather than channel-independent noise. The change adds no learned parameters, works with any coupling and any drift network, and uses the same three hyperparameters on every dataset.

Across eight EEG datasets and four flow-matching methods, the graph-Matérn source lowers the spectral discrepancy between generated and real signals in the five clinical bands (PSD-KL) on most datasets. PSD-KL falls by 12% to 17% in geometric mean over datasets depending on the method and by up to 40% on PhysioNet-MI, the densest montage. We show that the improvement stems from the spatial eigenvectors of the local graph of sensor positions, since randomizing the eigenvectors while preserving the eigenvalue spectrum eliminates the gain. Furthermore, a prior fitted directly to the empirical data covariance performs worse than isotropic noise. The same construction applies unchanged to MEG, intracranial EEG with patient-specific grids, and a traffic-sensor network, lowering PSD-KL for every method on each.

Highlights

  • A source prior from sensor coordinates alone. A graph-Matérn covariance on the Laplacian of the \(k\)-nearest-neighbor graph of sensor positions replaces the isotropic source. It adds no learned parameters, leaves the coupling and drift network unchanged, and uses \(k{=}4\), \(\tau{=}1\), and \(\alpha{=}2\) on every dataset.
  • Better spectra for every flow-matching method. Across SF2M, SI, OT-CFM, and rectified flow on eight EEG datasets with 16 to 64 channels, PSD-KL falls by 12% to 17% in geometric mean and by up to 40% on PhysioNet-MI.
  • Beyond scalp EEG. The identical construction lowers PSD-KL on whole-head MEG, on intracranial EEG with patient-specific grids (34% to 40%), and on a freeway traffic-sensor network.
  • The eigenvectors carry the gain. Keeping the spectrum but randomizing the eigenvectors eliminates the improvement, and a prior fitted to the empirical data covariance performs worse than isotropic noise.

Method

Given 3-D sensor coordinates \(\{p_c\}_{c=1}^{C}\), we connect sensors whenever either is among the other's \(k\) nearest neighbors, with weights \(W_{ij}=\exp(-\lVert p_i-p_j\rVert^2/h^2)\). The normalized Laplacian \(\mathbf{L}=\mathbf{I}-\mathbf{D}^{-1/2}\mathbf{W}\mathbf{D}^{-1/2}=\mathbf{U}\boldsymbol{\Lambda}\mathbf{U}^\top\) has eigenvectors that range from smooth patterns across the head to patterns that flip sign between neighboring sensors. We give each eigenvector an independent Gaussian with variance \(\phi(\lambda_m)=(1+\tau\lambda_m)^{-\alpha}\), normalized so that the total variance matches an isotropic source, and draw the flow's starting point as

$$x_0 = \mathbf{U}\,\mathrm{diag}\big(\sqrt{\phi(\boldsymbol{\Lambda})}\big)\,z_0,\qquad z_0\sim\mathcal{N}(0,\mathbf{I}_{C\times T}).$$

This is the only line of training that changes. The prior shapes the spatial covariance and leaves all temporal structure to the drift network. In our experiments the drift is a 1D U-Net that treats the sensors as input channels, so the network itself has no spatial inductive bias and the sensor layout enters only through the source.

Real vs. Generated EEG

Qualitative effect of the graph prior on TUAB. 16-channel waveforms from one real test window and one sample each from GP-SI (SI trained with the graph prior) and SI with the isotropic source. Real EEG shows slow oscillations that are coherent across channels. GP-SI samples retain both properties, whereas SI samples resemble channel-independent high-frequency noise.

Power Spectrum on Eight EEG Datasets

PSD-KL compares the distribution of log band power in the five clinical bands (\(\delta,\theta,\alpha,\beta,\gamma\)) between real and generated signals, per channel (lower is better). Its scale depends on the dataset, so comparisons are made within each row. Values are means over five training seeds on the test split; the paper reports standard deviations and statistical tests. Bold denotes the better source within each method, and the prefix GP- marks the graph prior.

\(\sigma{>}0\), Hungarian \(\sigma{>}0\), independent \(\sigma{=}0\), Hungarian \(\sigma{=}0\), indep.+reflow
Dataset (channels) SF2MGP-SF2MSIGP-SI OT-CFMGP-OT-CFMRFGP-RF
TUAB (16)1.681.041.431.161.671.402.272.53
TUEV (16)2.033.852.552.155.303.803.193.09
Mumtaz (19)4.423.893.913.914.975.516.063.79
BCI-IV 2a (22)3.783.563.303.503.063.614.274.24
FACED (32)32.931.833.329.033.631.532.126.4
SHU (32)11381.411482.811185.411682.9
SEED-V (62)32.828.132.928.032.928.133.328.3
PhysioNet-MI (64)34.820.733.419.933.019.933.322.4
Geometric mean of GP/iso 0.880.830.860.83

The largest reductions shared by all four methods occur on the densest montages, PhysioNet-MI (33–35 to 20–22) and SHU (111–116 to 81–85). The graph prior also raises the correlation between real and generated phase-lag coupling (wPLI) by 0.01 to 0.03 on average and by 0.05 to 0.12 on SEED-V and PhysioNet-MI. Where the prior does not help, the paper gives a likely reason, for example that the 22 electrodes of BCI-IV 2a cover only the centro-parietal region, leaving the graph little spatial structure to encode.

MEG, Intracranial EEG, and Traffic Sensors

The identical construction, using coordinates alone, lowers PSD-KL for every method on whole-head MEG (16 subjects, 102 magnetometers), on intracranial EEG (five patients with individual 72–112-channel ECoG grids), and on the PEMS-BAY freeway network (325 loop sensors, graph from GPS coordinates). The reduction is largest on iEEG (34% to 40%), where every patient has a different grid. The traffic network involves no tissue or volume conduction, which indicates that the benefit comes from the spatial geometry encoded in the graph rather than from electrophysiology.

Modality SF2MGP-SF2MSIGP-SI OT-CFMGP-OT-CFMRFGP-RF
MEG, Wakeman–Henson (16 subjects)28902715289127152893271229022726
iEEG, CCEP ECoG (5 patients)31.319.231.218.730.518.937.624.9
PEMS-BAY traffic (325 sensors)62.155.662.155.562.155.363.956.8

Which Part of the Prior Carries the Gain?

The graph prior differs from isotropic noise in its eigenvectors, its spectrum, and the graph that defines them. We train SI on TUAB with priors that replace one of these parts while keeping the rest fixed (validation split, mean ± std over seeds; \(d_{\mathrm{eff}}\) is the prior's effective dimension). Replacing the eigenvectors, even with the eigenvectors of the data covariance, is worse than isotropic noise. A smooth spectrum that covers every mode suffices, and the graph must be sparse and local.

PriorReplaces\(d_{\mathrm{eff}}\)PSD-KL
isotropic—16.01.85 ± 0.18
graph-Matérn (ours)—10.11.08 ± 0.41
shuffled sensor positionseigenvectors10.13.70 ± 2.79
uniformly random eigenvectorseigenvectors10.14.42 ± 2.97
empirical eigenvectors, graph spectrumeigenvectors10.12.76 ± 2.93
empirical covarianceeigenvectors and spectrum5.24.61 ± 2.40
Ledoit–Wolf-shrunk empirical covarianceeigenvectors and spectrum6.63.56 ± 2.43
heat-kernel spectrumspectrum12.01.38 ± 0.86
hard low-pass spectrumspectrum8.071.4 ± 12.5
fully connected Gaussian graphgraph11.22.53 ± 1.81
correlation graph (\(k\)-NN on |corr|)graph11.01.39 ± 1.70

Data Augmentation for EEG Classification

On Mumtaz-MDD, a binary depression-detection task with a subject-disjoint split, we add one generated window per real training window and train an EEGNet-8,2 classifier. Synthetic windows raise balanced accuracy from 60.8% (real data only) to 81.9% with the isotropic source and 83.3% with the graph prior, mainly through minority-class recall (0.284 to 0.794). Class-weighted training on real data alone reaches only 54.5%.

BibTeX

@article{hwang2026sensor,
    author  = {Hwang, Jaedong},
    title   = {Sensor Geometry as a Flow-Matching Prior for Multi-Channel Brain Signals},
    journal = {arXiv preprint},
    year    = {2026},
}