
The audit spans machine learning and deep learning models that ingest structural brain MRI to forecast disability trajectories. Performance varies widely. Reproducibility does not.
Where the models stand
Multimodal radiomics and deep learning architectures consistently outperform unimodal baselines across the reviewed cohorts. The reported predictive utility — driven largely by convolutional and transformer-based feature extractors applied to T1, T2-FLAIR, and diffusion-weighted inputs — degrades sharply the moment a model leaves its training distribution. Vendor-specific acquisition parameters remain the dominant uncontrolled variable. A model trained on 3T GE data tolerates poorly the same sequence acquired on a Siemens 1.5T scanner. Gradient slew rates, echo train lengths, and b-value selection all constrain the feature space in ways that no preprocessing normalization fully corrects.
The bias audit
The reviewers identified widespread risk of bias across the included studies. Small cohorts, retrospective splits, and absent external validation degrade the reported metrics. Selection bias — patients scanned at tertiary centers on modern hardware — inflates accuracy and suppresses failure modes that any deployed system will eventually encounter. The authors emphasize that without standardized evaluation benchmarks, cross-study comparison yields noise rather than signal. AUC numbers from different papers cannot be stacked. They do not share denominators, scanner hardware, or label definitions.
What to watch
The bottleneck is no longer the network architecture. It is the absence of a shared evaluation protocol — fixed acquisition constraints, held-out multi-vendor test cohorts, prospective validation windows. Until the field commits to a common benchmark, claimed predictive accuracy for MS progression will continue to function as a marketing artifact rather than a clinical constraint. The mathematics is mature. The audit layer is not.