The warning sign is not necessarily a failed quality-control report or an ICC that collapses below an accepted threshold. In fact, the more dangerous case is often the opposite: agreement between scanners remains high while the average measurement moves in one direction. Subjects appear to track each other well, but the biomarker has acquired a systematic offset.
That is the central lesson from the 2021 PLOS ONE analysis of 116 healthy young adults scanned before and after a Siemens Verio-to-Skyra transition at 3T. The change involved different scanner models from the same vendor, along with the syngo MR E11 software platform and a Tim 4G body coil. The intraclass correlation coefficient for cortical and subcortical volumes remained above 0.8. Yet approximately 67% of the regions examined showed a systematic shift of about 2–4%.
The figure matters because it is not an interchangeable property of every upgrade. It belongs specifically to that Verio-to-Skyra study. It should not be presented as evidence that an identical-model hardware replacement, a coil swap, or every vendor migration produces the same magnitude of change.
An ICC above 0.8 can coexist with a meaningful bias. Reliability tells you that subjects remain ordered; it does not tell you that the scanner has kept the same scale.
A shift of this size can sit in the same general range as the longitudinal changes that neurodegeneration studies are trying to detect. That does not mean every 2–4% change is clinically meaningful, or that the Verio-to-Skyra result can be transferred directly to another platform. It means that a scanner transition can be large enough to compete with the biological signal in a quantitative MRI endpoint.
When a vendor describes an upgrade as equivalent, the word usually refers to operational continuity, protocol availability, or broad image quality. It does not automatically establish longitudinal equivalence for cortical thickness, gray-matter volume, FA, mean diffusivity, susceptibility, perfusion, or an automated segmentation output.
The Upgrade That Wasn’t “Drop-In”
The Verio-to-Skyra result deserves careful reading precisely because it is easy to overgeneralize.
The study compared measurements acquired at the same field strength and from the same vendor, but it did not compare two units of the exact same scanner model. That distinction matters. A same-model replacement may behave differently from a Verio-to-Skyra transition, and the published 2–4% gray-matter-volume shift should not be used as a universal estimate for all hardware changes.
Still, the design exposes the underlying problem. A high ICC does not imply zero bias. ICC evaluates how consistently subjects retain their relative position across measurements. If participants with larger volumes on one scanner also tend to have larger volumes on another, the coefficient can remain high even when the second scanner reads systematically higher or lower.
For a clinical trial, that distinction creates two separate questions:
- Can the measurement rank subjects consistently?
- Does the measurement remain on the same quantitative scale over time?
The first question is about reliability. The second is about agreement and bias. Both matter, but they are not interchangeable.
A scanner transition can therefore leave a familiar pattern in the data. The between-subject ordering looks stable. Group-level plots remain visually tidy. The ICC survives the methods section. Yet a pre-upgrade scan and a post-upgrade scan are no longer directly comparable without accounting for the change in acquisition and reconstruction conditions.
In a single-site study, the problem may be manageable. The site can define a new baseline, repeat a reference scan, or model the transition as a study event. In a multicenter trial, the timing becomes part of the endpoint architecture. Sites upgrade on different schedules. One center may change the scanner platform while another changes only the software. A third may replace a coil or alter the reconstruction pathway during service work. The same nominal protocol can then represent different measurement conditions at different sites and timepoints.
That is how multicenter imaging trial endpoint bias enters the analysis. The bias does not need to be dramatic at any one site. A small site-specific shift can become consequential when it aligns with enrollment, treatment exposure, or follow-up timing.
Why the Mean Can Move While the ICC Holds
Suppose a new scanner produces values that are closely related to the old values but consistently higher. The subjects still occupy nearly the same order from smallest to largest. Correlation and ICC may remain strong. The mean difference, however, is not zero.
This is particularly dangerous in longitudinal designs because the scanner event can resemble a biological transition. If the upgrade occurs between baseline and follow-up, the post-upgrade value may look like progression or treatment response. If the upgrade is concentrated at sites with a particular recruitment profile, it can also resemble a site effect or a treatment-by-site interaction.
The statistical remedy is not to select one reassuring reliability coefficient and stop there. A proper assessment should examine agreement, paired differences, site effects, temporal effects, and the distribution of the shift across regions and tissue classes. Bland–Altman-style analyses, repeat scans, reference subjects, and phantom data can all contribute to that picture, but none replaces a clear record of what changed.
The basic metadata must answer more than whether the scan was performed on a Siemens, GE, or Philips system. It should identify the scanner model, software version, coil configuration, sequence parameters, reconstruction pathway, and the date on which each configuration became active. Without that timeline, the analyst may be forced to infer an instrument change from the biomarker itself.
When Only the Software Changes
Hardware replacement is easy to notice. Software drift is easier to miss because the scanner may remain in the same room, with the same field strength, the same bore, and a protocol name that has not changed.
Routine updates can affect reconstruction kernels, distortion correction, eddy-current correction, motion handling, gradient-imperfection compensation, denoising, and the way multiecho or multiband data are combined. A patch may be introduced as a service improvement rather than as a change to the measurement system. From the perspective of a quantitative endpoint, that distinction is not sufficient.
The 2012 diffusion tensor imaging work in Human Brain Mapping reported detectable changes in fractional anisotropy, axial diffusivity, and radial diffusivity after a software update on the same physical scanner. A 2013 follow-up extended the concern to voxel-based morphometry. The point is not that every software patch produces the same effect. The point is that a hardware-stable scanner can still change its quantitative behavior.
For a traumatic axonal injury study, the implications are direct. If diffusion data collected during the first part of the study were processed under one reconstruction or correction implementation, and later data under another, the FA trajectory may contain a step associated with the update. The subject has not changed because of the patch. The disease has not changed because of the patch. The instrument has changed.
That step may be visible in raw images, but it may also emerge only after tensor fitting, registration, smoothing, tract extraction, or region-of-interest summarization. A change that looks negligible at the image level can become more visible after several processing stages have amplified or redistributed it.
A software release is not automatically a new scanner, but it can still be a new measurement condition.
The same principle applies to automated pipelines. A vendor may update the reconstruction, then the segmentation software may receive its own update, and the core laboratory may revise preprocessing between analysis releases. If the study treats all of these as ordinary maintenance, it loses the ability to separate biological change from pipeline change.
The phrase mri biomarker drift scanner software update describes a practical governance problem as much as a technical one. The trial needs a change-control process that treats software versions as part of the instrument definition. Release notes, installation dates, build strings, configuration files, and validation results should be retained with the imaging record rather than left in a service mailbox.
The Complexity of QSM and DTI Pipeline Sensitivity
Quantitative MRI does not produce a biomarker through acquisition alone. The final number is the output of an acquisition protocol, reconstruction software, preprocessing sequence, registration strategy, segmentation method, and statistical summarization. A change at any stage can alter the value assigned to the same anatomy.
Diffusion tensor imaging is already sensitive to choices about susceptibility correction, eddy-current correction, motion correction, outlier handling, tensor fitting, and the treatment of signal dropouts. A software update can affect one of these steps without changing the DICOM series description in a way that is obvious to the downstream analyst.
QSM is even more exposed because it is a chain of transformations rather than a single reconstruction. A typical pipeline may include phase unwrapping, background-field removal, and dipole inversion. Different implementations can use different assumptions, regularization strategies, masks, and boundary conditions. The result is a susceptibility map whose regional value depends on the stability of the entire chain.
That matters when the endpoint is a subtle difference in iron-sensitive signal in the substantia nigra, putamen, or other deep-gray-matter structures. A vendor-side modification to phase reconstruction, a change in echo combination, a protocol adjustment affecting B0 homogeneity, or a pipeline update can propagate through the analysis. The resulting difference may be real as a measurement change even when the anatomy is unchanged.
The appropriate conclusion is not that QSM is unusable. It is that QSM requires stricter configuration control than a workflow based only on visual image review. A radiologist can judge whether images remain diagnostically acceptable while missing a shift in a regional quantitative value. Clinical adequacy and biomarker equivalence are different validation targets.
The Uncertainty Around AI Reconstruction
Deep-learning reconstruction and denoising make the problem more difficult because their behavior may not be easy to describe with a single correction factor. The effect of an AI reconstruction tool can depend on signal level, anatomy, motion, acquisition parameters, coil sensitivity, and the distribution of images used during development and validation.
It is not established that such changes are always uniform across tissue types or scanner conditions. They may be approximately stable for one protocol and less predictable for another. They may alter image texture without producing a large change in a conventional radiological assessment, while still affecting a quantitative algorithm downstream. At present, the possibility of uniform or non-linear drift should remain explicitly uncertain rather than being presented as a settled property of all reconstruction AI.
This uncertainty is especially relevant when a trial changes from conventional reconstruction to a deep-learning denoiser during enrollment. The absence of a visible artifact is not evidence of longitudinal equivalence. The validation question must be tied to the actual endpoint: a segmentation volume, a diffusion metric, a susceptibility value, a perfusion parameter, or another quantitative output.
| Potential source of drift | What can be stated responsibly | How it should be handled |
|---|---|---|
| Verio-to-Skyra hardware transition | A 2021 study reported approximately 2–4% shifts in many gray-matter regions after this specific same-vendor, different-model transition | Use the finding as study-specific evidence, not as a universal upgrade factor |
| Same-model hardware replacement | The magnitude of any quantitative shift is protocol- and system-dependent | Validate the actual replacement rather than importing the Verio-to-Skyra estimate |
| Software-only update | DTI and morphometric measures can change detectably even when the physical scanner is unchanged | Record the build and treat the update as a measurement event |
| Coil change | Effects depend on the coil, sequence, loading, signal behavior, and reconstruction | Do not assign a universal percentage; compare before-and-after data under the study protocol |
| Reconstruction AI or deep-learning denoiser | The direction, uniformity, and linearity of drift are not established in general | Validate the specific implementation and endpoint before combining timepoints |
| Cross-site and cross-vendor variation | Differences can arise from hardware, protocols, preprocessing, and local operations together | Model site and configuration explicitly and harmonize where feasible |
Automated Segmentation and the Regulatory Boundary
Regulatory clearance is another place where workflow assumptions can outrun the evidence.
A 510(k) clearance is primarily a determination of substantial equivalence to a legally marketed predicate device. It is not a blanket certification that an automated measurement is longitudinally interchangeable with every earlier software version, scanner configuration, or analysis pipeline. Nor does clearance by itself answer whether a mid-trial change preserves the statistical properties of a registered endpoint.
That distinction becomes important when an automated volumetric tool adds capabilities, changes segmentation behavior, or modifies its processing implementation. A newer version may be appropriate for clinical use and still require a bridging exercise before its outputs are pooled with measurements generated by an earlier version.
For an Alzheimer’s disease study, an updated tool may support additional volumetric regions, lesion handling, or ARIA-related workflows. Those capabilities can be valuable. They can also change the segmentation boundary, quality-control rules, failure handling, or normative comparison. The relevant question for a longitudinal trial is not simply whether the new version is cleared. It is whether the output can be interpreted on the same scale as the prior output for the endpoint being analyzed.
A quiet transition from NeuroQuant v4.x to v5.0 in the middle of enrollment is therefore not automatically compliant merely because both products are cleared. The trial protocol, analysis plan, sponsor procedures, institutional controls, and applicable regulatory requirements determine whether the change is permissible. It may require documentation, validation, sponsor and site approvals, a predefined deviation process, or a formal bridging analysis. Product clearance does not replace those obligations.
The minimum operational record should identify:
- the exact software version and build used for each scan;
- whether processing was performed locally, centrally, or through a hosted service;
- the segmentation model, atlas, or processing configuration;
- any changes to quality-control thresholds or failure-handling rules;
- the date of transition and the scans affected;
- the decision about whether pre- and post-transition results may be pooled.
This is not bureaucracy added after the scientific work. It is part of defining the measurement.
A version change can affect not only the mean value but also missingness. If the newer software successfully processes scans that the older version rejected, the apparent improvement may alter the composition of the analyzable sample. If a new algorithm fails on a different set of scans, the endpoint can acquire a selection effect in addition to a measurement shift.
What Longitudinal Studies Need to Quantify
The phrase “calibration error” can be misleading when the change is not a simple offset. Some transitions may be approximated by an additive or multiplicative adjustment. Others may interact with age, tissue volume, motion, signal-to-noise ratio, disease burden, or site-specific acquisition conditions.
This is why a single global correction factor is rarely an adequate first response. The study needs to establish what changed and for whom.
A useful bridging design may include repeated scans of the same participants before and after the transition, a stable reference cohort, quantitative phantom measurements, or a combination of these. The exact design depends on the modality, endpoint, available time, and operational constraints. A phantom acquisition does not have one universal duration: scan time depends on the phantom, pulse sequences, number of repeats, calibration requirements, and the quality-control question being asked. The relevant requirement is sufficient coverage of the protocol, not adherence to a fixed number of minutes.
Reference participants can help estimate a transition effect when rescanning is feasible, but they are not a substitute for proper metadata. A paired sample can show that values moved; it cannot explain whether the cause was hardware, software, coil loading, protocol drift, or a processing change unless those factors were recorded.
The analysis should also distinguish among:
1. A constant shift, in which most observations move by a similar amount.
2. A scale change, in which the difference grows with the underlying measurement.
3. A tissue- or region-specific effect, in which some structures are affected more than others.
4. A site-dependent effect, in which the same nominal update behaves differently across centers.
5. A time-varying effect, in which the system changes gradually rather than at one identifiable transition.
6. An interaction with subject characteristics, such as motion, atrophy, lesion burden, or image quality.
The last three are particularly difficult because they can defeat a simple pre/post indicator in the statistical model. A scanner event recorded only as “before” or “after” may be insufficient if the implementation was staged, if different reconstruction options were enabled at different sites, or if a vendor update changed only part of the processing chain.
For quantitative MRI endpoints, reproducibility is therefore a property of the full measurement system. It is not guaranteed by scanner brand, field strength, protocol name, or a high correlation coefficient alone.
Mitigating Systematic Drift in Multicenter Clinical Trials
The practical mitigation strategy is operational before it is algorithmic. Harmonization tools, statistical adjustment, and machine-learning correction can help, but only after the trial has established what changed and preserved enough information to model it.
Build a change-control record
Every scanner and processing event should have a timestamped record. That includes hardware replacement, software installation, firmware changes, coil changes, sequence edits, reconstruction options, local preprocessing, and central-pipeline releases.
The record should connect the event to actual subject scans. A service log that says an update occurred at a site is not enough if the study cannot identify which participants were scanned before and after the change.
Validate the configuration that the trial actually uses
A phantom or reference scan is useful only when it represents the endpoint’s acquisition and analysis conditions. A generic quality-control scan may show that the scanner is functioning while missing a drift in diffusion, QSM, segmentation, or another derived measure.
Validation should therefore include the sequence, coil, reconstruction pathway, and processing software relevant to the biomarker. If the trial uses multiple endpoints, one modality’s stability should not be treated as evidence that all other modalities remained stable.
Preserve pre- and post-event comparability
When a change is unavoidable, the trial should decide in advance how it will handle affected data. Options may include:
- modeling scanner configuration as a covariate or fixed effect;
- including site-by-configuration interactions where justified;
- analyzing pre- and post-transition periods separately;
- using a bridging cohort to estimate the transition;
- reprocessing all data through a common version-controlled pipeline;
- excluding a specific endpoint from pooled longitudinal analysis if equivalence cannot be established.
No single option is automatically correct. The choice depends on the evidence generated by the bridging work and on the estimand defined by the protocol.
Keep processing versioned and reproducible
A vendor black box can be clinically convenient, but it makes longitudinal interpretation dependent on the vendor’s release cycle. A centralized, version-controlled pipeline gives the core laboratory a clearer record of what was done and makes reprocessing possible when an implementation changes.
That does not mean local or vendor software is inherently unsuitable. It means the study must know which implementation produced each value and must retain the ability to reproduce or audit the result. A biomarker that cannot be regenerated from documented inputs is difficult to defend when the trial encounters a scanner transition.
Do not treat “minor” changes as automatically minor
A coil swap may have a limited effect in one sequence and a larger effect in another. A reconstruction change may matter for denoised diffusion data but not for a conventional anatomical endpoint. A protocol edit that appears harmless in one region may affect a downstream segmentation boundary.
The correct response is not to assume that every service action invalidates the trial. It is to classify the action, test the relevant endpoint, and document the decision. The burden is proportional to the measurement risk, but the decision should be explicit.
In a multicenter trial, configuration history is not background metadata. It is part of the biomarker definition.
The Cost of Ignoring the Transition
The most expensive time to discover scanner drift is after the database has been locked and the primary endpoint has been calculated.
At that stage, the team may have to reconstruct service histories, identify affected scans, locate archived software builds, repeat processing, and explain why a transition was not handled prospectively. If the event coincides with treatment milestones or a change in recruitment, the statistical ambiguity becomes harder to resolve.
The problem can also remain hidden when investigators focus on group-level averages. A site with a positive shift and another with a negative shift may appear stable in the pooled mean while increasing variance and weakening power. Conversely, a transition concentrated in one arm or one enrollment period can create an apparent treatment effect.
The same issue applies to normative tools. A segmentation output may be compared with a reference database whose own software, scanner distribution, or preprocessing history is not transparent. A volume can be technically precise and still be difficult to interpret if the measurement scale changed between the patient scan and the reference population.
This is why longitudinal qMRI calibration errors are not merely a scanner-room concern. They can affect sample-size assumptions, endpoint variance, treatment-effect estimates, missing-data patterns, and the credibility of a biomarker claim.
The Bottom Line
MRI biomarker drift after a scanner or software change is not a hypothetical edge case. Quantitative measures can shift even when the scanner remains clinically usable and reliability coefficients remain high. The approximately 2–4% result belongs specifically to the Siemens Verio-to-Skyra study; it is a warning about the scale of possible bias, not a universal conversion factor for every hardware upgrade.
The same caution applies to coils, software-only updates, automated segmentation, QSM pipelines, and AI reconstruction. Coil effects are protocol-dependent. Phantom acquisition time depends on the phantom and protocol. The direction and linearity of AI-related drift remain uncertain. A 510(k) clearance establishes substantial equivalence to a predicate device; it does not certify longitudinal equivalence across software versions or make an undocumented mid-trial transition automatically compliant.
The practical posture is not paranoia. It is configuration control.
Record the change. Identify the affected scans. Validate the endpoint that matters. Preserve a bridge between old and new conditions where possible. Model the transition when the data support it, and do not pool measurements simply because the scanner brand and protocol name stayed the same.
If the scanner changed and there is no baseline record showing how the quantitative output behaved across that change, the longitudinal number carries an unresolved measurement question. Vendor assurances may support an operational decision, but they are not a calibration log—and they are not a substitute for evidence of biomarker equivalence.
