Clinical Research & Biomarkers

QSM versus R2* mapping: competing paths for brain iron analysis

The choice between quantitative susceptibility mapping and R2* mapping can change the apparent trajectory of brain iron in a clinical trial, even when both metrics are derived from the same family of multi-echo gradient-echo MRI acquisitions.

QSM versus R2* mapping: competing paths for brain iron analysis

That is the central difficulty: the scanner is not merely observing iron, but translating its magnetic effects through a chain of sequence design, reconstruction, tissue selection, background-field correction, reference definition, and statistical interpretation.

For studies of Parkinson’s disease and related neurodegenerative conditions, this distinction is especially consequential in deep grey matter structures where iron accumulation may be anatomically localized and biologically meaningful. QSM has shown greater sensitivity and specificity than R2* mapping for distinguishing patients with Parkinson’s disease from normal controls, with particularly clear changes in the substantia nigra pars compacta. Yet greater sensitivity does not remove the need for methodological discipline. In cases of extreme iron overload, the QSM value can shift substantially depending on the reference region selected, which means that a technically sophisticated map may still produce a fragile endpoint if its calibration is not carried consistently across time and sites.

The practical question, then, is not whether QSM has replaced R2*. It has not. The more useful question is which metric better preserves the biological signal a study is trying to follow, and under what acquisition and analysis conditions that signal remains interpretable.

The same susceptibility effect, measured through different contrasts

QSM and R2* mapping are often discussed as though they were interchangeable readouts of iron concentration. They are better understood as two different ways of expressing how tissue iron alters the local magnetic environment.

Both approaches commonly use multi-echo gradient-echo data. In R2* mapping, the principal observation is how rapidly transverse signal decays across echo times. A higher R2* value indicates faster dephasing, which may occur when iron creates microscopic magnetic field variations within the voxel. The measurement is relatively direct from the signal-decay curve, and its units are typically expressed in inverse seconds.

QSM takes a different route. It begins with phase information, estimates the local magnetic field generated by susceptibility sources, and then reconstructs a susceptibility map, usually reported in parts per million. This reconstruction requires several processing decisions that do not arise in quite the same form for R2*: phase unwrapping, background-field removal, and an inversion step that estimates the underlying susceptibility distribution.

That distinction matters because the two metrics do not respond identically to the same anatomical pattern. R2* is sensitive to local field inhomogeneity and signal dephasing, but its contrast can be influenced by factors beyond iron alone, including tissue microstructure and the geometry of magnetic sources. QSM attempts to recover a property more closely related to magnetic susceptibility, which can improve the anatomical specificity of iron-sensitive imaging, although the reconstruction itself introduces additional sources of variability.

Consider the implications for a longitudinal trial. If a treatment is expected to alter iron handling in the substantia nigra, the desired endpoint is not simply a number that changes between baseline and follow-up. It is a number whose direction and magnitude can be connected, with reasonable confidence, to the biological process under investigation rather than to a different reference region, scanner calibration, or reconstruction choice.

A quantitative MRI metric becomes a biomarker only when its change remains biologically interpretable across time, observers, scanners, and analytical decisions.

Sensitivity and specificity in subcortical iron detection

The strongest comparative evidence in the available material concerns the ability of QSM and R2* mapping to identify localized iron differences in the deep grey matter of patients with Parkinson’s disease.

QSM demonstrated higher sensitivity and specificity than R2* mapping when separating patients with Parkinson’s disease from normal controls. The distinction was particularly apparent in the substantia nigra pars compacta, a region whose vulnerability is central to the motor pathology of Parkinson’s disease. QSM also detected significantly increased susceptibility values in the substantia nigra and red nucleus, whereas R2* mapping showed increased values only in the substantia nigra.

This is not a minor difference in anatomical coverage. When a metric identifies change in one structure but not another, the result can affect how researchers formulate the disease model. The red nucleus is not simply an extra region added to a table of results; it may alter the interpretation of how iron distribution relates to network-level degeneration, compensatory pathways, or disease stage. A biomarker with greater spatial sensitivity can therefore support a more precise biological hypothesis, provided the additional signal is reproducible and not an artifact of processing.

The comparison can be summarized as follows:

ParameterQSMR2* mapping
Primary signal basisReconstructed magnetic susceptibility from phase-derived field informationTransverse signal decay across echo times
Typical reported unitParts per million, or ppmInverse seconds, or s⁻¹
Detection of localized nigral iron changesHigher sensitivity and specificity in the cited Parkinson’s disease comparisonDetected increased values in the substantia nigra, with less extensive regional detection
Findings in deep grey matterIncreased susceptibility in the substantia nigra and red nucleusIncreased R2* values in the substantia nigra
Dependence on reference definitionSubstantial, particularly in extreme iron overloadDoes not use QSM-style susceptibility referencing in the same way
Main translational attractionAnatomically specific susceptibility contrast that may reveal subtle regional differencesEstablished iron-sensitive measure with comparatively straightforward signal-decay interpretation
Principal cautionReconstruction and reference choices can change measured susceptibilityThe relationship between R2* and iron is indirect and not universally convertible to absolute ppm

The table should not be read as a verdict in which one technique is clinically useful and the other is obsolete. R2* remains widely used, and its relative simplicity can be valuable in multicentre research environments where harmonized acquisition and robust signal modeling are priorities. The more nuanced conclusion is that QSM may provide a sharper view of localized iron changes, while R2* can remain a stable and practical companion measure, particularly when interpreted alongside anatomical and clinical data.

Why localized sensitivity matters in neurodegeneration

Neurodegenerative pathology rarely progresses as a uniform process across the brain. Iron accumulation may be concentrated in particular nuclei, may appear at different stages of disease, and may coexist with neuronal loss, gliosis, vascular changes, or alterations in tissue composition. A metric that averages or blurs these processes can conceal the trajectory a trial is trying to measure.

This is where QSM’s regional performance becomes attractive for clinical research. Increased susceptibility in the substantia nigra and red nucleus may allow investigators to examine whether a treatment-related change is anatomically constrained or whether it reflects a broader shift in tissue composition. It can also support more careful stratification of participants, especially when a trial is designed around a mechanistic hypothesis rather than a purely symptomatic endpoint.

Still, sensitivity should not be confused with specificity for a single pathological mechanism. Susceptibility is influenced by magnetic materials and tissue conditions that extend beyond one molecule or one disease pathway. The proper claim is that QSM can be more sensitive to localized iron-related susceptibility changes in the examined settings, not that every change in ppm represents a direct measurement of neuronal iron burden.

Reference regions are part of the measurement, not a formatting choice

QSM is often described as though it produces an absolute tissue property. In practice, susceptibility values are commonly interpreted relative to a reference region, because the reconstructed map is sensitive to the chosen baseline. This makes the reference region part of the biomarker definition.

The issue becomes especially visible in extreme brain iron overload, such as aceruloplasminemia. In the cited comparison, whole-brain referencing produced a median susceptibility value of 0.147 ppm, while cerebrospinal-fluid referencing produced a median value of 0.279 ppm. The ranges also differed: 0.527 ppm for whole-brain referencing and 0.593 ppm for cerebrospinal-fluid referencing.

These values should not be treated as two competing measurements in which one is automatically correct and the other is merely noisy. They demonstrate that the same underlying acquisition can yield materially different susceptibility estimates when the reference framework changes. Whole-brain referencing led to a lower median value than cerebrospinal-fluid referencing, consistent with systematic underestimation in the setting of diffuse or extreme iron burden.

This has direct consequences for longitudinal imaging. If a protocol uses one reference at baseline and another at follow-up, the resulting difference may reflect the analytical framework rather than a biological change. Even when the same reference is nominally used, changes in segmentation quality, tissue involvement, field strength, background-field correction, or the composition of the reference region can influence the endpoint.

A robust QSM protocol therefore needs to specify, before analysis begins:

  • which anatomical or fluid compartment serves as the reference;
  • how that reference is segmented and quality-controlled;
  • whether the reference is expected to remain biologically stable in the disease under study;
  • which background-field removal and dipole-inversion methods are applied;
  • how susceptibility values are harmonized across scanners, sites, and field strengths;
  • whether the study reports relative susceptibility rather than presenting ppm as an unconditional absolute concentration.

The question of a universal reference region remains unresolved across neurodegenerative pathologies. A reference that is reasonable for one disease may be vulnerable to biological alteration in another. Cerebrospinal fluid may offer a useful framework in some contexts, but it is not a universal answer simply because it is anatomically distinct from deep grey matter. Whole-brain referencing may be convenient, yet diffuse iron loading can make it systematically misleading.

In QSM, the reference region is not a footnote beneath the result; it is one of the conditions that gives the result its meaning.

Reliability is not the same as biological validity

One of the more reassuring findings in the comparison is the high inter-rater reliability reported for both techniques in deep grey matter. The intra-class correlation coefficient was 0.977 for QSM susceptibility values and 0.945 for R2* values, with p < 0.001 for both.

These figures indicate strong agreement between raters under the conditions of the analysis. That is valuable, because a clinical trial biomarker cannot depend on one expert’s ability to recognize a region or draw a boundary in a way that another trained observer cannot reproduce. High reliability reduces one important source of measurement error and strengthens confidence in the segmentation and reading workflow.

But reliability answers a narrower question than validity. It tells us whether observers agree, not whether the value corresponds to the biological quantity the trial intends to measure. Two raters can reproducibly obtain the same susceptibility value from a poorly chosen reference region. A highly consistent R2* measurement can still reflect several contributors to signal decay rather than iron alone. Conversely, a QSM method can provide excellent regional contrast while remaining vulnerable to reconstruction-related bias.

For longitudinal studies, at least three forms of consistency need to be kept separate:

1. Observer consistency: whether different raters identify and measure the same structures similarly.

2. Technical consistency: whether repeated acquisitions and processing pipelines produce comparable values.

3. Biological interpretability: whether a measured change can reasonably be connected to iron-related pathology rather than to a change in acquisition, reference, or tissue composition.

The ICC values support confidence in the first category. They do not, on their own, settle the second or third.

This distinction is particularly important when a trial seeks subtle degradation or stabilization over time. A small apparent treatment effect may be smaller than the variation introduced by scanner upgrades, coil changes, echo-time differences, or algorithmic updates. If the study’s endpoint is built around a regional change in ppm or s⁻¹, the analysis plan should preserve the acquisition and reconstruction pathway as carefully as the clinical follow-up schedule.

From susceptibility to iron concentration: where translation becomes difficult

Researchers often want to move from MRI metrics to a more tangible statement about iron concentration. That ambition is understandable. A value in ppm or s⁻¹ is useful for imaging science, but the clinical and biological question usually concerns tissue iron, its distribution, and its relationship to neuronal injury.

The difficulty is that neither QSM nor R2* should be treated as a universally calibrated direct assay of iron concentration in the brain. The relationship depends on tissue composition, iron compartmentalization, magnetic field strength, sequence parameters, and the mathematical model used to reconstruct or fit the signal. There is also no single unified conversion formula that directly equates R2* values in s⁻¹ with absolute QSM values in ppm across different MRI field strengths.

Evidence from liver iron quantification shows why the translational direction is promising but technically conditional. In comparisons with SQUID-based biomagnetic liver susceptibility measurements, QSM-BLS demonstrated a strong linear relationship, with an r² of 0.88. Magnetic susceptibility also correlated strongly with iron concentration, with a Spearman correlation of 0.918. These findings support the broader concept that magnetic susceptibility can carry meaningful information about iron burden.

However, liver validation cannot simply be transferred to the substantia nigra. The brain is a structurally heterogeneous organ, and the distribution of iron across neuromelanin-containing neurons, oligodendrocytes, blood products, and other tissue compartments may not reproduce the physical conditions of hepatic iron accumulation. The biological organization of iron matters as much as the total amount. A regional susceptibility value may reflect the magnetic behavior of iron deposits without providing a one-to-one estimate of total iron mass.

For clinical trials, this suggests a more defensible hierarchy of interpretation:

  • use QSM or R2* as quantitative imaging biomarkers of iron-sensitive tissue change;
  • validate their regional and longitudinal behavior against clinical measures, pathology where available, or an independent imaging or biochemical reference;
  • avoid presenting a ppm or s⁻¹ value as an absolute iron concentration unless the conversion has been established for that specific tissue, sequence, field strength, and population;
  • report the acquisition and reconstruction details needed for another group to reproduce the endpoint.

This shift allows the imaging biomarker to remain clinically useful without asking it to claim more than it can support. A biomarker does not become stronger by being described as an absolute concentration when its calibration is conditional. It becomes stronger when its limits are explicit and its change over time is reproducible.

The practical limitations of R2* in neurodegeneration

R2* mapping has a significant advantage in that its interpretation begins with a familiar signal-decay model. The method is not free from confounding, but it does not carry the same dependence on a susceptibility reference region. For studies that need a relatively efficient and repeatable iron-sensitive sequence, this can make R2* attractive.

Its limitation is that faster signal decay is not a unique signature of one pathological process. R2* may respond to iron-related microscopic field variation, but the observed value can also be shaped by tissue microstructure, orientation effects, partial-volume averaging, and sequence-specific factors. In a small deep grey matter nucleus, these influences can become meaningful because the structure may occupy only a limited number of voxels and may border tissues with very different magnetic properties.

The result is a metric that can be highly reproducible yet less anatomically discriminating in some disease contexts. In the Parkinson’s disease comparison summarized here, R2* detected increased values in the substantia nigra, while QSM identified increased susceptibility in both the substantia nigra and red nucleus. That does not make the R2* finding uninformative. It means that the two contrasts may support different levels of regional inference.

R2* also remains difficult to convert into QSM-like susceptibility units. Values expressed in s⁻¹ and ppm arise from different measurement models, and no universal conversion should be assumed across field strengths or protocols. A study that combines R2* data from several centres therefore needs to harmonize the sequence and fitting procedure rather than attempting to erase the distinction through a simple mathematical transformation.

For a neurodegenerative trial, the most appropriate use of R2* may depend on the endpoint:

  • If the study is focused on a well-defined region such as the substantia nigra and has a standardized acquisition, R2* may offer a practical longitudinal marker.
  • If the hypothesis concerns spatially selective susceptibility changes across several deep grey matter structures, QSM may provide greater anatomical sensitivity.
  • If the treatment is expected to produce a small effect, collecting both metrics may help distinguish a shared iron-sensitive signal from a method-specific change.
  • If the study spans multiple scanners or sites, the QSM reference and reconstruction pipeline require particularly detailed harmonization, while R2* requires equally careful control of echo sampling, fitting, and signal quality.

The choice should therefore be made at the level of the biological question, not at the level of whichever map looks more visually compelling.

Designing a longitudinal endpoint that survives contact with reality

A cross-sectional comparison can show that one metric separates groups more effectively than another. A clinical trial asks a harder question: can the metric detect change within individuals while preserving a meaningful relationship with disease trajectory?

That requires the imaging endpoint to be designed as part of the trial rather than added after the acquisition protocol has already been fixed. The following decisions are especially consequential:

1. Define the anatomical target before examining treatment effects.

If the primary hypothesis concerns the substantia nigra pars compacta, the segmentation strategy should be specified in advance, with a plan for partial-volume effects and failed or ambiguous segmentations.

2. Preserve the acquisition across visits.

Echo times, echo spacing, spatial resolution, field strength, scan geometry, and reconstruction software can all influence the apparent trajectory. A scanner or sequence change should be treated as a methodological event, not an invisible background detail.

3. Lock the QSM reference framework.

The reference region, preprocessing sequence, background-field correction, and inversion algorithm should remain consistent across baseline and follow-up. In extreme iron overload, the reference choice can move the median susceptibility from 0.147 ppm to 0.279 ppm in the reported comparison, a difference large enough to overwhelm a subtle biological effect if the framework is changed casually.

4. Separate observer agreement from endpoint precision.

High ICC values are encouraging, but repeated scans, test–retest data, and site-level variance are also needed to establish how much change is required before a trajectory can be interpreted as more than measurement noise.

5. Use multimodal context rather than a single decisive number.

Clinical severity, cognitive measures, structural imaging, diffusion metrics, or other biomarkers may help determine whether a QSM or R2* shift follows the expected biological pathway. No single iron-sensitive metric should be used to claim disease modification on its own.

6. Prespecify the interpretation of negative findings.

A failure to detect a change with R2* does not prove that iron biology is unchanged, just as a QSM change does not automatically establish progressive iron deposition. The sensitivity and specificity of the selected metric define what can be concluded from a null result.

A useful protocol may therefore retain both QSM and R2* as complementary measures, particularly during method development or early-phase trials. QSM can provide increased sensitivity to regional susceptibility differences, while R2* offers an independent iron-sensitive contrast that may help reveal whether the result is robust to the measurement model. This paired strategy adds acquisition and analysis burden, but it can be worthwhile when the expected treatment effect is subtle and the biological stakes are high.

Which metric should a clinical researcher choose?

For the narrow question of localized brain iron detection in Parkinson’s disease, QSM has the stronger case in the evidence summarized here. It showed higher sensitivity and specificity than R2* mapping, and it detected susceptibility increases in both the substantia nigra and red nucleus. Those findings make it particularly attractive for studies in which anatomical specificity and the detection of subtle regional change are central.

R2* remains a credible and useful option. It is not obsolete, and its lack of QSM-style reference dependence can simplify certain longitudinal workflows. It may be especially valuable when the protocol has already been extensively standardized, when the target anatomy is clearly defined, or when investigators want an independent iron-sensitive measure alongside QSM.

The decision can be framed in three questions:

  • Is the primary objective to detect the most spatially specific susceptibility change possible?
  • Can the study team maintain a stable reference region and reconstruction pipeline over time?
  • Would a second, differently modeled contrast improve confidence in a subtle treatment effect?

If the first two answers are yes, QSM may be the preferred primary endpoint. If reference stability is uncertain or the study infrastructure favors a simpler signal-decay approach, R2* may be the more practical choice. If the trial is exploratory and the biological effect is expected to be small, collecting both can provide a stronger basis for later endpoint selection.

The mature position is not that QSM wins every comparison. It is that QSM and R2* answer related but non-identical questions, and that their differences become most visible precisely where clinical research needs the greatest precision: in small nuclei, over long follow-up periods, and in disease states where iron distribution is neither uniform nor static.

A biomarker is a trajectory, not a colour map

Quantitative susceptibility mapping has moved brain iron imaging toward greater regional sensitivity, but it has also made the analytical assumptions more visible. The reference region, the inversion algorithm, the acquisition, and the tissue under examination all participate in the final number. R2* mapping carries fewer of these particular referencing decisions, yet its signal remains an indirect measure whose relationship to iron depends on tissue and sequence context.

For clinical trials, this is less a reason for hesitation than a reason for precision. The most valuable endpoint will not necessarily be the metric with the highest contrast on a single scan. It will be the one that preserves a coherent biological trajectory across visits, sites, observers, and changing clinical states.

QSM offers a compelling path when localized susceptibility changes in structures such as the substantia nigra and red nucleus are central to the hypothesis. R2* offers a complementary path with continuing practical value, particularly when reproducibility and protocol stability are paramount. Consider the implications carefully: the future of iron-sensitive MRI is unlikely to be a simple replacement of one metric by another, but a more disciplined use of both, with each interpreted according to the biology it can genuinely support.

FAQ

Is QSM better than R2* for detecting Parkinson’s disease?
QSM has shown higher sensitivity and specificity than R2* mapping when distinguishing patients with Parkinson’s disease from normal controls, particularly regarding changes in the substantia nigra pars compacta.
Why does the choice of reference region matter in QSM?
QSM values are interpreted relative to a chosen baseline, meaning the reference region is a fundamental part of the measurement. Changes in the reference framework can lead to significantly different susceptibility estimates, which may be mistaken for biological changes in longitudinal studies.
Can R2* and QSM be used interchangeably?
No, they are not interchangeable. They measure different properties—R2* tracks transverse signal decay, while QSM reconstructs magnetic susceptibility—and they do not respond identically to the same anatomical patterns.
Are QSM and R2* direct measures of brain iron concentration?
No. Both are iron-sensitive imaging biomarkers, but their relationship to actual iron concentration depends on factors like tissue composition, magnetic field strength, and the specific mathematical models used for reconstruction.
Does high inter-rater reliability guarantee a valid biomarker?
High reliability indicates that observers can consistently measure the same structures, but it does not confirm biological validity. A measurement can be highly reproducible while still being influenced by processing artifacts or an inappropriate reference region.

Also interesting