The reconstruction then converts those measurements into an estimate of tissue iron. The shared sequence does not produce shared biomarkers.
R2* measures the transverse relaxation rate, 1/T2*. QSM estimates magnetic susceptibility from phase data. One is a decay-rate measurement. The other is an inversion problem constrained by the scanner’s field geometry, background-field correction, reference choice, and reconstruction model. Treating them as interchangeable metrics is not a harmless simplification. It changes what the trial is measuring.
For clinical iron quantification, the practical question is not whether QSM is newer or R2* is more established. The question is narrower and more consequential: which metric preserves biological meaning when fat, fibrosis, focal lesions, cellular pathology, scanner differences, and longitudinal acquisition drift enter the dataset?
The physics of iron quantification: R2* versus QSM
Iron changes the local magnetic field. That field perturbation affects the gradient-echo signal in two observable ways.
The first is magnitude decay. In an idealized model, the signal falls as echo time increases, and the fitted decay rate is R2*. Higher iron concentration generally produces stronger microscopic field variation and faster signal loss. R2* therefore acts as an indirect iron-sensitive metric. It is computationally efficient. It is widely implemented. It has substantial vendor and clinical familiarity.
The second is phase accumulation. Local susceptibility differences alter the phase of the measured MR signal. QSM uses those phase shifts to estimate the underlying tissue magnetic susceptibility, typically expressed in parts per million relative to a selected reference. In principle, this is closer to the physical source of the contrast. In practice, the reconstruction is not a direct readout. The measured phase contains tissue effects, background fields, phase wraps, noise, and geometry-dependent terms. QSM must remove and invert them.
That distinction determines the failure modes.
R2* is often described as robust because its pipeline is comparatively short: acquire multi-echo magnitude data, fit signal decay, and report a rate. But the rate is not specific to iron. It responds to any factor that changes effective transverse relaxation. Fat, collagen, fibrosis, focal lesions, and gadolinium-related effects can alter the relationship between tissue iron and R2*. The dependency is nonlinear. A single R2* value can therefore represent different tissue states under different pathological conditions.
QSM has the opposite profile. It carries more physical information, but it requires more mathematical control. The reconstruction must account for phase unwrapping, removal of background fields, and dipole inversion. The dipole kernel is ill-conditioned around its magic-angle cone. Noise and residual background field can be amplified during inversion. A susceptibility value without a specified algorithm and reference region is not a fully defined biomarker.
| Parameter | R2* mapping | Quantitative susceptibility mapping |
|---|---|---|
| Primary input | Multi-echo magnitude signal, usually with optional phase support | Multi-echo phase data from gradient-echo acquisition |
| Reported quantity | Transverse relaxation rate, 1/T2* | Relative tissue magnetic susceptibility |
| Relationship to iron | Indirect and pathology-dependent | More directly linked to magnetic susceptibility from paramagnetic iron |
| Main confounders | Fat, fibrosis, collagen, lesions, gadolinium, signal-model mismatch | Reference region, background-field removal, phase wraps, dipole inversion, SNR |
| Reconstruction burden | Lower; curve fitting dominates | Higher; several mathematically coupled processing stages |
| Multicenter maturity | More established and vendor-standardized | Still dependent on harmonized acquisition and post-processing |
| Best role in trials | Established quantitative endpoint or comparator | Mechanistically informative endpoint when reconstruction is controlled |
The distinction between acquisition and biomarker is central. A standardized multi-echo sequence does not guarantee a standardized QSM result. Conversely, a standardized R2* fit does not eliminate pathology-related bias.
The scanner is a mathematical engine. It does not measure “iron” as a clinical noun. It records complex signal samples under a pulse sequence and gradient trajectory. The biomarker exists only after a model assigns biological meaning to those samples.
R2* is easier to deploy. QSM is closer to the underlying field perturbation. Neither survives an uncontrolled pipeline.
Why pathology breaks a simple R2* interpretation
R2* mapping performs well when the tissue environment is sufficiently stable and the dominant source of dephasing is the parameter under investigation. Liver iron assessment is the obvious use case, but even there the signal model is not trivial. Fat and iron occupy the same magnitude data but do not produce identical spectral behavior. Fibrosis changes tissue microstructure. Focal lesions add regional heterogeneity. Contrast agents can alter relaxation and susceptibility. The fitted R2* rate becomes a composite response.
This is not a minor correction term. It affects endpoint validity.
Suppose a clinical trial tracks hepatic iron across serial scans. A decrease in R2* may indicate reduced iron. It may also reflect a change in fat fraction, tissue composition, lesion burden, or acquisition parameters. If the confounding variables move with treatment, the biomarker can produce a plausible but biologically misassigned treatment effect.
The problem becomes sharper at high iron concentration. Signal decay can become very rapid. Later echoes fall into a low-SNR regime, while early echoes may be affected by dead time, imperfect excitation, or insufficient temporal resolution. A mono-exponential fit can then become structurally wrong even if the residuals appear acceptable. The fitting routine yields a number. The number is not automatically a valid estimate.
A robust R2* protocol must therefore constrain several interacting variables:
- Echo spacing must resolve the early decay without discarding usable signal.
- The first echo must be early enough to capture heavily iron-loaded tissue.
- The final echoes must retain adequate SNR rather than contribute noise-dominated curvature.
- Fat-water behavior must be represented by an appropriate signal model where required.
- The same fitting method must be applied across scanners and time points.
- Regions of interest must avoid necrosis, large vessels, focal lesions, and severe partial-volume contamination.
- Acquisition geometry and motion handling must remain stable across longitudinal visits.
These are not administrative details. They determine the mapping from signal to endpoint.
QSM reduces some of the pathology interference because it estimates susceptibility from phase rather than relying only on magnitude decay. Fat and fibrosis can distort R2* through their effects on relaxation and microstructure. QSM is less sensitive to those confounding cellular pathologies in the estimation of susceptibility. That advantage is meaningful, but it should not be overstated. QSM does not make tissue composition irrelevant. It relocates the technical burden from the decay model to phase processing and spatial inversion.
In a clinical trial, that trade can be favorable when pathology is expected to vary independently of iron. A susceptibility estimate may preserve iron-related information that R2* compresses or misinterprets. But the reconstruction must be frozen, documented, and validated before the endpoint is analyzed.
QSM is a reconstruction chain, not a single measurement
The phrase “QSM value” conceals a sequence of operations. Each operation can shift the final susceptibility distribution.
A typical QSM pipeline begins with phase images from several echo times. The phase must be unwrapped because the scanner records phase modulo 2π. The background field must then be estimated and removed. Sources outside the tissue of interest can create smooth but substantial field variation. If that field remains in the data, the inversion assigns external contributions to internal susceptibility.
The local field is then related to susceptibility through the dipole kernel. In Fourier space, the kernel contains a conical region where its values approach zero. Direct inversion becomes unstable there. Regularization, morphology constraints, total-variation penalties, nonlinear formulations, or other priors are used to suppress noise amplification and streaking. Different choices yield different spatial behavior.
The result is not merely a technical nuisance. Regularization changes the balance between spatial sharpness, noise, and quantitative bias. A method that looks cleaner may suppress small deposits. A method that preserves high-frequency detail may generate stronger streak artifacts. Two pipelines can process identical raw data and report different regional susceptibility values.
The core processing layers include:
1. Complex signal formation. Multi-echo gradient-echo data must preserve phase fidelity. Coil combination and vendor reconstruction can affect the phase input before any research algorithm begins.
2. Phase unwrapping. Wrapped phase creates discontinuities unrelated to tissue susceptibility. Errors here propagate into the local field map and can remain invisible after smoothing.
3. Background-field removal. The algorithm must separate local tissue-generated fields from fields generated outside the target volume. Boundary conditions and mask geometry matter.
4. Dipole inversion. The ill-conditioned kernel requires regularization. The selected solver defines the trade-off between noise, resolution, and bias.
5. Referencing. QSM values are relative. The chosen reference region establishes the zero point. A whole-brain reference and a cerebrospinal-fluid reference do not produce the same reported susceptibility.
6. Quality control. Residual phase artifacts, mask erosion, streaking, signal voids, and registration errors must be tracked rather than discarded silently.
R2* also requires quality control, but its mathematical dependence is narrower. QSM’s added physical specificity comes with a larger analytical state space. In a single-center research study, that may be manageable. In a multicenter trial, it becomes a governance problem.
The sequence protocol must specify more than field strength and voxel size. It should define echo train, bandwidth, flip angle, acquisition orientation, receiver handling, phase storage, motion correction, segmentation, preprocessing, and reconstruction version. If one site exports vendor-filtered phase and another exports raw or differently scaled phase, the downstream comparison is already compromised.
Validation: QSM-BLS against SQUID-BLS
Validation against an independent reference is where the QSM argument becomes more concrete.
Biomagnetic liver susceptometry using QSM, or QSM-BLS, has been compared with SQUID-based biomagnetic liver susceptometry. In patients with liver iron overload, linear regression produced a coefficient of determination of r² = 0.88. The reported relationship was:
QSM-BLS = (-0.22 ± 0.11) + (0.49 ± 0.05) · SQUID-BLS
That result is substantial. It indicates that QSM-derived susceptibility tracks an established external susceptometry benchmark across iron-loading conditions. It does not mean that QSM and SQUID are numerically interchangeable. The slope is not unity. The intercept is not zero. The relationship is a calibration, not a declaration that both systems measure on the same native scale.
The distinction matters for trial design. A biomarker can correlate strongly with a reference while still requiring population-specific calibration, scanner harmonization, and a prespecified analysis model. An r² of 0.88 does not establish universal cutoffs. It does not resolve the effect of field strength, sequence design, coil configuration, reference region, or reconstruction method. It does establish that the physical information contained in MRI phase can support clinically relevant iron estimation.
R2* has a longer history as an operational metric. It is integrated into established workflows and is more readily standardized across vendors. That maturity matters to regulators and trial operators. A metric that can be deployed consistently across sites may outperform a theoretically more direct metric whose implementation changes from laboratory to laboratory.
This is the central comparison:
- R2* offers procedural stability with biological ambiguity.
- QSM offers stronger physical interpretability with computational variability.
The correct endpoint depends on which risk the protocol can control.
For a trial with a narrow, well-characterized population and stable tissue composition, R2* may remain the more defensible primary endpoint. Its established status reduces operational uncertainty. QSM can serve as a mechanistic secondary endpoint, particularly when the study aims to characterize spatial iron distribution or separate iron-related susceptibility from pathology that distorts relaxation.
For a study involving heterogeneous liver disease, neurological iron accumulation, or longitudinal tissue remodeling, QSM may provide information that R2* cannot preserve reliably. But that benefit should be tested against an external standard, not inferred from visual map quality.
The validation package should include, at minimum:
- Repeatability across the intended acquisition interval.
- Reproducibility across scanners and sites.
- Sensitivity to changes in echo timing and SNR.
- Stability under the selected segmentation and registration workflow.
- Agreement with an accepted external or biochemical reference where available.
- Prespecified handling of failed phase unwrapping, incomplete echoes, and severe signal dropout.
- Frozen reconstruction software with versioned parameters and audit logs.
A map that looks anatomically plausible is not enough. Quantitative imaging fails quietly. The image remains smooth. The endpoint shifts.
A strong correlation validates a relationship, not a universal cutoff. Calibration remains part of the biomarker.
Reference regions become decisive in extreme iron overload
QSM does not produce an absolute susceptibility value without a reference. The reported number is relative to whatever region is assigned as the baseline. This is usually acceptable only when the reference is stable, biologically appropriate, and consistently segmented.
Extreme iron overload exposes the problem.
In patients with aceruloplasminemia, QSM brain susceptibility values differed substantially depending on the reference region. Using a whole-brain reference produced a median value of 0.147 ppm. Using cerebrospinal fluid produced a median of 0.279 ppm. The difference is not a biological contradiction. It is a consequence of the reference definition.
That magnitude is large enough to alter interpretation across cohorts. If one site references the whole brain and another uses cerebrospinal fluid, the resulting values cannot be compared as though they were measurements on one common scale. Even within a single site, a reference region can become unstable if disease affects the tissue assumed to be normal.
Whole-brain referencing can be problematic in diffuse disease. If iron deposition is widespread, the reference distribution itself shifts. The measurement becomes relative to an altered brain. Cerebrospinal fluid may offer a different baseline, but it introduces its own segmentation and partial-volume constraints. Small errors at tissue boundaries can matter because susceptibility estimates are sensitive to mask geometry and background-field handling.
Reference selection must therefore be treated as part of the endpoint definition, not as a post-processing preference. The protocol should specify:
- The reference tissue and anatomical boundaries.
- Whether the reference is global or regional.
- How the reference mask is eroded to reduce partial volume.
- How pathology within the reference region is handled.
- Whether the same reference is used for all visits and scanners.
- Whether susceptibility values are reported relative to the reference or transformed through a calibration model.
In neurological trials, this issue is especially severe when disease affects multiple deep-gray nuclei, white matter, and surrounding structures. A global reference can improve numerical stability while reducing biological specificity. A local reference can preserve regional contrast while increasing sensitivity to segmentation and disease spread.
The choice cannot be optimized after inspecting treatment results. That would allow the reference definition to absorb the very effect the trial is meant to measure.
R2* validation remains relevant
The rise of QSM does not invalidate R2*. R2* remains the regulatory-established primary metric in many clinical contexts, and its implementation is supported by mature vendor workflows. The practical advantage is not merely historical. A well-controlled R2* protocol can provide reliable longitudinal contrast when the tissue environment and signal model are known.
R2* is also useful as a companion metric. QSM and R2* respond to related but nonidentical aspects of the same iron-induced field behavior. Discordance between them can be informative.
A high R2* with modest QSM may indicate that relaxation is being affected by factors beyond bulk susceptibility. A strong QSM signal with unstable R2* can indicate rapid decay, low late-echo SNR, fat-related model error, or pathology-driven relaxation changes. The disagreement should not be averaged away. It is a diagnostic signal about the tissue and the acquisition.
For brain iron, the interpretation is further constrained by anatomy. Susceptibility is influenced by iron oxidation state, microstructural organization, myelin, calcium, and local geometry. R2* can be sensitive to iron but remains an indirect relaxation measure. Neither metric should be treated as a standalone molecular assay.
For liver iron, QSM can reduce interference from fat and fibrosis relative to R2* estimation. That is precisely where a dual-endpoint design can be useful: retain R2* for continuity with established evidence, while use QSM to test whether susceptibility provides a more stable biological relationship across pathological strata.
The trial should not ask which map is prettier. It should ask which metric maintains rank order, effect size, and repeatability under the conditions that matter clinically.
Multicenter reproducibility is the actual barrier
The most difficult question in QSM versus R2* mapping is not whether either metric can be computed. Both can. The difficult question is whether a value from Site A remains commensurate with a value from Site B after the scanner, coil, sequence implementation, reconstruction software, and reference definition have changed.
R2* has an advantage because the computation is less dependent on a long chain of inversion choices. Vendor-standardized implementations still differ, but the space of plausible outputs is narrower. QSM requires deeper harmonization. The field-to-susceptibility inversion is sensitive to acquisition orientation, echo sampling, phase quality, mask generation, and regularization.
The ISMRM Electro-Magnetic Tissue Properties Study Group has issued consensus recommendations intended to standardize QSM implementation in clinical brain research. That consensus work is important because it identifies the problem correctly: QSM adoption depends on methodological transparency and reproducibility, not only on physical plausibility.
Consensus guidance does not erase site variability. It establishes a common technical language. A trial still needs phantom testing, traveling-subject or repeat-scan data where feasible, and a central analysis pipeline. Raw-data retention is preferable when permitted. If only processed maps are stored, the study loses the ability to re-evaluate phase scaling, background-field correction, or inversion behavior after the protocol is locked.
A defensible multicenter QSM workflow should include:
- A common multi-echo gradient-echo protocol with controlled echo spacing and bandwidth.
- Scanner-specific calibration or cross-platform harmonization.
- A shared phase-scaling and coil-combination strategy.
- Centralized or identically containerized reconstruction.
- Fixed reference-region definitions.
- Phantom data acquired at every site.
- Automated artifact detection followed by expert review.
- Full recording of software version, parameter values, masks, and exclusions.
- Prespecified repeatability thresholds before unblinding treatment results.
R2* also benefits from centralized fitting and phantom controls. The difference is degree, not category. No quantitative MRI endpoint should rely on undocumented vendor defaults.
Choosing the endpoint by trial objective
A clinical trial should select the metric according to its estimand.
If the objective is continuity with existing regulatory evidence, broad deployment, and operational simplicity, R2* remains the safer primary choice. It has established use, recognizable failure modes, and greater vendor standardization.
If the objective is spatially resolved susceptibility, mechanistic characterization, or improved resistance to fat and fibrosis-related interference, QSM deserves a prominent role. It may be primary in a tightly controlled protocol, but only when the reconstruction and reference framework are validated before enrollment data are interpreted.
If the objective is biological confidence rather than procedural minimalism, the strongest design may use both. R2* supplies continuity. QSM supplies a susceptibility-based comparator. Their agreement strengthens interpretation. Their divergence identifies confounding and should trigger investigation rather than automatic exclusion.
A practical decision sequence is:
1. Define the biological target. Is the endpoint total iron burden, regional deposition, treatment-associated change, or mechanistic tissue characterization?
2. Characterize pathology. Fat, fibrosis, lesions, calcification, edema, and heterogeneous cellular composition all alter the error profile.
3. Select the signal model. Do not force a mono-exponential R2* fit onto data with rapid decay, fat-water interference, or severe low-SNR truncation.
4. Freeze the QSM pipeline. Specify phase handling, unwrapping, background removal, inversion, regularization, mask generation, and reference region.
5. Validate externally. Use an accepted benchmark where available. The QSM-BLS relationship with SQUID-BLS, including r² = 0.88, supports validation feasibility but does not remove the need for local calibration.
6. Test multicenter behavior. A single-site correlation is not a multicenter endpoint qualification.
7. Prespecify discordance analysis. If QSM and R2* disagree, the protocol should define how that disagreement will be examined.
The regulatory question is still unresolved
QSM has not fully replaced R2* mapping in clinical trials. That is not a failure of QSM. It reflects the difference between a physically attractive metric and a regulatory-ready endpoint.
Universal multicenter cutoff values for QSM susceptibility are not established across scanner manufacturers without standardized post-processing pipelines. The missing component is not simply more data. It is controlled comparability. A cutoff is meaningful only when the measurement scale is stable across the population, scanner platform, acquisition protocol, reconstruction version, and reference definition.
R2* is closer to that operational condition, although it is not immune to model error or pathology-related bias. Its established status makes it easier to defend. QSM may offer stronger biological specificity in selected settings, but its evidence must include the entire pipeline, from complex signal acquisition through final referenced map.
That requirement is demanding by design. A quantitative biomarker should be difficult to fool.
The most defensible current position is therefore conditional:
- Use R2* when established deployment, continuity, and simpler cross-site implementation dominate.
- Use QSM when susceptibility specificity and reduced interference from fat or fibrosis address a defined weakness of R2*.
- Use both when the trial can support central processing and wants to separate relaxation effects from susceptibility effects.
- Do not combine values from different QSM references or reconstruction pipelines without explicit harmonization.
- Do not convert correlation into interchangeability.
QSM versus R2* for MRI iron quantification is not a contest between an old metric and a superior successor. It is a choice between two different estimators of iron-related magnetic behavior. R2* compresses complex tissue effects into a relaxation rate and yields a mature, deployable endpoint. QSM solves a harder inverse problem and can yield a more direct susceptibility measure, but only under tighter control.
For clinical trials, the winner is not the metric with the stronger theoretical claim. It is the metric whose bias, variance, reference scale, and reconstruction behavior remain known after the data leave the originating scanner. Until QSM reaches that level of multicenter standardization, R2* remains the operational baseline. QSM remains the more informative instrument when the protocol is rigorous enough to deserve it.
