The difference is large enough to make a biostatistician pause before approving a pooled analysis.
It is also consistent with the kind of vendor- and protocol-dependent variability that complicates cardiac T1 mapping in multi-center workflows. The scanners may operate at the same field strength and still produce materially different baseline values. Sequence implementation, acquisition timing, reconstruction, motion handling, fitting algorithms, and local protocol choices all contribute to the number that eventually appears in the report.
That is the uncomfortable position of T1 mapping in clinical research. The metric has been used in the literature for more than a decade, but clinical trials still cannot reliably pool raw native T1 values across sites without documenting the conditions under which those values were acquired. The unit is the same. The measurement is not necessarily interchangeable.
The vendors’ sequence families may look comparable on a protocol sheet. That does not make their outputs equivalent. The Society for Cardiovascular Magnetic Resonance has been documenting these caveats in consensus statements since 2013, usually in language that radiologists read carefully once and then have to explain to a principal investigator during a trial meeting.
Why native T1 won’t sit still across scanners
The uncomfortable truth about native T1 is that it is not a fixed tissue property in the operational sense. It is a sequence-specific estimate of a physical relaxation property.
T1 relaxation is a property of myocardial tissue at a given field strength and physiological state. But the way it is measured matters. The pulse sequence, inversion or saturation preparation, readout strategy, flip angle, heart-rate dependence, breath-hold performance, B1 homogeneity, motion correction, and fitting model all enter the final estimate. So does the post-processing pipeline supplied by the vendor or used by the research core lab.
That is why two values can both be internally valid and still fail to be directly comparable.
A map can shift because the acquisition samples the recovery curve differently. It can shift because the subject’s heart rate changes the effective timing of the readouts. It can shift because a patient does not reproduce the intended breath-hold, because the basal slice is affected by partial volume, or because the region of interest includes a slightly different amount of blood pool or epicardial fat. None of those sources of variation has to represent a scanner malfunction.
Field strength is the obvious lever. A 1.5T scanner and a 3T scanner produce systematically different baseline T1 values in the same patient, even when the sequence design is otherwise similar. But field strength is rarely the only variable in a real-world imaging network. It is often the variable a trial inherits from procurement decisions made years before enrollment begins.
The practical problem is therefore not simply that one vendor has a higher number and another has a lower one. The problem is that the offset may interact with the sequence, the scanner configuration, the reconstruction software, the heart rate, and the population being scanned. A value that looks abnormal against a published reference range may be perfectly ordinary for the local protocol. Conversely, a value that falls inside a familiar textbook range may not be directly comparable with a value acquired elsewhere.
What contributes to the observed value?
For a multi-center study, the most consequential sources of variation usually include:
- Field strength. Native T1 values at 1.5T and 3T should not be treated as if they came from a single scale.
- Pulse-sequence design. MOLLI, ShMOLLI, SASHA, and related approaches sample magnetization recovery differently.
- Timing parameters. Inversion times, saturation recovery timing, readout spacing, and the number of images used for fitting affect the estimate.
- Heart rate and rhythm. Inversion-recovery methods are particularly sensitive to the relationship between cardiac cycle length and acquisition timing.
- Motion and breath-hold quality. Misregistration can distort the map or alter the usable region of myocardium.
- B0 and B1 behavior. Imperfect field uniformity and regional transmit variation can affect the preparation pulse and readout.
- Segmentation and region-of-interest placement. A septal measurement is not automatically equivalent to a global myocardial measurement.
- Reconstruction and post-processing. Vendor-specific corrections and fitting procedures can move the reported value even when the source images appear similar.
This is why “harmonization” needs to be defined precisely. Harmonizing acquisition parameters is not the same as making the resulting raw T1 values numerically identical. A trial can standardize the protocol, centralize analysis, and reduce avoidable variation without eliminating the biological and technical differences between platforms.
If your trial’s reproducibility budget cannot absorb a 60-millisecond baseline spread, native T1 alone is not your endpoint. Pick a metric that survives the field.
MOLLI versus SASHA: the sequence you did not know was choosing
Cardiac T1 mapping has spent more than a decade balancing two methodological priorities: precision in routine clinical acquisition and fidelity to the underlying relaxation behavior. Inversion-recovery sequences, particularly MOLLI and its variants such as ShMOLLI, remain widely used because they offer strong signal-to-noise performance and can be incorporated into a familiar cardiac MRI examination. Saturation-recovery approaches, primarily SASHA, take a different route.
The difference is not cosmetic. MOLLI uses an inversion pulse followed by a series of single-shot readouts. The acquisition samples the recovery curve repeatedly, and a Look-Locker correction is applied during fitting. That correction helps account for the fact that the readouts themselves disturb magnetization. The resulting estimate is also influenced by heart rate and by the sequence’s sensitivity to other relaxation effects, including T2-related behavior.
SASHA uses a saturation preparation rather than the same inversion-recovery strategy. It is less affected by some of the biases associated with incomplete recovery and repeated inversion readouts, and validation studies have often treated it as closer to the underlying T1 under particular conditions. The trade-off is reduced precision in many in vivo settings. A method can be less biased in principle while producing a noisier measurement in practice.
That distinction matters in a clinical trial. “More accurate” and “more precise” are not interchangeable advantages. A sequence with tighter repeatability but a systematic offset may be useful for detecting change within a tightly controlled protocol. A sequence with a different bias profile may be preferable when the study prioritizes agreement with a reference measurement. Neither can simply be substituted for the other after data collection has begun.
The practical consequence is a methodological gap, not random noise. Differences on the order of tens of milliseconds, and sometimes roughly 50 to 100 milliseconds in comparisons of healthy myocardium, can emerge between MOLLI and SASHA depending on field strength, implementation, timing, and analysis. That is a design decision appearing on the scanner console as a different number.
| Parameter | MOLLI, such as 5(3)3 | SASHA |
|---|---|---|
| Magnetization preparation | Inversion recovery | Saturation recovery |
| Effect of repeated readouts | Present and addressed in the fitting model | Different recovery behavior with less dependence on the same Look-Locker correction |
| Heart-rate sensitivity | Relevant, particularly because acquisition timing follows the cardiac cycle | Generally lower for some sources of heart-rate dependence, but not absent |
| Typical in vivo precision | Often higher | Often lower |
| Main practical strength | Established clinical workflow and relatively precise maps | Reduced sensitivity to some inversion-recovery biases |
| Main limitation | Sequence- and heart-rate-dependent bias | Greater measurement variability in many clinical settings |
| Interchangeability across vendors | Not assumed without validation | Not assumed without validation |
The final row is the one that matters to the trial statistician. If one center uses MOLLI and another uses SASHA because of historical preference or local availability, the study does not have one native T1 endpoint in the strict sense. It has two related measurements with different error structures.
Even when both centers use MOLLI, the problem does not disappear. A sequence name is not a complete acquisition protocol. The implementation can differ in timing, readout, reconstruction, motion correction, fitting, and the way the vendor exports the map. “MOLLI” in a methods section tells the reader where to start, not whether the values are numerically interchangeable.
ECV is not a workaround for bad T1 data. It is a different measurement with a different error budget.
The 60-millisecond gap: quantifying inter-site variance at 3T
The 1163-versus-1225 comparison is a useful illustration of the problem because both sites operate at 3T, yet the healthy-myocardium baselines differ substantially. The difference is consistent with vendor- and protocol-dependent variability. It should not be presented as proof that a particular scanner age, calibration history, or phantom procedure caused — or did not cause — the discrepancy.
That distinction matters. Inter-site variation can be observed without knowing which individual component produced it. If the available data show different values from two sites, the responsible conclusion is that the measurement is not automatically portable across those sites. It is not that a specific calibration artifact has been ruled out, nor that the scanners’ manufacturing history explains the result.
In a multi-center cardiac MRI study, that uncertainty is not a semantic footnote. It changes the analysis plan.
Suppose Site A’s healthy reference population has a mean native T1 around 1163 ms and Site B’s is around 1225 ms. A patient-level value cannot be interpreted using the same raw cutoff at both locations unless the study has demonstrated that the two acquisition-and-analysis pipelines are equivalent. A disease-related change of 30 ms may be biologically meaningful, but it can be difficult to identify when the between-site baseline offset is of a similar or greater magnitude.
The statistical cost is equally straightforward. Uncontrolled site effects increase variance. Increased variance widens confidence intervals and reduces the ability to detect a treatment effect. The issue is not limited to the mean value: different sites may also have different dispersion, different proportions of technically inadequate maps, and different relationships between T1 and covariates such as heart rate or hematocrit.
This is one reason quantitative MRI needs more than a scanner field-strength label. A clinical trial database should capture enough acquisition metadata to explain how the number was produced. At a minimum, that usually means recording:
- scanner manufacturer and model;
- field strength;
- sequence family and exact variant;
- acquisition timing and relevant readout parameters;
- contrast status for post-contrast measurements;
- map reconstruction and post-processing version;
- motion-correction and segmentation procedures;
- region-of-interest definition;
- heart rate and rhythm information where relevant;
- the timing of hematocrit and post-contrast acquisition for ECV.
Centralized analysis can reduce some sources of variation, especially segmentation and quality-control differences. It cannot retroactively make acquisitions identical. Nor can a core lab infer every sequence detail from a final exported map. If the raw images and acquisition metadata are not retained, later harmonization becomes an exercise in approximation.
The most defensible multi-center approach is therefore layered. Harmonize the acquisition as far as the hardware allows. Keep the sequence implementation stable during the trial. Use a consistent analysis pipeline. Predefine how technically inadequate maps and outliers will be handled. Then choose an endpoint whose interpretation does not depend entirely on the raw native T1 value being universal.
That is why many studies treat native T1 as a locally normalized measure, while using ECV or another derived metric as a complementary endpoint. The goal is not to pretend the platforms are identical. The goal is to make the remaining differences visible and manageable.
ECV fraction as the metric that travels better
Extracellular volume fraction is what native T1 wants to be when it grows up — not because it eliminates technical error, but because it combines information in a way that can reduce the effect of some baseline offsets.
The calculation uses pre-contrast and post-contrast T1 values together with the patient’s hematocrit. It produces a percentage intended to estimate the proportion of myocardial tissue occupied by extracellular space. In healthy myocardium, ECV is commonly around the mid-twenties, while diffuse fibrosis and other processes that expand the extracellular compartment can increase it. The biological interpretation is different from that of native T1, but it is often more directly connected to the tissue compartment of interest.
The cross-platform advantage comes from the paired design. If a scanner or sequence implementation shifts both pre-contrast and post-contrast T1 in a similar direction, the relationship between the two measurements may be more stable than either raw value alone. A site with generally higher native T1 does not necessarily produce a proportionally higher ECV. The calculation can compress part of the platform-related offset.
That is not a guarantee of equivalence. ECV still depends on acquisition timing, contrast administration, hematocrit measurement, the contrast agent and dose, renal and circulatory factors, and the post-processing method. The assumption that pre- and post-contrast shifts behave similarly can also fail. A protocol that produces a modest difference in native T1 may not produce the same difference after contrast, particularly if timing or recovery conditions differ between sites.
Still, cardiac MRI extracellular volume fraction often offers a more stable basis for cross-vendor comparisons than raw native T1. For a clinical trial, that can be the difference between an endpoint that can plausibly be pooled across sites and one that requires extensive site adjustment before interpretation.
Where ECV helps — and where it does not
ECV is particularly attractive when the biological question concerns diffuse interstitial expansion, fibrosis, or longitudinal tissue change. It can be useful when the trial includes multiple scanner platforms and when the acquisition protocol can be controlled sufficiently to keep contrast timing consistent.
It does not solve every operational problem:
- A contrast-enhanced examination is required.
- Hematocrit must be measured and linked correctly to the scan.
- The interval between contrast administration and post-contrast mapping needs to be defined.
- The patient’s circulation and contrast distribution can affect the result.
- The formula and units must be implemented consistently.
- Differences in segmentation can propagate into the final percentage.
- Sites need a procedure for incomplete or mistimed hematocrit samples.
The logistical burden is real. ECV requires a venous blood draw, a contrast bolus, and a post-contrast acquisition that is timed within the study’s protocol. In a single-center research study, that may be routine. In a consortium with different scanner platforms, field strengths, patient flows, and local laboratory processes, it becomes an operations problem that needs to be designed rather than improvised.
A trial should also resist the temptation to describe ECV as inherently “vendor-independent.” It is better described as less dependent on the raw native T1 scale under a defined protocol. That wording is less dramatic, but it is more accurate and more useful when writing a statistical analysis plan.
From universal cutoffs to site-specific reference ranges
The SCMR position on parametric mapping is cautious for a reason: native T1 values depend on sequence, vendor, field strength, and local implementation. There is no single universal native T1 cutoff that can be transferred across every scanner and patient population.
That does not make native T1 clinically useless. It changes how the measurement should be interpreted. Each center needs a reference range established with its own scanner, sequence, and analysis workflow. The reference population should be characterized well enough to reflect the population in which the measurement will be used, with relevant demographic and physiological factors taken into account.
The exact sample size and statistical method depend on the study’s purpose and available resources. A local reference cohort is not a decorative appendix to the protocol. It is the baseline against which the site’s patient measurements become interpretable.
In practice, the process has several linked parts:
1. Define the local acquisition. Fix the scanner, field strength, sequence variant, timing parameters, and reconstruction version. A reference range generated before a sequence change may not remain valid after the change.
2. Recruit an appropriate reference population. Healthy volunteers should be screened and characterized according to the study’s clinical context. Age, sex, heart rate, and other relevant factors may affect the distribution.
3. Use a consistent analysis method. The same segmentation rules, exclusions, and quality-control procedures should be applied to reference and patient scans.
4. Estimate the local distribution. Report the center, spread, and clinically useful limits rather than relying on a single average.
5. Document changes over time. Software updates, sequence revisions, coil changes, or scanner maintenance can alter the measurement and may require revalidation.
6. Normalize multi-center data before pooling. Raw values can be retained for transparency, but the primary analysis may need site-specific Z-scores or another prespecified adjustment.
For a multi-center trial, the implication is layered. Each site needs enough local information to understand its own baseline. The coordinating team needs a harmonized acquisition protocol. The core lab needs a consistent method for reading and adjudicating maps. The statistician needs the site effect represented in the analysis rather than discovered after unblinding.
A site-specific Z-score is one practical solution. It expresses a patient’s native T1 relative to the local reference distribution, making the value more interpretable across centers than the unadjusted millisecond measurement. This does not create a universal biological scale out of thin air. It makes the local measurement comparable in terms of deviation from that site’s expected baseline.
That distinction is important for clinical trials. A Z-score can support pooling when the trial has prespecified how reference data are generated and how uncertainty in those estimates is handled. It should not be used as a cosmetic conversion applied after the fact to rescue incompatible protocols.
What changes in the protocol
The first decision is biological: determine whether the disease process and trial question are better represented by native T1, ECV, or both. If the hypothesis concerns diffuse extracellular expansion, ECV may offer a more portable endpoint, provided the study can support the contrast and hematocrit workflow. If native T1 is central to the biology, it should not be discarded simply because it is technically demanding. It should be normalized and interpreted within a protocol that acknowledges its dependencies.
The second decision is operational: freeze what can be frozen. A trial should not casually change sequence variants, reconstruction software, or segmentation rules midway through enrollment. If a change is unavoidable, the study needs a bridging procedure that quantifies its effect rather than assuming continuity.
The third decision is statistical: define the site effect before the data arrive. That may involve site-specific reference ranges, site-adjusted models, Z-scores, stratification, or a combination of these approaches. The method should match the endpoint and the expected biological effect. A raw native T1 comparison is rarely defensible when the acquisition pipelines have not been shown to agree.
The practical documentation should include:
- scanner manufacturer, model, and field strength;
- exact sequence name and implementation;
- sequence timing and readout details;
- software and reconstruction versions;
- contrast agent and timing for ECV;
- hematocrit source and collection time;
- image-quality exclusions;
- segmentation rules;
- reference-population characteristics;
- procedures for scanner or software changes.
When a manufacturer describes its measurements as comparable with another platform, the right response is not to dismiss the claim or accept it as a guarantee. Ask what was compared, under which sequence conditions, at which field strength, using which reference object or patient cohort, and whether the clinical data support the proposed interchangeability. A phantom comparison may be useful, but it does not by itself establish that patient-level native T1 values will behave identically across sites.
The same caution applies to calibration language. A site-to-site difference should not automatically be labeled a phantom-calibration artifact, but neither should calibration effects be declared absent without evidence. The responsible interpretation is narrower: the observed difference is compatible with vendor- and protocol-dependent variability, and the trial should be designed accordingly.
Native T1 is not a universal number with a scanner attached. It is the output of a measurement system, and the system travels with the protocol.
The endpoint is only as strong as its comparability
Cardiac T1 mapping is not broken. It is simply not the vendor-agnostic, drop-in biomarker that a clean-looking map can suggest.
The field has moved beyond the question of whether T1 mapping is useful. The harder question is how to use it without confusing numerical precision with biological comparability. A map can be sharp, reproducible within one center, and clinically informative while still failing to share a universal scale with a map acquired elsewhere.
For clinical research, the answer is not to abandon quantitative MRI. It is to stop treating native T1 as self-interpreting. Sequence design matters. Vendor implementation matters. Field strength matters. Local reference data matter. So do contrast timing, hematocrit, segmentation, and the statistical model used to pool sites.
ECV can provide a more stable cross-platform endpoint when the disease biology and study logistics support it. Site-specific Z-scores can make native T1 more useful when the raw millisecond values cannot be pooled directly. Centralized analysis can reduce avoidable variation, but it cannot erase differences that were built into the acquisition.
The core principle is simple: cardiac MRI T1 mapping vendor variation is not a nuisance to be hidden in a footnote. It is part of the measurement model. Once the trial acknowledges that, the choices become clearer: standardize where possible, normalize where necessary, and never call two numbers interchangeable merely because they carry the same unit.
