Clinical Research & Biomarkers

Feature Tracking or Tagging for Cardiac Strain MRI

Cardiac MRI strain analysis is constrained by how the myocardium is encoded before reconstruction. Myocardial tagging writes a temporary spatial pattern into the tissue.

Feature Tracking or Tagging for Cardiac Strain MRI

Feature tracking infers motion from anatomical boundaries already present in routine cine images. One measures deformation from an imposed signal grid. The other estimates deformation from image texture and contour evolution.

That distinction is not cosmetic. It determines acquisition time, spatial and temporal resolution requirements, failure modes, and whether a reported strain value can be transported across scanners, vendors, and clinical trials.

For global longitudinal strain and global circumferential strain, cardiac MRI feature tracking and myocardial tagging often correlate strongly. Reported Pearson correlation coefficients range from 0.86 to 0.92 for GLS and from 0.85 to 0.94 for GCS. These are useful results. They are not evidence that the methods generate interchangeable regional measurements.

Feature tracking removes the tagging acquisition. It does not remove the measurement problem.

The evolution of myocardial deformation assessment

Myocardial tagging was introduced in the late 1980s. The seminal methodology published by Zerhouni and colleagues in 1988 established the basis for non-invasive assessment of intramyocardial deformation with MRI. The principle remains mechanically direct.

A spatial modulation of magnetization is applied to the myocardium. The resulting tag lines, grids, or other patterns move with the tissue during the cardiac cycle. Their displacement is tracked over time. The deformation field is then estimated from that displacement.

The scanner is not merely observing the myocardium. It is placing a measurable coordinate system inside it.

This is why myocardial tagging has long been treated as the non-invasive reference standard for regional myocardial strain. The method supplies an explicit motion feature. The tag pattern is tied to tissue at the time of preparation. Its evolution provides information about local deformation that does not depend entirely on endocardial or epicardial border visibility.

The cost is acquisition complexity. Tagging requires a dedicated pulse sequence. It adds scan time. It may reduce spatial resolution, constrain temporal sampling, and suffer from tag fading as the sequence progresses. In a clinical workflow already limited by breath-hold duration, every additional acquisition has operational consequences.

CMR feature tracking takes the opposite route. It uses routine cine steady-state free precession images, generally acquired for ventricular volume and function assessment. No dedicated tag pattern is applied. No additional tagging scan is required. The deformation analysis occurs during post-processing.

This changes the economics of the measurement. Existing cine data can be reanalyzed retrospectively. Archived studies become available for quantitative work. Longitudinal cohorts can be screened without repeating the MRI examination. Clinical research can obtain strain metrics without modifying the acquisition protocol.

The signal source, however, is less direct. Feature tracking software follows image features, contours, or intensity patterns through the cine series. It estimates myocardial motion from visible anatomy. That makes the result dependent on image quality, segmentation, temporal resolution, spatial resolution, and the tracking model implemented by the software vendor.

Tagging encodes motion into the tissue. Feature tracking reconstructs motion from the image.

Those are different inverse problems.

Dedicated pulse sequences versus post-processing cine SSFP

The practical difference between the methods begins in the acquisition layer.

Myocardial tagging

Tagging is a sequence-level intervention. The acquisition must be planned around the tag preparation and the intended deformation analysis. The operator accepts additional scan time in exchange for a more explicit tissue-tracking signal.

A tagging dataset can be analyzed in several ways, including measurement of displacement, strain, strain rate, and temporal features such as time to peak strain. The tag geometry provides a basis for estimating deformation throughout the myocardium rather than relying exclusively on the visibility of anatomical interfaces.

The method is not free of constraints:

  • Tag contrast can decay during the cardiac cycle, degrading late-diastolic tracking.
  • The sequence adds a dedicated acquisition to the examination.
  • Breath-hold demands increase.
  • Resolution and temporal sampling remain protocol-dependent.
  • Analysis may require specialized software and experienced quality control.
  • Short-axis and long-axis acquisitions can produce different sensitivity to through-plane motion and regional deformation.

Tagging is therefore physically informative but operationally expensive. It is a strong reference method, not automatically the best method for every cohort.

CMR feature tracking

Feature tracking begins with cine SSFP images. These images are already part of standard cardiac MRI in many protocols. The software identifies the endocardial and epicardial boundaries, propagates them across frames, and calculates myocardial deformation from the evolving contours or internal image features.

This is attractive for clinical research for an obvious reason: the acquisition burden is low. A trial can collect cine imaging first and derive strain afterward. A retrospective database can be mined without the missing tagging sequence becoming an absolute exclusion criterion.

The post-processing model introduces different constraints:

  • Poor endocardial definition can produce unstable contour propagation.
  • Papillary muscles and trabeculations complicate border placement.
  • Through-plane motion can make a feature appear to move because the imaging plane is no longer sampling the same tissue.
  • Low temporal resolution reduces the fidelity of peak strain timing.
  • Low spatial resolution can blur thin myocardial segments and distort local deformation.
  • Vendor-specific algorithms can produce systematically different values from the same cine dataset.

The software is not reading a universal physical quantity directly. It is estimating deformation through a particular segmentation and tracking pipeline.

This point is routinely diluted in clinical reporting. A GLS value looks like a simple scalar. In reality, it is the output of an acquisition protocol, reconstruction process, segmentation decision, tracking algorithm, temporal interpolation scheme, and quality-control threshold.

A strain number without its measurement chain is incomplete.

The hardware layer still matters

Feature tracking is post-processing, but it is not hardware-independent. Scanner field strength, gradient performance, coil configuration, parallel imaging, compressed sensing, and reconstruction choices affect the cine images supplied to the algorithm.

Gradient slew rate and gradient amplitude constrain how rapidly the sequence can traverse k-space. The selected k-space trajectory and acceleration factor influence spatial resolution, temporal footprint, and artifact behavior. These variables are not abstract engineering details. They determine whether the software receives sharply defined ventricular borders or a temporally efficient but spatially compromised cine series.

A feature-tracking pipeline may tolerate modest variation in image quality when calculating global strain. It will tolerate less when calculating segmental peak strain or regional timing. Averaging across the ventricle suppresses some local errors. Segmental analysis exposes them.

This is the first major separation between a clinically useful global biomarker and a fragile regional metric.

ParameterMyocardial taggingCMR feature tracking
AcquisitionDedicated tagging pulse sequenceRoutine cine SSFP images
Additional scan timeRequiredNot required for the analysis
Motion signalTag pattern moves with tissueImage features and contours are tracked
Workflow burdenHigher acquisition and analysis burdenLower acquisition burden; software-dependent analysis
Global strain agreementReference comparison methodStrong correlation reported for GLS and GCS
Segmental peak magnitudeMore direct deformation encodingMore sensitive to segmentation and image quality
Retrospective useLimited to datasets with taggingBroad, if suitable cine images exist
Cross-vendor portabilityStill protocol-dependentParticularly vulnerable to software variability

What the correlation numbers actually establish

The central result in the cardiac MRI feature tracking versus myocardial tagging literature is straightforward. At the global level, agreement can be strong.

Reported correlations for global longitudinal strain range from 0.86 to 0.92. For global circumferential strain, reported correlations range from 0.85 to 0.94. These values indicate that subjects with higher or lower global deformation by one method tend to rank similarly by the other.

That is useful for group comparisons, risk stratification research, and exploratory biomarker development. It supports the idea that routine cine images contain sufficient information to recover a clinically meaningful global deformation signal.

Correlation, however, is not interchangeability.

A correlation coefficient measures association. It does not prove that two methods produce the same absolute value. Two methods can correlate strongly while maintaining a systematic bias, a proportional bias, or a clinically meaningful spread of differences across the measurement range.

The distinction becomes more severe at the segmental level. Comparative studies report closer agreement for temporal strain metrics, such as segmental time to peak strain, than for absolute peak segmental strain magnitude.

That result is physically plausible. Timing can remain relatively stable even when local contour position or magnitude is imperfectly estimated. A tracking algorithm may identify the phase of maximal deformation with reasonable consistency while misestimating the exact amplitude of that deformation.

Reported mean differences illustrate the asymmetry. For peak strain magnitude, one comparative result reported a mean difference of approximately 1 ± 9%. For time to peak strain, the reported mean difference was approximately 1 ± 58 ms. These values should not be treated as universal conversion factors. They describe agreement within a specific comparison context, not a calibration rule for every scanner, sequence, software package, or patient population.

The important observation is not that one method wins every metric. It is that the error structure changes with the metric.

Global strain benefits from spatial averaging. Regional peak magnitude does not. Temporal measures may show a different agreement profile from amplitude measures. A clinical trial that treats all strain outputs as equally robust is already misclassifying its endpoints.

Strong global correlation supports substitution of acquisition burden. It does not authorize substitution of measurement definitions.

Why segmental strain is the harder problem

The left ventricle is not a rigid shell moving synchronously. Its deformation varies by location, wall thickness, loading condition, conduction pattern, scar distribution, and contractile state. Segmental strain therefore carries more biological information than a global average. It also carries more measurement noise.

For feature tracking, the algorithm must preserve a plausible contour across frames while dealing with changing myocardial appearance. The endocardial border may be clear in one frame and ambiguous in the next. The basal slices are affected by valve-plane motion. The apex is vulnerable to partial-volume effects. The right ventricle is thinner and more difficult to delineate than the left ventricle.

A small contour displacement can produce a large change in local strain when the segment itself is thin or the deformation gradient is steep. The numerical derivative amplifies positional noise. Temporal filtering can suppress that noise, but it can also shift or flatten the apparent peak.

Tagging has its own failure modes. Tag fading reduces late-cycle visibility. Tag spacing limits the spatial scale of measurable deformation. Through-plane motion can move tissue out of the tagged imaging plane. Local defects, artifacts, and inadequate tag contrast compromise the displacement field.

Neither method produces a ground-truth map of the myocardium. Tagging is closer to a direct physical encoding of tissue motion. Feature tracking is more accessible and often more scalable. Their outputs should be compared metric by metric, not method by method in the abstract.

For a global longitudinal strain endpoint, the question may be whether the software preserves subject ranking and group-level differences. For segmental peak circumferential strain, the question is stricter: does the pipeline preserve regional magnitude with sufficient repeatability and bias control?

Those are not equivalent validation requirements.

Directional strain adds another layer of risk

Longitudinal, circumferential, and radial strain are calculated along different anatomical directions. They do not have identical sensitivity to segmentation error or through-plane motion.

GLS is influenced by longitudinal shortening and by the tracking of long-axis boundaries. GCS depends heavily on short-axis geometry and circumferential contour propagation. Radial strain is particularly sensitive to wall-thickness estimation and opposing endocardial and epicardial boundary motion.

The available evidence does not establish universal reference cutoff values for global radial strain across all software vendors. That unknown should remain explicit. A threshold transported from one platform to another without calibration is not a biomarker. It is an unverified software assumption.

In research datasets, sign conventions also require control. Longitudinal and circumferential shortening are commonly represented as negative strain, while radial thickening is commonly positive. A pipeline that changes sign convention, smoothing, or peak-selection logic can create apparent biological differences that are actually computational.

Clinical trials and the inter-vendor problem

The principal weakness of CMR feature tracking in clinical research is not that it fails to produce useful strain. It is that different implementations can produce different strain from the same underlying cine data.

Inter-vendor software variability affects contour initialization, tracking constraints, regularization, temporal interpolation, definition of the region of interest, and exclusion of low-confidence segments. Two platforms may both produce a value labeled GLS while using different internal representations of the myocardium.

This matters immediately in multicenter trials.

A trial endpoint must be sensitive to biological change and stable against technical variation. If baseline scans are processed on one software platform and follow-up scans on another, an apparent change in strain may reflect the processing chain rather than myocardial remodeling. The problem becomes more difficult when sites use different scanners, field strengths, cine resolutions, acceleration factors, or breath-hold protocols.

A robust trial architecture therefore needs a locked analysis environment. The protocol should define:

  • The cine sequence family and acceptable spatial and temporal resolution.
  • The ventricular views and slice coverage required for analysis.
  • Endocardial and epicardial contour rules.
  • Papillary muscle handling.
  • The strain directions and sign conventions.
  • The method for rejecting or flagging poor tracking.
  • The software version and processing configuration.
  • The handling of missing segments and inadequate image quality.
  • Whether the endpoint is global, segmental, temporal, or magnitude-based.
  • The repeatability and cross-vendor validation procedure.

This is not administrative overhead. It is endpoint definition.

For longitudinal studies, the same raw images should be retained whenever possible. Reprocessing may be necessary as algorithms change, but version changes must be recorded. A future software update can alter the measured value without altering the patient or the scan.

The cleanest design is often a central core laboratory with a fixed software version and blinded analysis. That does not eliminate all variability. It constrains it. When multiple platforms must be used, cross-vendor calibration should be performed on representative datasets before the trial begins, not inferred after the primary analysis.

The phrase “vendor-neutral biomarker” should be used cautiously. A biomarker is not vendor-neutral merely because the input is DICOM cine data. Neutrality requires demonstrated analytical equivalence or a validated harmonization model.

Reproducibility is more than a correlation plot

A validation study that reports Pearson correlation alone is incomplete. Correlation answers whether two methods move together. It does not fully characterize bias, repeatability, heteroscedasticity, or clinically relevant limits of agreement.

A serious comparison should examine:

1. Systematic bias. Does feature tracking consistently overestimate or underestimate tagging-derived strain?

2. Range dependence. Does disagreement increase in patients with severe dysfunction or unusually high strain?

3. Repeatability. Does the same analyst, or the same software, reproduce the result?

4. Interobserver variability. How much does manual contour correction affect the output?

5. Segment failure rate. How often is a regional value unavailable or rejected?

6. Temporal stability. Are time-to-peak metrics preserved under different frame rates?

7. Cross-platform behavior. Does the same cine dataset yield comparable values across vendors?

8. Clinical sensitivity. Does the metric detect treatment response or disease progression beyond conventional ejection fraction?

A method can show a high global correlation and still fail a trial’s reproducibility requirement. Conversely, a modest numerical difference may be acceptable if it is stable, predictable, and incorporated into the endpoint model.

The correct question is not whether feature tracking matches tagging in every pixel. It is whether the selected metric is sufficiently valid for the clinical decision or research hypothesis being tested.

Choosing the biomarker instead of choosing the software

The first decision should be the biological question.

If the objective is a scalable global measure of ventricular deformation from routine cine data, CMR feature tracking is a rational choice. It avoids dedicated tagging acquisitions. It supports retrospective analysis. It can extend quantitative cardiac MRI biomarkers into cohorts where tagging was never acquired.

If the objective is precise regional deformation mapping, especially when local amplitude is the primary endpoint, myocardial tagging retains a stronger physical basis. The additional acquisition burden may be justified by the measurement requirement.

The choice changes again for temporal biomarkers. Segmental time to peak strain may show closer agreement between feature tracking and tagging than absolute segmental peak strain magnitude. A study focused on mechanical dyssynchrony or regional timing should validate the temporal endpoint directly rather than assuming amplitude validation transfers automatically.

For clinical trial design, the decision can be organized around the following hierarchy:

  • Global versus regional. Global metrics are generally more tolerant of local tracking error.
  • Magnitude versus timing. Peak amplitude and time to peak do not share the same agreement profile.
  • Prospective versus retrospective. Feature tracking has a major advantage when the cine dataset already exists.
  • Single-vendor versus multicenter. Inter-vendor variability becomes a primary design issue as platform diversity increases.
  • Exploratory versus confirmatory. Exploratory endpoints can tolerate more methodological uncertainty; confirmatory endpoints require locked acquisition and analysis.
  • Detection versus quantification. Detecting a group difference is not the same as producing an interchangeable patient-level measurement.

A practical comparison for research programs

A research program can assign the two methods different roles rather than forcing a binary choice.

Tagging can serve as the reference acquisition in a validation substudy. Feature tracking can then be evaluated on routine cine images across the wider cohort. This design measures the real translation problem: whether the lower-burden method preserves the biomarker signal that matters.

The validation cohort should include the disease spectrum relevant to the trial. Healthy volunteers alone are insufficient. Reduced ejection fraction, regional scar, hypertrophy, dyssynchrony, tachycardia, and suboptimal breath-holds can all alter tracking behavior.

The analysis should preserve the distinction between global and segmental outputs. A platform may be acceptable for GLS while remaining unsuitable for segmental peak strain magnitude. One cannot promote the entire software package based on its strongest metric.

The same logic applies to quantitative cardiac MRI more broadly. A biomarker is not validated because it is numerically sophisticated. It is validated when its acquisition, reconstruction, analysis, repeatability, and biological association are controlled well enough for the intended use.

The endpoint must survive the scanner

Feature tracking has earned a serious place in cardiac MRI because it converts ordinary cine SSFP images into deformation data without requiring a dedicated tagging sequence. The global results are strong enough to support clinical research use, particularly for GLS and GCS. The reported correlations are not trivial: 0.86 to 0.92 for GLS and 0.85 to 0.94 for GCS.

But the method remains an inference pipeline. It depends on image quality and software behavior. Regional peak strain magnitudes show weaker agreement with tagging. Temporal metrics can behave differently from magnitude metrics. Cross-vendor outputs are not automatically interchangeable. Universal radial strain cutoffs remain unresolved.

Myocardial tagging is still the more direct deformation reference. Feature tracking is the more deployable measurement system.

For most translational programs, the productive position is not to declare one method obsolete. It is to define the endpoint narrowly, validate it against the appropriate reference, lock the processing chain, and report the uncertainty honestly.

The scanner produces images. The software produces measurements. The clinical trial depends on the difference.

FAQ

Is cardiac MRI feature tracking as accurate as myocardial tagging?
While feature tracking shows strong correlation with tagging for global strain metrics, it is not interchangeable for all measurements. Tagging remains the reference standard for regional deformation because it provides a more direct physical encoding of tissue motion.
Why is feature tracking preferred for clinical research?
Feature tracking uses routine cine SSFP images, which eliminates the need for dedicated tagging pulse sequences. This reduces scan time, avoids additional breath-hold requirements, and allows for the retrospective analysis of existing clinical datasets.
Does feature tracking work on all cardiac MRI scanners?
Feature tracking is not hardware-independent; it is influenced by scanner field strength, gradient performance, and reconstruction choices. The quality of the cine images provided to the software directly affects the reliability of the strain measurements.
Can I use the same strain values across different software platforms?
No, strain values are not automatically interchangeable between vendors. Different software implementations use unique tracking constraints, segmentation rules, and interpolation schemes, which can produce systematically different results from the same cine data.
Why is segmental strain harder to measure than global strain?
Global strain benefits from spatial averaging, which suppresses local errors. Segmental strain is more vulnerable to issues like poor endocardial definition, through-plane motion, and the amplification of positional noise during numerical derivation.

Also interesting