Radiomics Feature Stability in Longitudinal Clinical Trials: What 8 of 1,596 Features Can—and Cannot—Tell Us
It is also a result that demands careful interpretation.
The finding does not mean that 99.5% of the biological signal was noise. A feature-selection experiment cannot establish that proportion. It tells us something narrower and more useful: under the acquisition conditions, preprocessing decisions, and robustness criteria used in that analysis, only eight features met both prespecified requirements for agreement and dynamic behavior. The other features may have been unstable, biased across sites, insufficiently variable, or simply unsuitable for those particular criteria. Their rejection does not prove that they contained no biological information.
That distinction matters in longitudinal clinical trials. A radiomic signature may appear precise on one workstation and still change when the scanner software is upgraded, the coil is replaced, the reconstruction pipeline is modified, or the patient moves differently at follow-up. The disease can progress while the imaging biomarker records a mixture of biology and acquisition history. The central task is therefore not to maximize the number of extracted features. It is to determine which changes can be defended as biological rather than technical.
The Fragility of Radiomic Signatures in Multi-Center Environments
Radiomics converts an image into a large collection of numerical descriptors: shape measurements, intensity summaries, texture statistics, and features derived from transformed images. The attraction is obvious. Medical images contain information that may not be captured by a single visual impression or a simple volumetric measurement. A radiomic model can examine heterogeneity, spatial organization, intensity distributions, and changes over time.
The difficulty is that these measurements are conditional on how the image was acquired and processed.
In a multicenter MRI trial, nominally similar protocols can still produce different voxel-level data. Echo time, repetition time, flip angle, bandwidth, acceleration, field strength, coil configuration, slice geometry, and sequence implementation all influence the intensity structure presented to the feature-extraction pipeline. Reconstruction software can alter noise texture and spatial sharpness. Vendor updates may change filtering or correction procedures without producing an obvious visual discontinuity in routine clinical review. Patient positioning and motion introduce another layer of variation, particularly in populations that may have difficulty remaining still during a high-resolution acquisition.
These effects are not equally distributed across feature classes. Shape features are calculated primarily from a segmentation mask. If the segmentation protocol is consistent and the anatomical boundary is reasonably stable, measurements such as volume or surface area may be less sensitive to scanner-specific intensity behavior. First-order features can also be comparatively tractable when the intensity scale has been standardized and the region of interest is defined consistently.
Texture features are more exposed to the acquisition chain. Gray-level co-occurrence matrices, gray-level run-length matrices, neighborhood-based descriptors, and related statistics depend on relationships between neighboring voxels. A small change in intensity distribution, spatial resolution, noise pattern, or quantization can alter those relationships. The effect then propagates into measures such as entropy, contrast, homogeneity, correlation, and run-length non-uniformity.
This does not make texture features useless. It changes the burden of proof. A texture feature proposed as a longitudinal biomarker must demonstrate that its behavior is sufficiently stable under the conditions in which the trial will use it. Stability is not an abstract property of the feature alone; it is a property of the feature together with the scanner, sequence, reconstruction, segmentation, preprocessing, and statistical definition.
The meaningful result is not that most extracted features were rejected. It is that the retained features were tested against a defined source of technical variation.
The 8-of-1,596 result should therefore be read as a robustness result within a particular experimental framework, not as a universal survival rate for radiomics. Change the MRI sequence, the patient population, the segmentation method, the gray-level discretization, or the agreement criteria, and the retained set may change. That is expected. Reproducibility is not demonstrated by finding the same number of features in every study; it is demonstrated by showing that the validation design matches the intended clinical use.
Why longitudinal trials expose the problem
Cross-sectional prediction can sometimes tolerate a feature that partly reflects site or protocol, especially if the training and deployment environments are nearly identical. A longitudinal trial has less room for that shortcut. The same patient is compared with themselves across visits, and the interpretation of change depends on the assumption that the measurement process is sufficiently consistent over time.
A scanner replacement can create an apparent shift. So can a software update, a changed motion-correction method, a different segmentation version, or a small alteration in the acquisition prescription. If the study endpoint is based on a delta-radiomic feature, the technical discontinuity may be interpreted as treatment response or disease progression.
This is why longitudinal imaging biomarker validation must begin before the first formal model is trained. The trial needs a record of scanner models, software versions, sequence parameters, preprocessing releases, segmentation rules, and deviations from protocol. Without that provenance, a later feature shift may be impossible to interpret.
Quantifying Feature Robustness: Beyond Simple Statistical Thresholds
No single reliability statistic answers every question in radiomics. Different metrics describe different failure modes, and a robust workflow uses them as complementary evidence rather than interchangeable gates.
ICC: repeatability without complete agreement
The Intraclass Correlation Coefficient (ICC) is commonly used to assess the consistency of repeated measurements. In broad terms, an ICC near 1 indicates that differences between subjects account for much more of the observed variation than differences between repeated measurements of the same subject.
That is useful for test-retest analysis, but ICC is not a complete description of measurement quality. Its interpretation depends on the ICC formulation, the study design, the population variance, and the variance components included in the calculation. A feature can show a high ICC because subjects are widely separated from one another, even if the absolute disagreement between repeated measurements is clinically meaningful.
ICC also does not, by itself, guarantee agreement without bias. Two scanners may rank patients similarly while producing systematically different values. For longitudinal analysis, that offset can matter as much as random variation. A feature that is consistently higher on one platform may appear to change when a participant moves between scanners.
CCC: agreement and bias in the same assessment
The Concordance Correlation Coefficient addresses both correlation and agreement. It is therefore useful when the question is not merely whether two measurements preserve the same ranking, but whether they occupy approximately the same measurement scale.
The multicenter brain MRI analysis used a joint requirement of CCC above 0.9 and Dynamic Range above 0.9. Only eight of the 1,596 evaluated features met both criteria. That result indicates that the combined filter was highly selective in that dataset. It does not establish a field-wide benchmark, nor does it show that the remaining features were biologically meaningless.
Dynamic Range adds a different concern: a feature that is technically stable but nearly constant across subjects may have limited value for discrimination or tracking. A useful biomarker must be stable enough to measure and variable enough to carry information. The exact interpretation of DR depends on how it is defined and normalized in the study, so the statistic should always be reported with its calculation method and reference population.
Coefficient of variation and absolute error
The Coefficient of Variation can help describe relative dispersion under repeated acquisition. It is often easier to interpret alongside absolute differences, confidence intervals, Bland–Altman plots, or limits of agreement than as a standalone pass/fail rule.
A relative error that looks small for a large-valued feature may be substantial for a feature whose clinically relevant changes are also small. Conversely, a high percentage variation in a feature with wide biological separation may not automatically make it unusable. The question is whether the measurement error is acceptable relative to the effect the trial is trying to detect.
Robustness is conditional, not universal
A feature does not have one permanent robustness label. It may be repeatable on the same scanner but not transferable across vendors. It may be stable in a phantom but sensitive to patient motion. It may behave well with one segmentation algorithm and poorly with another. It may be suitable for group-level comparisons but too noisy for individual patient monitoring.
A more informative reporting framework records the conditions under which the feature was assessed:
- same-session or separate-session repeatability;
- same scanner, same vendor, or cross-vendor comparison;
- identical or deliberately varied acquisition parameters;
- manual, semiautomatic, or automated segmentation;
- fixed or adaptive gray-level discretization;
- raw, normalized, or harmonized images;
- absolute agreement or rank-based consistency;
- and the intended clinical endpoint.
That context turns a reliability coefficient into evidence that a reader can evaluate.
Delta-Radiomics and the Reliable Change Index Framework
Longitudinal radiomics, often described as delta-radiomics when it focuses on temporal differences, asks a harder question than ordinary feature reproducibility. It is not enough to know whether a feature can be measured twice. The trial must determine whether the observed change between visits is larger than the change expected from measurement error.
A simple difference score can be written as:
\[
\Delta X = X_{\text{follow-up}} - X_{\text{baseline}}
\]
But the difference alone says little about reliability. A large delta may reflect a true biological event, a scanner transition, motion at one visit, or an unstable segmentation boundary. The interpretation requires an estimate of the feature’s error under relevant repeat-measurement conditions.
The Reliable Change Index (RCI) provides one way to formalize that comparison. Depending on the formulation, the RCI incorporates the standard error of measurement, test-retest variability, and the reliability coefficient of the feature. The resulting score expresses the observed change in units of expected measurement error. Values beyond a prespecified boundary, often ±1.96 in a conventional two-sided framework, may be treated as changes unlikely to be explained by random error alone.
That conclusion remains conditional. An RCI beyond the threshold does not prove that the disease has changed. It indicates that the observed change is large relative to the error model used to calculate the index. If the error model does not represent the trial’s actual scanner mix, motion profile, preprocessing, or segmentation variability, the RCI may give a false sense of certainty.
A hypothetical example
Consider a hypothetical feature measured at baseline and follow-up. Its value changes by 8% between visits. If repeat scans performed under the relevant conditions commonly produce differences of a similar magnitude, the change should not automatically be treated as a biological response. If the change is substantially larger than the estimated measurement error and the acquisition record shows no technical discontinuity, the evidence for a real change becomes stronger.
The example is deliberately generic. The acceptable boundary cannot be imported from one dataset without checking whether the same feature definition, modality, disease context, and acquisition conditions apply. A threshold derived from a controlled repeatability study may not transfer unchanged to a multicenter oncology trial.
RCI is most useful when it is treated as part of a chain of evidence:
1. Establish repeatability under the acquisition conditions that matter.
2. Define the feature and preprocessing pipeline before evaluating longitudinal change.
3. Estimate the measurement error from an appropriate test-retest or calibration dataset.
4. Calculate the RCI using a prespecified formulation.
5. Investigate changes that exceed the threshold alongside clinical, laboratory, and imaging evidence.
6. Document scanner, protocol, segmentation, and software events that could explain the shift.
RCI can show that a change exceeds an estimated measurement-error boundary. It cannot, by itself, identify the biological cause of that change.
This distinction is particularly important in early-phase trials. Small cohorts and short follow-up windows increase the temptation to treat every large delta as meaningful. A reliability-based framework helps prevent that mistake, but it does not replace clinical interpretation or endpoint validation.
Nor should RCI-validated delta-features be presented as demonstrated early-warning systems unless a study has actually shown that capability. A feature that exceeds measurement error may be useful for describing temporal change. That is different from proving that it detects perfusion, cellular, or structural deterioration before a conventional clinical or volumetric endpoint.
Mitigating Scanner-Induced Variance in Longitudinal Datasets
Feature filtering is necessary, but it cannot repair every defect introduced earlier in the imaging pipeline. The most effective strategy is to reduce avoidable variance at acquisition and to preserve enough metadata to identify the variance that remains.
Protocol control and change management
A trial protocol should specify more than the sequence name. It should define the parameters that materially influence the image, acceptable deviation ranges, positioning requirements, motion-management procedures, reconstruction settings, and procedures for repeat acquisition. The protocol should also include a change-control process for scanner software, coils, reconstruction packages, and preprocessing code.
A scanner upgrade is not merely an administrative event. It may be a measurement-system change. Before and after such a change, the trial team should consider whether calibration scans, phantom data, repeat scans, or bridging analyses are needed. If the change cannot be avoided, its timing and affected participants should be recorded for later sensitivity analysis.
Image-level preprocessing
Common preprocessing steps include resampling to a defined voxel geometry, intensity normalization, bias-field correction, and standardized masking. Each can improve comparability, but each can also introduce assumptions.
Resampling changes the spatial relationship among voxels and may affect texture values. Interpolation choices matter, particularly for small lesions or thin anatomical structures. Intensity normalization can reduce between-scan scale differences, yet MRI intensity is not intrinsically standardized in the way many laboratory assays are. Histogram matching, reference-region normalization, and z-score procedures should therefore be evaluated against the intended biological interpretation rather than applied as automatic repairs.
Gray-level discretization deserves explicit attention. Texture features may change when the number of bins, bin width, or discretization range changes. A pipeline that leaves this choice undocumented makes replication difficult and can turn a supposedly stable feature into a moving target.
Feature-level harmonization
Methods such as ComBat and related batch-correction approaches are often used to reduce site-associated differences in feature distributions. They can be valuable when site effects are separable from the biological signal, but that assumption requires scrutiny.
If disease severity is distributed differently across centers, a correction model may mistake biology for batch. If one treatment group is concentrated at one site, harmonization can also interact with treatment assignment. The correction should be trained and applied within a design that respects the trial structure, and the analysis should include sensitivity checks with and without harmonization.
Harmonization is not a substitute for protocol control. It may align distributions while leaving individual-level measurement error, nonlinear scanner effects, or clinically important interactions unresolved. Its role is to reduce a documented source of variation, not to make heterogeneous data interchangeable by declaration.
Phantoms, repeat scans, and traveling subjects
Physical phantoms can provide a stable reference for selected quantitative properties and help identify changes in scanner behavior. They are especially useful for monitoring drift over the life of a multicenter study. Their limitation is equally important: a phantom does not reproduce all the anatomical, physiological, and motion-related conditions of a patient scan.
Repeat scans and traveling subjects add realism. A participant or volunteer scanned at multiple sites can help estimate the combined effect of scanner, protocol, positioning, and physiological variation. These designs require time and coordination, but they provide evidence that a feature behaves consistently beyond a single instrument.
The most defensible approach is layered:
- control the acquisition protocol and document deviations;
- monitor scanner and software changes;
- standardize segmentation and feature extraction;
- evaluate repeatability with appropriate datasets;
- use harmonization only after examining what the site effect represents;
- and test the final feature definitions under conditions resembling deployment.
No layer is sufficient on its own. A high ICC cannot compensate for an undocumented protocol change, and a harmonization algorithm cannot create biological validity where the endpoint has not been defined.
Strategic Feature Selection for Clinical Endpoint Validation
Feature selection for a clinical trial should begin with the endpoint, not with the largest available feature library. The relevant question is not which descriptors can be extracted, but which measurements have a plausible relationship to the endpoint and can remain interpretable under the trial’s operating conditions.
A staged evaluation can be useful, provided it is treated as a study-specific design rather than a universal funnel.
| Evaluation stage | Question it addresses | Evidence to document |
|---|---|---|
| Feature definition | Is the feature precisely specified? | Image type, segmentation, discretization, software version, and calculation method |
| Repeatability assessment | Does the feature remain consistent under repeated measurement? | Test-retest design, variance estimates, ICC or related metrics |
| Agreement assessment | Does it preserve an appropriate scale across measurements or sites? | CCC, bias analysis, limits of agreement, and site comparisons |
| Longitudinal change assessment | Does the observed delta exceed estimated measurement error? | Error model, RCI formulation, and prespecified interpretation |
| Endpoint association | Is the feature related to the clinical or biological endpoint? | Analysis plan, covariate handling, and validation strategy |
| External confirmation | Does the finding persist in an independent setting? | Independent cohort, protocol differences, and performance estimates |
The eight features retained in the multicenter brain MRI analysis are therefore candidates for further study, not automatically validated biomarkers. Their retention means they satisfied the specified CCC and DR requirements in that analysis. It does not establish clinical utility, treatment-response prediction, regulatory acceptability, or generalization to another disease, modality, or scanner population.
The same caution applies to shape and first-order features. Their relative stability can make them sensible starting points, but stability alone is not enough. A feature may be highly reproducible and have no useful relationship to the endpoint. Conversely, a biologically informative feature may require a carefully controlled acquisition or a narrower claim than the original discovery analysis proposed.
Avoiding circular validation
One of the most common weaknesses in radiomics studies is allowing the same data to define the feature, select the feature, tune the model, and evaluate performance. Robustness filtering should be separated from endpoint modeling wherever possible. If the test-retest data used to select features also determine the final predictive performance, the uncertainty around that performance must be acknowledged.
The clinical endpoint should also be defined before exploratory associations are interpreted. A feature can correlate with tumor volume, treatment arm, site, or disease severity for reasons that do not make it a useful biomarker. Prespecified covariates and sensitivity analyses help distinguish an imaging signal from a proxy for trial logistics.
Feature panels versus single features
A small panel can be more useful than a single feature when its members capture complementary aspects of the image. But adding features also increases the opportunity for redundancy, instability, and overfitting. Panel construction should therefore consider correlation structure, missingness, scanner effects, and the number of observations available for model development.
A panel should not be described as robust merely because every member passed the same threshold. The combined model has its own reliability properties. Changes in one feature may be amplified or canceled by the others, especially after normalization or harmonization. The panel needs evaluation as a unit under the same longitudinal conditions in which it will be used.
A More Precise and More Useful Field
The most important lesson from the 8-of-1,596 result is not that radiomics has failed. It is that feature extraction is cheap compared with feature validation. A pipeline can generate thousands of descriptors in minutes, while demonstrating that a small subset is stable across scanners, visits, segmentations, and clinically relevant populations requires deliberate experimental work.
For longitudinal clinical trials, radiomics feature stability is therefore a design property, not a decorative statistic added after model development. The study must define what kind of stability is required: same-session repeatability, cross-scanner agreement, resistance to segmentation variation, sensitivity to genuine change, or some combination of these. The choice of ICC, CCC, Dynamic Range, RCI, and harmonization method should follow that question.
The multicenter finding gives the field a defensible starting point: eight features met both CCC and DR criteria under the reported conditions. It does not tell us that the other 1,588 features were noise, and it does not provide a universal retention rate for future studies. It tells us that strict agreement criteria can substantially narrow a candidate feature set—and that such narrowing may be exactly what a clinical endpoint program needs before it makes stronger claims.
A reliable imaging biomarker is not the feature that produces the most impressive association in a single dataset. It is the feature whose definition, measurement error, longitudinal behavior, and biological interpretation remain clear when the scanner, site, visit, and analysis context are no longer perfectly controlled. That is a smaller promise than radiomics once made. It is also a promise that clinical research can test.
