Neuroimaging & Brain Mapping

Cortical thickness pipeline errors and how to fix them

A cortical thickness pipeline can produce a complete, visually convincing dataset while still misrepresenting the cortex in a clinically meaningful way.

Cortical thickness pipeline errors and how to fix them

The problem often becomes visible only after the measurements have been converted into regional statistics: an unexpectedly thin occipital cortex, an implausible thickening near the frontal pole, or a group of apparently extreme subjects whose values seem to confirm the hypothesis too neatly.

This is the uncomfortable point at which cortical thickness pipeline segmentation errors become more than a technical nuisance. A surface that has been drawn around the wrong tissue can still generate valid-looking numbers, and those numbers can travel quietly through quality control, statistical modelling, and longitudinal interpretation before anyone returns to the original MRI volume.

The remedy is not to treat every outlier as an unusable scan, nor to assume that a popular software suite has made visual inspection unnecessary. It is to follow the error backward—from the regional estimate, through the reconstructed surface or voxel-wise thickness map, and finally to the anatomy and image contrast that allowed the mistake to occur.

Where cortical thickness errors begin: the boundary is not always a boundary

Most cortical thickness pipelines depend on a chain of anatomical decisions. The software must separate brain from non-brain tissue, identify the white matter surface, locate the pial surface, establish anatomical topology, and then calculate the distance between the relevant boundaries. Each step is influenced by the same basic material: the intensity pattern and spatial resolution of the MRI acquisition.

That means an error in the first stage can become a plausible-looking measurement several stages later.

Skull stripping is one of the most consequential points of failure. If dura, blood vessels, or other non-brain structures remain inside the brain mask, the pial surface may be placed too far outward. The resulting cortex appears thicker because the distance between the white matter boundary and the incorrectly expanded pial boundary has increased. The reverse problem is equally serious: an aggressive brain extraction can remove genuine cortical tissue, particularly near the posterior cortex or cerebellum, and create an artificially thin or truncated anatomical surface.

These errors do not always announce themselves as dramatic failures. A scan may show a broadly reasonable brain mask while retaining a narrow rim of dura along the convexity. That rim can be enough to distort local thickness estimates, especially where the cortex is already folded and the gray–white or gray–CSF contrast is less distinct.

Intensity normalization introduces a related vulnerability. Cortical reconstruction assumes that the relationship between tissue classes is sufficiently stable for the algorithm to identify boundaries across the entire brain. Bias fields, motion-related signal variation, coil sensitivity effects, and sequence-specific contrast can weaken that assumption. The software may then interpret a region of altered intensity as a tissue boundary, or fail to recognize the boundary where the anatomy is intact but the signal is not.

A cortical thickness value is never only a property of the cortex; it is also a record of how confidently the pipeline identified the two surfaces that define it.

This is why cortical thickness measurement accuracy cannot be separated from the acquisition and preprocessing history. A segmentation algorithm does not see anatomy in the abstract. It sees a three-dimensional field of intensities, and it makes anatomical inferences from that field.

FreeSurfer segmentation errors: what visual inspection should look for

FreeSurfer remains widely used because it provides a detailed surface-based workflow, extensive anatomical parcellation, and a mature ecosystem for cortical morphometry. It also makes the logic of quality control relatively concrete: inspect the brain mask, white matter surface, and pial surface in multiple planes, and determine whether each follows the anatomy rather than merely producing a continuous contour.

The commonly cited estimate for manual visual inspection and editing is around 30 minutes per subject. In practice, comprehensive review across hundreds of slices and multiple viewing planes can take an hour or more per scan, particularly when the subject has motion degradation, atypical anatomy, vascular structures near the cortex, or several regions requiring intervention. This time requirement is not evidence that the pipeline is poorly designed. It reflects the anatomical complexity of the problem and the fact that a reviewer is checking a three-dimensional reconstruction rather than a single image.

A useful review follows the surfaces rather than jumping from one suspicious region to another. Start with the brain mask, because a downstream pial error may simply be the visible consequence of an earlier extraction failure. Then examine the white matter surface, where inclusion of gray matter or exclusion of white matter can alter the inner boundary. Finally, inspect the pial surface along the convexities, sulcal depths, medial surfaces, and regions adjacent to ventricles and deep nuclei.

The following patterns deserve particular attention:

  • Dura included in the brain mask. The pial surface may extend beyond the cortical ribbon, producing local overestimation of thickness.
  • Posterior or inferior cortex removed. Over-stripping can truncate the surface and create regional underestimation or missing parcels.
  • Blood vessels interpreted as cortical tissue. A vessel crossing the cortical surface may pull the pial boundary outward or interrupt it.
  • Motion-induced intensity irregularity. The algorithm may create a patchwork of local boundary errors rather than one obvious global failure.
  • Partial-volume effects. Thin cortex near CSF or closely apposed sulcal walls can be represented by mixed voxels, making the surface placement uncertain.
  • Implausible continuity. A surface that jumps across a sulcus, bridges a narrow gap, or cuts through a clearly visible anatomical structure is not rescued by the fact that the final mesh is closed.

In a large dataset, it is tempting to review only subjects with extreme thickness values. That is a sensible way to prioritize effort, but it is not a complete quality-control strategy. A study of the NKI-Rockland Sample using FreeSurfer 5.3 found that groups defined as statistical outliers at z-scores of ±3 standard deviations had 21% more visually identified segmentation errors than non-outlier groups. The finding is useful because it confirms that statistical screening can direct attention. It is not proof that non-outlier scans are reliable: segmentation errors also occurred among subjects whose measurements fell within the expected distribution.

This distinction matters in longitudinal studies. A subtle error that is consistent across time points may not appear as a statistical outlier, yet a change in acquisition, head positioning, motion, or preprocessing can create an artificial trajectory. Apparent cortical thinning may then reflect a shift in segmentation behaviour rather than biological change.

Why topological defects corrupt an otherwise plausible surface

Surface-based pipelines do more than trace tissue boundaries. They must also construct a mesh with the expected topology, so that the reconstructed cortical surface can be represented, inflated, registered, and mapped across subjects. When non-cortical structures are misclassified, the mesh can develop artificial handles, holes, or other defects.

The fornix, basal ganglia, and lateral ventricles are common anatomical sources of confusion because their intensity and geometry can resemble neighbouring tissue in ways that are difficult for a generic segmentation rule to resolve. A misclassification near the ventricles may cause the white matter surface to intrude into a non-cortical structure. An error involving the basal ganglia can create a connection that should not exist. In the fornix, a thin elongated structure can be absorbed into the wrong tissue class and alter the topology of the reconstructed hemisphere.

The impact is not limited to the small region where the defect begins. Surface inflation and spherical registration depend on a coherent mesh. A local topological problem can therefore influence the representation of nearby folds, affect anatomical correspondence, and complicate the interpretation of regional measurements. In other words, the mesh is not just a display layer placed over the data after analysis; it is part of the measurement framework.

A practical troubleshooting sequence is more informative than asking whether a subject has a topological defect in isolation:

1. Locate the defect in native anatomical space. Determine which structure has been misclassified and whether the problem is confined to one slice or persists across adjacent sections.

2. Inspect the corresponding brain mask. If non-brain tissue is present or true brain tissue has been removed, surface correction alone may not address the origin of the error.

3. Compare white and pial surfaces separately. The two surfaces can fail for different reasons, and correcting one does not guarantee that the other is anatomically plausible.

4. Review the neighbouring sulci and gyri. A local repair should not create a new discontinuity or push the surface across a sulcal boundary.

5. Re-run the appropriate reconstruction stage. A corrected input must propagate through the downstream step that generated the faulty surface; otherwise, the visible edit may not be reflected in the final measurements.

This is also where the distinction between a segmentation error and a topological error becomes clinically relevant. A pial surface extending into the dura is a boundary-placement problem. A mesh that contains an artificial handle is a structural representation problem. They may coexist, but they are not interchangeable, and they should not be repaired with the same assumption.

Manual editing is slower, but it preserves anatomical accountability

Manual correction is sometimes discussed as though it were an embarrassing remnant of pre-automation imaging. That framing is unhelpful. In cortical morphometry, manual editing is a controlled way to restore anatomical information when the automated model has made a specific, inspectable error.

For FreeSurfer pial surface errors caused by skull voxels being misclassified as cortex, the established correction involves editing the brainmask volume with the Recon Edit tool and then rerunning the pial reconstruction stage with recon-all -autorecon-pial. The important principle is that the correction is made at the level of the volume that influenced the surface, rather than by attempting to paint over the final measurement.

The same principle applies beyond this particular workflow: edit the earliest representation that is demonstrably wrong, then regenerate the dependent output. Directly manipulating a derived surface without understanding its input can make the result look cleaner while leaving the underlying segmentation logic unresolved.

Manual editing does, however, create its own methodological obligations. The operator needs a consistent protocol, a record of which regions were edited, and a clear distinction between correction and subjective anatomical preference. If two reviewers apply different thresholds for what counts as an acceptable pial deviation, the dataset may acquire an invisible operator effect. This can be especially consequential in developmental, ageing, or neurodegenerative cohorts, where small changes in cortical thickness are interpreted as part of a longitudinal trajectory.

A useful editing record includes:

  • the subject and processing version;
  • the anatomical region affected;
  • the original failure mode, such as dura inclusion or tissue truncation;
  • the volume or surface stage edited;
  • the reconstruction stage rerun;
  • the reviewer and date;
  • whether the final output passed a second visual inspection.

That level of documentation may feel excessive for a small exploratory analysis, but it becomes essential when data are pooled across scanners, sites, or study waves. Consider the implications for reproducibility: an edited surface without an edit history is difficult to distinguish from an automated result, and an automated result without visual confirmation is difficult to interpret when the measurement later becomes clinically consequential.

Automated error maps and deep-learning pipelines

The time required for manual review has encouraged several strategies for prioritizing or reducing quality-control burden. One approach is to generate automated error maps that identify regions where a segmentation is unusual relative to a reference distribution. Such tools can help reviewers direct their attention, and Gaussian smoothing parameters—for example, a sigma of 1 mm in automated error-map frameworks—may be used to create spatially coherent representations of local abnormalities.

These methods should be understood as triage rather than adjudication. An error map can identify an unusual pattern, but unusual does not necessarily mean incorrect. Conversely, a segmentation can be wrong in a way that is not sufficiently rare to trigger an outlier flag. The 21% increase in visually identified errors among ±3 SD outlier groups illustrates the value of statistical prioritization, while also showing why outlier filtering cannot replace anatomical inspection.

Deep-learning pipelines such as FastSurfer offer another route to reducing processing time and, in some settings, improving scalability. Their value is particularly evident in large cohorts where a conventional processing schedule would make full manual inspection difficult to sustain. Faster inference can allow teams to process more scans, repeat analyses during development, and allocate human review to cases with the greatest uncertainty.

Yet a faster segmentation is not automatically a more trustworthy segmentation. A deep-learning model may fail differently from a conventional pipeline, especially when the training distribution does not represent the scanner, sequence, age range, pathology, or anatomical variation in the target cohort. The question is therefore not whether a deep-learning model eliminates quality control, but whether it changes the pattern of errors and provides sufficiently transparent signals for detecting them.

A robust large-scale workflow can combine several layers:

Quality-control layerWhat it can detectWhat it cannot establish alone
Global image and preprocessing checksMissing data, severe motion, orientation or intensity problemsAnatomically correct cortical boundaries
Automated outlier screeningUnusual regional thickness or surface measuresThat every non-outlier scan is error-free
Error maps or uncertainty indicatorsSpatially concentrated deviations that merit reviewWhether an unusual region reflects pathology or segmentation failure
Visual inspection in multiple planesDura inclusion, truncation, surface crossings, topological defectsPerfect biological validity of the measurement
Manual correction and rerunSpecific, documented boundary or mask errorsThat the acquisition itself was adequate
Test–retest or longitudinal consistencyInstability across repeated measurements or time pointsWhether a stable bias is biologically correct

This shift allows quality control to become a managed scientific process rather than an all-or-nothing decision made at the end of a pipeline. The goal is not to force every image into a clean statistical distribution. It is to identify which measurements retain an anatomically defensible relationship to the underlying brain.

ANTs and DiReCT: avoiding mesh defects without avoiding segmentation questions

ANTs takes a different methodological route from surface-based pipelines such as FreeSurfer and CIVET. Its DiReCT method estimates cortical thickness directly in voxel space using diffeomorphic registration-based calculations. Because it does not depend on constructing the same type of cortical surface mesh, it avoids the mesh-topology defects that can arise in surface-based reconstruction.

That distinction is important, but it should not be mistaken for immunity from error. A voxel-based method still depends on image contrast, tissue classification, registration quality, and the biological plausibility of the resulting thickness field. If the cortical ribbon is poorly represented in the input image, changing the mathematical representation of thickness does not restore information that was never captured reliably.

The choice between approaches should therefore follow the study question and the failure mode. Surface-based methods are powerful when the analysis requires explicit cortical geometry, folding patterns, surface registration, or parcel-wise interpretation aligned to the cortical sheet. A voxel-based method can be attractive when mesh topology is a recurrent source of instability or when the analysis benefits from direct volumetric estimation.

For comparative studies, it is unwise to treat measurements from different pipelines as though they were interchangeable simply because both are labelled cortical thickness. Differences in tissue definitions, registration, smoothing, surface placement, and regional correspondence can create systematic offsets. If a study compares methods, the methods themselves become part of the analysis and should be evaluated with the same attention given to the clinical cohort.

Changing the pipeline can remove one class of failure while exposing another; the correct question is not which software is flawless, but which error model the study can see and control.

Building a troubleshooting workflow that respects the biological trajectory

Cortical thickness is often used to describe subtle degradation, developmental change, cognitive reserve, or disease-related trajectories. Those interpretations require more than a single plausible image. They require confidence that the apparent direction and magnitude of change are not produced by unstable preprocessing.

A practical brain morphometry pipeline troubleshooting process should therefore connect image-level inspection to the final scientific claim. If the study examines neurodegenerative disease, review the regions that drive the hypothesis, but do not restrict quality control to those regions. If the study is longitudinal, inspect the same subject across time points and look for changes in surface behaviour that coincide with acquisition or processing changes. If the analysis involves multiple sites, examine whether one scanner or sequence produces a different distribution of failures.

The following sequence is deliberately conservative:

1. Preserve the original outputs. Never overwrite the first automated result. The initial mask and surfaces are evidence of how the pipeline behaved and may be needed to understand a later correction.

2. Use statistics to prioritize, not to absolve. Begin with extreme values, unusual regional combinations, and abrupt longitudinal changes, while retaining a plan to sample non-outlier scans.

3. Inspect native-space anatomy. Overlay the brain mask, white matter surface, and pial surface on the original structural MRI rather than relying only on rendered surfaces.

4. Classify the failure. Distinguish skull stripping, intensity normalization, tissue misclassification, topology, motion, and registration problems.

5. Correct the earliest defensible source. If the mask is wrong, edit the mask; if the pial boundary is wrong because of that mask, rerun the dependent reconstruction.

6. Record and verify the edit. A correction is incomplete until the regenerated output has been inspected again.

7. Assess group-level consequences. Compare results before and after quality control, and determine whether the principal finding depends on a small number of corrected or excluded subjects.

8. Report the workflow transparently. State how visual inspection was performed, which stages were edited, how outliers were handled, and whether the process differed across sites or time points.

This approach also protects against a common interpretive error: assuming that a biologically credible finding must have come from biologically credible measurements. A result can align with the literature and still be influenced by a systematic segmentation tendency. Conversely, an unusual cortical pattern may be real and should not be discarded merely because it is statistically inconvenient.

The most reliable studies hold both possibilities in view. They ask whether the signal follows anatomy, whether it persists under reasonable quality-control decisions, and whether its trajectory remains coherent when the pipeline is inspected at the level of individual subjects.

The clinical meaning of a repaired surface

A corrected pial surface is not the same as a corrected diagnosis. It is a repaired measurement boundary, one step that makes the relationship between image and anatomy more defensible. That distinction matters particularly in clinical research, where cortical thickness may contribute to a broader assessment of disease burden, cognitive change, or treatment response but rarely provides an independent explanation for the patient’s condition.

The value of quality control is therefore cumulative. Better skull stripping reduces one source of bias. More careful inspection of topology reduces another. Automated outlier maps make large cohorts more manageable. Deep-learning methods may shorten processing time. Voxel-based alternatives can avoid particular mesh failures. None of these measures replaces the others, because each addresses a different point at which the image can be translated incorrectly into a biological quantity.

Cortical thickness pipeline segmentation errors are best understood as failures of translation: from signal intensity to tissue class, from tissue class to boundary, from boundary to surface or thickness field, and from that field to a claim about human biology. At every transition, a small technical error can become a large interpretive problem if it is allowed to pass without inspection.

The final standard is not visual perfection, nor is it the removal of every statistical outlier. It is a documented and anatomically reasoned account of why the measurement deserves to be believed. When the pipeline preserves that accountability, cortical thickness becomes more than a derived number. It becomes a traceable observation of the brain across time—one that can be connected, with appropriate humility, to cognition, disease, and the changing reserve of human biology.

FAQ

Why does a cortical thickness pipeline produce errors even if the final output looks reasonable?
The pipeline may misclassify non-brain tissue or fail to identify boundaries correctly, resulting in a surface that appears visually plausible but misrepresents the underlying anatomy.
How can I identify if my cortical thickness measurements are affected by segmentation errors?
You should inspect the brain mask, white matter surface, and pial surface in multiple planes against the original MRI volume to ensure they follow the actual anatomy.
Is it sufficient to only review subjects identified as statistical outliers?
No, while outliers are more likely to contain errors, segmentation failures also occur among subjects whose measurements fall within the expected distribution.
What is the correct way to fix a segmentation error in a pipeline like FreeSurfer?
You should edit the earliest representation that is demonstrably wrong, such as the brain mask, and then rerun the dependent reconstruction stages to propagate the correction.
Do voxel-based methods like ANTs DiReCT avoid all cortical thickness pipeline errors?
While voxel-based methods avoid mesh-topology defects, they remain dependent on image contrast, registration quality, and the biological plausibility of the input data.

Also interesting