Abstract
Abstract
Background: Four-grade lumbar foraminal stenosis labels are ordered, and automated four-grade assessment has an established sagittal T1-weighted precedent. Whether ordinal modelling and structured supervision behave consistently in a sagittal T2-specific setting remains unclear. Reliable grading from routine sagittal T2-weighted imaging may also inform future streamlined lumbar MRI workflows.
Methods: We analysed 2,978 expert-defined sagittal T2-weighted foraminal regions of interest (ROIs) from 468 patients using patient-disjoint outer five-fold evaluation with a prespecified validation fold for each outer split. Source-canonical Original CNN, Original DeiT, and fixed 0.40 CNN + 0.60 DeiT fusion were retained for task provenance. CORAL provided structure-matched ordinal modelling; matched moderate-to-severe auxiliary-supervision (MSaux) variants and ConvNeXt-Tiny cross-entropy/CORAL controls tested whether objective-level effects persisted across model contexts. Compatible branches used the same prespecified fixed fusion. Primary metrics were accuracy, balanced accuracy, macro-F1, weighted F1, quadratic weighted kappa (QWK), and mean absolute error (MAE); final fixed-fusion comparisons used 20,000 paired patient-cluster bootstrap replicates.
Results: Ordinal modelling was context dependent. In the source CNN family, CORAL improved all six primary point estimates relative to Original CNN, including macro-F1 from 0.5182 to 0.5689 and QWK from 0.7596 to 0.7861, whereas ConvNeXt-CORAL was weaker than ConvNeXt cross-entropy (macro-F1 0.4325 vs 0.6006; QWK 0.7067 vs 0.7759). MSaux likewise showed no universal benefit: adding it to standalone ConvNeXt-CORAL was adverse on most primary metrics. Fixed DeiT fusion improved both modern ordinal branches to accuracy 0.7482/0.7495 and QWK 0.7939/0.7924. CORAL-MSaux + fixed DeiT achieved macro-F1 0.6155 and QWK 0.7974; versus the two modern fixed-fusion controls, macro-F1 was higher by 0.0170 (95% CI 0.0038 to 0.0303) and 0.0168 (95% CI 0.0028 to 0.0309), while displayed accuracy, QWK, and MAE intervals included zero.
Conclusions: Neither ordinal modelling nor MSaux was universally beneficial across model contexts. Held-fixed DeiT fusion materially changed weak standalone ordinal performance, consistent with heterogeneous system-level complementarity. These results support evaluating label-structure supervision, auxiliary supervision, and fusion as distinct hypotheses rather than assuming a monotonic performance staircase.