In this research for JMIR Formative Research, our team evaluated 11 AutoML frameworks, 6 general-purpose and 5 built specifically for radiomics, across 10 public and private datasets covering CT and MRI, different anatomies, sample sizes, and clinical endpoints. All frameworks were tested under the same standardized cross-validation, and evaluated on predictive performance (AUC), runtime, and qualitative aspects including software status, accessibility, and interpretability.
Most radiomics-specific frameworks were excluded from the final performance comparison: they were obsolete, required extensive programming, or were computationally impractical. General-purpose frameworks proved more accessible to implement. Simplatab, a radiomics-specific tool with a no-code interface, achieved the best overall balance between performance and computational efficiency (mean AUC 78.46%, SD 12.22%, runtime 1.1h), however this result was not statistically superior to the most computationally intensive general-purpose frameworks.
No single framework demonstrated absolute predictive superiority across the datasets tested. The study concludes that AutoML tools for radiomics still require further development to address the field’s specific needs: support for survival analysis, and the integration of upstream steps such as feature extraction, harmonization, and reproducibility control across the full radiomics workflow.