Skip to content

Latest commit

 

History

History
3 lines (2 loc) · 1.52 KB

File metadata and controls

3 lines (2 loc) · 1.52 KB

TrustQueryNet Paper Abstract

We present TrustQueryNet, an externally validated dermatoscopic classification study under simulated class-dependent label corruption, budgeted trusted-label repair under simulated oracle supervision, post-hoc calibration, and selective prediction. The pipeline combines lesion-level HAM10000 splits, persistent noise manifests, explicit best-checkpoint evaluation, multi-seed aggregation, and external testing on the official ISIC 2019 test set. With the corrected ConvNeXt-Tiny recipe, repair reached 0.8350 ± 0.0059 internal calibrated accuracy and 0.7152 ± 0.0216 internal macro-F1 across five seeds. However, strong baselines remained close: no repair reached 0.7145 ± 0.0130 internal macro-F1, random repair reached 0.7105 ± 0.0239, and a clean-label upper bound reached 0.7330 ± 0.0288. On external ISIC 2019, all methods degraded sharply. Repair reached 0.5692 ± 0.0145 calibrated accuracy and 0.4427 ± 0.0117 macro-F1, versus 0.5630 ± 0.0078 and 0.4288 ± 0.0123 for no repair, and 0.5591 ± 0.0203 and 0.4311 ± 0.0328 for random repair. Internal temperature scaling did not remove external calibration fragility, and generalized cross-entropy collapsed under the chosen corruption regime. An overlap audit found zero exact duplicate images between HAM10000 and the mapped ISIC 2019 external slice. Overall, the study shows that modest internal gains from trusted-label repair do not guarantee external trustworthiness and must be judged against strong baselines, including random repair.