Deferral Policies for Medical Vision-Language Question Answering Under Limited Radiologist Review Capacity in Emergency Radiology Workflows
- Authors
-
-
Samer Khoury
Department of Computer Science, City University, Al Kafaat Intersection, Aley District, Mount Lebanon, LebanonAuthor -
Tariq Hamdan
Department of Information Systems, Arts, Sciences and Technology University in Lebanon, Tal Abbas West, Akkar Governorate, LebanonAuthor
-
- Abstract
-
Emergency radiology services increasingly face workloads in which automated question answering systems may be used to prioritize, summarize, or pre-screen imaging findings. A central difficulty is that model performance is usually reported as if every automated answer must be accepted, although clinical use may instead require selective deferral when the answer appears unreliable. This paper develops and evaluates a risk-calibrated deferral framework for medical vision-language question answering under constrained radiologist review capacity. The study uses a simulation-based emergency chest imaging benchmark with 18,900 question-image cases, four urgency strata, heterogeneous label prevalence, and delayed-review cost. Three automated policies are compared: confidence thresholding, entropy-based abstention, and a proposed harm-weighted conformal deferral method that combines calibrated nonconformity scores with queue-aware review allocation. The primary outcome is expected clinical penalty, defined as a weighted sum of automated error harm, review burden, and delay harm. Across 250 Monte Carlo replications, the proposed policy reduced expected penalty by 18.6% compared with confidence thresholding and by 14.2% compared with entropy-based abstention at a fixed 30\% review budget. It also improved sensitivity for high-urgency positive cases from 86.1% to 92.7% while maintaining radiologist workload constraints. Mixed-effects regression, paired permutation testing, and calibration diagnostics showed that the benefit was largest when model confidence was poorly aligned with case urgency. The results indicate that deployment evaluation should examine selective automation, review capacity, and harm-weighted outcomes rather than treating answer accuracy as a stand-alone property.
- Downloads
- Published
- 2026-04-04
- Section
- Articles
- License
-
Copyright (c) 2026 authors

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
