Deferral Policies for Medical Vision-Language Question Answering Under Limited Radiologist Review Capacity in Emergency Radiology Workflows

Authors
  • Samer Khoury

    Department of Computer Science, City University, Al Kafaat Intersection, Aley District, Mount Lebanon, Lebanon
    Author
  • Tariq Hamdan

    Department of Information Systems, Arts, Sciences and Technology University in Lebanon, Tal Abbas West, Akkar Governorate, Lebanon
    Author
Abstract

Emergency radiology services increasingly face workloads in which automated question answering systems may be used to prioritize, summarize, or pre-screen imaging findings. A central difficulty is that model performance is usually reported as if every automated answer must be accepted, although clinical use may instead require selective deferral when the answer appears unreliable. This paper develops and evaluates a risk-calibrated deferral framework for medical vision-language question answering under constrained radiologist review capacity. The study uses a simulation-based emergency chest imaging benchmark with 18,900 question-image cases, four urgency strata, heterogeneous label prevalence, and delayed-review cost. Three automated policies are compared: confidence thresholding, entropy-based abstention, and a proposed harm-weighted conformal deferral method that combines calibrated nonconformity scores with queue-aware review allocation. The primary outcome is expected clinical penalty, defined as a weighted sum of automated error harm, review burden, and delay harm. Across 250 Monte Carlo replications, the proposed policy reduced expected penalty by 18.6% compared with confidence thresholding and by 14.2% compared with entropy-based abstention at a fixed 30\% review budget. It also improved sensitivity for high-urgency positive cases from 86.1% to 92.7% while maintaining radiologist workload constraints. Mixed-effects regression, paired permutation testing, and calibration diagnostics showed that the benefit was largest when model confidence was poorly aligned with case urgency. The results indicate that deployment evaluation should examine selective automation, review capacity, and harm-weighted outcomes rather than treating answer accuracy as a stand-alone property.

Downloads
Published
2026-04-04
Section
Articles
License

Copyright (c) 2026 authors

Creative Commons License

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.

How to Cite

Khoury, Samer, and Tariq Hamdan. 2026. “Deferral Policies for Medical Vision-Language Question Answering Under Limited Radiologist Review Capacity in Emergency Radiology Workflows”. Transactions on General Science, Evidence Synthesis, and Interdisciplinary Methods 16 (4): 1-18. https://grovesocieties.com/index.php/TGSESIM/article/view/Deferral-Policies-for.