Auditability Gaps in ICU Deterioration Models Through Slice-Stable Evaluation Attribution Reliability and Deployment-Facing Error Discovery
- Authors
-
-
Santiago Andrés Rojas
Departamento de Ingeniería de Sistemas, Universidad de la Amazonia, Calle 17 Diagonal 17 con Carrera 3F, Barrio Porvenir, Florencia, Caquetá, ColombiaAuthor -
Daniel Esteban Cárdenas
Programa de Ingeniería Informática, Universidad de Córdoba, Carrera 6 No. 76–103, Barrio Mocarí, Montería, Córdoba, ColombiaAuthor -
Miguel Ángel Herrera
Facultad de Ingeniería y Tecnología, Universidad de los Llanos, Kilómetro 12 Vía Puerto López, Vereda Barcelona, Villavicencio, Meta, ColombiaAuthor
-
- Abstract
-
Evaluation in intensive care forecasting is usually organized around a small set of scalar metrics such as AUROC, AUPRC, calibration error, and thresholded sensitivity. These summaries are useful, but they often fail to reveal where a warning system is brittle, which patient-time regimes dominate its apparent success, and how model behavior changes when evidence becomes sparse, delayed, or semantically inconsistent. As a result, systems that appear strong in retrospective benchmarking can remain poorly understood in the exact regimes where clinicians would need reliability most. This paper develops an audit-centered framework for ICU deterioration modeling in which evaluation is treated not as a final scoring step, but as a structured inference problem over hidden failure regions. The central idea is that error should be decomposed over temporally local, semantically coherent, and operationally meaningful slices rather than averaged across the full population of prediction windows. To support this goal, the paper introduces mathematical tools for slice discovery, temporal attribution stability, confidence-conditioned review utility, and deployment-facing risk decomposition under alert budgets. The framework also examines how repeated-window sampling, severe class imbalance, and institutional workflow patterns can create misleading impressions of generalization when only aggregate metrics are reported. By linking evaluation to hidden-state structure, support quality, and review constraints, the paper argues for a shift from score-centric benchmarking toward auditability as a first-class property of warning systems. Under this view, a useful ICU predictor is not simply one that ranks positives above negatives on average, but one whose failures are legible, whose explanations are stable, and whose deployment thresholds remain defensible across clinically distinct regions of the data.
- Downloads
- Published
- 2025-12-04
- Section
- Articles
- License
-
Copyright (c) 2025 authors

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
