Investigating Data Usage for Inductive Conformal Predictors
Publication status: arXiv preprint, first submitted June 18, 2024. Read the source and version history.
Research question
How should limited development data be allocated between training and calibration, and what happens when those sets overlap?
Method
The experiments use an inductive conformal predictor around a neural-network classifier on Covtype. They vary training/calibration allocation, development-set size, and overlap, using repeated randomized splits to examine coverage and prediction-set size.
Results and scope
Small calibration sets can increase variability. Overlapping training and calibration data can produce smaller prediction sets at the cost of undercoverage. These empirical findings concern one classification dataset; they do not establish a universally optimal split or validate overlapping calibration data in general.