SafetyCage is a Python package for detecting misclassified predictions from machine learning models in classification tasks. It provides a unified interface for multiple statistical detection methods, enabling users to quantify prediction reliability and flag potentially incorrect outputs across different models and datasets easily.
Available on PyPI: https://pypi.org/project/safetycage/.
Full documentation is available at https://safetycage.readthedocs.io/.
The idea behind safetycage is that we can find statistics on each predicted sample and compare that statistic to some statistic threshold to predict whether the sample prediction was incorrectly classified.
Machine learning models can produce incorrect predictions with high confidence. SafetyCage addresses this by providing post-hoc misclassification detection methods that operate on model outputs or internal representations.
The package includes several methods:
- MSP (Maximum Softmax Probability)
- DOCTOR (Error probability estimation)
- Mahalanobis (Distance-based statistical testing)
- SPARDACUS (Projection + density estimation approach)
Each method outputs a statistic or p-value that reflects how likely a prediction is to be incorrect.
Alternatively, you can implement your own method by initializing a base class from the safetycage abstract base class, which defines how methods should be implemented.
safetycage requires Python 3.13 or later.
Core dependencies (installed automatically): joblib, matplotlib, numpy.
pip install safetycage
# or: uv add safetycageSome methods need extra dependencies, installed via pip install safetycage[extra] (or uv add "safetycage[extra]"):
| Extra | Uses | Adds |
|---|---|---|
red |
RED |
torch, gpytorch |
spardacus |
SPARDACUS |
statsmodels, scipy, scikit-learn, tqdm |
mahalanobis |
Mahalanobis |
statsmodels, scipy |
torch |
TorchModelModule |
torch |
To learn how to use safetycage, check out the examples/ directory in this repository. It contains complete integrations with runnable notebooks, including scripts to train models to test the safetycage methods on.
See the CHANGELOG.MD for details on versioning.
If you encounter issues or have questions:
- Open an issue on the repository: https://github.com/SINTEF/safetycage/issues.
If you would like to contribute, please reach out to our safetycage team, listed below!
- Pål Vegard Bun Johnsen (palVJ)
- Joel Bjervig (joelbjervig)
- Julia Qiu (jq11)
The MSP method was introduced by Hendrycks and Gimpel in A Baseline for Detecting Misclassified and Out-of-Distribution Examples in Neural Networks.
The DOCTOR method was introduced by Granese et al. in DOCTOR: A Simple Method for Detecting Misclassification Errors.
The Mahalanobis method is described in Johnsen et al..
The SPARDACUS method is described in Johnsen et al..
A proper citation for these methods is provided in the docstring of the code using these methods.
A special thank you goes to previous co-authors of the methods we have built, Filippo Remonato, Shawn Benedict, and Albert Ndur-Osei.
This project is licensed under the MIT License - see the LICENSE file for details.
Active and under development!
If you use safetycage, please cite us!
