Uncertainty based selection of compatible inputs
-
Updated
Aug 3, 2026 - Python
Uncertainty based selection of compatible inputs
An editable, auditable 807K-param byte-level LLM: CRUD single facts with provable per-edit locality, and abstain when unsure instead of guessing. CPU, offline.
Behavioral Trust Clustering a thermodynamic governance layer that reduces LLM hallucination by 52% on HumanEval. Drop-in wrapper for any decoder. MIT.
We show that a model owner can artificially introduce uncertainty into their model and provide a corresponding detection mechanism.
Confidence-aware support intent classification with reproducible evaluation, selective prediction and a typed FastAPI inference boundary.
Safer continual learning for PyTorch models — research alpha
Uncertainty-aware reliability monitor for safety-critical ML models — calibration (ECE), risk-coverage (AURC), selective prediction with human review, and drift detection. React, FastAPI, PostgreSQL, Kafka.
A comprehensive library for uncertainty quantification in machine learning.
Trust layer for document→JSON extraction & AI agents: calibrated per-field confidence + source grounding + accept/review abstention on any OCR/VLM. Ships VerifyDocBench, a novel grounding-conditioned conformal method, and an MCP server.
Reliable medical QA with Mistral-7B, QLoRA, selective prediction, and learned abstention via warm-start SFT + DPO.
Calibrated trust and abstention for MS/MS molecular annotations
DegradeRisk-Seg: risk-controlled semantic segmentation under degraded multi-modal remote-sensing observations
Prompt-only boundary prediction for IFEval-style instruction-checker pass/fail behavior.
Investigation of how sampling strategies affect Selective Prediction performance in Multi Task Learning
Reproducible MEDAI deferral simulation (AIRI 2026). Synthetic research code.
Tsetlin Machines with a certificate on every answer: the exact number of feature flips a prediction survives, computed per sample, with predict-or-abstain when the radius is too small.
Code Repository for SCoRE paper
Proper scoring rules, reduces LLM overconfidence in multiple-choice QA.
Neuro-symbolic selective verification for LLM mathematical reasoning using SMT solvers and conformal prediction.
Selective prediction with calibrated uncertainty for contactless arrhythmia detection from facial video (rPPG). Benchmarks five UQ methods on AF / NSR / Other classification.
Add a description, image, and links to the selective-prediction topic page so that developers can more easily learn about it.
To associate your repository with the selective-prediction topic, visit your repo's landing page and select "manage topics."