AI Safety
Evaluation
Accepted
Unreliability or Disagreement?
What an interpercentile decomposition of multi-turn evaluation results can and cannot show. Accepted as a poster (non-archival) at the NeurIPS 2026 Trust-AI-Eval workshop.
Highlights
- Single author
- Accepted as a poster (non-archival), NeurIPS 2026 Trust-AI-Eval (TAE) workshop, to be presented Dec 2026
Overview
"Unreliability or Disagreement? What an Interpercentile Decomposition Can and Cannot Show" is a single-author paper accepted as a poster (non-archival) at the NeurIPS 2026 Trust-AI-Eval (TAE) workshop. It will be presented in Dec 2026.
What it covers
What an interpercentile decomposition of multi-turn evaluation results can and cannot separate, tested on conversations generated across seven models.


