Skip to content
AI Safety
Evaluation
Accepted

Unreliability or Disagreement?

What an interpercentile decomposition of multi-turn evaluation results can and cannot show. Accepted as a poster (non-archival) at the NeurIPS 2026 Trust-AI-Eval workshop.

Highlights

  • Single author
  • Accepted as a poster (non-archival), NeurIPS 2026 Trust-AI-Eval (TAE) workshop, to be presented Dec 2026

Overview

"Unreliability or Disagreement? What an Interpercentile Decomposition Can and Cannot Show" is a single-author paper accepted as a poster (non-archival) at the NeurIPS 2026 Trust-AI-Eval (TAE) workshop. It will be presented in Dec 2026.

What it covers

What an interpercentile decomposition of multi-turn evaluation results can and cannot separate, tested on conversations generated across seven models.

Related projects

Like what you see? Let's talk.

I am always open to discussing new projects, collaborations, or opportunities.