
AI Safety
Research
Active
Emergent Misalignment Research
Sole-author BlueDot Impact sprint study of which fine-tuning domain drives behavioural drift at 7B scale, scored by LLM judges. Manuscript.
LLM judges
4
Highlights
- 4-judge LLM panel across 3 model families
- Cross-architecture replication on Qwen 2.5 7B and Llama 3.1 8B
- BlueDot Rapid Grant recipient ($250)
Overview
This research investigates how the domain of fine-tuning data influences emergent misalignment in large language models. Building on Betley et al.'s finding that fine-tuning on insecure code can cause misaligned outputs on unrelated tasks, it extends the investigation to political content.
What it studies
- Which fine-tuning domain drives behavioural drift, comparing political content against insecure code at 7B scale
- Replication across Qwen 2.5 7B and Llama 3.1 8B
- A 2026 replication adds a valence-matched control
Research context
Sole-authored manuscript from the BlueDot Impact Technical AI Safety Project Sprint, Group 07, 2026, funded by a BlueDot rapid grant. All behavioural scoring is done by a 4-judge LLM panel, not by human raters.


