Skip to content
https://github.com/ascender1729/emergent-misalignment-political
AI Safety
Research
Active

Emergent Misalignment Research

Sole-author BlueDot Impact sprint study of which fine-tuning domain drives behavioural drift at 7B scale, scored by LLM judges. Manuscript.

LLM judges

4

Highlights

  • 4-judge LLM panel across 3 model families
  • Cross-architecture replication on Qwen 2.5 7B and Llama 3.1 8B
  • BlueDot Rapid Grant recipient ($250)

Overview

This research investigates how the domain of fine-tuning data influences emergent misalignment in large language models. Building on Betley et al.'s finding that fine-tuning on insecure code can cause misaligned outputs on unrelated tasks, it extends the investigation to political content.

What it studies

  • Which fine-tuning domain drives behavioural drift, comparing political content against insecure code at 7B scale
  • Replication across Qwen 2.5 7B and Llama 3.1 8B
  • A 2026 replication adds a valence-matched control

Research context

Sole-authored manuscript from the BlueDot Impact Technical AI Safety Project Sprint, Group 07, 2026, funded by a BlueDot rapid grant. All behavioural scoring is done by a 4-judge LLM panel, not by human raters.

Related projects

Like what you see? Let's talk.

I am always open to discussing new projects, collaborations, or opportunities.