Alignment
Does redemption-narrative steering move emergently misaligned LLMs back toward alignment? Tested behaviorally, on self-rated harmfulness, and geometrically against a derived misalignment direction.
Theoretical framing: emergent misalignment as moral injury. The repository has the full experimental design, cross-scale results (Llama / Qwen, 0.5B–8B), and the derived canonical misalignment direction.