Can a Warmth-Trained Model Learn When Not to Agree?
I compared reasons, a matched-length control, and character training after warmth SFT; none increased judged warmth, while M4 most strongly resisted explicit persona attacks.
Showing 3 of 17 posts tagged AI Safety
I compared reasons, a matched-length control, and character training after warmth SFT; none increased judged warmth, while M4 most strongly resisted explicit persona attacks.
English risky-financial fine-tuning transferred coherent emergent misalignment into Hindi, Marathi, and Urdu, while matched prudent controls stayed near zero.
I fine-tuned a 3.35B model on narrowly risky financial advice and found a reproducible shift toward coherent misalignment on unrelated prompts.