Can a Warmth-Trained Model Learn When Not to Agree?
I compared reasons, a matched-length control, and character training after warmth SFT; none increased judged warmth, while M4…
Showing 3 of 18 posts tagged AI Safety
I compared reasons, a matched-length control, and character training after warmth SFT; none increased judged warmth, while M4…
English risky-financial fine-tuning transferred coherent emergent misalignment into Hindi, Marathi, and Urdu, while matched…
I fine-tuned a 3.35B model on narrowly risky financial advice and found a reproducible shift toward coherent misalignment on…