back to ai
Hugging Face
home/ai/tags/ai-safety

AI Safety

Showing 3 of 19 posts tagged AI Safety

All19AI SafetyAlignmentAutoresearchContinual LearningICMLImage GenerationInferenceInterpretabilityMulti-Agent SystemsMultilingualPre-TrainingRLSFTSparse Autoencoders

Can a Warmth-Trained Model Learn When Not to Agree?

I compared reasons, a matched-length control, and character training after warmth SFT; none increased judged warmth, while M4…

August 20, 2026

When Emergent Misalignment Crosses Languages

English risky-financial fine-tuning transferred coherent emergent misalignment into Hindi, Marathi, and Urdu, while matched…

August 17, 2026

Narrow Fine-Tuning, Broad Misalignment in a 3B Model

I fine-tuned a 3.35B model on narrowly risky financial advice and found a reproducible shift toward coherent misalignment on…

August 10, 2026
August 20, 2026reads

Can a Warmth-Trained Model Learn When Not to Agree?

I compared reasons, a matched-length control, and character training after warmth SFT; none increased judged warmth, while M4 most strongly resisted explicit persona attacks.

August 17, 2026reads

When Emergent Misalignment Crosses Languages

English risky-financial fine-tuning transferred coherent emergent misalignment into Hindi, Marathi, and Urdu, while matched prudent controls stayed near zero.

August 10, 2026reads

Narrow Fine-Tuning, Broad Misalignment in a 3B Model

I fine-tuned a 3.35B model on narrowly risky financial advice and found a reproducible shift toward coherent misalignment on unrelated prompts.