Narrow Fine-Tuning, Broad Misalignment in a 3B Model
I fine-tuned a 3.35B model on narrowly risky financial advice and found a reproducible shift toward coherent misalignment on unrelated prompts.
August 10, 2026
Showing 2 of 15 posts tagged Trustworthy AI
I fine-tuned a 3.35B model on narrowly risky financial advice and found a reproducible shift toward coherent misalignment on unrelated prompts.
I used Anthropic's new Jacobian lens to test whether Qwen3.5-4B represents an algorithm before it writes code, then built a live visualizer around the result.