back to ai
Hugging Face
home/ai/tags/trustworthy-ai

Trustworthy AI

Showing 2 of 15 posts tagged Trustworthy AI

View as Markdown
All15AlignmentAutoresearchICMLImage GenerationMechanistic InterpretabilityMulti-Agent SystemsPre-TrainingRLSFTSparse AutoencodersTrustworthy AI

Narrow Fine-Tuning, Broad Misalignment in a 3B Model

I fine-tuned a 3.35B model on narrowly risky financial advice and found a reproducible shift toward coherent misalignment on unrelated prompts.

August 10, 2026

Before It Codes: Catching Qwen3.5-4B Planning With J-Lens

I used Anthropic's new Jacobian lens to test whether Qwen3.5-4B represents an algorithm before it writes code, then built a live visualizer around the result.

July 30, 2026
August 10, 2026reads

Narrow Fine-Tuning, Broad Misalignment in a 3B Model

I fine-tuned a 3.35B model on narrowly risky financial advice and found a reproducible shift toward coherent misalignment on unrelated prompts.

July 30, 2026reads

Before It Codes: Catching Qwen3.5-4B Planning With J-Lens

I used Anthropic's new Jacobian lens to test whether Qwen3.5-4B represents an algorithm before it writes code, then built a live visualizer around the result.