back to ai
Hugging Face
home/ai/tags/mechanistic-interpretability

Mechanistic Interpretability

Showing 2 of 15 posts tagged Mechanistic Interpretability

View as Markdown
All15AlignmentAutoresearchICMLImage GenerationMechanistic InterpretabilityMulti-Agent SystemsPre-TrainingRLSFTSparse AutoencodersTrustworthy AI

Before It Codes: Catching Qwen3.5-4B Planning With J-Lens

I used Anthropic's new Jacobian lens to test whether Qwen3.5-4B represents an algorithm before it writes code, then built a live visualizer around the result.

July 30, 2026

Looking Inside Qwen3-4B With Sparse Autoencoders

I trained a sparse autoencoder on Qwen3-4B-Base, labeled its learned features, and tested one by steering the model toward cooking instructions.

June 12, 2026
July 30, 2026reads

Before It Codes: Catching Qwen3.5-4B Planning With J-Lens

I used Anthropic's new Jacobian lens to test whether Qwen3.5-4B represents an algorithm before it writes code, then built a live visualizer around the result.

June 12, 2026reads

Looking Inside Qwen3-4B With Sparse Autoencoders

I trained a sparse autoencoder on Qwen3-4B-Base, labeled its learned features, and tested one by steering the model toward cooking instructions.