back to ai
Hugging Face
home/ai/tags/interpretability

Interpretability

Showing 2 of 18 posts tagged Interpretability

All18AI SafetyAlignmentAutoresearchContinual LearningICMLImage GenerationInterpretabilityMulti-Agent SystemsMultilingualPre-TrainingRLSFTSmall ModelsSparse Autoencoders

Before It Codes: Catching Qwen3.5-4B Planning With J-Lens

I used Anthropic's new Jacobian lens to test whether Qwen3.5-4B represents an algorithm before it writes code, then built a live…

July 30, 2026

Looking Inside Qwen3-4B With Sparse Autoencoders

I trained a sparse autoencoder on Qwen3-4B-Base, labeled its learned features, and tested one by steering the model toward…

June 12, 2026
July 30, 2026reads

Before It Codes: Catching Qwen3.5-4B Planning With J-Lens

I used Anthropic's new Jacobian lens to test whether Qwen3.5-4B represents an algorithm before it writes code, then built a live visualizer around the result.

June 12, 2026reads

Looking Inside Qwen3-4B With Sparse Autoencoders

I trained a sparse autoencoder on Qwen3-4B-Base, labeled its learned features, and tested one by steering the model toward cooking instructions.