---
title: Interpretability
description: AI notes by Mrinaal Arora tagged "Interpretability"
updated: 2026-07-30
canonical: https://aroramrinaal.com/ai/tags/interpretability
source: https://aroramrinaal.com/ai/tags/interpretability.md
alternate: https://aroramrinaal.com/ai/tags/interpretability/markdown
---

# Interpretability

AI notes, small builds, and writeups by Mrinaal Arora tagged "Interpretability".

- Page: https://aroramrinaal.com/ai/tags/interpretability
- Markdown: https://aroramrinaal.com/ai/tags/interpretability/markdown
- Markdown (.md): https://aroramrinaal.com/ai/tags/interpretability.md
- All tags: https://aroramrinaal.com/ai/tags/markdown
- Full AI index: https://aroramrinaal.com/ai/markdown
- Posts with this tag: 2 of 16
- Co-occurring tags: Sparse Autoencoders, Trustworthy AI

## Posts

- [Before It Codes: Catching Qwen3.5-4B Planning With J-Lens](https://aroramrinaal.com/ai/j-lens/markdown) — 2026-07-30; tags: Interpretability, Trustworthy AI. I used Anthropic's new Jacobian lens to test whether Qwen3.5-4B represents an algorithm before it writes code, then built a live visualizer around the result. [HTML](https://aroramrinaal.com/ai/j-lens)
- [Looking Inside Qwen3-4B With Sparse Autoencoders](https://aroramrinaal.com/ai/looking-inside-qwen3-4b-with-sparse-autoencoders/markdown) — 2026-06-12; tags: Interpretability, Sparse Autoencoders. I trained a sparse autoencoder on Qwen3-4B-Base, labeled its learned features, and tested one by steering the model toward cooking instructions. [HTML](https://aroramrinaal.com/ai/looking-inside-qwen3-4b-with-sparse-autoencoders)
