Event Dime

SAE It Across Models: Explaining Features With Foreign NLA Verbalizers

Music & Concerts
When:
June 5, 2026 · 6:36 PM
Where:
A review of Where Is My Flying Car? by J. Storrs Hall — LessWrong
Source:
A review of Where Is My Flying Car? by J. Storrs Hall — LessWrong

TLDR: I show that a foreign model's Natural Language Autoencoder (NLA) Activation Verbalizer (AV) can produce plausible explanations for SAE features from a model it was never trained on. It is currently assumed that these tools only work for the exact model and layer they were trained for. I show that is not the case. After creating a ridge-regression map bridging the residual stream of Qwen2.5-7B-IT at layer 20 and Gemma-3-27B-IT at layer 41, I mapped 45 SAE decoder directions from a Qwen SAE

Why this event appears here

This listing is shown because it is tied to the current city coverage and carries structured metadata that can be indexed and compared with other events.

  • City coverage

    A review of Where Is My Flying Car? by J. Storrs Hall — LessWrong

  • Primary source

    A review of Where Is My Flying Car? by J. Storrs Hall — LessWrong

  • Categories

    Music & Concerts