Gemma 3 Natural Language Autoencoders

Recently, Anthropic released a suite of Natural Language Autoencoders (NLAs) for Gemma 3. What is an NLA? A language model like ChatGPT converts words into lists of numbers and outputs words. These intermediate numbers are activation vectors, and are generally not interpretable by humans. An NLA is comprised of a language model and a copy of the same model that has been trained to map activation vectors to English (Activation Verbalizer). The model and activation verbalizer are trained together to minimize the reconstruction error between the original activation vector and an activation vector that has been verbalized and passed through the model again.NLA diagramOriginally devised as a tool to understand what LLMs are "thinking" at each text token position, we can apply the same technique to image tokens, since Gemma 3 is a multimodal model with a vision tower.NLA diagramAlthough suffering from limitations like hallucinations and rigid prompt structure, I find the results quite fun and interesting, as you can see below.
You can make your own image explanations with the code here

Hover on a patch for its explanation. Heatmap shades each image token by its L2 norm. Click to copy explanations
low norm
high norm

Hyangwonjeong pond

Hyangwonjeong pond, Gyeongbokgung palace
An image of the beautiful Hyangwonjeong pond in the north section of Gyeongbukgong palace, which I took on a recent trip to Korea. The king Gojong built this pond and an extension to the palace called Geoncheonggung to assert his independence from his father.

ICML 2026 GenBio Workshop

Conference

Korea Tim Hortons has Lobster Rolls!

Lobster