Gemma 3 Natural Language Autoencoders
Recently, Anthropic released a suite of Natural Language Autoencoders (NLAs) for Gemma 3. What is an NLA? A language model like ChatGPT converts words into lists of numbers and outputs words. These intermediate numbers are activation vectors, and are generally not interpretable by humans. An NLA is comprised of a language model and a copy of the same model that has been trained to map activation vectors to English (Activation Verbalizer). The model and activation verbalizer are trained together to minimize the reconstruction error between the original activation vector and an activation vector that has been verbalized and passed through the model again.
Originally devised as a tool to understand what LLMs are "thinking" at each text token position, we can apply the same technique to image tokens, since Gemma 3 is a multimodal model with a vision tower.
Although suffering from limitations like hallucinations and rigid prompt structure, I find the results quite fun and interesting, as you can see below.
You can make your own image explanations with the code here
Hyangwonjeong pond

ICML 2026 GenBio Workshop

Korea Tim Hortons has Lobster Rolls!
