Token compression
VideoChat-Flash, F-16
Decided once, before the LLM
Visual tokens are pooled or pruned up front, so fine details are blurred or gone for every later prediction.
When a multimodal LLM writes the word “dog”, does it really need to look at the entire image?
A black and white dog sits attentively on a porch, its gaze fixed on the camera, while a red bicycle leans against the railing behind it.
The routing is learned without any gaze supervision. Click a
word to jump to it.
Attention of Cambrian-4B (Qwen3-4B + SigLIP2-So/16-512) while
each word is generated.
Select before attend
Visual KV entries per query: 384 px image · video
Efficient MLLMs already reduce the visual information that the LLM attends to. But they fix it during prefill, before it is known which regions each prediction will need.
VideoChat-Flash, F-16
Decided once, before the LLM
Visual tokens are pooled or pruned up front, so fine details are blurred or gone for every later prediction.
ReKV, HERMES, InfiniPot-V, StreamMem
Decided once, from the question
Visual entries are scored once and the rest is evicted. The retained set stays fixed, even when later words need different evidence.
Ours
Decided at every decoding step
The full KV cache is kept and each query is routed to its own regions. A region skipped now can be attended to later.
Inside every LLM layer, Gaze Attention replaces dense attention over visual tokens with two steps: a lightweight routing step picks a few regions, and attention is computed only over them.
The cached keys of the visual tokens are arranged in their spatial and temporal layout and split into regions of \(r^h \times r^w \times r^t\) tokens (6 × 6 × 1 by default). A descriptor \(d_g^u\), the mean of the keys of a region, summarizes it without extra parameters.
At decoding step \(j\), the query scores every descriptor, \(s_g^u = q_j^{\top} d_g^u\), and attends only to the key–value pairs of its Top-K regions. Nothing is evicted, so a region skipped now can be selected later.
A few learnable tokens (4 by default) are appended to each image or frame and attend only to it, which yields a compact summary \(k_C^u\). It stays visible to every query and keeps the global scene in view while the routed regions change.
Routing is trained with the next-token prediction loss only. Because queries and descriptors are barely related early in training, a progressive Top-K schedule starts dense and gradually reduces the number of selected regions.
An 8-frame video has 8 × 576 = 4,608 visual tokens. With regions of 36 tokens, 20 routed regions and 4 context tokens per frame, each query attends to 20 × 36 + 4 × 8 = 752 entries: 16.3% of dense attention.
No gaze supervision. Routing is learned through next-token prediction alone. The examples below show the regions selected by the model.
InfiniPot-V and HERMES keep one set of visual tokens for the whole answer. Gaze Attention selects regions again for every generated token. Bright areas are the visual tokens attended to at layers 10 and 23.
Q. What is the dog holding in its mouth, and what is the cat doing?
The dog is holding a piece of paper in its mouth, and the cat is looking at the dog.
KV size 24 / 121, unchanged during decoding
The dog is holding a bunch of white flowers. The cat is standing, looking at it with an attentive gaze.
KV size 39 / 196, unchanged during decoding
The dog is holding a bunch of white flowers in its mouth, and the cat is sitting on the grass.
Q. Where are the people and the dog, and what are they doing?
The people and the dog are in a park. The dog is running towards the people.
The people and the dog are in a grassy area. The woman is holding a frisbee.
The people and the dog are in a grassy area. They are playing with a frisbee.
Q. Where are the man, the dog, and the sheep, and what are they doing?
The man is standing in the middle of a field, holding a leash attached to a black dog.
KV size 24 / 121, unchanged during decoding
The man is standing on the left side of the image, holding a red leash attached to a dog.
KV size 39 / 196, unchanged during decoding
The man, the dog, and the sheep are in a field. The man is holding a stick and the dog is herding the sheep.
Q. What is the dog doing, and how is the bicycle positioned?
The dog is sitting on the ground, looking at the bicycle. The bicycle is leaning against a tree.
The dog is sitting on a porch, and the bicycle is leaning against it.
The dog is sitting on the porch. The bicycle is leaning against the porch railing.
In videos, a query selects frame–region pairs. Red boxes mark the regions attended to while the highlighted word is generated.
Q. What is the man helping the woman with, and what is she enjoying?
A. The man is helping the woman to climb up the ship's ladder. She is enjoying the view of the sunset.
Q. What happens in this video? Describe the actions of the man and the boy briefly.
A. A man climbs down a staircase and falls on the floor. A boy sits on the stairs and laughs.
We build Cambrian-4B (SigLIP2-So/16 with Qwen3-4B) and Cambrian-7B (with Qwen2.5-7B) and train dense-attention and Gaze Attention models with an identical architecture and token-matched training. They are evaluated on 13 image and 6 video understanding benchmarks.
Average over the 13 image benchmarks, relative to the dense-attention baseline of each model
| Model | Visual KV entries | Avg. | Change |
|---|---|---|---|
| Qwen2.5-VL-3B | 256 | 63.3 | – |
| + InfiniPot-V | 128 | 60.3 | −3.0 |
| + InfiniPot-V | 50 | 52.8 | −10.5 |
| LLaVA-OV-7B | 196 | 62.4 | – |
| + HERMES | 100 | 57.9 | −4.5 |
| + HERMES | 40 | 52.3 | −10.1 |
| Cambrian-4B | 576 | 59.7 | – |
| + Gaze Attention | 288 + 4 | 61.1 | +1.4 |
| + Gaze Attention | 144 + 4 | 61.2 | +1.5 |
| + Gaze Attention | 72 + 4 | 60.4 | +0.7 |
| Cambrian-4B + token compression | 144 | 57.8 | – |
| + Gaze Attention | 54 + 4 | 59.2 | +1.4 |
| + Gaze Attention | 36 + 4 | 58.3 | +0.5 |
Cambrian-4B, averaged over six video benchmarks, relative to dense attention with 20K visual KV entries
| Model | KV | VideoMME | MLVU | EgoSchema | LongVideoB. | NExT-QA | TempCom. | Change |
|---|---|---|---|---|---|---|---|---|
| Cambrian-4B | 20K | 58.8 | 65.9 | 53.5 | 57.5 | 77.1 | 60.5 | – |
| Cambrian-4B | 4K | 55.7 | 61.4 | 53.2 | 54.0 | 75.3 | 63.7 | −1.7 |
| + HERMES | 4K | 58.1 | 65.0 | 53.4 | 56.2 | 76.5 | 62.7 | −0.2 |
| + Gaze Attention | 4K | 60.4 | 67.5 | 53.8 | 56.8 | 80.4 | 62.9 | +1.4 |
| + Gaze Attention | 2K | 59.4 | 67.1 | 53.2 | 56.6 | 79.9 | 63.8 | +1.1 |
| Cambrian-7B | 10K | 60.0 | 66.5 | 55.8 | 58.9 | 78.4 | 60.9 | – |
| + HERMES | 1.5K | 60.1 | 66.3 | 55.6 | 58.5 | 78.1 | 64.2 | +0.4 |
| + Gaze Attention | 1.5K | 60.4 | 67.0 | 56.2 | 58.6 | 80.4 | 63.7 | +1.0 |
GFLOPs, all attention layers
GB on the GPU
Generated tokens whose routed region overlaps the object they describe (%)
Mean ± standard deviation over three seeds.
Average score relative to the default of 4 tokens
Average score relative to the default of 6 × 6 tokens
@inproceedings{song2026gazeattention,
title = {Gaze Attention: Query-Adaptive Visual Routing for Efficient Multimodal {LLM}s},
author = {Song, Junha and Heo, Byeongho and Gu, Geonmo and Choo, Jaegul and Han, Dongyoon and Yun, Sangdoo},
booktitle = {Third Conference on Language Modeling (CoLM)},
year = {2026}
}