Gaze Attention: Query-Adaptive Visual Routing for Efficient Multimodal LLMs

CoLM 2026 Main
CoLM 2026 Workshop on Efficient Reasoning Spotlight 💫

1NAVER AI Lab 2KAIST

*Project lead †Corresponding author

TL;DR

When a multimodal LLM writes the word “dog”, does it really need to look at the entire image?

  • No! Humans fixate on the region they are about to describe.
  • Current multimodal LLMs (MLLMs) with dense attention look at every visual token at every generation step.
  • Gaze Attention shifts its focus as it generates text.
A black and white dog sits on a porch in front of a red bicycle.

A black and white dog sits attentively on a porch, its gaze fixed on the camera, while a red bicycle leans against the railing behind it.

The routing is learned without any gaze supervision. Click a word to jump to it.
Attention of Cambrian-4B (Qwen3-4B + SigLIP2-So/16-512) while each word is generated.

Dense Attention Looks Everywhere, All the Time

  • An image or a video turns into hundreds to tens of thousands of visual tokens. With dense attention, every generated token attends to all of them, whether or not they matter for the current word.
  • Human perception is selective: we shift our gaze to the region we are about to describe.
  • The mismatch dilutes the model's focus and makes attention over visual tokens the dominant cost for high-resolution images and long videos, where the visual KV cache keeps growing.
generating “dog” generating “bicycle”
Dense attention all 1,024 visual tokens
Gaze Attention 128 tokens in 2 regions
low high
Attention maps of one head at layer 15 over the 1,024 visual tokens of a 512 px image (log-normalized). Dense attention spreads over the whole image. Gaze Attention first selects two regions and attends only inside them.

Deciding Before Generation Is Too Early

Efficient MLLMs already reduce the visual information that the LLM attends to. But they fix it during prefill, before it is known which regions each prediction will need.

Four diagrams: token compression reduces vision tokens before the LLM; frame-level and token-level KV-cache eviction discard vision KV-caches during prefill; Gaze Attention keeps all vision KV-caches and selects the regions each generated word needs.
Token compression and KV-cache eviction reduce the visual KV cache during prefill.
Gaze Attention keeps the cache and selects query-relevant gaze regions at each generation step.

Token compression

VideoChat-Flash, F-16

Decided once, before the LLM

Visual tokens are pooled or pruned up front, so fine details are blurred or gone for every later prediction.

KV-cache eviction

ReKV, HERMES, InfiniPot-V, StreamMem

Decided once, from the question

Visual entries are scored once and the rest is evicted. The retained set stays fixed, even when later words need different evidence.

Gaze Attention

Ours

Decided at every decoding step

The full KV cache is kept and each query is routed to its own regions. A region skipped now can be attended to later.

See what a fixed selection misses during generation ↓

Gaze Attention: Select, Then Attend

Inside every LLM layer, Gaze Attention replaces dense attention over visual tokens with two steps: a lightweight routing step picks a few regions, and attention is computed only over them.

Overview of Gaze Attention: visual keys are grouped into regions and summarized by descriptors, each decoding query routes its attention to the relevant regions, and learnable context tokens preserve the global context. The attention masks of dense attention and Gaze Attention are compared on the right.
  1. 1

    Gaze regions and descriptors

    The cached keys of the visual tokens are arranged in their spatial and temporal layout and split into regions of \(r^h \times r^w \times r^t\) tokens (6 × 6 × 1 by default). A descriptor \(d_g^u\), the mean of the keys of a region, summarizes it without extra parameters.

  2. 2

    Visual routing

    At decoding step \(j\), the query scores every descriptor, \(s_g^u = q_j^{\top} d_g^u\), and attends only to the key–value pairs of its Top-K regions. Nothing is evicted, so a region skipped now can be selected later.

  3. 3

    Learnable context tokens

    A few learnable tokens (4 by default) are appended to each image or frame and attend only to it, which yields a compact summary \(k_C^u\). It stays visible to every query and keeps the global scene in view while the routed regions change.

Stable training

Routing is trained with the next-token prediction loss only. Because queries and descriptors are barely related early in training, a progressive Top-K schedule starts dense and gradually reduces the number of selected regions.

What a query attends to

An 8-frame video has 8 × 576 = 4,608 visual tokens. With regions of 36 tokens, 20 routed regions and 4 context tokens per frame, each query attends to 20 × 36 + 4 × 8 = 752 entries: 16.3% of dense attention.

Where the Model Looks

No gaze supervision. Routing is learned through next-token prediction alone. The examples below show the regions selected by the model.

A fixed selection can miss what later words need

InfiniPot-V and HERMES keep one set of visual tokens for the whole answer. Gaze Attention selects regions again for every generated token. Bright areas are the visual tokens attended to at layers 10 and 23.

A corgi holding white flowers in its mouth next to a black and white cat on grass.

Q. What is the dog holding in its mouth, and what is the cat doing?

InfiniPot-V

The dog is holding a piece of paper in its mouth, and the cat is looking at the dog.

Tokens kept by InfiniPot-V at layer 10: scattered patches across the image.
Tokens kept by InfiniPot-V at layer 23: scattered patches across the image.

KV size 24 / 121, unchanged during decoding

HERMES

The dog is holding a bunch of white flowers. The cat is standing, looking at it with an attentive gaze.

Tokens kept by HERMES at layer 10: mostly the bottom strip of grass and a few patches.
Tokens kept by HERMES at layer 23: scattered patches on the dog and the cat.

KV size 39 / 196, unchanged during decoding

Gaze Attention

KV size 18 / 144 at every step

The dog is holding a bunch of white flowers in its mouth, and the cat is sitting on the grass.

While generating “flowers”
Regions routed at layer 10 for the word flowers: the flowers in the dog's mouth.
Regions routed at layer 23 for the word flowers: the dog's head and the flowers.
While generating “cat”
Regions routed at layer 10 for the word cat: the cat.
Regions routed at layer 23 for the word cat: the cat.
While generating “grass”
Regions routed at layer 10 for the word grass: grass regions.
Regions routed at layer 23 for the word grass: the grass at the bottom of the image.

Routing across space and time

In videos, a query selects frame–region pairs. Red boxes mark the regions attended to while the highlighted word is generated.

Q. What is the man helping the woman with, and what is she enjoying?

A. The man is helping the woman to climb up the ship's ladder. She is enjoying the view of the sunset.

While generating “helping”
Seven video frames. The routed regions cover the man's hand holding the woman's hand and the two people standing at the bow of the ship.
While generating “She”
The same seven frames. The routed regions now cover the woman and the sunset sky in four frames.

Q. What happens in this video? Describe the actions of the man and the boy briefly.

A. A man climbs down a staircase and falls on the floor. A boy sits on the stairs and laughs.

While generating “falls”
Eight video frames. The routed regions cover the man on the staircase in the first frame and the man lying on the floor in the third frame.
While generating “boy”
The same eight frames. The routed regions now cover the boy in the fourth and the seventh frame.

Results

We build Cambrian-4B (SigLIP2-So/16 with Qwen3-4B) and Cambrian-7B (with Qwen2.5-7B) and train dense-attention and Gaze Attention models with an identical architecture and token-matched training. They are evaluated on 13 image and 6 video understanding benchmarks.

Up to 90% fewer visual KV entries Video: 2K instead of 20K entries, still +1.1 average points
13 image benchmarks +1.5 average points over dense attention with 148 instead of 576 visual KV entries per query
6 video benchmarks +1.4 average points over dense attention with 4K instead of 20K visual KV entries per query
Attention FLOPs −83.9% with 1.5K instead of 10K entries and 79.8% less KV-cache memory on the GPU

Results Summary

  • Selecting visual regions for each decoding query matches or improves the accuracy of dense attention while attending to up to 90% fewer visual KV entries.
  • Under matched visual KV budgets, Gaze Attention achieves higher average performance than KV-cache eviction and than dense attention over fewer tokens.
  • Routing learned from next-token prediction alone follows the content being generated: the model shifts its gaze without any gaze supervision.
1

Image understanding: accuracy holds as the visual KV budget shrinks

Average score change vs. visual KV budget

Average over the 13 image benchmarks, relative to the dense-attention baseline of each model

    Show the numbers
    Model Visual KV entries Avg. Change
    Qwen2.5-VL-3B 256 63.3 –
    + InfiniPot-V 128 60.3 −3.0
    + InfiniPot-V 50 52.8 −10.5
    LLaVA-OV-7B 196 62.4 –
    + HERMES 100 57.9 −4.5
    + HERMES 40 52.3 −10.1
    Cambrian-4B 576 59.7 –
    + Gaze Attention 288 + 4 61.1 +1.4
    + Gaze Attention 144 + 4 61.2 +1.5
    + Gaze Attention 72 + 4 60.4 +0.7
    Cambrian-4B + token compression 144 57.8 –
    + Gaze Attention 54 + 4 59.2 +1.4
    + Gaze Attention 36 + 4 58.3 +0.5
    • KV-cache eviction degrades quickly on single images: keeping about 20% of the visual entries costs InfiniPot-V 10.5 and HERMES 10.1 average points.
    • Gaze Attention stays above its dense baseline at every budget: +1.5 points with 144 of 576 entries, and +0.7 with only 72.
    • It also combines with input-level token compression: with 36 routed entries it scores 58.3, comparable to dense attention over 144 entries (57.8).
    2

    Video understanding: better than dense attention with a fraction of the cache

    Average score change vs. visual KV budget

    Cambrian-4B, averaged over six video benchmarks, relative to dense attention with 20K visual KV entries

      Show the numbers
      Model KV VideoMME MLVU EgoSchema LongVideoB. NExT-QA TempCom. Change
      Cambrian-4B 20K 58.8 65.9 53.5 57.5 77.1 60.5 –
      Cambrian-4B 4K 55.7 61.4 53.2 54.0 75.3 63.7 −1.7
      + HERMES 4K 58.1 65.0 53.4 56.2 76.5 62.7 −0.2
      + Gaze Attention 4K 60.4 67.5 53.8 56.8 80.4 62.9 +1.4
      + Gaze Attention 2K 59.4 67.1 53.2 56.6 79.9 63.8 +1.1
      Cambrian-7B 10K 60.0 66.5 55.8 58.9 78.4 60.9 –
      + HERMES 1.5K 60.1 66.3 55.6 58.5 78.1 64.2 +0.4
      + Gaze Attention 1.5K 60.4 67.0 56.2 58.6 80.4 63.7 +1.0
      • On Cambrian-4B, attending to 4K instead of 20K visual KV entries improves the average of the six benchmarks by 1.4 points. With 2K entries, 90% fewer, the gain is still 1.1.
      • Under the same 4K budget, a dense model that simply receives fewer tokens averages 60.6 and HERMES stays 0.2 below the baseline, while Gaze Attention reaches 63.6.
      • The trend holds for Cambrian-7B: +1.0 points with 1.5K of 10K entries.
      3

      Efficiency: attention FLOPs and KV-cache memory

      Attention FLOPs

      GFLOPs, all attention layers

      −83.9%
      Dense 5.52
      Gaze 0.89 +0.34
      • Attention 0.89
      • Routing 0.19
      • Context tokens 0.15

      KV-cache memory

      GB on the GPU

      −79.8%
      Dense 1.37
      Gaze 0.27
      • Dense attention, 10K visual KV entries
      • Gaze Attention, 1.5K
      • Each query attends to 1.5K instead of 10K visual KV entries, which reduces the attention FLOPs by 83.9%. Routing and the context tokens add 0.19 and 0.15 GFLOPs.
      • With KV-cache offloading, the full cache stays in CPU memory or on disk and only the selected regions are loaded onto the GPU: 79.8% less KV-cache memory.
      4

      Analysis: does routing find the right regions?

      Routing accuracy

      Generated tokens whose routed region overlaps the object they describe (%)

      Random 11.2 ±1.4
      Layer 9 56.6 ±2.3
      Layer 15 57.7 ±2.1
      Layer 21 51.3 ±2.7

      Mean ± standard deviation over three seeds.

      Learnable context tokens

      Average score relative to the default of 4 tokens

      Gaze region size

      Average score relative to the default of 6 × 6 tokens

      • For 200 images, each generated caption is split into phrases and each token is aligned to one of them. A token counts as correct when its routed region overlaps the segmentation mask of its phrase. Gaze Attention reaches 51.3–57.7% against 11.2% for random routing, without any supervision of where to look.
      • Context tokens add global information that the routed regions miss: removing them costs 0.9 points, and more than 4 tokens give diminishing returns while their cost grows with the number of frames.
      • Finer regions score slightly higher but multiply the routing candidates, and 12 × 12 regions are too coarse. Regions of 6 × 6 tokens balance both.
      • The progressive Top-K schedule matters: enforcing sparse routing from the first training step, before reliable routing signals emerge, lowers the average image score.

      Citation

      @inproceedings{song2026gazeattention,
        title     = {Gaze Attention: Query-Adaptive Visual Routing for Efficient Multimodal {LLM}s},
        author    = {Song, Junha and Heo, Byeongho and Gu, Geonmo and Choo, Jaegul and Han, Dongyoon and Yun, Sangdoo},
        booktitle = {Third Conference on Language Modeling (CoLM)},
        year      = {2026}
      }

      Acknowledgements

      This page follows the style of our PIVOT and MM-SeR pages, which were developed with reference to Web-SSL.