Skip to main content

Write a PREreview

Visual Tokenizers: A Comprehensive Survey

Posted
Server
Preprints.org
DOI
10.20944/preprints202607.1756.v1

Visual tokenization bridges the gap between high-dimensional visual data and sequence modeling by transforming images, videos, and 3D content into compact token sequences. Recent advances in Multimodal Large Language Models (LLMs) have further underscored the critical role of visual tokenizers. Despite this rapid progress, the literature in this area remains fragmented. This work addresses this gap by presenting a comprehensive taxonomy of visual tokenizers. We review the evolution of task-specific tokenizers along two complementary dimensions: high-level tokenizers that emphasize semantic representations and low-level tokenizers that focus on pixel reconstruction and compression. Motivated by the emergence of Multimodal Large Language Models capable of supporting both understanding and generation tasks, we highlight recent advances in unified tokenizers. These models aim to integrate semantic abstraction with fine-grained visual details within a single representation. We categorize unified tokenizers into single-encoder and dual-encoder architectures and further organize them according to their semantic preservation strategies and fusion mechanisms. We also summarize empirical results on downstream understanding and generation benchmarks, comparing unified tokenizers with both individual task-specific models and combinations of separate tokenizers. Finally, we discuss open challenges and future directions, including the trade-off between semantic abstraction and reconstruction fidelity, scalable token vocabularies, and extensions to video and 3D domains. We further examine the evolving role of visual tokens in next-generation multimodal systems. We maintain a curated list of related works at https://github.com/Shi-qingyu/Awesome-Visual-Tokenizer.

You can write a PREreview of Visual Tokenizers: A Comprehensive Survey. A PREreview is a review of a preprint and can vary from a few sentences to a lengthy report, similar to a journal-organized peer-review report.

Before you start

We will ask you to log in with your ORCID iD. If you don’t have an iD, you can create one.

What is an ORCID iD?

An ORCID iD is a unique identifier that distinguishes you from everyone with the same or similar name.

Start now