Saltar al contenido principal

Escribe una PREreview

Visual Tokenizers: A Comprehensive Survey

Publicada
Servidor
Preprints.org
DOI
10.20944/preprints202607.1756.v1

Visual tokenization bridges the gap between high-dimensional visual data and sequence modeling by transforming images, videos, and 3D content into compact token sequences. Recent advances in Multimodal Large Language Models (LLMs) have further underscored the critical role of visual tokenizers. Despite this rapid progress, the literature in this area remains fragmented. This work addresses this gap by presenting a comprehensive taxonomy of visual tokenizers. We review the evolution of task-specific tokenizers along two complementary dimensions: high-level tokenizers that emphasize semantic representations and low-level tokenizers that focus on pixel reconstruction and compression. Motivated by the emergence of Multimodal Large Language Models capable of supporting both understanding and generation tasks, we highlight recent advances in unified tokenizers. These models aim to integrate semantic abstraction with fine-grained visual details within a single representation. We categorize unified tokenizers into single-encoder and dual-encoder architectures and further organize them according to their semantic preservation strategies and fusion mechanisms. We also summarize empirical results on downstream understanding and generation benchmarks, comparing unified tokenizers with both individual task-specific models and combinations of separate tokenizers. Finally, we discuss open challenges and future directions, including the trade-off between semantic abstraction and reconstruction fidelity, scalable token vocabularies, and extensions to video and 3D domains. We further examine the evolving role of visual tokens in next-generation multimodal systems. We maintain a curated list of related works at https://github.com/Shi-qingyu/Awesome-Visual-Tokenizer.

Puedes escribir una PREreview de Visual Tokenizers: A Comprehensive Survey. Una PREreview es una revisión de un preprint y puede variar desde unas pocas oraciones hasta un extenso informe, similar a un informe de revisión por pares organizado por una revista.

Antes de comenzar

We will ask you to log in with your ORCID iD. If you don’t have an iD, you can create one.

What is an ORCID iD?

An ORCID iD is a unique identifier that distinguishes you from everyone with the same or similar name.

Comenzar ahora