Visual Tokenizers: A Comprehensive Survey
Escrito por Qingyu Shi, Haobo Yuan, Yikang Zhou, Xiu Li, Jason Li, Size Wu, Haochen Wang, Ye Tian, Tao Zhang, Jinbin Bai, Yujing Wang, Lu Qi, Yunhai Tong y Ming-Hsuan Yang
- Publicado
- Servidor
- Preprints.org
Resumen
Visual tokenization bridges the gap between high-dimensional visual data and sequence modeling by transforming images, videos, and 3D content into compact token sequences. Recent advances in Multimodal Large Language Models (LLMs) have further underscored the critical role of visual tokenizers. Despite this rapid progress, the literature in this area remains fragmented. This work addresses this gap by presenting a comprehensive taxonomy of visual tokenizers. We review the evolution of task-specific tokenizers along two complementary dimensions: high-level tokenizers that emphasize semantic representations and low-level tokenizers that focus on pixel reconstruction and compression. Motivated by the emergence of Multimodal Large Language Models capable of supporting both understanding and generation tasks, we highlight recent advances in unified tokenizers. These models aim to integrate semantic abstraction with fine-grained visual details within a single representation. We categorize unified tokenizers into single-encoder and dual-encoder architectures and further organize them according to their semantic preservation strategies and fusion mechanisms. We also summarize empirical results on downstream understanding and generation benchmarks, comparing unified tokenizers with both individual task-specific models and combinations of separate tokenizers. Finally, we discuss open challenges and future directions, including the trade-off between semantic abstraction and reconstruction fidelity, scalable token vocabularies, and extensions to video and 3D domains. We further examine the evolving role of visual tokens in next-generation multimodal systems. We maintain a curated list of related works at https://github.com/Shi-qingyu/Awesome-Visual-Tokenizer.
Leer el preprint