Source
COLING SMM
DATE OF PUBLICATION
10/17/2022
Authors
Alexander Panchenko Anton Razzhigaev Anton Voronov Andrey Kaznacheev Denis Dimitrov
Share

Pixel-Level BPE for Auto-Regressive Image Generation

Abstract

Pixel-level autoregression with Transformer models (Image GPT or iGPT) is one of the recent approaches to image generation that has not received massive attention and elaboration due to quadratic complexity of attention as it imposes huge memory requirements and thus restricts the resolution of the generated images.
In this paper, we propose to tackle this problem by adopting Byte-Pair-Encoding (BPE) originally proposed for text processing to the image domain to drastically reduce the length of the modeled sequence. The obtained results demonstrate that it is possible to decrease the amount of computation required to generate images pixel-by-pixel while preserving their quality and the expressiveness of the features extracted from the model. Our results show that there is room for improvement for iGPT-like models with more thorough research on
the way to the optimal sequence encoding techniques for images

Join AIRI