Write a PREreview

Data Efficient Training of a U-Net Based Architecture for Structured Documents Localization

by Anastasiia Kabeshova, Guillaume Betmont, Julien Lerouge, Evgeny Stepankevich, and Alexis Bergès

Posted: October 2, 2023
Server: arXiv
DOI: 10.48550/arxiv.2310.00937

Structured documents analysis and recognition are essential for modern online on-boarding processes, and document localization is a crucial step to achieve reliable key information extraction. While deep-learning has become the standard technique used to solve document analysis problems, real-world applications in industry still face the limited availability of labelled data and of computational resources when training or fine-tuning deep-learning models. To tackle these challenges, we propose SDL-Net: a novel U-Net like encoder-decoder architecture for the localization of structured documents. Our approach allows pre-training the encoder of SDL-Net on a generic dataset containing samples of various document classes, and enables fast and data-efficient fine-tuning of decoders to support the localization of new document classes. We conduct extensive experiments on a proprietary dataset of structured document images to demonstrate the effectiveness and the generalization capabilities of the proposed approach.

You can write a PREreview of Data Efficient Training of a U-Net Based Architecture for Structured Documents Localization. A PREreview is a review of a preprint and can vary from a few sentences to a lengthy report, similar to a journal-organized peer-review report.

Before you start

We will ask you to log in with your ORCID iD. If you don’t have an iD, you can create one.