Structured Tender Entities Extraction from Complex Tables with Few-short Learning

Asim Abbas, Mark Lee, Niloofer Shanavas, Venelin Kovatchev, Mubashir Ali

Published: 2025, Last Modified: 28 May 2026COLING Workshops 2025EveryoneRevisionsBibTeXCC BY-SA 4.0
Abstract: Extracting structured text from complex tables in PDF tender documents remains a challenging task due to the loss of structural and positional information during the extraction process. AI-based models often require extensive training data, making development from scratch both tedious and time-consuming. Our research focuses on identifying tender entities in complex table formats within PDF documents. To address this, we propose a novel approach utilizing few-shot learning with large language models (LLMs) to restore the structure of extracted text. Additionally, handcrafted rules and regular expressions are employed for precise entity classification. To evaluate the robustness of LLMs with few-shot learning, we employ data-shuffling techniques. Our experiments show that current text extraction tools fail to deliver satisfactory results for complex table structures. However, the few-shot learning approach significantly enhances the structural integrity of extracted data and improves the accuracy of tender entity identification.
Loading