TextTIGER: Text-based Intelligent Generation with Entity Prompt Refinement for Text-to-Image Generation

TMLR Paper8063 Authors

24 Mar 2026 (modified: 20 May 2026)Under review for TMLREveryoneRevisionsBibTeXCC BY 4.0
Abstract: When generating images from prompts that include specific entities, the model must retain as much entity-specific knowledge as possible. However, the number of entities is enormous, and new entities emerge; memorizing all of them completely is not realistic. To bridge this gap, our work proposes Text-based Intelligent Generation with Entity Prompt Refinement (\textsc{TextTIGER}). \textsc{TextTIGER} strengthens knowledge about entities that appear in the prompt by augmenting external information and then summarizes the expanded descriptions with large language models, preventing performance degradation that arises from excessively long inputs. To evaluate our method, we construct a new dataset consisting of captions, images, detailed descriptions, and lists of entities. Experiments with multiple image generation models show that \textsc{TextTIGER} improves image generation performance on widely used evaluation metrics compared with prompts that use captions alone. In addition, using Multimodal LLM (MLLM)-as-a-judge, which shows a strong correlation with human evaluation, we demonstrate that our method consistently achieves higher scores, which underscores its effectiveness. These results show that strengthening entity-related descriptions, summarizing them, and refining prompts to an appropriate length leads to substantial improvements in image generation performance. We will release the created dataset and code upon acceptance.
Submission Type: Regular submission (no more than 12 pages of main content)
Assigned Action Editor: ~Hongyang_Zhang1
Submission Number: 8063
Loading