Toward Reliable Ad-hoc Scientific Information Extraction: A Case Study on Two Materials Dataset

Published: 01 Jan 2024, Last Modified: 20 May 2025ACL (Findings) 2024EveryoneRevisionsBibTeXCC BY-SA 4.0
Abstract: We explore the ability of GPT-4 to perform ad-hoc schema-based information extraction from scientific literature. We assess specifically whether it can, with a basic one-shot prompting approach over the full text of the included manusciprts, replicate two existing material science datasets, one pertaining to multi-principal element alloys (MPEAs), and one to silicate diffusion. We collaborate with materials scientists to perform a detailed manual error analysis to assess where and why the model struggles to faithfully extract the desired information, and draw on their insights to suggest research directions to address this broadly important task.
Loading