`Don't Get Too Technical with Me': A Discourse Structure-Based Framework for Automatic Science Journalism

Published: 07 Oct 2023, Last Modified: 01 Dec 2023EMNLP 2023 MainEveryoneRevisionsBibTeX
Submission Type: Regular Long Paper
Submission Track: Summarization
Submission Track 2: Natural Language Generation
Keywords: automatic scientific journalism, summarization, style transfer
TL;DR: We introduce a new dataset for `scientific journalism', where a summary in journalistic style is generated conditioned on a scientific article. The proposed systems leverage a paper's discourse structure and its metadata to guide generation.
Abstract: Science journalism refers to the task of reporting technical findings of a scientific paper as a less technical news article to the general public audience. We aim to design an automated system to support this real-world task (i.e., automatic science journalism ) by 1) introducing a newly-constructed and real-world dataset (SciTechNews), with tuples of a publicly-available scientific paper, its corresponding news article, and an expert-written short summary snippet; 2) proposing a novel technical framework that integrates a paper's discourse structure with its metadata to guide generation; and, 3) demonstrating with extensive automatic and human experiments that our model outperforms other baseline methods (e.g. Alpaca and ChatGPT) in elaborating a content plan meaningful for the target audience, simplify the information selected, and produce a coherent final report in a layman's style.
Submission Number: 2154
Loading