A Corpus and Evaluation for Predicting Semi-Structured Human Annotations

Published: 07 Dec 2022, Last Modified: 02 Apr 20252nd Workshop on Natural Language Generation, Evaluation, and Metrics (GEM)EveryoneRevisionsCC BY 4.0
Abstract: A wide variety of tasks have been framed as text-to-text tasks to allow processing by sequence-to-sequence models. We propose a new task of generating a semi-structured interpretation of a source document. The interpretation is semi-structured in that it contains mandatory and optional fields with free-text information. This structure is surfaced by human annotations, which we standardize and convert to text format. We then propose an evaluation technique that is generally applicable to any such semi-structured annotation, called equivalence classes evaluation. The evaluation technique is efficient and scalable; it creates a large number of evaluation instances from a comparably cheap clustering of the free-text information by domain experts. For our task, we release a dataset about the monetary policy of the Federal Reserve. On this corpus, our evaluation shows larger differences between pretrained models than standard text generation metrics.
Loading