Sentence-Level Resampling for Named Entity RecognitionDownload PDF


08 Mar 2022 (modified: 05 May 2023)NAACL 2022 Conference Blind SubmissionReaders: Everyone
Paper Link:
Paper Type: Long paper (up to eight pages of content + unlimited references and appendices)
Abstract: As a fundamental task in natural language processing, named entity recognition (NER) aims to locate and classify named entities in unstructured text. However, named entities are always the minority among all tokens in the text. This data imbalance problem presents a challenge to machine learning models as their learning objective is usually dominated by the majority of non-entity tokens. To alleviate data imbalance, we propose a set of sentence-level resampling methods where the importance of each training sentence is computed based on its tokens and entities. We study the generalizability of these resampling methods on a wide variety of NER models (CRF, Bi-LSTM, and BERT) across corpora from diverse domains (general, social, and medical texts). Extensive experiments show that the proposed methods improve span-level macro F1-scores of the evaluated NER models on multiple corpora, frequently outperforming sub-sentence-level resampling, data augmentation, and special loss functions such as focal and Dice loss.
Presentation Mode: This paper will be presented virtually
Virtual Presentation Timezone: UTC-4
Copyright Consent Signature (type Name Or NA If Not Transferrable): Xiaochen Wang
Copyright Consent Name And Address: University of North Carolina at Chapel Hill, Chapel Hill, NC, USA
0 Replies
