Advancing Multi-Criteria Chinese Word Segmentation Through Criterion Classification and DenoisingDownload PDF

Anonymous

16 Dec 2022 (modified: 05 May 2023)ACL ARR 2022 December Blind SubmissionReaders: Everyone
Abstract: Recent research on multi-criteria Chinese word segmentation (MCCWS) mainly focuses on building complex private structures, adding more handcrafted features, or introducing complex optimization processes.In this work, we show that through a simple yet elegant input-hint-based MCCWS model, we can achieve state-of-the-art (SoTA) performances on several datasets simultaneously.We further propose a novel criterion-denoising objective that hurts slightly on F1 score but acheives SoTA recall on out-of-vocabulary words.Our result establishes a simple yet strong baseline for future MCCWS research.
Paper Type: long
Research Area: Syntax: Tagging, Chunking and Parsing / ML
0 Replies

Loading