Research Area: Data, Engineering for large LMs
Keywords: syntax, word order, postposition, case marker, syntactically-incomplete data
TL;DR: We confirmed the syntactic flexibility of Korean, examined if LLMs capture this feature, and applied it as a data augmentation method to enhance performance.
Abstract: Syntactic elements, such as word order and case markers, are fundamental in natural language processing. Recent studies show that syntactic information boosts language model performance and offers clues for people to understand their learning mechanisms. Unlike languages with a fixed word order such as English, Korean allows for varied word sequences, despite its canonical structure, due to case markers that indicate the functions of sentence components. This study explores whether Korean language models can accurately capture this flexibility. We note that incomplete word orders and omitted case markers frequently appear in ordinary Korean communication. To investigate this further, we introduce the Syntactically Incomplete Korean (SIKO) dataset. Through SIKO, we assessed Korean language models’ flexibility with incomplete syntax and confirmed the dataset’s training value. Results indicate these models reflect Korean’s inherent flexibility, accurately handling incomplete inputs. Moreover, fine-tuning with SIKO enhances the ability to handle common incomplete Korean syntactic forms. The dataset’s simple construction process, coupled with significant performance enhancements, solidifies its standing as an effective data augmentation technique. The SIKO will become accessible post-publication.
Code Of Ethics: I acknowledge that I and all co-authors of this work have read and commit to adhering to the COLM Code of Ethics on https://colmweb.org/CoE.html
Author Guide: I certify that this submission complies with the submission instructions as described on https://colmweb.org/AuthorGuide.html
Submission Number: 298
Loading