Abstract: This study introduces a pretrained large language model-based annotation methodology
for the first dependency treebank in Ottoman
Turkish. Our experimental results show that, iteratively, i) pseudo-annotating data using a multilingual BERT-based parsing model, ii) manually correcting the pseudo-annotations, and
iii) fine-tuning the parsing model with the corrected annotations, we speed up and simplify
the challenging dependency annotation process.
The resulting treebank, that will be a part of the
Universal Dependencies (UD) project, will facilitate automated analysis of Ottoman Turkish
documents, unlocking the linguistic richness
embedded in this historical heritage.
Loading