Introducing the Djinni recruitment dataset: A corpus of anonymized CVs and job postings

Nazarii Drushchak, Mariana Romanyshyn

Published: 24 May 2024, Last Modified: 27 Sept 2025Third Ukrainian Natural Language Processing Workshop (UNLP)@ LREC-COLING 2024EveryoneRevisionsCC BY 4.0

Abstract: This paper introduces the Djinni Recruitment Dataset, a large-scale open-source corpus of candidate profiles and job descriptions. With over 150,000 jobs and 230,000 candidates, the dataset includes samples in English and Ukrainian, thereby facilitating advancements in the recruitment domain of natural language processing (NLP) for both languages. It is one of the first open-source corpora in the recruitment domain, opening up new opportunities for AI-driven recruitment technologies and related fields. Notably, the dataset is accessible under the MIT license, encouraging widespread adoption for both scientific research and commercial projects.