Optimizing Code Retrieval: High-Quality and Scalable Dataset Annotation through Large Language Models

Optimizing Code Retrieval: High-Quality and Scalable Dataset Annotation through Large Language Models

ACL ARR 2024 June Submission4759 Authors

16 Jun 2024 (modified: 02 Jul 2024)ACL ARR 2024 June SubmissionEveryoneRevisionsBibTeXCC BY 4.0

Abstract: Code retrieval aims to identify code from extensive codebases that semantically aligns with a given query code snippet. Collecting a broad and high-quality set of query and code pairs is crucial to the success of this task. However, existing data collection methods struggle to effectively balance scalability and annotation quality. In this paper, we first analyze the factors influencing the quality of function annotations generated by Large Language Models (LLMs). We find that the invocation of intra-repository functions and third-party APIs plays a significant role. Building on this insight, we propose a novel annotation method that enhances the annotation context by incorporating the content of functions called within the repository and information on third-party API functionalities. Additionally, we integrate LLMs with a novel sorting method to address the multi-level function call relationships within repositories. Furthermore, by applying our proposed method across a range of repositories, we have developed the Query4Code dataset. The quality of this synthesized dataset is validated through both model training and human evaluation, demonstrating high-quality annotations. Moreover, cost analysis confirms the scalability of our annotation method.

Paper Type: Long

Research Area: Information Retrieval and Text Mining

Research Area Keywords: dense retrieval, code generation and understanding

Contribution Types: NLP engineering experiment, Data resources

Languages Studied: English, Code

Submission Number: 4759

Loading