(Updated submission 11/20/2020) MISIM: A Novel Code Similarity System

Fangke Ye; Shengtian Zhou; Anand Venkat; Ryan Marcus; Nesime Tatbul; Jesmin Jahan Tithi; Niranjan Hasabnis; Paul Petersen; Timothy G Mattson; Tim Kraska; Pradeep Dubey; Vivek Sarkar; Justin Gottschlich

(Updated submission 11/20/2020) MISIM: A Novel Code Similarity System

Fangke Ye, Shengtian Zhou, Anand Venkat, Ryan Marcus, Nesime Tatbul, Jesmin Jahan Tithi, Niranjan Hasabnis, Paul Petersen, Timothy G Mattson, Tim Kraska, Pradeep Dubey, Vivek Sarkar, Justin Gottschlich

28 Sept 2020 (modified: 05 May 2023)ICLR 2021 Conference Blind SubmissionReaders: Everyone

Keywords: Machine Programming, Machine Learning, Code Similarity, Code Representation

Abstract: Semantic code similarity systems are integral to a range of applications from code recommendation to automated software defect correction. Yet, these systems still lack the maturity in accuracy for general and reliable wide-scale usage. To help address this, we present Machine Inferred Code Similarity (MISIM), a novel end-to-end code similarity system that consists of two core components. First, MISIM uses a novel context-aware semantic structure (CASS), which is designed to aid in lifting semantic meaning from code syntax. We compare CASS with the abstract syntax tree (AST) and show CASS is more accurate than AST by up to 1.67x. Second, MISIM provides a neural-based code similarity scoring algorithm, which can be implemented with various neural network architectures with learned parameters. We compare MISIM to four state-of-the-art systems: (i) Aroma, (ii) code2seq, (iii) code2vec, and (iv) Neural Code Comprehension. In our experimental evaluation across 328,155 programs (over 18 million lines of code), MISIM has 1.5x to 43.4x better accuracy across all four systems.

One-sentence Summary: We present a new state-of-the-art code similarity system that includes a novel code structure and a flexible neural back-end to learn the code similarity algorithm for different code corpi.

Code Of Ethics: I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics

Supplementary Material: zip

Reviewed Version (pdf): https://openreview.net/references/pdf?id=L9-_ngwU1D

10 Replies

Loading