Skip to main navigation Skip to search Skip to main content

A systematic mapping study on the research landscape of LLM-based code clone detection

  • Stevens Institute of Technology

Research output: Contribution to journalReview articlepeer-review

Abstract

Context: Code clone refactoring is the practice of identifying and manipulating code duplicates to enhance code quality. Code duplication can pose some difficulties in programming, making code clone refactoring a prominent topic of research. Initial models of code clone detection were introduced using textual analysis in NiCad. Soon after using token-based models, we are now using models with large language models (LLMs). LLMs are AI models trained on large-scale datasets, enabling them to understand programming problems at a high level. LLMs such as GPT, Llama, Starcoder, and Falcon have led new code clone detection ambitions. Because of the complexity and difficulty of many coding problems, LLMs are able to comprehend them and create results at an extremely high level. This has led to the usage of LLMs such as Claude, ChatGPT, and Gemini. Objective: The aim of our research is to provide a comprehensive analysis of existing literature on LLM-based code clone detection practices. Our research will provide a comparison between primary research studies, giving us an overall evaluation of the landscape of clone detection, with the rise of LLM models. Methods: Our research focused on 28 primary studies published within the past two years. These studies were selected through a combination of keyword-based database searches targeting terms related to LLMs, code clone detection, and refactoring. Human validation was used to confirm that each study directly addressed the core themes of LLM-driven software engineering practices. Results: Our analysis of 28 studies shows that LLMs can improve code clone detection, particularly for more complex clones such as Type-3 and Type-4. ChatGPT-4 outperformed earlier versions across multiple benchmarks. Few-shot and context-specific prompting were the most common and effective techniques, appearing in 65% of the studies and contributing to higher precision and recall. However, prompt sensitivity remains a major issue because subtle wording changes can lead to substantial variation in LLM responses. Most studies were limited in the programming languages tested, as Java was used in 83% of the studies and Python in 50%. Type-4 clone detection remains a challenge, with LLMs having difficulty. Conclusion: Future research on code clone detection and refactoring requires specific research directions. Specifically, multilingual testing, prompt sensitivity, and generalization. As research progresses, examining these methods will enable the effective, repeated use of LLM models in coding problems. Existing studies focus on detection rather than downstream refactoring applications, highlighting a gap in actionable tooling. Addressing these challenges will be essential for deploying LLMs in reliable, repeatable software engineering workflows.

Original languageEnglish
Article number108096
JournalInformation and Software Technology
Volume195
DOIs
StatePublished - Jul 2026

Keywords

  • Code clone
  • LLMs
  • Literature review
  • Quality

Fingerprint

Dive into the research topics of 'A systematic mapping study on the research landscape of LLM-based code clone detection'. Together they form a unique fingerprint.

Cite this