Efficient and Universal Watermarking for LLM-Generated Code Detection

Name: Efficient and Universal Watermarking for LLM-Generated Code Detection
Keywords: FOS: Computer and information sciences, Cryptography and Security, Cryptography and Security (cs.CR)

descriptionPublicationkeyboard_double_arrow_right Article , Preprint 01 Jan 2024Embargo end date: 01 Jan 2024Publisher:arXiv

Authors: Li, Boquan; Fu, Zirui; Zhang, Mengdi; Zhang, Peixin; Sun, Jun; Wang, Xingmei;

doi: 10.48550/arxiv.2402.07518

arXiv: 2402.07518

Efficient and Universal Watermarking for LLM-Generated Code Detection

- Summary
- Subjects
- Related research
  (1)
- Metrics

Abstract

Large language models (LLMs) have significantly enhanced the usability of AI-generated code, providing effective assistance to programmers. This advancement also raises ethical and legal concerns, such as academic dishonesty or the generation of malicious code. For accountability, it is imperative to detect whether a piece of code is AI-generated. Watermarking is broadly considered a promising solution and has been successfully applied to identify LLM-generated text. However, existing efforts on code are far from ideal, suffering from limited universality and excessive time and memory consumption. In this work, we propose a plug-and-play watermarking approach for AI-generated code detection, named ACW (AI Code Watermarking). ACW is training-free and works by selectively applying a set of carefully-designed, semantic-preserving and idempotent code transformations to LLM code outputs. The presence or absence of the transformations serves as implicit watermarks, enabling the detection of AI-generated code. Our experimental results show that ACW effectively detects AI-generated code, preserves code utility, and is resilient against code optimizations. Especially, ACW is efficient and is universal across different LLMs, addressing the limitations of existing approaches.

This work has been submitted to IEEE for possible publication

Keywords

FOS: Computer and information sciences, Cryptography and Security, Cryptography and Security (cs.CR)

1 Research products, page 1 of 1

PCWA software on GitHub
IsRelatedTo

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	0
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Average
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Average
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Average

Found an issue? Give us feedback

Average

Green

Efficient and Universal Watermarking for LLM-Generated Code Detection

Efficient and Universal Watermarking for LLM-Generated Code Detection

1 Research products, page 1 of 1

PCWA software on GitHub