Name: VeriLeaky: Navigating IP Protection vs Utility in Fine-Tuning for LLM-Driven Verilog Coding
Keywords: Hardware Architecture, Machine Learning, FOS: Computer and information sciences, Cryptography and Security, Hardware Architecture (cs.AR), Cryptography and Security (cs.CR), Machine Learning (cs.LG)

descriptionPublicationkeyboard_double_arrow_right Article , Preprint 26 Jun 2025Embargo end date: 01 Jan 2025Publisher:IEEEJournal:2025 IEEE International Conference on LLM-Aided Design (ICLAD)

Authors: Wang, Zeng; Shao, Minghao; Nabeel, Mohammed; Roy, Prithwish Basu; Mankali, Likhitha; Bhandari, Jitendra; Karri, Ramesh; +3 Authors

doi: 10.1109/iclad65226.2025.00018 , 10.48550/arxiv.2503.13116

arXiv: 2503.13116

VeriLeaky: Navigating IP Protection vs Utility in Fine-Tuning for LLM-Driven Verilog Coding

- Summary
- Subjects
- Metrics

Abstract

Large language models (LLMs) offer significant potential for coding, yet fine-tuning (FT) with curated data is essential for niche languages like Verilog. Using proprietary intellectual property (IP) for FT presents a serious risk, as FT data can be leaked through LLM inference. This leads to a critical dilemma for design houses: seeking to build externally accessible LLMs offering competitive Verilog coding, how can they leverage in-house IP to enhance FT utility while ensuring IP protection? For the first time in the literature, we study this dilemma. Using LLaMA 3.1-8B, we conduct in-house FT on a baseline Verilog dataset (RTLCoder) supplemented with our own in-house IP, which is validated through multiple tape-outs. To rigorously assess IP leakage, we quantify structural similarity (AST/Dolos) and functional equivalence (Synopsys Formality) between generated codes and our in-house IP. We show that our IP can indeed be leaked, confirming the threat. As defense, we evaluate logic locking of Verilog codes (ASSURE). This offers some level of protection, yet reduces the IP's utility for FT and degrades the LLM's performance. Our study shows the need for novel strategies that are both effective and minimally disruptive to FT, an essential effort for enabling design houses to fully utilize their proprietary IP toward LLM-driven Verilog coding.

Keywords

Hardware Architecture, Machine Learning, FOS: Computer and information sciences, Cryptography and Security, Hardware Architecture (cs.AR), Cryptography and Security (cs.CR), Machine Learning (cs.LG)

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	0
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Average
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Average
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Average

Found an issue? Give us feedback

Average

Green