Name: Revisiting Deep Learning for Variable Type Recovery
Keywords: FOS: Computer and information sciences, Computer Science - Machine Learning, Machine Learning (cs.LG)

descriptionPublicationkeyboard_double_arrow_right Article , Preprint 01 May 2023Embargo end date: 01 Jan 2023Publisher:IEEEJournal:2023 IEEE/ACM 31st International Conference on Program Comprehension (ICPC)

Authors: Cao, Kevin; Leach, Kevin;

doi: 10.1109/icpc58990.2023.00042 , 10.48550/arxiv.2304.03854

arXiv: 2304.03854

Revisiting Deep Learning for Variable Type Recovery

- Summary
- Subjects
- Related research
  (4)
- Metrics

Abstract

Compiled binary executables are often the only available artifact in reverse engineering, malware analysis, and software systems maintenance. Unfortunately, the lack of semantic information like variable types makes comprehending binaries difficult. In efforts to improve the comprehensibility of binaries, researchers have recently used machine learning techniques to predict semantic information contained in the original source code. Chen et al. implemented DIRTY, a Transformer-based Encoder-Decoder architecture capable of augmenting decompiled code with variable names and types by leveraging decompiler output tokens and variable size information. Chen et al. were able to demonstrate a substantial increase in name and type extraction accuracy on Hex-Rays decompiler outputs compared to existing static analysis and AI-based techniques. We extend the original DIRTY results by re-training the DIRTY model on a dataset produced by the open-source Ghidra decompiler. Although Chen et al. concluded that Ghidra was not a suitable decompiler candidate due to its difficulty in parsing and incorporating DWARF symbols during analysis, we demonstrate that straightforward parsing of variable data generated by Ghidra results in similar retyping performance. We hope this work inspires further interest and adoption of the Ghidra decompiler for use in research projects.

In The 31st International Conference on Program Comprehension(ICPC 2023 RENE)

Related Organizations

Vanderbilt University
United States

Keywords

FOS: Computer and information sciences, Computer Science - Machine Learning, Machine Learning (cs.LG)

4 Research products, page 1 of 1

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	0
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Average
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Average
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Average

Found an issue? Give us feedback

Average

Green

Revisiting Deep Learning for Variable Type Recovery

Revisiting Deep Learning for Variable Type Recovery

4 Research products, page 1 of 1

remill software on GitHub

ghidra software on GitHub

ghcc software on GitHub

radare2 software on GitHub