Downloads provided by UsageCounts
arXiv: 2010.07754
handle: 11573/1487082
Reliable identification of encrypted file fragments is a requirement for several security applications, including ransomware detection, digital forensics, and traffic analysis. A popular approach consists of estimating high entropy as a proxy for randomness. However, many modern content types (e.g. office documents, media files, etc.) are highly compressed for storage and transmission efficiency. Compression algorithms also output high-entropy data, thus reducing the accuracy of entropy-based encryption detectors. Over the years, a variety of approaches have been proposed to distinguish encrypted file fragments from high-entropy compressed fragments. However, these approaches are typically only evaluated over a few, select data types and fragment sizes, which makes a fair assessment of their practical applicability impossible. This paper aims to close this gap by comparing existing statistical tests on a large, standardized dataset. Our results show that current approaches cannot reliably tell apart encryption and compression, even for large fragment sizes. To address this issue, we design EnCoD, a learning-based classifier which can reliably distinguish compressed and encrypted data, starting with fragments as small as 512 bytes. We evaluate EnCoD against current approaches over a large dataset of different data types, showing that it outperforms current state-of-the-art for most considered fragment sizes and data types.
19 pages, 6 images, 2 tables. Accepted for publication at the 14th International Conference on Network and System Security (NSS2020)
FOS: Computer and information sciences, Computer Science - Machine Learning, Computer Science - Cryptography and Security, Machine Learning, Security, Machine Learning (cs.LG), Cryptography and Security (cs.CR), MAG: Traffic analysis, MAG: Computer science, MAG: Digital forensics, MAG: Encryption, MAG: Identification (information), MAG: Ransomware, MAG: Entropy (information theory), MAG: Data mining, MAG: Randomness, MAG: Data compression
FOS: Computer and information sciences, Computer Science - Machine Learning, Computer Science - Cryptography and Security, Machine Learning, Security, Machine Learning (cs.LG), Cryptography and Security (cs.CR), MAG: Traffic analysis, MAG: Computer science, MAG: Digital forensics, MAG: Encryption, MAG: Identification (information), MAG: Ransomware, MAG: Entropy (information theory), MAG: Data mining, MAG: Randomness, MAG: Data compression
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 18 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Top 10% | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Top 10% | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Top 10% |
| views | 5 | |
| downloads | 13 |

Views provided by UsageCounts
Downloads provided by UsageCounts