Compressed Index for Dictionary Matching

descriptionPublicationkeyboard_double_arrow_right Article , Conference object 01 Mar 2008 China (People's Republic of) Publisher:IEEEJournal:Data Compression Conference (dcc 2008)

Authors: Wing-Kai Hon; Tak Wah Lam; Rahul Shah 0001; Siu-Lung Tam; Jeffrey Scott Vitter;

doi: 10.1109/dcc.2008.62

handle: 10722/57233

Compressed Index for Dictionary Matching

- Summary
- Subjects
- Metrics

Abstract

The past few years have witnessed several exciting results on compressed representation of a string T that supports efficient pattern matching, and the space complexity has been reduced to |T| Hk (T) + o (|T| log sigma) bits, where Hk(T) denotes the kth-order empirical entropy of T, and sigma is the size of the alphabet. In this paper we study compressed representation for another classical problem of string indexing, which is called dictionary matching in the literature. Precisely, a collection D of strings (called patterns) of total length n is to be indexed so that given a text T, the occurrences of the patterns in T can be found efficiently. In this paper we show how to exploit a sampling technique to compress the existing O(n)-word index to an (n Hk (D) + o(n log sigma))-bit index with only a small sacrifice in search time.

Country

China (People's Republic of)

Related Organizations

University of Hong Kong (香港大學)
China (People's Republic of)
University of Hong Kong
China (People's Republic of)
Purdue University System
United States
National Tsing Hua University
Taiwan
Purdue University West Lafayette
United States

View all View all

Keywords

Computers, Information science and information theory and abstracting and indexing services

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	18
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Average
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Top 10%
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Top 10%