
The massive volumes of data generated in diverse domains, suchas scientific computing, finance and environmental monitoring,hinder our ability to perform multidimensional analysis at highspeeds and also yield significant storage and egress costs. Applying compression algorithms to reduce these costs is particularlysuitable for column-oriented DBMSs, as the values of individualcolumns are usually similar and thus, allow for effective compression. However, this has not been the case for binary floating point numbers, as the space savings achieved by respectivecompression algorithms are usually very modest. We presenthere two lossless compression algorithms for floating point data,termed Chimp and Patas, that attain impressive compression ratios and greatly outperform state-of-the-art approaches. We focuson how these two algorithms impact the performance of DuckDB,a purpose-built embeddable database for interactive analytics.Our demonstration will showcase how our novel compressionapproaches a) reduce storage requirements, and b) improve thetime needed to load and query data using DuckDB.
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
