Powered by OpenAIRE graph
Found an issue? Give us feedback
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/ ZENODOarrow_drop_down
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
ZENODO
Other literature type . 2024
License: CC BY
Data sources: ZENODO
ZENODO
Presentation . 2024
License: CC BY
Data sources: Datacite
ZENODO
Presentation . 2024
License: CC BY
Data sources: Datacite
versions View all 2 versions
addClaim

VirtualiZarr and DMR++

Authors: Nag, Ayush; Gallagher, James;

VirtualiZarr and DMR++

Abstract

The Challenge of Big Data: Scientists working with massive datasets, like those from the SWOT satellite, face a significant hurdle: processing speed. With tens of thousands of individual data files, traditional methods of reading, combining, and analyzing data are simply too slow. This bottleneck hampers research and innovation. A New Approach: To tackle this problem, researchers have developed a new strategy involving a combination of technologies. At the core is DMR++, a system that efficiently stores metadata about data chunks. This metadata is then used by VirtualiZarr to create virtual Zarr datasets, which offer a more efficient way to access and manipulate data. To streamline the process, the team has integrated VirtualiZarr with earthaccess, a tool that quickly finds and accesses data. For even faster processing, they've incorporated dask, which allows for parallel computing. Together, these technologies create a powerful pipeline for handling vast amounts of data. The Benefits: This new approach promises several advantages. By processing data in parallel and using optimized metadata, scientists can dramatically reduce the time it takes to analyze data. Additionally, the ability to create virtual datasets without duplicating data saves storage space and computational resources. Looking Ahead: While this solution is already showing promise, there's still work to be done. Improving the accessibility of DMR++ and optimizing its structure for performance are key priorities. The team is also working to expand compatibility with different data formats and to finalize the VirtualiZarr API and specification. The ultimate goal is to create a standardized system that can be used by a wide range of researchers, accelerating scientific discovery. By addressing the challenges of big data, this innovative approach has the potential to revolutionize how scientists work with massive datasets.

Related Organizations
Keywords

Web accessibility, Big data, Data exchange

  • BIP!
    Impact byBIP!
    selected citations
    These citations are derived from selected sources.
    This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    0
    popularity
    This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
    Average
    influence
    This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    Average
    impulse
    This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
    Average
Powered by OpenAIRE graph
Found an issue? Give us feedback
selected citations
These citations are derived from selected sources.
This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Citations provided by BIP!
popularity
This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
BIP!Popularity provided by BIP!
influence
This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Influence provided by BIP!
impulse
This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
BIP!Impulse provided by BIP!
0
Average
Average
Average
Green