Powered by OpenAIRE graph
Found an issue? Give us feedback
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/ ZENODOarrow_drop_down
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
ZENODO
Software . 2026
Data sources: ZENODO
ZENODO
Software . 2026
Data sources: Datacite
ZENODO
Software . 2026
Data sources: Datacite
versions View all 2 versions
addClaim

OWLer Crawler

Authors: Dinzinger, Michael; Zerhoudi, Saber; Granitzer, Michael;

OWLer Crawler

Abstract

The Open Web Crawler (OWLer) is a distributed web crawling system developed within the OpenWebSearch.eu initiative, contributing to the construction of an open, transparent, and publicly accessible Web Index. This Zenodo release contains the source code of the OWLer Crawler and its associated file transfer service, published as two complementary components: owler-crawler: The main crawling engine responsible for fetching, parsing, and storing web content. owler-crawler-filetransfer: A supporting service responsible for exchanging URL files and crawl log files with the URL Frontier. owler-crawler The OWLer Crawler is designed to be polite, efficient, and horizontally scalable. It respects robots.txt directives, supports configurable crawl delays, and enforces host-based politeness policies. The crawler retrieves web content, parses HTML pages to extract metadata and outlinks, and stores results in structured log files and WARC format for archival purposes . Its modular architecture consists of the following core components: Reader / WARCReader: Ingest URL lists or WARC records into the crawling topology. Fetcher: Retrieves web pages via HTTP/HTTPS with configurable limits and timeouts. Parser: Extracts metadata and outgoing links from HTML content. Writer / WARCWriter: Persists crawl results as structured log files and WARC archives. The crawler can be deployed as a standalone Java application or as a containerized service. Horizontal scaling is achieved by running multiple crawler instances connected to the same URL Frontier service . owler-crawler-filetransfer The owler-crawler-filetransfer module provides controlled file-based communication between crawler instances and the URL Frontier. It is responsible for: Retrieving URL batches from the URL Frontier, Delivering discovered URLs and crawl logs back to the Frontier, Managing input/output directory synchronization. This module depends on the separate Maven artifact owler-urlfrontier-api, which defines the gRPC interface used for communication with the URL Frontier service. By separating the API definition into its own repository, the architecture ensures modularity, clearer dependency management, and reuse across services. System Context Within the OWLer architecture, the crawler operates as a distributed worker component connected to a centralized URL Frontier. The URL Frontier manages crawl state and scheduling, while crawler instances perform fetching and content processing. This separation of concerns allows scalable and robust crawling in distributed environments. This release corresponds to a stable development snapshot used within the OpenWebSearch.eu project. It supports transparency, reproducibility, and long-term archival of the crawling components contributing to the Open Web Index infrastructure.

Related Organizations
  • BIP!
    Impact byBIP!
    selected citations
    These citations are derived from selected sources.
    This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    0
    popularity
    This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
    Average
    influence
    This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    Average
    impulse
    This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
    Average
Powered by OpenAIRE graph
Found an issue? Give us feedback
selected citations
These citations are derived from selected sources.
This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Citations provided by BIP!
popularity
This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
BIP!Popularity provided by BIP!
influence
This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Influence provided by BIP!
impulse
This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
BIP!Impulse provided by BIP!
0
Average
Average
Average