Powered by OpenAIRE graph
Found an issue? Give us feedback
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/ ZENODOarrow_drop_down
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
ZENODO
Report
Data sources: ZENODO
addClaim

Comparative Analysis of Semantic File Routing and Chunk-Based Retrieval on Homogeneous Financial Documents Using MSR-VTT

Authors: SOVEREIGN Research Kernel;

Comparative Analysis of Semantic File Routing and Chunk-Based Retrieval on Homogeneous Financial Documents Using MSR-VTT

Abstract

Retrieval-Augmented Generation (RAG) systems for financial document question answering typically follow a chunk-based paradigm: documents are split into fragments, embedded into vector space, and retrieved via similarity search. While effective in general settings, this approach suffers from cross-document chunk confusion in structurally homogeneous corpora such as regulatory filings. Semantic File Routing (SFR), which uses LLM structured output to route queries to whole documents, reduces catastrophic failures but sacrifices the precision of targeted chunk retrieval. We identify this robustneResearch goal: How does the Semantic File Routing (SFR) method compare to traditional chunk-based retrieval in terms of precision and recall on the MSR-VTT benchmark when applied to structurally homogeneous financial documents?Autonomous synthesis report generated by SOVEREIGN Research Kernel. Tribunal consensus score: 8.9/10.

Powered by OpenAIRE graph
Found an issue? Give us feedback