
This report synthesises findings from 9 peer-reviewed papers addressing the following research question: How do cross-domain multi-hop QA benchmarks like HotPotQA reveal robustness gaps in RAG systems when evaluated against adversarial distractor contexts. 11 claims were extracted from source literature; 11 were independently verified against retrieved documents. An automated multi-reviewer quality assessment produced a score of 7.6/10. This report is a machine-generated literature synthesis and does not constitute original research. Research goal: How do cross-domain multi-hop QA benchmarks like HotPotQA reveal robustness gaps in RAG systems when evaluated against adversarial distractor contexts? Autonomous literature synthesis. Automated review score: 7.6/10. Full text and citation available at Assignee Research.
reveal, like, gaps, HotPotQA, benchmarks, robustness, cross-domain, multi-hop
reveal, like, gaps, HotPotQA, benchmarks, robustness, cross-domain, multi-hop
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
