
This paper presents a small-scale empirical evaluation of hallucination and citation accuracy in retrieval-augmented generation (RAG) question answering. Using a 50-question sample derived from HotpotQA, the study compares a gold-evidence condition, where the model receives the benchmark supporting facts, with a vector-retrieval condition, where evidence is retrieved automatically using sentence-level embeddings and cosine similarity. The evaluation measures answer accuracy, citation accuracy, hallucination rate, and retrieval completeness across multiple large language models under a strict citation prompt. Results show that while vector top-8 retrieval finds at least one gold supporting sentence for every question, it retrieves the complete evidence chain for only 52% of questions. This partial evidence gap is associated with lower answer accuracy, weaker citation grounding, and increased hallucination. The study highlights the importance of evaluating retrieval completeness and citation correctness separately from answer accuracy in multi-hop RAG systems.
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
