
A copyright-safe Korean long-context benchmark for evaluating long-context LLMs, RAG systems, and table/tool pipelines over real housing announcements, public tabular data, and housing statutes. The public release contains QA labels, evidence locators, deterministic predicates, answerability labels, split and provider/region metadata, and long-context-bundle references. It does not redistribute raw PDF/HWP/HWPX documents, bundle text, API keys, or hidden gold answers. v0.8 is a human-review repair build (1,997 QA): all positional-cloze probes were regenerated into natural source-grounded questions or removed, and location/answer errors were fixed (LLM-assisted gpt-5.4 + Claude cross-model review).
table reasoning, benchmark, question answering, long-context, Korean, large language models, retrieval-augmented-generation, housing
table reasoning, benchmark, question answering, long-context, Korean, large language models, retrieval-augmented-generation, housing
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
