
DisasterVQA is a benchmark dataset for evaluating Vision-Language Models (VLMs) on disaster-response visual question answering. It includes Binary, Multiple-Choice, and Open-Ended questions, and contains 1,395 real-world disaster images and 4,405 expert-curated question–answer pairs covering floods, wildfires, and earthquakes. Questions span situational awareness and operational decision-making tasks, grounded in humanitarian frameworks (FEMA ESF, OCHA MIRA). We benchmark seven state-of-the-art vision–language models, revealing performance gaps in fine-grained quantitative reasoning, object counting, and context-sensitive interpretation — especially for underrepresented disaster scenarios. Files in this release: disastervqa_annotations.jsonl: the benchmark annotations and metadata (question text, ground-truth answers, image paths, and taxonomy labels). disastervqa_model_outputs.jsonl: model predictions for each question (join with annotations using question_id). For Open-Ended questions, some records include a judge-LLM decision label (Right/Wrong). taxonomy.json: final taxonomy definitions and references for each crisis_info_code. Paper:Please cite the accompanying paper: “DisasterVQA: A Visual Question Answering Benchmark Dataset for Disaster Scenes”, arXiv:2601.13839.
disaster response, benchmark, humanitarian, DisasterVQA, crisis informatics, VQA, VLM, Damage assessment, vision-language models
disaster response, benchmark, humanitarian, DisasterVQA, crisis informatics, VQA, VLM, Damage assessment, vision-language models
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
