Learning to Answer Visual Questions From Web Videos

Name: Learning to Answer Visual Questions From Web Videos
Keywords: Question Generation, [INFO.INFO-AI] Computer Science [cs]/Artificial Intelligence [cs.AI], FOS: Computer and information sciences, Computer Science - Machine Learning, Computer Science - Computation and Language, Computer Vision and Pattern Recognition (cs.CV), Computer Science - Computer Vision and Pattern Recognition, Cross-Modal Supervision, Video Question Answering, [INFO.INFO-LG] Computer Science [cs]/Machine Learning [cs.LG]

Antoine Yang; Antoine Miech; Josef Sivic; Ivan Laptev; Cordelia Schmid

Found an issue? Give us feedback

arXiv.org e-Print Ar...arrow_drop_down

arXiv.org e-Print Archive

Preprint . 2022

Data sources: arXiv.org e-Print Archive

INRIA2

Article . 2022

Data sources: INRIA2

INRIA a CCSD electronic archive server

Article . 2022

Data sources: INRIA a CCSD electronic archive server

IEEE Transactions on Pattern Analysis and Machine Intelligence

Article . 2025 . Peer-reviewed

License: IEEE Copyright

Data sources: Crossref

https://dx.doi.org/10.48550/ar...

Article . 2022

License: arXiv Non-Exclusive Distribution

Data sources: Datacite

https://pubmed.ncbi.nlm.nih.go...

Article

Data sources: Europe PubMed Central

Learning to Answer Visual Questions From Web Videos

descriptionPublicationkeyboard_double_arrow_right Article , Preprint 01 May 2025Embargo end date: 01 Jan 2022Publisher:Institute of Electrical and Electronics Engineers (IEEE)Journal:IEEE Transactions on Pattern Analysis and Machine Intelligence, volume 47, pages 3,202-3,218 (issn: 0162-8828, eissn: 1939-3539,

Copyright policy )Funded by:ANR | PRAIRIE

Authors: Antoine Yang; Antoine Miech; Josef Sivic; Ivan Laptev; Cordelia Schmid;

doi: 10.1109/tpami.2022.3173208 , 10.48550/arxiv.2205.05019

pmid: 35533174

arXiv: 2205.05019

Learning to Answer Visual Questions From Web Videos

- Summary
- Subjects
- Related research
  (5)
- Metrics

Abstract

Recent methods for visual question answering rely on large-scale annotated datasets. Manual annotation of questions and answers for videos, however, is tedious, expensive and prevents scalability. In this work, we propose to avoid manual annotation and generate a large-scale training dataset for video question answering making use of automatic cross-modal supervision. We leverage a question generation transformer trained on text data and use it to generate question-answer pairs from transcribed video narrations. Given narrated videos, we then automatically generate the HowToVQA69M dataset with 69M video-question-answer triplets. To handle the open vocabulary of diverse answers in this dataset, we propose a training procedure based on a contrastive loss between a video-question multi-modal transformer and an answer transformer. We introduce the zero-shot VideoQA task and the VideoQA feature probe evaluation setting and show excellent results, in particular for rare answers. Furthermore, our method achieves competitive results on MSRVTT-QA, ActivityNet-QA, MSVD-QA and How2QA datasets. We also show that our VideoQA dataset generation approach generalizes to another source of web video and text data. We use our method to generate the WebVidVQA3M dataset from the WebVid dataset, i.e., videos with alt-text annotations, and show its benefits for training VideoQA models. Finally, for a detailed evaluation we introduce iVQA, a new VideoQA dataset with reduced language bias and high-quality manual annotations. Code, datasets and trained models are available at https://antoyang.github.io/just-ask.html

Accepted at the TPAMI Special Issue on the Best Papers of ICCV 2021. Journal extension of the conference paper arXiv:2012.00451. 16 pages, 13 figures

Related Organizations

French National Centre for Scientific Research
France
PSL Research University
France
French Institute for Research in Computer Science and Automation
France
Czech Technical University in Prague
Czech Republic
Centre Inria de Paris
France

Keywords

Question Generation, [INFO.INFO-AI] Computer Science [cs]/Artificial Intelligence [cs.AI], FOS: Computer and information sciences, Computer Science - Machine Learning, Computer Science - Computation and Language, Computer Vision and Pattern Recognition (cs.CV), Computer Science - Computer Vision and Pattern Recognition, Cross-Modal Supervision, Video Question Answering, [INFO.INFO-LG] Computer Science [cs]/Machine Learning [cs.LG], [INFO] Computer Science [cs], Machine Learning (cs.LG), [INFO.INFO-CV] Computer Science [cs]/Computer Vision and Pattern Recognition [cs.CV], [INFO.INFO-CL] Computer Science [cs]/Computation and Language [cs.CL], Zero-Shot Learning, Computation and Language (cs.CL)

5 Research products, page 1 of 1

Just Ask: Learning to Answer Questions from Millions of Narrated Videos
2021IsAmongTopNSimilarDocuments
Look Before you Speak: Visually Contextualized Utterances
2021IsAmongTopNSimilarDocuments
Long-Term Video Question Answering via Multimodal Hierarchical Memory Attentive Networks
2021IsAmongTopNSimilarDocuments
Zero-Shot Video Question Answering via Frozen Bidirectional Language Models
2022IsAmongTopNSimilarDocuments
punctuator2 software on GitHub
IsRelatedTo

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	9
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Top 10%
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Average
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Top 10%