Visual Storytelling

Preprint English OPEN
Ting-Hao; Huang; Ferraro, Francis; Mostafazadeh, Nasrin; Misra, Ishan; Agrawal, Aishwarya; Devlin, Jacob; Girshick, Ross; He, Xiaodong; Kohli, Pushmeet; Batra, Dhruv; Zitnick, C. Lawrence; Parikh, Devi; Vanderwende, Lucy; Galley, Michel; Mitchell, Margaret;
  • Subject: Computer Science - Computation and Language | Computer Science - Computer Vision and Pattern Recognition | Computer Science - Artificial Intelligence

We introduce the first dataset for sequential vision-to-language, and explore how this data may be used for the task of visual storytelling. The first release of this dataset, SIND v.1, includes 81,743 unique photos in 20,211 sequences, aligned to both descriptive (capt... View more
Share - Bookmark