A Free Lunch in Generating Datasets: Building a VQG and VQA System with Attention and Humans in the Loop

Lee, Jihyeon; Arora, Sho

Computer Science > Computer Vision and Pattern Recognition

arXiv:1912.00124 (cs)

[Submitted on 30 Nov 2019 (v1), last revised 28 Aug 2020 (this version, v2)]

Title:A Free Lunch in Generating Datasets: Building a VQG and VQA System with Attention and Humans in the Loop

Authors:Jihyeon Lee, Sho Arora

View PDF

Abstract:Despite their importance in training artificial intelligence systems, large datasets remain challenging to acquire. For example, the ImageNet dataset required fourteen million labels of basic human knowledge, such as whether an image contains a chair. Unfortunately, this knowledge is so simple that it is tedious for human annotators but also tacit enough such that they are necessary. However, human collaborative efforts for tasks like labeling massive amounts of data are costly, inconsistent, and prone to failure, and this method does not resolve the issue of the resulting dataset being static in nature. What if we asked people questions they want to answer and collected their responses as data? This would mean we could gather data at a much lower cost, and expanding a dataset would simply become a matter of asking more questions. We focus on the task of Visual Question Answering (VQA) and propose a system that uses Visual Question Generation (VQG) to produce questions, asks them to social media users, and collects their responses. We present two models that can then parse clean answers from the noisy human responses significantly better than our baselines, with the goal of eventually incorporating the answers into a Visual Question Answering (VQA) dataset. By demonstrating how our system can collect large amounts of data at little to no cost, we envision similar systems being used to improve performance on other tasks in the future.

Comments:	9 pages, 12 figures
Subjects:	Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Human-Computer Interaction (cs.HC); Machine Learning (cs.LG)
MSC classes:	I.2, I.4, I.6, I.7
ACM classes:	I.2; I.4; I.6; I.7
Cite as:	arXiv:1912.00124 [cs.CV]
	(or arXiv:1912.00124v2 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.1912.00124

Submission history

From: Jihyeon Lee [view email]
[v1] Sat, 30 Nov 2019 03:45:17 UTC (1,795 KB)
[v2] Fri, 28 Aug 2020 17:52:03 UTC (1,795 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:A Free Lunch in Generating Datasets: Building a VQG and VQA System with Attention and Humans in the Loop

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:A Free Lunch in Generating Datasets: Building a VQG and VQA System with Attention and Humans in the Loop

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators