Understanding More about Human and Machine Attention in Deep Neural Networks

Lai, Qiuxia; Khan, Salman; Nie, Yongwei; Shen, Jianbing; Sun, Hanqiu; Shao, Ling

Computer Science > Computer Vision and Pattern Recognition

arXiv:1906.08764 (cs)

[Submitted on 20 Jun 2019 (v1), last revised 6 Jul 2020 (this version, v3)]

Title:Understanding More about Human and Machine Attention in Deep Neural Networks

Authors:Qiuxia Lai, Salman Khan, Yongwei Nie, Jianbing Shen, Hanqiu Sun, Ling Shao

View PDF

Abstract:Human visual system can selectively attend to parts of a scene for quick perception, a biological mechanism known as Human attention. Inspired by this, recent deep learning models encode attention mechanisms to focus on the most task-relevant parts of the input signal for further processing, which is called Machine/Neural/Artificial attention. Understanding the relation between human and machine attention is important for interpreting and designing neural networks. Many works claim that the attention mechanism offers an extra dimension of interpretability by explaining where the neural networks look. However, recent studies demonstrate that artificial attention maps do not always coincide with common intuition. In view of these conflicting evidence, here we make a systematic study on using artificial attention and human attention in neural network design. With three example computer vision tasks, diverse representative backbones, and famous architectures, corresponding real human gaze data, and systematically conducted large-scale quantitative studies, we quantify the consistency between artificial attention and human visual attention and offer novel insights into existing artificial attention mechanisms by giving preliminary answers to several key questions related to human and artificial attention mechanisms. Overall results demonstrate that human attention can benchmark the meaningful `ground-truth' in attention-driven tasks, where the more the artificial attention is close to human attention, the better the performance; for higher-level vision tasks, it is case-by-case. It would be advisable for attention-driven tasks to explicitly force a better alignment between artificial and human attention to boost the performance; such alignment would also improve the network explainability for higher-level computer vision tasks.

Comments:	Q. Lai, S. Khan, Y. Nie, J. Shen, H. Sun, L. Shao, TMM, in press, 2020
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
MSC classes:	62H35
Cite as:	arXiv:1906.08764 [cs.CV]
	(or arXiv:1906.08764v3 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.1906.08764

Submission history

From: Qiuxia Lai [view email]
[v1] Thu, 20 Jun 2019 17:41:57 UTC (3,382 KB)
[v2] Mon, 24 Jun 2019 06:17:11 UTC (3,382 KB)
[v3] Mon, 6 Jul 2020 05:13:58 UTC (3,436 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Understanding More about Human and Machine Attention in Deep Neural Networks

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Understanding More about Human and Machine Attention in Deep Neural Networks

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators