Multi-modal dialog for browsing large visual catalogs using exploration-exploitation paradigm in a joint embedding space

Bhattacharya, Indrani; Chowdhury, Arkabandhu; Raykar, Vikas

Computer Science > Computer Vision and Pattern Recognition

arXiv:1901.09854 (cs)

[Submitted on 28 Jan 2019 (v1), last revised 29 Jan 2019 (this version, v2)]

Title:Multi-modal dialog for browsing large visual catalogs using exploration-exploitation paradigm in a joint embedding space

Authors:Indrani Bhattacharya, Arkabandhu Chowdhury, Vikas Raykar

View PDF

Abstract:We present a multi-modal dialog system to assist online shoppers in visually browsing through large catalogs. Visual browsing is different from visual search in that it allows the user to explore the wide range of products in a catalog, beyond the exact search matches. We focus on a slightly asymmetric version of the complete multi-modal dialog where the system can understand both text and image queries but responds only in images. We formulate our problem of "showing $k$ best images to a user" based on the dialog context so far, as sampling from a Gaussian Mixture Model in a high dimensional joint multi-modal embedding space, that embed both the text and the image queries. Our system remembers the context of the dialog and uses an exploration-exploitation paradigm to assist in visual browsing. We train and evaluate the system on a multi-modal dialog dataset that we generate from large catalog data. Our experiments are promising and show that the agent is capable of learning and can display relevant results with an average cosine similarity of 0.85 to the ground truth. Our preliminary human evaluation also corroborates the fact that such a multi-modal dialog system for visual browsing is well-received and is capable of engaging human users.

Comments:	10 pages including reference, 8 figures. First two authors are equal contributors
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:1901.09854 [cs.CV]
	(or arXiv:1901.09854v2 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.1901.09854

Submission history

From: Arkabandhu Chowdhury [view email]
[v1] Mon, 28 Jan 2019 17:49:55 UTC (3,261 KB)
[v2] Tue, 29 Jan 2019 21:13:03 UTC (3,260 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Multi-modal dialog for browsing large visual catalogs using exploration-exploitation paradigm in a joint embedding space

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Multi-modal dialog for browsing large visual catalogs using exploration-exploitation paradigm in a joint embedding space

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators