De-Anonymizing Text by Fingerprinting Language Generation

Sun, Zhen; Schuster, Roei; Shmatikov, Vitaly

Computer Science > Cryptography and Security

arXiv:2006.09615 (cs)

[Submitted on 17 Jun 2020 (v1), last revised 3 Nov 2020 (this version, v2)]

Title:De-Anonymizing Text by Fingerprinting Language Generation

Authors:Zhen Sun, Roei Schuster, Vitaly Shmatikov

View PDF

Abstract:Components of machine learning systems are not (yet) perceived as security hotspots. Secure coding practices, such as ensuring that no execution paths depend on confidential inputs, have not yet been adopted by ML developers. We initiate the study of code security of ML systems by investigating how nucleus sampling---a popular approach for generating text, used for applications such as auto-completion---unwittingly leaks texts typed by users. Our main result is that the series of nucleus sizes for many natural English word sequences is a unique fingerprint. We then show how an attacker can infer typed text by measuring these fingerprints via a suitable side channel (e.g., cache access times), explain how this attack could help de-anonymize anonymous texts, and discuss defenses.

Comments:	NeurIPS 2020
Subjects:	Cryptography and Security (cs.CR); Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as:	arXiv:2006.09615 [cs.CR]
	(or arXiv:2006.09615v2 [cs.CR] for this version)
	https://doi.org/10.48550/arXiv.2006.09615

Submission history

From: Zhen Sun [view email]
[v1] Wed, 17 Jun 2020 02:49:15 UTC (2,024 KB)
[v2] Tue, 3 Nov 2020 04:47:25 UTC (1,202 KB)

Full-text links:

Access Paper:

view license

Current browse context:

cs.CR

< prev | next >

new | recent | 2020-06

Change to browse by:

cs
cs.CL
cs.LG

References & Citations

DBLP - CS Bibliography

listing | bibtex

Roei Schuster
Vitaly Shmatikov

export BibTeX citation

Computer Science > Cryptography and Security

Title:De-Anonymizing Text by Fingerprinting Language Generation

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Cryptography and Security

Title:De-Anonymizing Text by Fingerprinting Language Generation

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators