The Deployment of End-to-End Audio Language Models Should Take into Account the Principle of Least Privilege

He, Luxi; Qi, Xiangyu; Liao, Michel; Cheong, Inyoung; Mittal, Prateek; Chen, Danqi; Henderson, Peter

Computer Science > Sound

arXiv:2503.16833 (cs)

[Submitted on 21 Mar 2025]

Title:The Deployment of End-to-End Audio Language Models Should Take into Account the Principle of Least Privilege

Authors:Luxi He, Xiangyu Qi, Michel Liao, Inyoung Cheong, Prateek Mittal, Danqi Chen, Peter Henderson

View PDF HTML (experimental)

Abstract:We are at a turning point for language models that accept audio input. The latest end-to-end audio language models (Audio LMs) process speech directly instead of relying on a separate transcription step. This shift preserves detailed information, such as intonation or the presence of multiple speakers, that would otherwise be lost in transcription. However, it also introduces new safety risks, including the potential misuse of speaker identity cues and other sensitive vocal attributes, which could have legal implications. In this position paper, we urge a closer examination of how these models are built and deployed. We argue that the principle of least privilege should guide decisions on whether to deploy cascaded or end-to-end models. Specifically, evaluations should assess (1) whether end-to-end modeling is necessary for a given application; and (2), the appropriate scope of information access. Finally, We highlight related gaps in current audio LM benchmarks and identify key open research questions, both technical and policy-related, that must be addressed to enable the responsible deployment of end-to-end Audio LMs.

Subjects:	Sound (cs.SD); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY); Audio and Speech Processing (eess.AS)
Cite as:	arXiv:2503.16833 [cs.SD]
	(or arXiv:2503.16833v1 [cs.SD] for this version)
	https://doi.org/10.48550/arXiv.2503.16833

Submission history

From: Luxi He [view email]
[v1] Fri, 21 Mar 2025 04:03:59 UTC (994 KB)

Computer Science > Sound

Title:The Deployment of End-to-End Audio Language Models Should Take into Account the Principle of Least Privilege

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Sound

Title:The Deployment of End-to-End Audio Language Models Should Take into Account the Principle of Least Privilege

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators