Audio-CoT: Exploring Chain-of-Thought Reasoning in Large Audio Language Model

Ma, Ziyang; Chen, Zhuo; Wang, Yuping; Chng, Eng Siong; Chen, Xie

Computer Science > Sound

arXiv:2501.07246 (cs)

[Submitted on 13 Jan 2025]

Title:Audio-CoT: Exploring Chain-of-Thought Reasoning in Large Audio Language Model

Authors:Ziyang Ma, Zhuo Chen, Yuping Wang, Eng Siong Chng, Xie Chen

View PDF HTML (experimental)

Abstract:Large Audio-Language Models (LALMs) have demonstrated remarkable performance in tasks involving audio perception and understanding, such as speech recognition and audio captioning. However, their reasoning capabilities - critical for solving complex real-world problems - remain underexplored. In this work, we conduct the first exploration into integrating Chain-of-Thought (CoT) reasoning into LALMs to enhance their reasoning ability across auditory modalities. We evaluate representative CoT methods, analyzing their performance in both information extraction and reasoning tasks across sound, music, and speech domains. Our findings reveal that CoT methods significantly improve performance on easy and medium tasks but encounter challenges with hard tasks, where reasoning chains can confuse the model rather than improve accuracy. Additionally, we identify a positive correlation between reasoning path length and accuracy, demonstrating the potential of scaling inference for advanced instruction-following and reasoning. This study not only highlights the promise of CoT in enhancing LALM reasoning capabilities but also identifies key limitations and provides actionable directions for future research.

Subjects:	Sound (cs.SD); Computation and Language (cs.CL); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
Cite as:	arXiv:2501.07246 [cs.SD]
	(or arXiv:2501.07246v1 [cs.SD] for this version)
	https://doi.org/10.48550/arXiv.2501.07246

Submission history

From: Ziyang Ma [view email]
[v1] Mon, 13 Jan 2025 11:54:40 UTC (113 KB)

Computer Science > Sound

Title:Audio-CoT: Exploring Chain-of-Thought Reasoning in Large Audio Language Model

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Sound

Title:Audio-CoT: Exploring Chain-of-Thought Reasoning in Large Audio Language Model

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators