Typing to Listen at the Cocktail Party: Text-Guided Target Speaker Extraction

Hao, Xiang; Wu, Jibin; Yu, Jianwei; Xu, Chenglin; Tan, Kay Chen

Electrical Engineering and Systems Science > Audio and Speech Processing

arXiv:2310.07284 (eess)

[Submitted on 11 Oct 2023 (v1), last revised 15 Oct 2023 (this version, v3)]

Title:Typing to Listen at the Cocktail Party: Text-Guided Target Speaker Extraction

Authors:Xiang Hao, Jibin Wu, Jianwei Yu, Chenglin Xu, Kay Chen Tan

View PDF

Abstract:Humans possess an extraordinary ability to selectively focus on the sound source of interest amidst complex acoustic environments, commonly referred to as cocktail party scenarios. In an attempt to replicate this remarkable auditory attention capability in machines, target speaker extraction (TSE) models have been developed. These models leverage the pre-registered cues of the target speaker to extract the sound source of interest. However, the effectiveness of these models is hindered in real-world scenarios due to the unreliable or even absence of pre-registered cues. To address this limitation, this study investigates the integration of natural language description to enhance the feasibility, controllability, and performance of existing TSE models. Specifically, we propose a model named LLM-TSE, wherein a large language model (LLM) extracts useful semantic cues from the user's typed text input. These cues can serve as independent extraction cues, task selectors to control the TSE process or complement the pre-registered cues. Our experimental results demonstrate competitive performance when only text-based cues are presented, the effectiveness of using input text as a task selector, and a new state-of-the-art when combining text-based cues with pre-registered cues. To our knowledge, this is the first study to successfully incorporate LLMs to guide target speaker extraction, which can be a cornerstone for cocktail party problem research.

Comments:	Under review, this https URL
Subjects:	Audio and Speech Processing (eess.AS); Computation and Language (cs.CL)
Cite as:	arXiv:2310.07284 [eess.AS]
	(or arXiv:2310.07284v3 [eess.AS] for this version)
	https://doi.org/10.48550/arXiv.2310.07284

Submission history

From: Xiang Hao [view email]
[v1] Wed, 11 Oct 2023 08:17:54 UTC (14,179 KB)
[v2] Thu, 12 Oct 2023 01:40:37 UTC (7,172 KB)
[v3] Sun, 15 Oct 2023 03:58:29 UTC (7,187 KB)

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:Typing to Listen at the Cocktail Party: Text-Guided Target Speaker Extraction

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:Typing to Listen at the Cocktail Party: Text-Guided Target Speaker Extraction

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators