Many-to-Many Voice Transformer Network

Kameoka, Hirokazu; Huang, Wen-Chin; Tanaka, Kou; Kaneko, Takuhiro; Hojo, Nobukatsu; Toda, Tomoki

Electrical Engineering and Systems Science > Audio and Speech Processing

arXiv:2005.08445 (eess)

[Submitted on 18 May 2020 (v1), last revised 6 Nov 2020 (this version, v4)]

Title:Many-to-Many Voice Transformer Network

Authors:Hirokazu Kameoka, Wen-Chin Huang, Kou Tanaka, Takuhiro Kaneko, Nobukatsu Hojo, Tomoki Toda

View PDF

Abstract:This paper proposes a voice conversion (VC) method based on a sequence-to-sequence (S2S) learning framework, which enables simultaneous conversion of the voice characteristics, pitch contour, and duration of input speech. We previously proposed an S2S-based VC method using a transformer network architecture called the voice transformer network (VTN). The original VTN was designed to learn only a mapping of speech feature sequences from one speaker to another. The main idea we propose is an extension of the original VTN that can simultaneously learn mappings among multiple speakers. This extension called the many-to-many VTN makes it able to fully use available training data collected from multiple speakers by capturing common latent features that can be shared across different speakers. It also allows us to introduce a training loss called the identity mapping loss to ensure that the input feature sequence will remain unchanged when the source and target speaker indices are the same. Using this particular loss for model training has been found to be extremely effective in improving the performance of the model at test time. We conducted speaker identity conversion experiments and found that our model obtained higher sound quality and speaker similarity than baseline methods. We also found that our model, with a slight modification to its architecture, could handle any-to-many conversion tasks reasonably well.

Comments:	submitted to IEEE/ACM Trans. ASLP. Please also refer to our related article: arXiv:1811.01609
Subjects:	Audio and Speech Processing (eess.AS); Sound (cs.SD); Machine Learning (stat.ML)
Cite as:	arXiv:2005.08445 [eess.AS]
	(or arXiv:2005.08445v4 [eess.AS] for this version)
	https://doi.org/10.48550/arXiv.2005.08445

Submission history

From: Hirokazu Kameoka [view email]
[v1] Mon, 18 May 2020 04:02:08 UTC (4,293 KB)
[v2] Sun, 7 Jun 2020 13:14:07 UTC (4,314 KB)
[v3] Fri, 9 Oct 2020 14:48:19 UTC (4,284 KB)
[v4] Fri, 6 Nov 2020 22:46:52 UTC (4,201 KB)

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:Many-to-Many Voice Transformer Network

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:Many-to-Many Voice Transformer Network

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators