Sound

Authors and titles for recent submissions, skipping first 8

Tue, 4 Jun 2024
Mon, 3 Jun 2024
Fri, 31 May 2024
Thu, 30 May 2024
Wed, 29 May 2024

[ total of 45 entries: 1-50 | 9-45 ]
[ showing up to 50 entries per page: fewer | more ]

Tue, 4 Jun 2024 (continued, showing last 11 of 19 entries)

[9] arXiv:2406.00146 [pdf, other]: Title: A Survey of Deep Learning Audio Generation Methods

Authors: Matej Božić, Marko Horvat

Comments: 14 pages, 2 figures

Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[10] arXiv:2406.01205 (cross-list from eess.AS) [pdf, other]: Title: ControlSpeech: Towards Simultaneous Zero-shot Speaker Cloning and Zero-shot Language Style Control With Decoupled Codec

Authors: Shengpeng Ji, Jialong Zuo, Minghui Fang, Siqi Zheng, Qian Chen, Wen Wang, Ziyue Jiang, Hai Huang, Xize Cheng, Rongjie Huang, Zhou Zhao

Subjects: Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Sound (cs.SD)
[11] arXiv:2406.01018 (cross-list from eess.AS) [pdf, other]: Title: Accent Conversion in Text-To-Speech Using Multi-Level VAE and Adversarial Training

Authors: Jan Melechovsky, Ambuj Mehrish, Berrak Sisman, Dorien Herremans

Comments: Under review

Subjects: Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Sound (cs.SD)
[12] arXiv:2406.00976 (cross-list from cs.CL) [pdf, other]: Title: Generative Pre-trained Speech Language Model with Efficient Hierarchical Transformer

Authors: Yongxin Zhu, Dan Su, Liqiang He, Linli Xu, Dong Yu

Comments: Accept in ACL2024-main

Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[13] arXiv:2406.00901 (cross-list from cs.MM) [pdf, other]: Title: Robust Multi-Modal Speech In-Painting: A Sequence-to-Sequence Approach

Authors: Mahsa Kadkhodaei Elyaderani, Shahram Shirani

Subjects: Multimedia (cs.MM); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[14] arXiv:2406.00899 (cross-list from cs.CL) [pdf, other]: Title: YODAS: Youtube-Oriented Dataset for Audio and Speech

Authors: Xinjian Li, Shinnosuke Takamichi, Takaaki Saeki, William Chen, Sayaka Shiota, Shinji Watanabe

Comments: ASRU 2023

Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[15] arXiv:2406.00654 (cross-list from cs.CL) [pdf, other]: Title: Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback

Authors: Chen Chen, Yuchen Hu, Wen Wu, Helin Wang, Eng Siong Chng, Chao Zhang

Comments: 19 pages, Preprint

Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[16] arXiv:2406.00626 (cross-list from cs.MM) [pdf, other]: Title: Intelligent Text-Conditioned Music Generation

Authors: Zhouyao Xie, Nikhil Yadala, Xinyi Chen, Jing Xi Liu

Subjects: Multimedia (cs.MM); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[17] arXiv:2406.00356 (cross-list from eess.AS) [pdf, other]: Title: AudioLCM: Text-to-Audio Generation with Latent Consistency Models

Authors: Huadai Liu, Rongjie Huang, Yang Liu, Hengyuan Cao, Jialei Wang, Xize Cheng, Siqi Zheng, Zhou Zhao

Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[18] arXiv:2406.00022 (cross-list from cs.CL) [pdf, other]: Title: Multilingual Prosody Transfer: Comparing Supervised & Transfer Learning

Authors: Arnav Goel, Medha Hira, Anubha Gupta

Comments: 7 pages, Accepted to ICLR 2024 - Tiny Track

Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[19] arXiv:2406.00021 (cross-list from cs.CL) [pdf, other]: Title: CrossVoice: Crosslingual Prosody Preserving Cascade-S2ST using Transfer Learning

Authors: Medha Hira, Arnav Goel, Anubha Gupta

Comments: 8 pages, Accepted at ICLR 2024 - Tiny Track

Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)

Mon, 3 Jun 2024

[20] arXiv:2405.20887 [pdf, other]: Title: On the Condition Monitoring of Bolted Joints through Acoustic Emission and Deep Transfer Learning: Generalization, Ordinal Loss and Super-Convergence

Authors: Emmanuel Ramasso, Rafael de O. Teloli, Romain Marcel

Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[21] arXiv:2405.20884 [pdf, other]: Title: Effects of Dataset Sampling Rate for Noise Cancellation through Deep Learning

Authors: Brandon Colelough, Andrew Zheng

Comments: 16 pages, 8 pictures, 3 tables

Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[22] arXiv:2405.20410 (cross-list from cs.CL) [pdf, other]: Title: SeamlessExpressiveLM: Speech Language Model for Expressive Speech-to-Speech Translation with Chain-of-Thought

Authors: Hongyu Gong, Bandhav Veluri

Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[23] arXiv:2405.20402 (cross-list from eess.AS) [pdf, other]: Title: Cross-Talk Reduction

Authors: Zhong-Qiu Wang, Anurag Kumar, Shinji Watanabe

Comments: in International Joint Conference on Artificial Intelligence (IJCAI), 2024

Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD); Signal Processing (eess.SP)

Fri, 31 May 2024

[24] arXiv:2405.20289 [pdf, other]: Title: DITTO-2: Distilled Diffusion Inference-Time T-Optimization for Music Generation

Authors: Zachary Novack, Julian McAuley, Taylor Berg-Kirkpatrick, Nicholas Bryan

Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
[25] arXiv:2405.20172 [pdf, other]: Title: Iterative Feature Boosting for Explainable Speech Emotion Recognition

Authors: Alaa Nfissi, Wassim Bouachir, Nizar Bouguila, Brian Mishara

Comments: Published in: 2023 International Conference on Machine Learning and Applications (ICMLA)

Journal-ref: 2023 International Conference on Machine Learning and Applications (ICMLA), Jacksonville, FL, USA, 2023, pp. 543-549

Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[26] arXiv:2405.20101 [pdf, other]: Title: Fill in the Gap! Combining Self-supervised Representation Learning with Neural Audio Synthesis for Speech Inpainting

Authors: Ihab Asaad, Maxime Jacquelin, Olivier Perrotin, Laurent Girin, Thomas Hueber

Subjects: Sound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
[27] arXiv:2405.20059 [pdf, other]: Title: Spectral Mapping of Singing Voices: U-Net-Assisted Vocal Segmentation

Authors: Adam Sorrenti

Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[28] arXiv:2405.19796 [pdf, other]: Title: Explainable Attribute-Based Speaker Verification

Authors: Xiaoliang Wu, Chau Luu, Peter Bell, Ajitha Rajan

Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[29] arXiv:2405.19343 [pdf, other]: Title: Luganda Speech Intent Recognition for IoT Applications

Authors: Andrew Katumba, Sudi Murindanyi, John Trevor Kasule, Elvis Mugume

Comments: Presented as a conference paper at ICLR 2024/AfricaNLP

Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
[30] arXiv:2405.19342 [pdf, other]: Title: Sonos Voice Control Bias Assessment Dataset: A Methodology for Demographic Bias Assessment in Voice Assistants

Authors: Chloé Sekkat, Fanny Leroy, Salima Mdhaffar, Blake Perry Smith, Yannick Estève, Joseph Dureau, Alice Coucke

Subjects: Sound (cs.SD); Computation and Language (cs.CL); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[31] arXiv:2405.20336 (cross-list from cs.CV) [pdf, other]: Title: RapVerse: Coherent Vocals and Whole-Body Motions Generations from Text

Authors: Jiaben Chen, Xin Yan, Yihang Chen, Siyuan Cen, Qinwei Ma, Haoyu Zhen, Kaizhi Qian, Lie Lu, Chuang Gan

Comments: Project website: this https URL

Subjects: Computer Vision and Pattern Recognition (cs.CV); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[32] arXiv:2405.20064 (cross-list from eess.AS) [pdf, other]: Title: 1st Place Solution to Odyssey Emotion Recognition Challenge Task1: Tackling Class Imbalance Problem

Authors: Mingjie Chen, Hezhao Zhang, Yuanchao Li, Jiachen Luo, Wen Wu, Ziyang Ma, Peter Bell, Catherine Lai, Joshua Reiss, Lin Wang, Philip C. Woodland, Xie Chen, Huy Phan, Thomas Hain

Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[33] arXiv:2405.19497 (cross-list from eess.AS) [pdf, other]: Title: Gaussian Flow Bridges for Audio Domain Transfer with Unpaired Data

Authors: Eloi Moliner, Sebastian Braun, Hannes Gamper

Comments: Submitted to IWAENC 2024

Subjects: Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Sound (cs.SD)
[34] arXiv:2405.19426 (cross-list from cs.CL) [pdf, other]: Title: Deep Learning for Assessment of Oral Reading Fluency

Authors: Mithilesh Vaidya, Binaya Kumar Sahoo, Preeti Rao

Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)

Thu, 30 May 2024

[35] arXiv:2405.18726 [pdf, other]: Title: Reverse the auditory processing pathway: Coarse-to-fine audio reconstruction from fMRI

Authors: Che Liu, Changde Du, Xiaoyu Chen, Huiguang He

Subjects: Sound (cs.SD); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[36] arXiv:2405.18503 [pdf, other]: Title: SoundCTM: Uniting Score-based and Consistency Models for Text-to-Sound Generation

Authors: Koichi Saito, Dongjun Kim, Takashi Shibuya, Chieh-Hsin Lai, Zhi Zhong, Yuhta Takida, Yuki Mitsufuji

Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[37] arXiv:2405.19041 (cross-list from cs.CL) [pdf, other]: Title: BLSP-KD: Bootstrapping Language-Speech Pre-training via Knowledge Distillation

Authors: Chen Wang, Minpeng Liao, Zhongqiang Huang, Jiajun Zhang

Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[38] arXiv:2405.18639 (cross-list from q-bio.NC) [pdf, other]: Title: Improving Speech Decoding from ECoG with Self-Supervised Pretraining

Authors: Brian A. Yuan, Joseph G. Makin

Subjects: Neurons and Cognition (q-bio.NC); Computation and Language (cs.CL); Machine Learning (cs.LG); Sound (cs.SD); Audio and Speech Processing (eess.AS)

Wed, 29 May 2024

[39] arXiv:2405.18386 [pdf, other]: Title: Instruct-MusicGen: Unlocking Text-to-Music Editing for Music Language Models via Instruction Tuning

Authors: Yixiao Zhang, Yukara Ikemiya, Woosung Choi, Naoki Murata, Marco A. Martínez-Ramírez, Liwei Lin, Gus Xia, Wei-Hsiang Liao, Yuki Mitsufuji, Simon Dixon

Comments: Code and demo are available at: this https URL

Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[40] arXiv:2405.18213 [pdf, other]: Title: NeRAF: 3D Scene Infused Neural Radiance and Acoustic Fields

Authors: Amandine Brunetto, Sascha Hornauer, Fabien Moutarde

Comments: Project Page: this https URL

Subjects: Sound (cs.SD); Computer Vision and Pattern Recognition (cs.CV); Audio and Speech Processing (eess.AS)
[41] arXiv:2405.18153 [pdf, other]: Title: Practical aspects for the creation of an audio dataset from field recordings with optimized labeling budget with AI-assisted strategy

Authors: Javier Naranjo-Alcazar, Jordi Grau-Haro, Ruben Ribes-Serrano, Pedro Zuccarello

Comments: Submitted to ICML 2024 Workshop on Data-Centric Machine Learning Research

Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[42] arXiv:2405.17615 [pdf, other]: Title: Listenable Maps for Zero-Shot Audio Classifiers

Authors: Francesco Paissan, Luca Della Libera, Mirco Ravanelli, Cem Subakan

Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS); Signal Processing (eess.SP)
[43] arXiv:2405.17842 (cross-list from cs.CV) [pdf, other]: Title: Discriminator-Guided Cooperative Diffusion for Joint Audio and Video Generation

Authors: Akio Hayakawa, Masato Ishii, Takashi Shibuya, Yuki Mitsufuji

Subjects: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Multimedia (cs.MM); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[44] arXiv:2405.17809 (cross-list from cs.CL) [pdf, other]: Title: TransVIP: Speech to Speech Translation System with Voice and Isochrony Preservation

Authors: Chenyang Le, Yao Qian, Dongmei Wang, Long Zhou, Shujie Liu, Xiaofei Wang, Midia Yousefi, Yanmin Qian, Jinyu Li, Sheng Zhao, Michael Zeng

Comments: Work in progress

Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[45] arXiv:2405.17569 (cross-list from cs.LG) [pdf, other]: Title: Discriminant audio properties in deep learning based respiratory insufficiency detection in Brazilian Portuguese

Authors: Marcelo Matheus Gauy, Larissa Cristina Berti, Arnaldo Cândido Jr, Augusto Camargo Neto, Alfredo Goldman, Anna Sara Shafferman Levin, Marcus Martins, Beatriz Raposo de Medeiros, Marcelo Queiroz, Ester Cerdeira Sabino, Flaviane Romani Fernandes Svartman, Marcelo Finger

Comments: 5 pages, 2 figures, 1 table. Published in Artificial Intelligence in Medicine (AIME) 2023

Journal-ref: Artificial Intellingence in Medicine Proceedings 2023, page 271-275

Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Sound (cs.SD); Audio and Speech Processing (eess.AS)

Tue, 4 Jun 2024
Mon, 3 Jun 2024
Fri, 31 May 2024
Thu, 30 May 2024
Wed, 29 May 2024

[ total of 45 entries: 1-50 | 9-45 ]
[ showing up to 50 entries per page: fewer | more ]

Disable MathJax (What is MathJax?)

Links to: arXiv, form interface, find, cs, new, 2406, contact, help (Access key information)

> cs > cs.SD

Sound

Authors and titles for recent submissions, skipping first 8

Tue, 4 Jun 2024 (continued, showing last 11 of 19 entries)

Mon, 3 Jun 2024

Fri, 31 May 2024

Thu, 30 May 2024

Wed, 29 May 2024