We gratefully acknowledge support from
the Simons Foundation and member institutions.

Sound

Authors and titles for recent submissions, skipping first 8

[ total of 45 entries: 1-50 | 9-45 ]
[ showing up to 50 entries per page: fewer | more ]

Tue, 4 Jun 2024 (continued, showing last 11 of 19 entries)

[9]  arXiv:2406.00146 [pdf, other]
Title: A Survey of Deep Learning Audio Generation Methods
Comments: 14 pages, 2 figures
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[10]  arXiv:2406.01205 (cross-list from eess.AS) [pdf, other]
Title: ControlSpeech: Towards Simultaneous Zero-shot Speaker Cloning and Zero-shot Language Style Control With Decoupled Codec
Subjects: Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Sound (cs.SD)
[11]  arXiv:2406.01018 (cross-list from eess.AS) [pdf, other]
Title: Accent Conversion in Text-To-Speech Using Multi-Level VAE and Adversarial Training
Comments: Under review
Subjects: Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Sound (cs.SD)
[12]  arXiv:2406.00976 (cross-list from cs.CL) [pdf, other]
Title: Generative Pre-trained Speech Language Model with Efficient Hierarchical Transformer
Comments: Accept in ACL2024-main
Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[13]  arXiv:2406.00901 (cross-list from cs.MM) [pdf, other]
Title: Robust Multi-Modal Speech In-Painting: A Sequence-to-Sequence Approach
Subjects: Multimedia (cs.MM); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[14]  arXiv:2406.00899 (cross-list from cs.CL) [pdf, other]
Title: YODAS: Youtube-Oriented Dataset for Audio and Speech
Comments: ASRU 2023
Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[15]  arXiv:2406.00654 (cross-list from cs.CL) [pdf, other]
Title: Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback
Comments: 19 pages, Preprint
Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[16]  arXiv:2406.00626 (cross-list from cs.MM) [pdf, other]
Title: Intelligent Text-Conditioned Music Generation
Subjects: Multimedia (cs.MM); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[17]  arXiv:2406.00356 (cross-list from eess.AS) [pdf, other]
Title: AudioLCM: Text-to-Audio Generation with Latent Consistency Models
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[18]  arXiv:2406.00022 (cross-list from cs.CL) [pdf, other]
Title: Multilingual Prosody Transfer: Comparing Supervised & Transfer Learning
Comments: 7 pages, Accepted to ICLR 2024 - Tiny Track
Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[19]  arXiv:2406.00021 (cross-list from cs.CL) [pdf, other]
Title: CrossVoice: Crosslingual Prosody Preserving Cascade-S2ST using Transfer Learning
Comments: 8 pages, Accepted at ICLR 2024 - Tiny Track
Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)

Mon, 3 Jun 2024

[20]  arXiv:2405.20887 [pdf, other]
Title: On the Condition Monitoring of Bolted Joints through Acoustic Emission and Deep Transfer Learning: Generalization, Ordinal Loss and Super-Convergence
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[21]  arXiv:2405.20884 [pdf, other]
Title: Effects of Dataset Sampling Rate for Noise Cancellation through Deep Learning
Comments: 16 pages, 8 pictures, 3 tables
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[22]  arXiv:2405.20410 (cross-list from cs.CL) [pdf, other]
Title: SeamlessExpressiveLM: Speech Language Model for Expressive Speech-to-Speech Translation with Chain-of-Thought
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[23]  arXiv:2405.20402 (cross-list from eess.AS) [pdf, other]
Title: Cross-Talk Reduction
Comments: in International Joint Conference on Artificial Intelligence (IJCAI), 2024
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD); Signal Processing (eess.SP)

Fri, 31 May 2024

[24]  arXiv:2405.20289 [pdf, other]
Title: DITTO-2: Distilled Diffusion Inference-Time T-Optimization for Music Generation
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
[25]  arXiv:2405.20172 [pdf, other]
Title: Iterative Feature Boosting for Explainable Speech Emotion Recognition
Comments: Published in: 2023 International Conference on Machine Learning and Applications (ICMLA)
Journal-ref: 2023 International Conference on Machine Learning and Applications (ICMLA), Jacksonville, FL, USA, 2023, pp. 543-549
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[26]  arXiv:2405.20101 [pdf, other]
Title: Fill in the Gap! Combining Self-supervised Representation Learning with Neural Audio Synthesis for Speech Inpainting
Subjects: Sound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
[27]  arXiv:2405.20059 [pdf, other]
Title: Spectral Mapping of Singing Voices: U-Net-Assisted Vocal Segmentation
Authors: Adam Sorrenti
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[28]  arXiv:2405.19796 [pdf, other]
Title: Explainable Attribute-Based Speaker Verification
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[29]  arXiv:2405.19343 [pdf, other]
Title: Luganda Speech Intent Recognition for IoT Applications
Comments: Presented as a conference paper at ICLR 2024/AfricaNLP
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
[30]  arXiv:2405.19342 [pdf, other]
Title: Sonos Voice Control Bias Assessment Dataset: A Methodology for Demographic Bias Assessment in Voice Assistants
Subjects: Sound (cs.SD); Computation and Language (cs.CL); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[31]  arXiv:2405.20336 (cross-list from cs.CV) [pdf, other]
Title: RapVerse: Coherent Vocals and Whole-Body Motions Generations from Text
Comments: Project website: this https URL
Subjects: Computer Vision and Pattern Recognition (cs.CV); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[32]  arXiv:2405.20064 (cross-list from eess.AS) [pdf, other]
Title: 1st Place Solution to Odyssey Emotion Recognition Challenge Task1: Tackling Class Imbalance Problem
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[33]  arXiv:2405.19497 (cross-list from eess.AS) [pdf, other]
Title: Gaussian Flow Bridges for Audio Domain Transfer with Unpaired Data
Comments: Submitted to IWAENC 2024
Subjects: Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Sound (cs.SD)
[34]  arXiv:2405.19426 (cross-list from cs.CL) [pdf, other]
Title: Deep Learning for Assessment of Oral Reading Fluency
Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)

Thu, 30 May 2024

[35]  arXiv:2405.18726 [pdf, other]
Title: Reverse the auditory processing pathway: Coarse-to-fine audio reconstruction from fMRI
Subjects: Sound (cs.SD); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[36]  arXiv:2405.18503 [pdf, other]
Title: SoundCTM: Uniting Score-based and Consistency Models for Text-to-Sound Generation
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[37]  arXiv:2405.19041 (cross-list from cs.CL) [pdf, other]
Title: BLSP-KD: Bootstrapping Language-Speech Pre-training via Knowledge Distillation
Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[38]  arXiv:2405.18639 (cross-list from q-bio.NC) [pdf, other]
Title: Improving Speech Decoding from ECoG with Self-Supervised Pretraining
Subjects: Neurons and Cognition (q-bio.NC); Computation and Language (cs.CL); Machine Learning (cs.LG); Sound (cs.SD); Audio and Speech Processing (eess.AS)

Wed, 29 May 2024

[39]  arXiv:2405.18386 [pdf, other]
Title: Instruct-MusicGen: Unlocking Text-to-Music Editing for Music Language Models via Instruction Tuning
Comments: Code and demo are available at: this https URL
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[40]  arXiv:2405.18213 [pdf, other]
Title: NeRAF: 3D Scene Infused Neural Radiance and Acoustic Fields
Comments: Project Page: this https URL
Subjects: Sound (cs.SD); Computer Vision and Pattern Recognition (cs.CV); Audio and Speech Processing (eess.AS)
[41]  arXiv:2405.18153 [pdf, other]
Title: Practical aspects for the creation of an audio dataset from field recordings with optimized labeling budget with AI-assisted strategy
Comments: Submitted to ICML 2024 Workshop on Data-Centric Machine Learning Research
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[42]  arXiv:2405.17615 [pdf, other]
Title: Listenable Maps for Zero-Shot Audio Classifiers
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS); Signal Processing (eess.SP)
[43]  arXiv:2405.17842 (cross-list from cs.CV) [pdf, other]
Title: Discriminator-Guided Cooperative Diffusion for Joint Audio and Video Generation
Subjects: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Multimedia (cs.MM); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[44]  arXiv:2405.17809 (cross-list from cs.CL) [pdf, other]
Title: TransVIP: Speech to Speech Translation System with Voice and Isochrony Preservation
Comments: Work in progress
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[45]  arXiv:2405.17569 (cross-list from cs.LG) [pdf, other]
Title: Discriminant audio properties in deep learning based respiratory insufficiency detection in Brazilian Portuguese
Comments: 5 pages, 2 figures, 1 table. Published in Artificial Intelligence in Medicine (AIME) 2023
Journal-ref: Artificial Intellingence in Medicine Proceedings 2023, page 271-275
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[ total of 45 entries: 1-50 | 9-45 ]
[ showing up to 50 entries per page: fewer | more ]

Disable MathJax (What is MathJax?)

Links to: arXiv, form interface, find, cs, new, 2406, contact, help  (Access key information)