Sound

Authors and titles for recent submissions, skipping first 14

[ total of 51 entries: 1-10 | 5-14 | 15-24 | 25-34 | 35-44 | 45-51 ]
[ showing 10 entries per page: fewer | more | all ]

Tue, 9 Dec 2025 (continued, showing last 5 of 19 entries)

[15] arXiv:2512.07226 (cross-list from eess.AS) [pdf, ps, other]: Title: Unsupervised Single-Channel Audio Separation with Diffusion Source Priors

Authors: Runwu Shi, Chang Li, Jiang Wang, Rui Zhang, Nabeela Khan, Benjamin Yen, Takeshi Ashizawa, Kazuhiro Nakadai

Comments: 15 pages, 31 figures, accepted by The 40th Annual AAAI Conference on Artificial Intelligence (AAAI 2026)

Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[16] arXiv:2512.07209 (cross-list from cs.MM) [pdf, ps, other]: Title: Coherent Audio-Visual Editing via Conditional Audio Generation Following Video Edits

Authors: Masato Ishii, Akio Hayakawa, Takashi Shibuya, Yuki Mitsufuji

Subjects: Multimedia (cs.MM); Machine Learning (cs.LG); Sound (cs.SD)
[17] arXiv:2512.06417 (cross-list from cs.LG) [pdf, ps, other]: Title: Hankel-FNO: Fast Underwater Acoustic Charting Via Physics-Encoded Fourier Neural Operator

Authors: Yifan Sun (1), Lei Cheng (1), Jianlong Li (1), Peter Gerstoft (2) ((1) College of Information Science and Electronic Engineering, Zhejiang University, Hangzhou, China, (2) Scripps Institution of Oceanography, University of California San Diego, La Jolla, USA)

Subjects: Machine Learning (cs.LG); Sound (cs.SD)
[18] arXiv:2512.06304 (cross-list from eess.AS) [pdf, ps, other]: Title: Degrading Voice: A Comprehensive Overview of Robust Voice Conversion Through Input Manipulation

Authors: Xining Song, Zhihua Wei, Rui Wang, Haixiao Hu, Yanxiang Chen, Meng Han

Subjects: Audio and Speech Processing (eess.AS); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR); Sound (cs.SD)
[19] arXiv:2512.05994 (cross-list from eess.AS) [pdf, ps, other]: Title: KidSpeak: A General Multi-purpose LLM for Kids' Speech Recognition and Screening

Authors: Rohan Sharma, Dancheng Liu, Jingchen Sun, Shijie Zhou, Jiayu Qin, Jinjun Xiong, Changyou Chen

Subjects: Audio and Speech Processing (eess.AS); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Sound (cs.SD)

Mon, 8 Dec 2025

[20] arXiv:2512.05592 [pdf, ps, other]: Title: The T12 System for AudioMOS Challenge 2025: Audio Aesthetics Score Prediction System Using KAN- and VERSA-based Models

Authors: Katsuhiko Yamamoto, Koichi Miyazaki, Shogo Seki

Comments: Accepted by IEEE ASRU 2025

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[21] arXiv:2512.05508 [pdf, ps, other]: Title: Lyrics Matter: Exploiting the Power of Learnt Representations for Music Popularity Prediction

Authors: Yash Choudhary, Preeti Rao, Pushpak Bhattacharyya

Comments: 8 pages

Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
[22] arXiv:2512.05528 (cross-list from q-bio.NC) [pdf, ps, other]: Title: Decoding Selective Auditory Attention to Musical Elements in Ecologically Valid Music Listening

Authors: Taketo Akama, Zhuohao Zhang, Tsukasa Nagashima, Takagi Yutaka, Shun Minamikawa, Natalia Polouliakh

Subjects: Neurons and Cognition (q-bio.NC); Machine Learning (cs.LG); Sound (cs.SD); Audio and Speech Processing (eess.AS); Signal Processing (eess.SP)
[23] arXiv:2512.05201 (cross-list from cs.NI) [pdf, ps, other]: Title: MuMeNet: A Network Simulator for Musical Metaverse Communications

Authors: Ali Al Housseini, Jaime Llorca, Luca Turchet, Tiziano Leidi, Cristina Rottondi, Omran Ayoub

Comments: To appear in 2025 IEEE 6th International Symposium on the Internet of Sounds (IS2) proceedings

Subjects: Networking and Internet Architecture (cs.NI); Sound (cs.SD)
[24] arXiv:2512.05126 (cross-list from eess.AS) [pdf, ps, other]: Title: SyncVoice: Towards Video Dubbing with Vision-Augmented Pretrained TTS Model

Authors: Kaidi Wang, Yi He, Wenhao Guan, Weijie Wu, Hongwu Ding, Xiong Zhang, Di Wu, Meng Meng, Jian Luan, Lin Li, Qingyang Hong

Subjects: Audio and Speech Processing (eess.AS); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM); Sound (cs.SD)

[ total of 51 entries: 1-10 | 5-14 | 15-24 | 25-34 | 35-44 | 45-51 ]
[ showing 10 entries per page: fewer | more | all ]

Disable MathJax (What is MathJax?)

Links to: arXiv, form interface, find, cs, new, 2512, contact, help (Access key information)

> cs > cs.SD

Sound

Authors and titles for recent submissions, skipping first 14

Tue, 9 Dec 2025 (continued, showing last 5 of 19 entries)

Mon, 8 Dec 2025