报告题目:Advanced Methods in Regional Speech Enhancement and Far-to-Near-Field Transformation
报告人: 于蒙 博士(腾讯混元大模型部门)
邀请人: 黄公平 教授
报告时间: 2026/8/27(周四)下午15:00
报告地点: 武汉大学电子信息学院 于刚·宋晓楼 A603会议室
报告简介:
Enhancing speech signals in complex acoustic environments remains a persistent challenge in audio processing. In recent work, the presenter introduces several innovative strategies to address core issues in this field. First, they propose a deep learning–based audio zooming technique that moves beyond conventional direction-dependent beamforming, instead enabling sound capture within a user-adjustable 3D region. This design allows for more precise and adaptable audio acquisition, supporting real-time use cases such as remote conferencing, education, and live streaming.
Building on this region-based capture paradigm, the work further aims to convert far-field audio into near-field quality. Using real-world acoustic data, the authors develop a novel framework that combines a Schrödinger Bridge–based diffusion model with generative adversarial networks. This hybrid approach achieves state-of-the-art performance in jointly suppressing noise and reverberation, leading to substantial improvements in speech quality. The proposed method sets a new benchmark for far-field to near-field enhancement in practical scenarios. Collectively, these advancements provide robust solutions to long-standing speech processing challenges, paving the way for high-fidelity audio experiences across a wide range of applications.
报告人简介:
Meng Yu has been a Principal Research Scientist at Tencent AI Lab, now at Tencent Hunyuan since 2016. From 2013 to 2016, he worked as a Staff Research Engineer at Audience, focusing on audio/speech enhancement for voice communication and improving speech recognition. Prior to that, he was a Software Engineer at Cisco from 2012 to 2013, specializing in speaker segmentation and recognition. He received B.S. in Mathematics from Peking University, Beijing, China in 2007, and a Ph.D. degree in Mathematics from University of California, Irvine, CA, USA in 2012. His research interests focus on audio and speech processing, with a particular emphasis on single and multi-channel far-field frontend speech enhancement applications.
欢迎感兴趣的老师和同学们积极参与!