Integrated M.S. & Ph.D. Student
Department of Artificial Intelligence, Korea University
Seoul, South Korea
Google Scholar • PRML Speech Team • Email
I am an Integrated M.S. & Ph.D. student at Korea University, advised by
Seong-Whan Lee, and a member of the PRML Speech Team.
My research focuses on speech generation and affective speech modeling, including expressive and controllable speech synthesis, neural vocoders, voice conversion, and generative models.
- Speech Synthesis
- Neural Vocoder
- Voice Conversion
- Generative Models
- Affective Computing
-
L-Proto: Language-Aware Episodic Prototypical Training for Multilingual Speaker Verification
H.-S. Oh, D.-H. Cho, S.-B. Kim, S.-W. Lee INTERSPEECH, 2026 -
Affectron: Emotional Speech Synthesis with Affective and Contextually Aligned Nonverbal Vocalizations
D.-H. Cho, H.-S. Oh, S.-B. Kim, S.-W. Lee
ACL Findings, 2026 -
Toward Complex-Valued Neural Networks for Waveform Generation
H.-S. Oh, D.-H. Cho, S.-B. Kim, S.-W. Lee
ICLR, 2026
-
FillerSpeech: Towards Human-Like Text-to-Speech Synthesis with Filler Injection and Filler Style Control
S.-B. Kim, J.-H. Cha, H.-S. Oh, H.-J. Choi, S.-W. Lee
EMNLP, 2025 -
EmoSphere-SER: Enhancing Speech Emotion Recognition through Spherical Representation with Auxiliary Classification
D.-H. Cho, H.-S. Oh, S.-B. Kim, S.-W. Lee
INTERSPEECH, 2025 -
VibE-Singer: Vibrato Extraction with High Frequency Pitch for Singing Voice Conversion
J.-S. Choi, D.-M. Byun, H.-S. Oh, S.-W. Lee
INTERSPEECH, 2025 -
DiEmo-TTS: Disentangled Emotion Representations via Self-Supervised Distillation for Cross-Speaker Emotion Transfer in Text-to-Speech
D.-H. Cho, H.-S. Oh, S.-B. Kim, S.-W. Lee
INTERSPEECH, 2025 -
EmoSphere++: Emotion-Controllable Zero-Shot Text-to-Speech via Emotion-Adaptive Spherical Vector
D.-H. Cho, H.-S. Oh, S.-B. Kim, S.-W. Lee
IEEE Transactions on Affective Computing, 2025 -
UnitCorrect: Unit-based Mispronunciation Correcting System with a DTW-based Detection
H.-W. Bae, H.-S. Oh, S.-B. Kim, S.-W. Lee
IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2025 -
DurFlex-EVC: Duration-Flexible Emotional Voice Conversion Leveraging Discrete Representations Without Text Alignment
H.-S. Oh, S.-H. Lee, D.-H. Cho, S.-W. Lee
IEEE Transactions on Affective Computing, 2025 -
JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis
J.-H. Cha, S.-B. Kim, H.-S. Oh, S.-W. Lee
ICASSP, 2025
-
EmoSphere-TTS: Emotional Style and Intensity Modeling via Spherical Emotion Vector for Controllable Emotional Text-to-Speech
D.-H. Cho, H.-S. Oh, S.-B. Kim, S.-H. Lee, S.-W. Lee
INTERSPEECH, 2024 -
DiffProsody: Diffusion-based Latent Prosody Generation for Expressive Speech Synthesis with Prosody Conditional Adversarial Training
H.-S. Oh, S.-H. Lee, S.-W. Lee
IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2024
-
HierVST: Hierarchical Adaptive Zero-shot Voice Style Transfer
S.-H. Lee, H.-Y. Choi, H.-S. Oh, S.-W. Lee
INTERSPEECH, 2023 -
Haze Removal Network Using Unified Function for Image Dehazing
H.-S. Oh, H.-N. Kim, Y.-J. Kim, C.-H. Yim
Electronics Letters, 2021
Machine Learning Researcher Intern
NAVER Cloud, Seongnam, South Korea
Mar 2023 – Jun 2023
- Korea University — Integrated M.S. & Ph.D. in Artificial Intelligence (2021–Present)
- Konkuk University — B.S. in Computer Science and Engineering (2017–2021)
- Excellence Award, 1st AI Frontier Challenge, Korean Society for Artificial Intelligence (2025)
Reviewer
- IEEE/ACM Transactions on Audio, Speech, and Language Processing (TASLP)
- IEEE Transactions on Affective Computing (TAFFC)
- ACL



