Locality-aware continual unlearning (LACU), a framework that stabilizes sequential concept removal
Theoretically analyzes the MaskGIT sampler, poviding a choose-then-sample (CTS) formulation
Training flow-map models in RAE latent with consistency mid-training for trajectory-aware initialization
A framework for Identify which training examples influenced specific concepts within the diffusion model
CMT reduces the training cost of diffusion-based flow map models by up to 90% while reaching SOTA performance
Learning conditional, unconditional, and matching-aware discriminator with adaptive weighting mechanism (cSAN)
Propose tensor-decomposition-based PEFT method, showing its effectiveness on T-to-I generation tasks
Theoretical analysis of limitation of current discrete diffusion and a method for effectively capturing element-wise dependency
A method efficiently leverages online human feedback to fine-tune Stable Diffusion for various range of tasks
An enhanced multimodal representation using weighted point clouds and its theoretical benefits
A 64x64 pre-trained diffusion model is all you need for 1-step high-resolution SOTA generation
Unified framework enables diverse samplers and 1-step generation SOTAs
Applications:
[SoundGen]
<div class="tile">
<h3>MCA</h3>
<a href=""><img src="./assets/mca.png"></a>
<h5>
[EMNLP]
<a href="https://arxiv.org/abs/2510.15543">[arXiv]</a>
[code]
</h5>
<p>MCA: Modality Composition Awareness for Robust Composed Multimodal Retrieval</p>
<div class="tile_venue">EMNLP26</div>
</div>
<div class="tile">
<h3>MLLMCLIP</h3>
<a href=""><img src="./assets/mllmclip.png"></a>
<h5>
[EMNLP]
[arXiv]
[code]
</h5>
<p>MLLMCLIP: Feature-Level Distillation of MLLM for Robust Vision-Language Representations</p>
<div class="tile_venue">EMNLP26</div>
</div>
<div class="tile">
<h3>Syn-Omni</h3>
<a href=""><img src="./assets/synomni.png"></a>
<h5>
[EMNLP]
[arXiv]
[code]
</h5>
<p>Syn-Omni: Structured Specialization and Progressive Collaboration for Omnimodal Embeddings</p>
<div class="tile_venue">EMNLP26</div>
</div>
<div class="tile">
<h3>DynaVieW</h3>
<a href=""><img src="./assets/dynaview.png"></a>
<h5>
<a href="https://icml.cc/virtual/2026/poster/65080">[ICML]</a>
<a href="https://arxiv.org/abs/2607.04112">[arXiv]</a>
<a href="https://github.com/Silin159/DynaVieW">[code]</a>
</h5>
<p>DynaVieW: Schema-Guided World Modeling for Understanding Hierarchical Visual Dynamics</p>
<div class="tile_venue">ICML26</div>
</div>
<div class="tile">
<h3>LRPO</h3>
<a href=""><img src="./assets/lrpo.png"></a>
<h5>
<a href="https://icml.cc/virtual/2026/poster/66718">[ICML]</a>
<a href="https://arxiv.org/abs/2605.25360">[arXiv]</a>
<a href="https://github.com/Guochry/LRPO">[code]</a>
</h5>
<p>Learning to Route Languages for Multilingual Policy Optimization</p>
<div class="tile_venue">ICML26</div>
</div>
<div class="tile">
<h3>DeepResonance</h3>
<a href=""><img src="./assets/deepresonance.png"></a>
<h5>
<a href="https://aclanthology.org/2025.emnlp-main.653/">[EMNLP]</a>
<a href="https://arxiv.org/abs/2502.12623">[arXiv]</a>
<a href="https://github.com/sony/DeepResonance">[code]</a>
</h5>
<p>DeepResonance: Enhancing Multimodal Music Understanding via Music-centric Multi-way Instruction Tuning</p>
<div class="tile_venue">EMNLP25</div>
</div>
<div class="tile">
<h3>CARE</h3>
<a href=""><img src="./assets/care.png"></a>
<h5>
<a href="https://aclanthology.org/2025.emnlp-main.1669/">[EMNLP]</a>
<a href="https://arxiv.org/abs/2504.05154">[arXiv]</a>
<a href="https://github.com/Guochry/CARE">[data]</a>
</h5>
<p>CARE: Assessing the Impact of Multilingual Human Preference Learning on Cultural Awareness</p>
<div class="tile_venue">EMNLP25</div>
</div>
<div class="tile">
<h3>BiAug</h3>
<a href=""><img src="./assets/biaug.png"></a>
<h5>
[MRR@ICCV25]
<a href="https://arxiv.org/abs/2310.01330">[arXiv]</a>
</h5>
<p>Towards reporting bias in visual-language datasets: bimodal augmentation by decoupling object-attribute association</p>
<div class="tile_venue">ICCV25 MRR Workshop</div>
</div>
<div class="tile">
<h3>GLOV</h3>
<a href=""><img src="./assets/glov.png"></a>
<h5>
[TMLR]
<a href="https://arxiv.org/abs/2410.06154">[arXiv]</a>
</h5>
<p>GLOV: Guided Large Language Models as Implicit Optimizers for Vision Language Models</p>
<div class="tile_venue">TMLR</div>
</div>
<div class="tile">
<h3>Music-to-MVD</h3>
<a href=""><img src="./assets/mvd.png"></a>
<h5>
<a href="https://aclanthology.org/2025.repl4nlp-1.4.pdf">[RepL4NLP@NAACL25]</a>
<a href="https://arxiv.org/abs/2503.11190">[arXiv]</a>
</h5>
<p>Cross-Modal Learning for Music-to-Music-Video Description Generation</p>
<div class="tile_venue">NAACL25 RepL4NLP Workshop</div>
</div>
<div class="tile">
<h3>VinaBench</h3>
<a href=""><img src="./assets/vinabench.png"></a>
<h5>
[CVPR]
<a href="https://arxiv.org/abs/2503.20871">[arXiv]</a>
<a href="https://silin159.github.io/Vina-Bench/">[data]</a>
</h5>
<p>VinaBench: Benchmark for Faithful and Consistent Visual Narratives</p>
<div class="tile_venue">CVPR25</div>
</div>
<div class="tile">
<h3>OpenMU</h3>
<a href=""><img src="./assets/openmu.png"></a>
<h5>
<a href="https://arxiv.org/abs/2410.15573">[arXiv]</a>
<a href="https://huggingface.co/datasets/Sony/OpenMU-Bench">[data]</a>
<a href="https://mzhaojp22.github.io/open_music_understanding/">[demo]</a>
</h5>
<p>OpenMU: Your Swiss Army Knife for Music Understanding</p>
<div class="tile_venue">ISMIR2024 Late Breaking Demos</div>
</div>
<div class="tile">
<h3>DiffuCOMET</h3>
<a href="https://arxiv.org/abs/2402.17011"><img src="./assets/diffcomet.png"></a>
<h5>
<a href="https://aclanthology.org/2024.acl-long.264/">[ACL]</a>
<a href="https://arxiv.org/abs/2402.17011">[arXiv]</a>
<a href="https://github.com/Silin159/DiffuCOMET">[code]</a>
</h5>
<p>DiffuCOMET: Contextual Commonsense Knowledge Diffusion</p>
<div class="tile_venue">ACL24</div>
</div>
<div class="tile">
<h3>CyCLIPs/CyCLAPs</h3>
<a href="https://arxiv.org/abs/2310.13267"><img src="./assets/cyclips.png"></a>
<h5>
<a href="https://aclanthology.org/2024.findings-acl.293/">[ACL]</a>
<a href="https://arxiv.org/abs/2310.13267">[arXiv]</a>
</h5>
<p>On the Language Encoder of Contrastive Cross-modal Models</p>
<div class="tile_venue">ACL24</div>
</div>
<div class="tile">
<h3>DIIR</h3>
<a href="https://arxiv.org/abs/2403.15737"><img src="./assets/diir.png"></a>
<h5>
<a href="https://aclanthology.org/2024.findings-acl.782/">[ACL]</a>
<a href="https://arxiv.org/abs/2403.15737">[arXiv]</a>
<a href="https://github.com/zhouhanxie/DIIR">[code]</a>
</h5>
<p>Few-shot Dialogue Strategy Learning for Motivational Interviewing via Inductive Reasoning</p>
<div class="tile_venue">ACL24</div>
</div>
<div class="tile">
<h3>PeaCok</h3>
<a href="https://arxiv.org/abs/2305.02364"><img src="./assets/personas.png"></a>
<h5>
<a href="https://aclanthology.org/2023.acl-long.362/">[ACL]</a>
<a href="https://arxiv.org/abs/2305.02364">[arXiv]</a>
<a href="https://github.com/Silin159/PeaCoK">[code]</a>
</h5>
<p>PeaCoK: Persona Commonsense Knowledge for Consistent and Engaging Narratives<br>(Outstanding Paper Award)</p>
<div class="tile_venue">ACL23</div>
</div>
<div class="tile">
<h3>ComFact</h3>
<a href="https://aclanthology.org/2022.findings-emnlp.120/"><img src="./assets/comfact.png"></a>
<h5>
<a href="https://aclanthology.org/2022.findings-emnlp.120/">[EMNLP]</a>
<a href="https://arxiv.org/abs/2210.12678">[arXiv]</a>
<a href="https://github.com/epfl-nlp/ComFact">[code]</a>
</h5>
<p>ComFact: A Benchmark for Linking Contextual Commonsense Knowledge</p>
<div class="tile_venue">EMNLP22 Findings</div>
</div>
<div class="tile" style="background-color: white;"></div>
<div class="tile" style="background-color: white;"></div>
Large-Scale Training Data Attribution for Music Generative Models via Unlearning
SOTA Fx representation: Extract instrument-wise audio effects representations from music mixtures
Reverse Engineering of Music Mixing Graphs with Differentiable Processors and Iterative Pruning
Supervised contrastive learning from weakly-labeled audio segments for musical version matching
DiffVox: A Differentiable Model for Capturing and Analysing Professional Effects Distributions
Improving Unsupervised Clean-to-Rendered Guitar Tone Transformation Using GANs and Integrated Unaligned Clean Data
DiffRoll: Diffusion-based Generative Music Transcription with Unsupervised Pretraining Capability
PAVAS: a framework for generating physically plausible audio from video, by integrating physics estimation
A diffusion-based post-processor for perceptually improving speech enhancement and separation outputs
MMDisCo: Multi-Modal Discriminator-Guided Cooperative Diffusion for Joint Audio and Video Generation
Mining Your Own Secrets: Diffusion Classifier Scores for Continual Personalization of Text-to-Image Diffusion Models
SpecMaskGIT: Masked Generative Modeling of Audio Spectrograms for Efficient Audio Synthesis and Beyond
STARSS23: An Audio-Visual Dataset of Spatial Recordings of Real Scenes with Spatiotemporal Annotations of Sound Events































































