Hi!
I'd like to recommend two recent works on multimodal agent memory for the awesome list:
PersonaVLM: Long-Term Personalized Multimodal LLMs (focuses on long-term personalization, CVPR 2026 Highlight) and
Light-Omni: Reflex over Reasoning in Agentic Video Understanding with Long-Term Memory (focuses on efficient reflexive memory for video understanding).
Here are the links:
PersonaVLM: https://github.com/MiG-NJU/PersonaVLM
Light-Omni: https://github.com/Clare-Nie/Light-Omni
Would be great if they could be included in the multimodal/video agent sections. Let me know if you'd prefer I submit a PR directly! Thanks!
Hi!
I'd like to recommend two recent works on multimodal agent memory for the awesome list:
PersonaVLM: Long-Term Personalized Multimodal LLMs (focuses on long-term personalization, CVPR 2026 Highlight) and
Light-Omni: Reflex over Reasoning in Agentic Video Understanding with Long-Term Memory (focuses on efficient reflexive memory for video understanding).
Here are the links:
PersonaVLM: https://github.com/MiG-NJU/PersonaVLM
Light-Omni: https://github.com/Clare-Nie/Light-Omni
Would be great if they could be included in the multimodal/video agent sections. Let me know if you'd prefer I submit a PR directly! Thanks!