Hi WeMM-Embedding team — congratulations on the release, and thank you for open-sourcing the models.
We are building Cerul.ai, a product focused on video understanding and video retrieval. The video results of WeMM-Embedding are especially interesting to us.
We would love to learn more about:
-
Compared with Qwen3-VL-Embedding and Gemini Embedding 2, which video retrieval capabilities improve most significantly—for example, fine-grained actions, temporal localization, motion understanding, OCR, long-video retrieval, or compositional queries?
-
How does WeMM-Embedding perform in latency- and throughput-sensitive scenarios? Is near-real-time video indexing practical, and how does its efficiency compare with Microsoft’s MAGE-VL?
-
Are official Apple Silicon/MLX support, quantized checkpoints, or a managed API planned? Complete multimodal vLLM or SGLang API examples would also be very helpful.
If some results or deployment details are not suitable for public discussion, we would be happy to communicate privately at jiaxi@cerul.ai.
If there is a good fit, we would also be very interested in exploring potential collaboration around product-oriented video understanding and retrieval.
Thank you again for the great work.
Hi WeMM-Embedding team — congratulations on the release, and thank you for open-sourcing the models.
We are building Cerul.ai, a product focused on video understanding and video retrieval. The video results of WeMM-Embedding are especially interesting to us.
We would love to learn more about:
Compared with Qwen3-VL-Embedding and Gemini Embedding 2, which video retrieval capabilities improve most significantly—for example, fine-grained actions, temporal localization, motion understanding, OCR, long-video retrieval, or compositional queries?
How does WeMM-Embedding perform in latency- and throughput-sensitive scenarios? Is near-real-time video indexing practical, and how does its efficiency compare with Microsoft’s MAGE-VL?
Are official Apple Silicon/MLX support, quantized checkpoints, or a managed API planned? Complete multimodal vLLM or SGLang API examples would also be very helpful.
If some results or deployment details are not suitable for public discussion, we would be happy to communicate privately at jiaxi@cerul.ai.
If there is a good fit, we would also be very interested in exploring potential collaboration around product-oriented video understanding and retrieval.
Thank you again for the great work.