Skip to content

Questions about video retrieval, deployment, and potential collaboration #2

Description

@JessyTsui

Hi WeMM-Embedding team — congratulations on the release, and thank you for open-sourcing the models.

We are building Cerul.ai, a product focused on video understanding and video retrieval. The video results of WeMM-Embedding are especially interesting to us.

We would love to learn more about:

  1. Compared with Qwen3-VL-Embedding and Gemini Embedding 2, which video retrieval capabilities improve most significantly—for example, fine-grained actions, temporal localization, motion understanding, OCR, long-video retrieval, or compositional queries?

  2. How does WeMM-Embedding perform in latency- and throughput-sensitive scenarios? Is near-real-time video indexing practical, and how does its efficiency compare with Microsoft’s MAGE-VL?

  3. Are official Apple Silicon/MLX support, quantized checkpoints, or a managed API planned? Complete multimodal vLLM or SGLang API examples would also be very helpful.

If some results or deployment details are not suitable for public discussion, we would be happy to communicate privately at jiaxi@cerul.ai.

If there is a good fit, we would also be very interested in exploring potential collaboration around product-oriented video understanding and retrieval.

Thank you again for the great work.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions