Skip to content
@multivon-ai

multivon-ai

Popular repositories Loading

  1. multivon-eval multivon-eval Public

    Practical LLM evaluation for teams that ship to production. Deterministic + LLM-as-judge evaluators, dataset support, CI/CD integration.

    Python 26

  2. eval-framework-benchmark eval-framework-benchmark Public

    Reproducible head-to-head benchmark of multivon-eval, DeepEval, and RAGAS on hallucination detection. Same judge, same dataset, same seed.

    Python

  3. eval-action eval-action Public

    GitHub Action wrapper for multivon-eval — runs LLM eval suites on PRs, posts diff comments, gates merges on regressions or safety-class failures.

    Python

  4. pdfhell pdfhell Public

    Adversarial PDFs that break AI document readers. Procedural ground truth, not LLM-as-judge.

    Python

  5. multivon-mcp multivon-mcp Public

    MCP server exposing multivon-eval + pdfhell as agent-callable tools. Drop into Claude Desktop, Cursor, Cline.

    Python

Repositories

Showing 5 of 5 repositories

People

This organization has no public members. You must be a member to see who’s a part of this organization.

Top languages

Loading…

Most used topics

Loading…