multivon-ai
Popular repositories Loading
-
multivon-eval
multivon-eval PublicPractical LLM evaluation for teams that ship to production. Deterministic + LLM-as-judge evaluators, dataset support, CI/CD integration.
Python 26
-
eval-framework-benchmark
eval-framework-benchmark PublicReproducible head-to-head benchmark of multivon-eval, DeepEval, and RAGAS on hallucination detection. Same judge, same dataset, same seed.
Python
-
eval-action
eval-action PublicGitHub Action wrapper for multivon-eval — runs LLM eval suites on PRs, posts diff comments, gates merges on regressions or safety-class failures.
Python
-
pdfhell
pdfhell PublicAdversarial PDFs that break AI document readers. Procedural ground truth, not LLM-as-judge.
Python
-
multivon-mcp
multivon-mcp PublicMCP server exposing multivon-eval + pdfhell as agent-callable tools. Drop into Claude Desktop, Cursor, Cline.
Python
Repositories
- multivon-eval Public
Practical LLM evaluation for teams that ship to production. Deterministic + LLM-as-judge evaluators, dataset support, CI/CD integration.
- multivon-mcp Public
MCP server exposing multivon-eval + pdfhell as agent-callable tools. Drop into Claude Desktop, Cursor, Cline.
- eval-action Public
GitHub Action wrapper for multivon-eval — runs LLM eval suites on PRs, posts diff comments, gates merges on regressions or safety-class failures.
- pdfhell Public
Adversarial PDFs that break AI document readers. Procedural ground truth, not LLM-as-judge.
- eval-framework-benchmark Public
Reproducible head-to-head benchmark of multivon-eval, DeepEval, and RAGAS on hallucination detection. Same judge, same dataset, same seed.
People
This organization has no public members. You must be a member to see who’s a part of this organization.
Top languages
Loading…
Most used topics
Loading…