AI annotation QA workflows for ranking, relevance scoring, factuality checks, and safety evaluation.
-
Updated
May 26, 2026 - Python
AI annotation QA workflows for ranking, relevance scoring, factuality checks, and safety evaluation.
Automated quality evaluation for RL agent discoveries using LLM as Judge, rule based validation, and closed loop reward feedback
Capability Schema Spec defines a shared semantic language for world model evaluation. Standardize capability definition, observation, and verification across models and benchmarks. Not a benchmark—a shared language. Define • Observe • Verify
LLM-as-a-Judge evaluation pipeline — score and compare LLM outputs with a separate judge LLM
🏥 Modular Agentic RAG Medical Assistant 🤖 built with LangGraph, Pinecone, Groq & Tavily. It intelligently routes queries between RAG, live web search, and direct answers. 🔬 Live tracing shows every agent step, while FastAPI + Streamlit deliver a production-ready full-stack experience. 🚀
Production-grade Safe GenAI Agent Orchestrator with intent routing, hallucination guard, tool orchestration, evaluation pipeline, and multi-screen Next.js dashboard.
End-to-end AI evals orchestration platform for comparing LLM outputs across providers with transcription, structured logging, human review, and Supabase-backed decision tracking.
AI Agent Validation Pipeline — agentic workflow tool that runs tasks through LLM, validates outputs via multi-rule harness, applies supervised handoff routing, and generates evaluation reports.
Production-ready automated evaluation pipeline for LLM prompts, prompt regressions, RAG systems, and autonomous agent workflows.
Generates prompt evaluation pipeline for an eCommerce website Agent.
🤖 An advanced, policy-compliant RAG-powered Internal Knowledge Assistant for TechNova. Built with LangChain, Chroma DB, and Llama 3.3 (Groq). Features strict source grounding, 100% retrieval precision, comprehensive factual citations, automated evaluation pipelines, and a dynamic multi-category Gradio chat interface. 🚀
To associate your repository with the evaluation-pipeline topic, visit your repo's landing page and select "manage topics."