Skip to content
#

evaluation-pipeline

Here are 11 public repositories matching this topic...

Capability Schema Spec defines a shared semantic language for world model evaluation. Standardize capability definition, observation, and verification across models and benchmarks. Not a benchmark—a shared language. Define • Observe • Verify

  • Updated Jul 3, 2026
  • Python

🏥 Modular Agentic RAG Medical Assistant 🤖 built with LangGraph, Pinecone, Groq & Tavily. It intelligently routes queries between RAG, live web search, and direct answers. 🔬 Live tracing shows every agent step, while FastAPI + Streamlit deliver a production-ready full-stack experience. 🚀

  • Updated Jul 10, 2026
  • Python

🤖 An advanced, policy-compliant RAG-powered Internal Knowledge Assistant for TechNova. Built with LangChain, Chroma DB, and Llama 3.3 (Groq). Features strict source grounding, 100% retrieval precision, comprehensive factual citations, automated evaluation pipelines, and a dynamic multi-category Gradio chat interface. 🚀

  • Updated Jun 14, 2026
  • Jupyter Notebook

Add this topic to your repo

To associate your repository with the evaluation-pipeline topic, visit your repo's landing page and select "manage topics."

Learn more