This fork applies CD4AI (Continuous Delivery for Artificial Intelligence) as a practical example for a MAC0499 capstone project. Its goal is to use problems found during real usage to improve regression tests for an AI system.
The application is Divination, a D&D assistant that uses retrieval-augmented generation (RAG). It was originally developed by Luis Carlos; see his original monograph. This fork adds testing, monitoring and curation to explore the CD4AI cycle.
| Stage | What this project does |
|---|---|
| Testing | Runs code checks and chatbot evaluations to detect regressions. Approval follows the criteria of each test. |
| Monitoring | Records interactions and user feedback, and flags possible failures with lightweight detectors. |
| Curation | Supports human review to distinguish real defects from noise. Developers manually turn confirmed defects into new regression cases. |
To close the cycle, a developer defines the expected behavior for a confirmed problem, adds a case to the evaluation dataset and runs the evaluations alongside the fix. That case becomes part of future regression runs. See the manual curation guide.
With Git, Python 3.10+, Poetry 1.6.1 and Make installed on Linux, macOS or WSL, run from the repository root:
make ci-setup
make checkmake check validates the dependency lock, runs Ruff and executes the CI runner
and monitoring tests. It needs no API keys and makes no paid API calls.
For chatbot evaluations, first configure backend/.env from
backend/.env.sample with your API keys:
| Command | Evaluation |
|---|---|
make eval |
Individual answers |
make eval-conversations |
Conversations with multiple turns |
make eval-conversations-if-changed |
Conversations, skipping the commit if it is already recorded as successfully evaluated |
These evaluations use paid APIs. The commands work independently of GitHub;
GitHub Actions runs the same commands for PRs and scheduled evaluations.
An optional pre-push hook runs the checks without paid evaluations.