Skip to content
 
 

Latest commit

 

History

89 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Divination — CD4AI case study

This fork applies CD4AI (Continuous Delivery for Artificial Intelligence) as a practical example for a MAC0499 capstone project. Its goal is to use problems found during real usage to improve regression tests for an AI system.

The application is Divination, a D&D assistant that uses retrieval-augmented generation (RAG). It was originally developed by Luis Carlos; see his original monograph. This fork adds testing, monitoring and curation to explore the CD4AI cycle.

How CD4AI works here

Stage What this project does
Testing Runs code checks and chatbot evaluations to detect regressions. Approval follows the criteria of each test.
Monitoring Records interactions and user feedback, and flags possible failures with lightweight detectors.
Curation Supports human review to distinguish real defects from noise. Developers manually turn confirmed defects into new regression cases.

To close the cycle, a developer defines the expected behavior for a confirmed problem, adds a case to the evaluation dataset and runs the evaluations alongside the fix. That case becomes part of future regression runs. See the manual curation guide.

Try the checks

With Git, Python 3.10+, Poetry 1.6.1 and Make installed on Linux, macOS or WSL, run from the repository root:

make ci-setup
make check

make check validates the dependency lock, runs Ruff and executes the CI runner and monitoring tests. It needs no API keys and makes no paid API calls.

For chatbot evaluations, first configure backend/.env from backend/.env.sample with your API keys:

Command Evaluation
make eval Individual answers
make eval-conversations Conversations with multiple turns
make eval-conversations-if-changed Conversations, skipping the commit if it is already recorded as successfully evaluated

These evaluations use paid APIs. The commands work independently of GitHub; GitHub Actions runs the same commands for PRs and scheduled evaluations. An optional pre-push hook runs the checks without paid evaluations.

Documentation

About

RAG assistant that answers Dungeons & Dragons rules questions straight from the official sourcebooks, built to minimise LLM hallucination. USP capstone project (MAC0499).

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages