This synthetic-data observatory turns AI interaction logs into inspectable session summaries of verification, revision, reflection, and answer adoption. Its intent labels use explicit lexical rules, and its agency index is a stated weighted heuristic rather than a psychological measure. The repository is useful for testing instrumentation and analysis workflows before collecting consented learner data.
Review scope: The existing suite requires unavailable dependencies; no full-suite pass is claimed. The bundled demonstration executed successfully in this review.
Category: AI in Education Observability for how learners actually use generative AI, not just whether they used it.
Research prototype. All bundled data and results are synthetic demonstrations. Nothing in this repository should be interpreted as evidence about real learners, teachers, or institutions.
Most dashboards reduce GenAI use to counts or time-on-task. This project treats learner–AI interaction as a process: what the learner asked for, whether they checked the answer, whether they revised it, and whether the interaction ended in reflection or simple adoption.
The implementation keeps interaction events, feature extraction, aggregation, and reporting separate so each learner-agency signal can be traced back to the underlying synthetic events. That makes it easier to challenge the assumptions behind a metric instead of treating the dashboard as a black box.
- How can we distinguish productive help-seeking from passive answer adoption?
- Which interaction patterns signal learner agency, verification, and reflection?
- How should an observability layer represent uncertainty without turning exploratory analytics into high-stakes labels?
The reference pipeline follows five stages:
- Event ingestion
- Intent tagging
- Behavioral feature extraction
- Session aggregation
- Agency-oriented reporting
The baseline is intentionally lightweight so the full pipeline can be audited before introducing real interaction logs, richer models, or institutional data.
verification_raterevision_ratereflection_rateanswer_adoption_rateintent_diversityagency_index
The dashboard above is generated from synthetic data and is included only to show what the analysis surface looks like. It is not a reported empirical result.
python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -e .[dev]
python examples/demo.py
pytest -qYou can also use Docker:
docker build -t genai_learning_observatory .
docker run --rm genai_learning_observatorygenai_learning_observatory/
├── src/genai_learning_observatory/ # core implementation and synthetic-data generator
├── examples/demo.py # end-to-end reproducible demo
├── tests/ # executable unit tests
├── docs/ # research design, data dictionary, references
│ └── images/ # original project diagrams and demo visualisations
├── results/ # synthetic demo outputs only
├── config/default.yaml
├── Dockerfile
├── Makefile
└── pyproject.toml
The fuller design rationale is in docs/research_design.md, including constructs, assumptions, validation steps, and a proposed empirical extension.
- Synthetic generation uses a fixed random seed.
- The core metrics are implemented as small, testable functions.
- The demo writes machine-readable results into
results/. - CI runs the tests on every push and pull request.
- No API keys, proprietary datasets, or external model calls are required for the baseline.
- Intent tagging is deliberately transparent and lightweight; it is not a validated psychological classifier.
- The included data are synthetic and cannot support claims about real learners.
- Agency-related features are descriptive research signals, not diagnostic labels or grading criteria.
- Replace rule-based intent tagging with a validated transformer classifier and report calibration.
- Add sequence models for interaction trajectories rather than only session aggregates.
- Run a learner/teacher co-design study to test whether the dashboard supports useful reflection.
See docs/references.md. The references are there to locate the project in current AIED, learning-analytics, human-centered AI, and instructional-design research. They do not imply endorsement or affiliation.
If you build on this research prototype, use the metadata in CITATION.cff.
MIT for the code in this repository. Research data from future studies should use a separate data-governance and consent process.