Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
18 changes: 9 additions & 9 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -12,16 +12,16 @@ jobs:
contents: read

steps:
- uses: actions/checkout@v5
- uses: actions/checkout@v6.0.2

- name: Setup Python 3.13 (with pip cache)
uses: actions/setup-python@v6
uses: actions/setup-python@v6.2.0
with:
python-version: "3.13"
cache: pip # ativa cache de dependencies pip :contentReference[oaicite:1]{index=1}

- name: Cache pre-commit environment
uses: actions/cache@v4
uses: actions/cache@v5.0.4
id: precommit-cache
with:
path: ~/.cache/pre-commit/
Expand All @@ -43,9 +43,9 @@ jobs:
contents: read
pull-requests: write
steps:
- uses: actions/checkout@v5
- uses: actions/checkout@v6.0.2

- uses: actions/setup-python@v6
- uses: actions/setup-python@v6.2.0
with:
python-version: "3.13"
cache: pip
Expand All @@ -62,7 +62,7 @@ jobs:
pytest tests/unit --cov=src --cov-report=xml --cov-fail-under=70

- name: Upload coverage report
uses: actions/upload-artifact@v4
uses: actions/upload-artifact@v7.0.1
with:
name: coverage-report
path: coverage.xml
Expand All @@ -81,8 +81,8 @@ jobs:
contents: read
pull-requests: write
steps:
- uses: actions/checkout@v5
- uses: actions/setup-python@v6
- uses: actions/checkout@v6.0.2
- uses: actions/setup-python@v6.2.0
with:
python-version: "3.13"
cache: pip
Expand All @@ -99,7 +99,7 @@ jobs:

- name: Add label for dependencies
id: fetch-metadata
uses: dependabot/fetch-metadata@v2
uses: dependabot/fetch-metadata@v3.0.0

- name: Label dependabot PR
if: steps.fetch-metadata.outputs.dependency-type == 'direct:production'
Expand Down
10 changes: 5 additions & 5 deletions .github/workflows/release.yml
Original file line number Diff line number Diff line change
Expand Up @@ -14,7 +14,7 @@ jobs:

steps:
- name: Checkout code
uses: actions/checkout@v5
uses: actions/checkout@v6.0.2
with:
fetch-depth: 0 # Required for setuptools_scm

Expand Down Expand Up @@ -45,9 +45,9 @@ jobs:
git push origin ${{ steps.get_next_version.outputs.NEXT_VERSION }}

- name: Set up Python
uses: actions/setup-python@v6
uses: actions/setup-python@v6.2.0
with:
python-version: "3.9"
python-version: "3.13"

- name: Install build dependencies
run: |
Expand All @@ -58,15 +58,15 @@ jobs:
run: python -m build

- name: Create GitHub Release
uses: ncipollo/release-action@v1
uses: ncipollo/release-action@v1.21.0
with:
tag: ${{ steps.get_next_version.outputs.NEXT_VERSION }}
name: Release ${{ steps.get_next_version.outputs.NEXT_VERSION }}
generateReleaseNotes: true
artifacts: "dist/*"

- name: Publish to PyPI
uses: pypa/gh-action-pypi-publish@v1.13.0
uses: pypa/gh-action-pypi-publish@v1.14.0

# - name: Set up Conda
# uses: conda-incubator/setup-miniconda@v3
Expand Down
14 changes: 14 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -17,6 +17,20 @@
AIUnitTest is a command-line tool that reads your `pyproject.toml` and test coverage
data (`.coverage`) to generate and update missing Python unit tests using AI.

## Status

The current public CLI and README still describe the v1 workflow.

AIUnitTest v2 is now being defined as a redesign around a tool-first test execution layer for coding agents.
It will ship first through its own CLI client and later expose the same core to external agents.
The product direction is to become the specialized testing layer that works with coding agents,
instead of competing with them as a general-purpose agent.
The v2 planning and architecture docs live in:

- `docs/v2/README.md`
- `docs/v2/architecture.md`
- `docs/v2/implementation-plan.md`

## How it Works

1. **Coverage Analysis**: The tool uses `coverage.py` to identify lines of code
Expand Down
3 changes: 3 additions & 0 deletions docs/initial.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,8 @@
# Initial Project Documentation

This document describes the original v1 thesis of the project.
For the current v2 redesign, see the documents under `docs/v2/`.

This section describes the initial idea, objectives, and architectural overview of the AIUnitTest project.

1. Objective
Expand Down
187 changes: 187 additions & 0 deletions docs/v2/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,187 @@
# AIUnitTest v2

AIUnitTest v2 is a redesign around a different product category.

v1 is a direct AI test generator driven by missing coverage.
v2 is a tool-first test execution layer for coding agents.

It should become the specialized testing layer that works with coding agents,
instead of competing with them as a general-purpose agent.

The goal is not to make the current generator slightly better.
The goal is to turn a testing target into a validated patch with explicit control
over target selection, context assembly, validator execution, retry policy,
and reporting.

## Official positioning

AIUnitTest v2 should be understood as:

- a test-focused orchestration engine
- a local CLI client built on top of that engine
- a future MCP-friendly tool surface for external agents
- a backend-agnostic system for reasoning
- a validation-first workflow that aims to produce a patch, not just generated text

AIUnitTest v2 is not:

- a generic coding agent
- a thin wrapper around a single model or provider
- a VS Code extension pretending to be the product
- a PR review bot for every possible concern
- a replacement for Copilot CLI, Gemini CLI, or similar agentic runtimes

## Delivery principle

The architecture is hybrid.
The delivery should be sequential.

That means:

- build the core engine first
- ship the first usable workflow through the CLI
- add CI and PR delivery after the local loop is reliable
- expose MCP only after the core contracts are stable

This avoids a common failure mode where "hybrid" is interpreted as
"build every surface at the same time."

## Product promise

Given one of these targets:

- uncovered code
- a git diff
- an explicit file or symbol
- a failing test

AIUnitTest v2 should:

1. decide what to fix first
2. assemble only the context needed for that target
3. request a patch proposal from a reasoning backend or external agent
4. apply the patch under explicit guardrails
5. run validators and collect structured feedback
6. retry when the first attempt fails
7. emit a report that can be reviewed locally or in CI

## Why this shape exists

Modern coding agents are getting better at reasoning, but that does not remove
the testing problem. It shifts the product boundary.

The scarce part is no longer raw code generation.
The scarce part is test-specific orchestration:

- selecting the right target
- choosing the minimum useful context
- enforcing validation
- coordinating retries
- producing audit-friendly diffs and reports

That is the part AIUnitTest v2 should own.

## Product surfaces

### Core engine

The core engine owns:

- target selection
- context building
- patch application
- validation
- retry policy
- report generation

This is the real product.

### CLI client

The CLI is the first shipping surface because it is the fastest way to prove value in:

- local development
- benchmark runs
- CI and headless execution
- demos

The CLI should call the core engine. It should not contain the core logic.

### Future MCP surface

The future MCP surface should expose stable, test-specific capabilities from the same core.
That lets Copilot, Gemini, Claude, or custom agents use AIUnitTest as a testing tool
instead of forcing AIUnitTest to compete as a general agent.

### CI and PR reporting

CI and PR integrations should reuse the same run artifacts and reporters produced by the core.
They are delivery surfaces, not separate products.

## Reasoning strategy

AIUnitTest v2 keeps reasoning backend-agnostic.

Initial backend targets:

- Copilot CLI
- Gemini CLI

Possible future backends:

- OpenAI API
- local models
- MCP-mediated or SDK-backed runtimes

This means AIUnitTest does not try to out-think the strongest general agent.
It delegates general reasoning and keeps ownership of test-focused execution.

## Risks and guardrails

The v2 strategy only works if the product enforces guardrails instead of trusting raw model output.

The first cut should treat these as non-negotiable:

- bounded retry limits
- test-file-first writes by default
- explicit opt-in for source edits
- persisted diffs and validator artifacts for every run
- no automatic commit, push, or merge behavior in the core workflow
- validation summaries that make failures reviewable instead of opaque

## First useful scope

The first useful scope for v2 is intentionally narrow:

- explicit file targeting
- one backend adapter
- one patch application path
- targeted pytest validation
- bounded retry loop
- terminal and JSON reporting

That is enough to prove the new thesis without widening into PR bots,
editor plugins, or generic review automation.

## Docs in this folder

- `architecture.md` defines the concrete layer split and interfaces
- `implementation-plan.md` defines the bounded MVP and expansion path

## Current status

The v2 MVP is implemented and tested. The following components are functional:

- **Core models:** `RunRequest`, `TargetSpec`, `ContextBundle`, `PatchCandidate`,
`PatchApplication`, `ValidationResult`, `RunReport`
- **Target selection:** `ExplicitFileSelector` for explicit file targeting
- **Context building:** `FileContextBuilder` reads source, related tests, and project config
- **Backend adapters:** `CopilotCliBackend` and `GeminiCliBackend` (subprocess-based, JSON-first parsing)
- **Patch application:** `PatchApplier` with test-file-first guardrails, rollback between retries
- **Validation:** `SyntaxValidator` (py_compile) and `PytestValidator` (subprocess, overrides global addopts)
- **Feedback:** `FeedbackSummarizer` for structured retry context
- **Reporting:** `RunStore` (persists report.json, summary.md, patch.diff), `TerminalRenderer`, `JsonRenderer`
- **Orchestrator:** `V2Orchestrator` wires the full loop with bounded retries
- **CLI:** `v2 run`, `v2 report`, `v2 backends`, `v2 doctor` registered under the `v2` namespace

Work stays isolated from v1. The v1 CLI remains fully functional.
Loading
Loading