Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 6 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,12 @@ and this project uses [Semantic Versioning](https://semver.org/).

## [Unreleased]

### Added
- `downshift report` shows the one-time analysis cost (eval model calls plus estimated judge
calls, at config prices) and the payback time. New `--audit-cost USD` option for an
assistant audit. `downshift export` adds the same numbers to `summary.json`.
- Web Overview shows the analysis cost and payback time.

## [0.1.0] - 2026-09-26

First public release.
Expand Down
Binary file added bob_sessions/downshift_task12_payback_a.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file added bob_sessions/downshift_task12_payback_b.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file added bob_sessions/downshift_task12_payback_c.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file added bob_sessions/downshift_task12_payback_d.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file added bob_sessions/downshift_task12_payback_e.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file added bob_sessions/downshift_task12_payback_f.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file added bob_sessions/downshift_task12_payback_g.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file added bob_sessions/downshift_task12_payback_h.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file added bob_sessions/downshift_task12_payback_i.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file added bob_sessions/downshift_task13_review_a.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file added bob_sessions/downshift_task13_review_b.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file added bob_sessions/downshift_task13_review_c.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file added bob_sessions/downshift_task13_review_d.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file added bob_sessions/downshift_task13_review_e.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file added bob_sessions/downshift_task13_review_f.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file added bob_sessions/downshift_task13_review_g.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
377 changes: 377 additions & 0 deletions docs/specs/audit-payback-plan.md

Large diffs are not rendered by default.

91 changes: 91 additions & 0 deletions docs/specs/audit-payback.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,91 @@
# Spec: analysis cost and payback in `downshift report`

## Why
The report shows monthly savings but not what it cost to find them. Add a one-time
"analysis cost" (eval model calls + judge calls, priced from the config) and a payback
time, so users can see the tool pays for itself.

## Scope (files)
- NEW `src/downshift/payback.py`
- EDIT `src/downshift/report.py` (Report field, build_report, render_markdown)
- EDIT `src/downshift/cli.py` (`report` gets `--audit-cost`)
- EDIT `src/downshift/export.py` (summary gets the new fields)
- NEW `tests/unit/test_payback.py`; extend existing report, export and CLI tests
- Do NOT edit the report snapshot file. The user regenerates it.

## payback.py
Constants:
- `JUDGE_EXTRA_PROMPT_TOKENS = 150` (rubric + instructions around the case)
- `JUDGE_COMPLETION_TOKENS = 200` (matches the judge's max_tokens default in scorer.py)

Dataclasses (frozen):
- `AnalysisCost`: `model_calls: int`, `model_cost: float`, `judge_calls: int`,
`judge_cost: float`, `judge_model: str | None`, `judge_priced_as: str | None`,
`audit_cost: float` (default 0.0). Property `total` = model_cost + judge_cost + audit_cost.
- `Payback`: `hours: float | None` (None when monthly savings <= 0).

Functions:
- `analysis_cost(sites, results_dir, config, *, audit_cost=0.0) -> AnalysisCost`
- For every call site and every model with a results file, read rows with the EXISTING
results loader in `runner.py` (last row per case wins). Do not write a new JSONL parser.
- Model cost: for each row, `decide.call_cost(price, row.prompt_tokens, row.completion_tokens)`
with `config.price_for(row_model)`. Count one call per row. Skip models with no price
(do not crash), and skip rows with an error.
- Judge cost: for each row with `judge_model` set, estimate
prompt = row.prompt_tokens + row.completion_tokens + JUDGE_EXTRA_PROMPT_TOKENS,
completion = JUDGE_COMPLETION_TOKENS. Price with the judge model's price if it is in
`config.pricing`, otherwise with the baseline model's price, and set `judge_priced_as`
to the model whose price was used.
- `audit_cost` is a user-supplied one-time amount in USD (e.g. what an AI-assistant audit
cost). Must be >= 0.
- `payback(total_one_time: float, monthly_savings: float) -> Payback`
- hours = total / (monthly_savings / cost.HOURS_PER_MONTH). None if monthly_savings <= 0.
- `format_payback(p: Payback) -> str`
- None -> "no payback (no projected savings)"
- < 1 hour -> "N minutes" (round up, minimum 1)
- < 48 hours -> "X.Y hours" (one decimal)
- otherwise -> "N days" (round up)

## report.py
- `Report` gets `analysis: AnalysisCost | None` and `payback: Payback | None`.
- `build_report` computes both from the same results folder and config it already uses,
and accepts `audit_cost: float = 0.0`.
- `render_markdown` adds this section right after the Summary section:

```
## What this analysis cost

| | One-time cost |
|---|---:|
| Eval model calls (N) | $X |
| Judge calls, estimated (M) | $Y |
| Assistant audit | $Z | <- only when audit_cost > 0
| **Total** | **$T** |

Pays back in **<format_payback>** of projected savings.

> Priced at the same illustrative prices as the rest of the report. Judge tokens are not
> recorded, so each judge call is estimated as (case prompt + output + 150) tokens in and
> 200 out, priced as `<judge_priced_as>`. Retries and warm-up calls are not counted.
> Local Ollama runs cost $0 in practice.
```
- Money uses the existing `_fmt_money`. With M = 0, drop the judge row and the judge sentence.

## cli.py
- `downshift report --audit-cost USD` (float, min 0, default 0). Help: "One-time cost of an
assistant audit, in USD, added to the analysis cost."

## export.py
- The summary JSON gets `analysis_cost_total`, `analysis_model_cost`, `analysis_judge_cost`,
`analysis_calls` (model + judge), `payback_hours` (float or null).

## Tests
- payback math with hand-computed values; format_payback for None, 0.2 h, 1.04 h, 47.9 h, 50 h.
- analysis_cost with a tmp results folder: two models, one judge-graded site, one row with an
error (skipped), a judge model with no price (priced as baseline), a model with no price (skipped).
- render_markdown: section present, judge row absent when there are no judge calls, audit row
only when audit_cost > 0.
- CLI: `--audit-cost 0.5` shows the audit row; negative value is rejected.
- export: new keys present.
- Follow repo standards: ruff (E,F,I,B,UP,SIM, line length 100), mypy clean, `zip(strict=True)`,
no network, FakeLLMClient not needed.
12 changes: 12 additions & 0 deletions examples/supportdesk/downshift.report.md
Original file line number Diff line number Diff line change
Expand Up @@ -15,6 +15,18 @@ Downgraded **3 of 8** call sites.

Rule: a cheaper model must keep at least 95% of the baseline pass rate and pass at least 80% of cases on its own. Decisions use pass rate, not mean score.

## What this analysis cost

| | One-time cost |
|---|---:|
| Eval model calls (724) | $0.15 |
| Judge calls, estimated (176) | $0.50 |
| **Total** | **$0.65** |

Pays back in **51 minutes** of projected savings.

> Priced at the same illustrative prices as the rest of the report. Judge tokens are not recorded, so each judge call is estimated as (case prompt + output + 150) tokens in and 200 out, priced as `qwen2.5:7b`. Retries and warm-up calls are not counted. Local Ollama runs cost $0 in practice.

## Decisions

| Call site | Grading | Decision | Model | Pass rate | Before / month | After / month |
Expand Down
7 changes: 7 additions & 0 deletions src/downshift/cli.py
Original file line number Diff line number Diff line change
Expand Up @@ -690,6 +690,12 @@ def report(
help="Minimum pass rate override [0,1]. Default: from config.",
),
out: Path | None = typer.Option(None, "--out", help="Write Markdown to this file."),
audit_cost: float = typer.Option(
0.0,
"--audit-cost",
min=0.0,
help="One-time cost of an assistant audit, in USD, added to the analysis cost.",
),
) -> None:
"""Render the cost and quality report."""
if threshold is not None and threshold <= 0:
Expand All @@ -713,6 +719,7 @@ def report(
results_dir,
threshold=threshold,
min_pass_rate=min_pass_rate,
audit_cost=audit_cost,
)
except ReportError as exc:
_fail(str(exc))
Expand Down
15 changes: 15 additions & 0 deletions src/downshift/export.py
Original file line number Diff line number Diff line change
Expand Up @@ -150,6 +150,21 @@ def summary_payload(report: Report, config: Config, rows: Rows, *, project: str)
},
"pricing": pricing,
"disclaimer": DISCLAIMER,
"analysis_cost_total": (
_round(report.analysis.total, 4) if report.analysis is not None else None
),
"analysis_model_cost": (
_round(report.analysis.model_cost, 4) if report.analysis is not None else None
),
"analysis_judge_cost": (
_round(report.analysis.judge_cost, 4) if report.analysis is not None else None
),
"analysis_calls": (
(report.analysis.model_calls + report.analysis.judge_calls)
if report.analysis is not None
else None
),
"payback_hours": (_round(report.payback.hours, 2) if report.payback is not None else None),
}


Expand Down
172 changes: 172 additions & 0 deletions src/downshift/payback.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,172 @@
"""Analysis cost and payback for one downshift report run.

`analysis_cost` prices the eval-model calls and judge calls that were made
to produce a downshift report. `payback` converts a one-time cost plus a
monthly savings figure into a payback period. `format_payback` renders the
period as a human-readable string.
"""

from __future__ import annotations

import math
from collections import Counter
from collections.abc import Sequence
from dataclasses import dataclass
from pathlib import Path

from downshift.config import Config
from downshift.cost import HOURS_PER_MONTH
from downshift.decide import call_cost
from downshift.runner import load_results, results_path
from downshift.schema import CallSite

JUDGE_EXTRA_PROMPT_TOKENS = 150
JUDGE_COMPLETION_TOKENS = 200


@dataclass(frozen=True)
class AnalysisCost:
"""One-time cost of the evaluation run that produced a downshift report."""

model_calls: int
model_cost: float
judge_calls: int
judge_cost: float
judge_model: str | None # most-frequent judge model seen in results
judge_priced_as: str | None # model whose price was used for judge calls
audit_cost: float = 0.0

@property
def total(self) -> float:
return self.model_cost + self.judge_cost + self.audit_cost


@dataclass(frozen=True)
class Payback:
"""Payback period expressed in hours (None when monthly savings <= 0)."""

hours: float | None


def analysis_cost(
sites: Sequence[CallSite],
results_dir: Path,
config: Config,
*,
audit_cost: float = 0.0,
) -> AnalysisCost:
"""Compute the one-time cost of running evals for *sites*.

For each call site × each model (baseline + candidates), the last result
row per case is used (that is how :func:`runner.load_results` works). Rows
with ``error`` set are skipped. Models without a price entry in
``config.pricing`` are also skipped silently.

The judge model is the most frequent ``judge_model`` value seen across all
non-error rows; ties are broken alphabetically. All judge rows are priced
with that model's price, or the baseline model's price when the judge model
has no price entry.

*audit_cost* is a user-supplied one-time amount in USD (e.g. the cost of an
AI-assisted audit session). It must be >= 0.
"""
if audit_cost < 0:
raise ValueError(f"audit_cost must be >= 0, got {audit_cost}")

models = list(config.models.all_models)
pricing = config.pricing
baseline = config.models.baseline

model_calls = 0
model_cost_total = 0.0
judge_calls = 0
judge_cost_total = 0.0
judge_model_counter: Counter[str] = Counter()

# First pass: accumulate model costs and count judge models.
# We also collect judge rows to price in a second pass once the dominant
# judge model is known.
judge_rows: list[tuple[int, int, str]] = [] # (prompt_tok, compl_tok, judge_model_name)

for site in sites:
for model in models:
path = results_path(results_dir, site.id, model)
if not path.is_file():
continue
rows = load_results(path)
for row in rows.values():
if row.error is not None:
continue
# Model call cost
price = pricing.get(row.model)
if price is not None:
model_cost_total += call_cost(price, row.prompt_tokens, row.completion_tokens)
model_calls += 1
# Judge call accounting
if row.judge_model is not None:
judge_model_counter[row.judge_model] += 1
prompt_est = (
row.prompt_tokens + row.completion_tokens + JUDGE_EXTRA_PROMPT_TOKENS
)
judge_rows.append((prompt_est, JUDGE_COMPLETION_TOKENS, row.judge_model))

# Resolve dominant judge model (most frequent; ties → alphabetically first).
dominant_judge: str | None = None
judge_priced_as: str | None = None
if judge_model_counter:
dominant_judge = min(
judge_model_counter,
key=lambda m: (-judge_model_counter[m], m),
)
# Determine which model's price to use.
if pricing.get(dominant_judge) is not None:
judge_priced_as = dominant_judge
elif pricing.get(baseline) is not None:
judge_priced_as = baseline
# else no price available at all; judge cost stays 0

if judge_priced_as is not None:
judge_price = pricing[judge_priced_as]
for prompt_est, compl_est, _jm in judge_rows:
judge_cost_total += call_cost(judge_price, prompt_est, compl_est)
judge_calls += 1

return AnalysisCost(
model_calls=model_calls,
model_cost=model_cost_total,
judge_calls=judge_calls,
judge_cost=judge_cost_total,
judge_model=dominant_judge,
judge_priced_as=judge_priced_as,
audit_cost=audit_cost,
)


def payback(total_one_time: float, monthly_savings: float) -> Payback:
"""Return the payback period for a one-time cost given a monthly saving.

Returns ``Payback(hours=None)`` when *monthly_savings* <= 0.
"""
if monthly_savings <= 0:
return Payback(hours=None)
hours = total_one_time / (monthly_savings / HOURS_PER_MONTH)
return Payback(hours=hours)


def format_payback(p: Payback) -> str:
"""Render a :class:`Payback` as a human-readable string.

- ``None`` → ``"no payback (no projected savings)"``
- < 1 hour → ``"N minutes"`` (ceil, minimum 1)
- < 48 hours → ``"X.Y hours"`` (one decimal)
- >= 48 hours → ``"N days"`` (ceil)
"""
if p.hours is None:
return "no payback (no projected savings)"
if p.hours < 1.0:
minutes = max(1, math.ceil(p.hours * 60))
return f"{minutes} minute" if minutes == 1 else f"{minutes} minutes"
if p.hours < 48.0:
return f"{p.hours:.1f} hours"
days = math.ceil(p.hours / 24)
return f"{days} days"
Loading
Loading