Skip to content

Repository files navigation

falsify

Deterministic checks for backtested trading strategies, run before you take them live — like npm audit for overfitting.

falsify check trades.csv --params 8

No AI/LLM calls. No dashboard, no billing. Free forever.

Install

pip install falsify-backtest

Quick start

TradingView (most common)

Export your strategy's "List of Trades" from TradingView, then:

falsify check ~/Downloads/BTCUSD_Strategy_Trades.csv --from tradingview --params 6

No manual CSV conversion needed. The adapter pairs Entry/Exit rows, maps Profit → pnl, and extracts symbol/timeframe automatically.

Canonical CSV

If you already have a 4-column CSV (entry_time,exit_time,pnl,side):

falsify check trades.csv --params 3

Sample CSVs live in the repo, not in the installed package — grab one to try the checks without a backtest of your own:

curl -O https://raw.githubusercontent.com/zecalis/falsify/main/examples/known_overfit.csv
falsify check known_overfit.csv --params 12
# → FAIL — param_overfit_ratio below the 10:1 floor

examples/known_good.csv is the counterpart that passes all three checks (2 parameters, 200 trades).

Checks

Check What it does Fail threshold
sample_size Binomial test — is the win rate distinguishable from 50/50? n < 30 trades
deflated_sharpe Deflates Sharpe ratio for multiple testing (Bailey & López de Prado, 2014) DSR < 0.5
param_overfit_ratio Trades per free parameter — heuristic overfitting signal ratio < 10

The overall verdict is the worst of the three checks.

JSON output

falsify check trades.csv --params 8 --json

Run record

--json outputs a structured run record — one JSON object per invocation. Append results over time to build a history:

falsify check trades.csv --params 8 --json >> runs.jsonl

With a strategy name for ledger grouping:

falsify check trades.csv --params 8 --strategy ema-cross --json >> runs.jsonl

Migration from v1: v1's --json had no schema_version and a top-level n_trades. v2 nests it under input. Records without schema_version are v1 (flat shape) — readers should branch on its presence.

Schema (version 1):

Field Description
schema_version Run record schema version (1)
engine_version Engine version (falsify --version)
input.path CSV path as provided
input.sha256 SHA-256 of the raw CSV file
input.canonical_sha256 SHA-256 of normalized trade content (sorted, consistent float format)
input.n_trades Number of trades
declared.params Free parameters (user-provided)
declared.trials Strategy variants tried (user-provided)
declared.trials_was_default Whether --trials was left at default
dataset.symbol Passthrough: symbol if present, null if absent, "<multiple>" if ambiguous
dataset.timeframe Passthrough: timeframe if present, null if absent, "<multiple>" if ambiguous
dataset.strategy Strategy name (passthrough, engine ignores it, null if not provided)
dataset.date_range {first_entry, last_exit} computed from trade timestamps
verdict Overall: pass, warn, or fail
checks[] Per-check results with name, verdict, value, threshold, explanation, and inputs
created_at UTC timestamp (ISO-8601)

Example:

{
  "schema_version": 1,
  "engine_version": "0.1.0",
  "input": {
    "path": "examples/known_good.csv",
    "sha256": "a1b2c3...",
    "canonical_sha256": "552af7e9b7ae20154f8dd70c34aa049c8fd51830d2e3785659f1e814900ca651",
    "n_trades": 200
  },
  "declared": {
    "params": 2,
    "trials": 1,
    "trials_was_default": true
  },
  "dataset": {
    "symbol": null,
    "timeframe": null,
    "strategy": "ema-cross",
    "date_range": {
      "first_entry": "2024-01-01T09:00:00+00:00",
      "last_exit": "2024-06-28T16:00:00+00:00"
    }
  },
  "verdict": "pass",
  "checks": [
    {
      "name": "sample_size",
      "verdict": "pass",
      "value": 200,
      "threshold": 30,
      "explanation": "n=200 trades, win rate 62% statistically distinguishable from chance at p=0.00",
      "inputs": {"n_trades": 200, "n_wins": 124}
    }
  ],
  "created_at": "2026-09-10T12:00:00+00:00"
}

Agent ledger (trials you actually tried)

--trials defaults to 1, but nobody remembers how many variants they tried — especially when an AI edits the strategy for you. The agent-side ledger counts for you, in the place you already work (the chat agent). Base install, stdlib only, no network, no account:

falsify check trades.csv --params 3 --strategy ema-cross --json | falsify-agent log
# → ema-cross: กรอก --trials 1 · ledger นับได้ 7

falsify-agent status ema-cross
# ema-cross — 9 runs · 2026-09-01 → 2026-09-24
#   trials: กรอก --trials 1 · ledger นับได้ 7
#   verdict: fail ×8 · pass ×1 (run #9, หลังลอง 7 แบบ)

falsify-agent report ema-cross
# → /Users/you/.falsify/reports/ema-cross-9.html (single file, opens offline)

Observed trials = distinct backtests (by content hash) under one strategy name, counted over the whole ledger in ~/.falsify/runs.jsonl (override with FALSIFY_HOME). Rules that surprise people exactly once:

  • Re-checking the same file does not add a trial — correct, it is the same backtest, not a new attempt.
  • Re-exporting the file (CRLF, column order) does not add a trial either — only content counts (canonical_sha256).
  • No --strategy → the run is logged but counted nowhere ("ไม่รู้" beats guessing).
  • Names match exactly: Ema-Cross and ema-cross are different strategies.

Claude Desktop / Code / Cursor (MCP)

pip install falsify-backtest[mcp]
falsify-agent install --client claude-desktop   # or: claude-code | cursor

Three tools: check (strategy required — it deflates with observed trials automatically and shows declared vs deflated verdicts side by side), history, report. A chat session looks like this:

user: ช่วยแก้ strategy ให้ sharpe ดีขึ้น แล้ว check ให้หน่อย agent → check(csv=..., strategy="ema-cross", params=3, trials=1) ← verdict: fail · กรอก trials 1 · ledger นับได้ 14 · report: ~/.falsify/reports/ema-cross-14.html

False-negative corpus

examples/corpus/ holds 10 backtests with known ground truth — 6 bad strategies the engine must not pass, 3 good ones it must pass, and 1 documented blind spot (in-regime backtest v1 cannot see — needs held-out data, out of engine scope). CI fails on any false negative or false positive:

python -m pytest tests/test_corpus.py -s   # prints the scoreboard
python3 examples/corpus/generate.py         # regenerates the synthetic CSVs

Current score: 6/6 bad caught, 3/3 good pass, 1 tracked gap.

Development

pip install -e .
pip install pytest
python -m pytest tests/

License

MIT

About

Deterministic checks for backtested trading strategies, run before you take them live — like npm audit for overfitting.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages