Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 7 additions & 0 deletions .env.example
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,13 @@ KAP_RATE_LIMIT_RPS=2
KAP_TIMEOUT_S=10
KAP_USER_AGENT=trailingedge/0.1 (research)

# Optional proxy pool for the KAP backfill. The WAF throttles per source IP (~50 requests,
# then a block that refills in ~2 min), so a rotating pool of IPs removes the block stall:
# a spent IP is parked to refill while another serves at full speed. Comma-separated URLs,
# or put one per line in a gitignored proxies.txt. Residential proxies survive far better
# than datacenter ones here. Leave unset for a single direct connection (default).
# KAP_PROXIES=http://user:pass@host1:port,http://user:pass@host2:port

# TSG (Ticaret Sicil Gazetesi) - semi-automatic scraping.
# The scraper opens a visible browser; you sign in and solve the login
# CAPTCHA by hand once, then the run proceeds automatically. No credentials
Expand Down
3 changes: 3 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -65,3 +65,6 @@ ehthumbs.db
.python-version
pip-log.txt
pip-delete-this-directory.txt

# Proxy list for the KAP scraper rotation pool - a credential, never commit.
proxies.txt*
58 changes: 36 additions & 22 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -21,19 +21,33 @@ disclosure is *public*, returns are measured in excess of XU100 over the same he
interval, and the round-trip cost is estimated per trade from that stock's own OHLC
(Abdi-Ranaldo 2017) rather than assumed as a flat fee.

Full 2015-2026 history, **2,279 survivorship-clean insider-cluster events**:

| Horizon | N | Gross AR | Cost | **Net AR** | t (net) |
|---|---:|---:|---:|---:|---:|
| 5d | 1,070 | +0.66% | 3.37% | **−2.71%** | −12.91 |
| 20d | 1,070 | +2.02% | 3.37% | **−1.35%** | −3.31 |
| 60d | 1,071 | +2.18% | 3.37% | **−1.20%** | −1.86 |

**The signal is real. The gross abnormal return is significantly positive at every
horizon** (20d: +2.07%, t = 5.36, N = 1,079). **And it is not tradeable**, because insider
clusters fire in illiquid small caps whose bid-ask spread is wider than the alpha: the
median round trip costs 1.93%, the upper quartile 4.34%. Nothing survives crossing it
twice. At 60 days the net loss is no longer statistically distinguishable from zero
(t = −1.86) - which buys nothing: the point estimate is still negative, and "you might
merely break even after three months" is not an edge either.
| 5d | 2,279 | +0.41% | 4.19% | **−3.78%** | −21.29 |
| 20d | 2,279 | +1.58% | 4.19% | **−2.61%** | −7.76 |
| 60d | 2,236 | +2.61% | 4.19% | **−1.58%** | −2.57 |

**The pooled gross signal is real** (20d: +1.58%, t = 5.2, `EDGE_DETECTED`) **and not
tradeable** — insider clusters fire in illiquid small caps whose bid-ask spread (median
round trip 2.33%) is wider than the alpha. Net negative at every horizon.

**But the full history says something sharper than "not tradeable".** Split by regime, the
gross signal was strong in 2015-2018 (+2.36% at 20d, t = 6.5) and has **decayed to nothing
in 2021-2026** (−0.67%, t = −1.0 — indistinguishable from zero, before costs):

| Regime | 20d Gross AR | t (gross) | 20d Net AR |
|---|---:|---:|---:|
| 2015-2018 | +2.36% | 6.46 | −1.00% |
| 2019-2020 | +3.34% | 3.34 | −5.01% |
| **2021-2026** | **−0.67%** | **−1.00** | −5.44% |

So the pooled number is carried entirely by the early era. Insiders' disclosed purchases
predicted abnormal returns in 2015-2018 — returns you still could not capture after the
spread — and in the regime that matters to a trader today they no longer predict them at
all. The edge was real, uncapturable, and has since decayed. (2019-2020 is a COVID
small-cap-mania artefact on a tiny, extremely illiquid sample, not a strategy.)

That is the whole finding, and it is why this repository exists. A gross number is not
an edge; an edge is what is left after the market takes its cut.
Expand Down Expand Up @@ -95,12 +109,12 @@ verdict standing.
`INSUFFICIENT_POWER` below ~784 events and `SURVIVORSHIP_BIASED` when too many
clusters cannot be priced. Both gates fired during this work, and both were right.

> **What is claimed, precisely:** a statistically strong *gross* abnormal return
> (20d: +2.07%, t = 5.36, N = 1,079, survivorship-clean) that does **not** survive a
> per-trade cost estimate. The window is 2015-2018 - a single regime - so the result is
> not yet regime-conditional, and that is stated rather than glossed. Remaining gaps are
> in [`docs/METHODOLOGY.md`](docs/METHODOLOGY.md#6-still-open), not left for a
> reader to discover.
> **What is claimed, precisely:** across the full 2015-2026 history (N = 2,279,
> survivorship-clean), a *gross* abnormal return that does **not** survive a per-trade cost
> estimate at any horizon — and that, split by regime, was statistically strong in 2015-2018
> (20d +2.36%, t = 6.5) and has **decayed to zero in 2021-2026** (−0.67%, t = −1.0). The edge
> was real, uncapturable, and is now gone. Remaining gaps are in
> [`docs/METHODOLOGY.md`](docs/METHODOLOGY.md), not left for a reader to discover.

## Türkçe özet

Expand Down Expand Up @@ -173,13 +187,13 @@ python scripts/net_of_cost.py # the one that decides it
```
=== Abnormal return, NET of round-trip cost (order 25,000 TRY) ===
spread: Abdi-Ranaldo (2017) from the stock's own OHLC, per trade
dropped (no cost estimate): 27
round-trip cost: median 1.93% p25 1.19% p75 4.34%
dropped (no cost estimate): 36
round-trip cost: median 2.33% p25 1.34% p75 4.76%

HORIZON N GROSS AR% COST% NET AR% HIT% 95% CI t VERDICT
5d 1070 0.66 3.37 -2.71 26.9 [24.3, 29.7] -12.91 LOSES MONEY (net)
20d 1070 2.02 3.37 -1.35 41.3 [38.4, 44.3] -3.31 LOSES MONEY (net)
60d 1071 2.18 3.37 -1.20 44.4 [41.5, 47.4] -1.86 NO EDGE (net)
5d 2280 0.41 4.19 -3.78 26.1 [24.4, 28.0] -21.29 LOSES MONEY (net)
20d 2279 1.58 4.19 -2.61 39.8 [37.8, 41.8] -7.76 LOSES MONEY (net)
60d 2236 2.61 4.19 -1.58 42.4 [40.4, 44.5] -2.57 LOSES MONEY (net)
```

The spread is not a parameter. It is estimated for each trade from the 30 sessions of
Expand Down
111 changes: 88 additions & 23 deletions docs/METHODOLOGY.md
Original file line number Diff line number Diff line change
Expand Up @@ -168,18 +168,22 @@ have flattered the answer, and not taken from a quote feed, which does not exist
delisted names the exchange bulletin carries. Impact is Kyle/Almgren square-root on the
same window; commission and BSMV are charged per side.

round-trip cost: median 1.93% p25 1.19% p75 4.34%
On the full 2015-2026 sample (N=2,279 priceable clusters, more than double the earlier cut):

horizon N=1070 gross AR net AR t (net)
5d +0.66% -2.71% -12.91
20d +2.02% -1.35% -3.31
60d +2.18% -1.20% -1.86 (not significant)
round-trip cost: median 2.33% mean 4.19% p75 4.76%

horizon N=2279 gross AR net AR t (net)
5d +0.41% -3.78% -21.29
20d +1.58% -2.61% -7.76
60d +2.61% -1.58% -2.57

**The signal does not survive the cost of trading it.** Insider clusters fire in illiquid
small caps, and the spread on those names is wider than the alpha. This is the project's
result, not a caveat on it. At 60 days the net loss stops being statistically
distinguishable from zero, which is not a reprieve: the point estimate is still negative,
and a signal that *may* break even over three months is not an edge either.
result, not a caveat on it - and on the full sample it is sharper than on the earlier cut:
the net loss is significant at every horizon, including 60 days. The pooled gross number
(+1.58%, t=5.2, still "EDGE_DETECTED" before costs) is nonetheless misleading on its own,
because it averages two different regimes - see §6, where the split shows the gross signal
was real in 2015-2018 and has since decayed to zero.

A second trap, found later and more serious than the first, because it sat under the
number that decides the answer. `price_history.close_try` is a **chained total-return
Expand Down Expand Up @@ -264,25 +268,86 @@ Worth recording without over-reading: the *gross* alpha collapses from +2.30% to
between the halves. That could be the market becoming more efficient, or it could be
sampling noise at N = 197. It is not interpreted here, because at that N it cannot be.

## 6. Still open

**Regime.** The window is 2015-2018. That spans the August 2018 currency crisis but not
the 2021-2023 negative-real-rate retail boom or the 2023+ normalisation. The result is
therefore **not regime-conditional**, and a signal that dies to the spread in one regime
could in principle survive in another where those names traded tighter. The honest position
is that this is untested, not that it is unaffected.

Closing it needs the KAP backfill to reach 2026. That is a data-collection problem, not a
methodological one, and it does not touch the mechanism: the spread eating the alpha is a
microstructure fact about illiquid names, not a regime phenomenon.
## 6. Closed by the full backfill, and what still remains

### Regime - CLOSED, and the signal itself decayed

The window used to be 2015-2018. The KAP backfill now reaches the present (2026-07), so the
question - does the result hold across regimes? - can finally be answered on the full
2,279-cluster sample instead of asserted. It does not hold. Split by the cluster's own date:

regime 20d gross t(gross) 20d net note
2015-2018 +2.36% 6.46 -1.00% real, uncapturable (as before)
2019-2020 +3.34% 3.34 -5.01% COVID small-cap mania, N=198, cost 8.4%
2021-2026 -0.67% -1.00 -5.44% the GROSS signal is gone

The finding is sharper than "not tradeable". In 2015-2018 the gross abnormal return was
real and strong (+2.36% at 20d, t=6.5) and died only to the spread. In **2021-2026 the gross
signal has decayed to nothing** (-0.67%, t=-1.0 - statistically indistinguishable from zero,
and if anything negative). Insider-cluster purchases no longer predict a positive abnormal
return at all in the recent regime, before costs are even considered.

So the pooled full-sample number (+1.61% gross at 20d, t=5.2, EDGE_DETECTED) is carried
entirely by the 2015-2018 era and is misleading on its own - it averages a real early edge
with a dead recent one. The honest statement is: **the signal was a 2015-2018 phenomenon
that has since decayed**, plausibly because the 2021+ retail boom changed who follows insider
filings and how fast. The 2019-2020 line is a genuine but regime-specific artefact: the COVID
small-cap mania produced huge gross returns (+16.7% at 60d) on a tiny, extremely illiquid
sample (round-trip cost 8.4%), and it clears cost only there and only at 60 days - not a
strategy, a curiosity.

This closes the last major open item. It does not rescue the signal - it removes even the
"real but uncapturable" consolation for the regime that matters most to a trader today.

### The ingest - CLOSED at 139/139, and the completion metric was not the obvious one

The full 2015-2026 archive is now in: **139 of 139 months SUCCESS, 18,415 disclosures,
15,316 transactions, 2,339 clusters** (up from the 1,079 the earlier analysis ran on). Two
things had to be got right to get here.

First, the completion metric. A chunk can lose filings to a WAF disconnect and still log
`chunk_done`; the ledger records the month `PARTIAL`, not `SUCCESS`. Counting `chunk_done`
overstated progress badly - at one point it read 92 of 139 months while `scraper_runs` held
20 SUCCESS, 82 PARTIAL, 44 FAILED, and 2016 and 2020 contained not one complete month. The
real criterion is `count(SUCCESS)`, and reaching it needed repeated sweeps, because
`backfill_kap_insider.py` fixes its `todo` list at startup and only collects a month's
WAF-dropped stragglers on a *later* run. A supervisor now re-runs until SUCCESS stops rising.

Second, the WAF itself, which made the difference between a 12-hour grind and a stall. It
throttles per source IP on cumulative volume (~50 disclosures, refilling in ~2 min), so the
fix was to pace under that budget and then to spread the load across a rotating pool of IPs
with concurrent requests. Throughput went from ~7 to ~100 disclosures/min with the block
rate driven to near zero. Direction of any residual loss: WAF drops are unrelated to what a
filing predicts, so they thin the sample rather than tilt it.

### Quarantine - a source ceiling, not a bug

On the complete archive, **3,795 of 18,415 disclosures (~21%) are quarantined** to
`reports/parse_failures/` - a DKB PDF that yields no transactions is set aside with its
bytes, never silently skipped. Sampling those files, **about three quarters carry no
extractable text at all**: they are scanned images with no text layer, and no parser will
ever read them. That is a hard ceiling of the same kind as the 250,000 TRY disclosure
threshold - a property of the source, not a defect to fix. Recovering them would need OCR,
which is not attempted here and would need its own accuracy audit before any number it
produced could be trusted.

The remaining quarter do carry text and fail for a different reason: they are **narrative
filings with no transaction table** ("... 3,63 - 3,64 TL fiyat aralığından 9.911 adet satış
işlemi gerçekleşmiştir"), so the arithmetic row-validation gate has nothing to bind to.
Those are recoverable in principle. `scripts/reparse_quarantine.py` re-runs the parser over
the whole quarantine and rescues the ones that now parse; the scanned images are the
residual floor.

### What still remains

**Cluster scoring is close to single-factor.** `cluster_score` blends insider count (0.50),
role seniority (0.30) and recency (0.20). In historical mode recency is pinned at 1.0, so
20% of the weight is a constant, and seniority falls back to its 0.5 default wherever the
scraped board roster does not cover an insider. When coverage is zero the score reduces to
a monotone function of `insider_count` alone - `detect_clusters` now logs `role_map_empty`
loudly in that case, where it used to happen silently. The score is not used to gate any
result reported here, so this is a latent defect rather than an active one.
scraped board roster does not cover an insider - and `person_company_roles` is unpopulated
until `graph scrape-management` runs, so in practice the score reduces to a monotone
function of `insider_count` alone. `detect_clusters` logs `role_map_empty` loudly in that
case. The score gates none of the results above, so this is a latent defect, not an active
one.

**Kyle's lambda is uncalibrated** (1.0). At retail order size the impact term is small
enough that the error changes no conclusion; at institutional size it would, and the number
Expand Down
49 changes: 44 additions & 5 deletions scripts/backfill_kap_insider.py
Original file line number Diff line number Diff line change
Expand Up @@ -71,10 +71,28 @@
# nothing had gone wrong: over 122 months that is ~2 hours of the run spent asleep on the
# happy path. That is not WAF protection, it is a tax on success.
#
# The escalation ladder itself is kept intact - if KAP's WAF does disconnect the warmup
# GET (httpx.RemoteProtocolError = IP throttled), back off 1 min, then 10, then 20. The
# protection is now reactive, which is the only thing a backoff can usefully be.
_WAF_BACKOFF_S = [60, 600, 1200]
# The ladder used to escalate 60s -> 600s -> 1200s on the assumption that a longer silence
# buys a bigger allowance from the WAF. It was never measured. Measured now, over 66
# windows in the run's own log - counting how many disclosures got through between one
# block and the next, bucketed by how long we had just slept:
#
# silence windows disclosures passed before the next block
# < 2 min 36 42.9
# 2-8 min 11 57.0
# 10 min 7 57.1
# 20 min 12 49.6
#
# It is flat. A 20-minute sleep buys the same ~50 disclosures as a 90-second one. The WAF's
# budget is roughly "~50 requests, then block", and it refills within about two minutes -
# so nearly all of the 10- and 20-minute sleeps were pure waste. Over the run that is many
# hours spent waiting for an allowance we already had.
#
# It is not shortened to zero. In the two observed runs of 3+ consecutive short sleeps, the
# 4th and 5th windows fell to 30 and 20 disclosures (n=1 each - thin, but pointing the
# wrong way), so hammering without pause may still draw a penalty. The ladder keeps its
# shape and loses its tail: escalate, but within the window where the budget is known to
# refill.
_WAF_BACKOFF_S = [60, 120, 240]

# A courtesy gap between chunks so a long backfill does not arrive as one unbroken
# burst. Small enough to be irrelevant to the runtime (122 x 2s = ~4 min), unlike the
Expand Down Expand Up @@ -254,7 +272,28 @@ async def main_async(
_log.info("chunk_dry_run", from_date=from_date, to_date=to_date)
continue

await _run_chunk(from_date, to_date)
# A month that exhausts its WAF backoffs must NOT kill the whole run. Its
# scraper_runs record is already written (PARTIAL, or FAILED if the warmup never
# got through), so a later sweep collects it - that is the entire point of the
# ledger. Before the ladder was shortened this rarely fired, because a 20-minute
# backoff usually outlasted the block; now that the tail is gone a genuinely
# blocked month exhausts faster, and letting its RuntimeError propagate abandoned
# every remaining month behind it. Log it and move on.
try:
await _run_chunk(from_date, to_date)
except Exception as exc:
# ANY unhandled per-chunk failure is contained, not just the RuntimeError from
# exhausted WAF backoffs. A dead proxy returning 402 Payment Required crashed a
# whole run once because it surfaced as httpx.ProxyError, which the narrower
# `except RuntimeError` did not catch. The month's ledger record is written by
# scraper.run() before it re-raises, so a sweep collects it; one bad month must
# never abandon the months behind it, whatever killed it.
_log.warning(
"chunk_abandoned",
from_date=from_date,
to_date=to_date,
error=f"{type(exc).__name__}: {exc}",
)
await asyncio.sleep(_CHUNK_GAP_S)

if dry_run:
Expand Down
Loading
Loading