Skip to content

The signal is real. It is not tradeable. (and the six silent data faults that hid both) - #8

Merged
caganco merged 29 commits into
masterfrom
fix/duy-era-silent-parse-failure
Jul 13, 2026
Merged

caganco merged 29 commits into
masterfrom
fix/duy-era-silent-parse-failure

Conversation

@caganco

@caganco caganco commented Jul 12, 2026 •

Copy link
Copy Markdown
Owner

The question this repo asks: do BIST insiders' disclosed purchases predict returns you can actually capture? They predict. You cannot capture them.

Horizon N Gross AR Cost Net AR t (net)
5d 1,070 +0.66% 3.37% −2.71% −12.91
20d 1,070 +2.02% 3.37% −1.35% −3.31
60d 1,071 +2.18% 3.37% −1.20% −1.86

The gross abnormal return is real and significant (20d: +2.07%, t = 5.36, N = 1,079, survivorship-clean). It does not survive the spread: insider clusters fire in illiquid small caps whose median round trip costs 1.93%.

Getting to a number worth trusting meant finding the faults that had been moving it. N was never 29 as a fact about the market — it was the sum of the bugs below.


Part I — the cost model was reading a return index as if it were a price

The one that sat directly under the answer.

price_history.close_try is a chained, corporate-action-adjusted total-return index. That is correct for returns — a bonus issue halves the print, and a raw series reads it as a 50% loss. It is not a price. But:

  • tick_floor_pct is 0.01 TRY / price on the exchange's grid — it was handed the index.
  • ADV is price × volume, and Kyle impact scales as sqrt(1/ADV) — same index.

BIST companies issue bonus shares constantly, so the two series pull apart. Measured on the 2018-12 bulletin: a median 0.98×, but a range of 0.60× to 118×, with 32% of ticker-days off by more than 10%. The factor falls on both sides of 1, so this was not a conservative error that happened to protect the conclusion — it mispriced trades in both directions, worst in the serial bonus-issuers, which are small caps, which is exactly where the tick floor binds and where tradeability is decided.

Migration 0008 keeps the raw print alongside the index. The Abdi-Ranaldo estimator is scale-free (it reads log(close) − log(high-low midpoint), so a constant factor cancels) and correctly stays on the index; only the floor and ADV moved to the price.

Correcting it raised N from 1,032 to 1,070 and left the verdict standing.

The left edge, found on the way

Prices started 2015-01-02, the same date as the KAP floor, so the earliest clusters had fewer than the 22 sessions the spread estimator needs and were silently dropped. That exclusion was not random: the 23 dropped clusters averaged +8.45% at 20 days against +1.93% for the rest. The exchange's bulletin archive goes back to 1988; loading from 2014-11 closes the edge.

relatedStocks is not always one ticker

A filing naming several stocks became a ticker like KRDMA, KRDMB, KRDMD — a key that joins to no price row, so 14 clusters left every result in silence (12 distinct strings, ~131 disclosures). split_related_tickers now splits it and deliberately refuses to choose: which class the insider bought is in the filing body, not this field, and a plausible wrong ticker is worse than a visible gap.

The report contradicted its own project

signal daily-report printed EDGE_DETECTED and a 54.7% hit rate and stopped there. Anyone running it would trade a signal measured to lose money. The net result now travels with the gross one — on stdout, and in a cost_adjusted block in the JSON, because a machine consumer reading verdict: EDGE_DETECTED never opens the README.

yfinance could silently overwrite the bulletin

prices backfill pulls every KAP ticker from yfinance and upserts close_try. Unguarded, it would replace a survivorship-clean chain (yfinance serves nothing for a delisted BIST ticker, not even the years it traded) with a biased one, drop the VBTS flags, and leave close_try and raw_close_try sourced from two different providers while the cost model reads both. The upsert now refuses any row the bulletin wrote.


Part II — the ingest was losing data and reporting SUCCESS

Four commits, one story: the ingest was losing or corrupting data and reporting SUCCESS the whole way. Each commit removes one silent failure; the last one makes the parser prove its output before storing it.

1. The date regex accepted slashes only

KAP writes transaction dates both ways: 14.06.2019 (dots, filings up to ~2020) and 07/06/2023 (slashes, ~2021+). _DATE_RE accepted slashes only, so every filing before 2021 parsed to zero transactions while the run logged success. Measured after the fix: 2016-2020 went from 0% to 75-100% parse yield.

2. A DKB filing that parses to zero transactions is a failure, not an empty filing

dkb_yielded_no_transactions / dkb_parse_yield_low now fire. Had they existed, bug #1 would have been caught on the first backfill.

3. The WAF cooldown was a tax on success

The retry ladder slept 60s before every chunk, including the first — ~2 hours of a 122-month run spent asleep before anything went wrong. The ladder (60s → 10m → 20m) now fires only after an actual RemoteProtocolError. The WAF is real (probing tripped it twice during this work); reactive backoff is what handles it.

4. Fixed column indices stored plausible-looking garbage — 21% of all rows

A missing or merged cell shifted every later column and the row still "parsed": share counts of 3.02 (a percentage), ownership percentages of 25,923,015 (a share count), transactions with share_count=0 (the table's totals line). The old policy nulled the one field that looked wrong and kept the rest — precisely backwards.

Canonical rows must now prove their layout arithmetically before anything is stored:

start_nominal + (buy − sell) == end_nominal
|buy − sell| == |net|
every percentage ∈ [0, 100]

A shifted layout essentially cannot satisfy the first identity by accident. Failing rows are logged (dkb_row_rejected, with reason) and dropped — recoverable from logs, never silently corrupt in the database.

And the 2015-2020 document is now actually parsed. It is a different form — a per-trade blotter, not the netted one-row-per-day table. pdfminer emits its cells in an unstable order (the 2015 fixture interleaves column-major and row-major in one table), so the parser uses only the two order-independent, self-checking anchors:

  • TOPLAM ALIŞ / TOPLAM SATIŞ → (buy qty, sell qty, buy amount, sell amount)
  • the narrative price range → implied average amount/qty must fall inside it

Verified to the kuruş on both committed era fixtures: TEKFEN 2015 (six buys → 120.000 @ 4,39125, range 4,36–4,41 ✓) and METRO 2018 (479.593 bought @ 1,1458, narrated 1,11–1,16 ✓; 69.368 sold @ exactly 1,13 ✓). Holdings columns are not order-independently recoverable in this form and stay NULL rather than guessed.

The floor question from the earlier revision — settled

An earlier revision set the floor at 2016-06 blaming the DUY→ODA label. That was a confound: the date-regex bug alone zeroed those years. Re-probed with the fixed parser: 2015 filings parse fine (the label was a red herring); 2013/2014 genuinely have no table (free-form mailed-in PDFs — OCR territory, out of scope). DEFAULT_START = 2015-01-01, measured.

Net effect

before after
Usable history ~14 months (truncated) ~11.5 years
Pre-2021 filings 0 transactions parsed (canonical + blotter)
Corrupt rows 21% silent rejected + logged
Expected sample N≈29 clusters ~1,000–1,900 clusters (power gate: 784)

Tests: 175 passing. Both era PDFs are committed as fixtures, so none of this ever has to be developed against the live WAF-guarded endpoint again.


Part III — what the measurement then said

With the data trustworthy, the remaining escape routes for the signal were closed by measuring them, not by arguing:

  • Is the signal real? Yes. +2.07% at 20d, t = 5.36, N = 1,079.
  • Is it tradeable? No. Median round trip 1.93%, p75 4.34%. Net negative at every horizon.
  • Does VBTS gross-settlement block it? No — only 16 of 1,079 entries (1.5%) land on a restricted day. Measured, then dismissed.
  • Does stripping routine insiders rescue it? No — and not because the test failed. Zero clusters classify as routine. SPK II-15.1's 250,000 TRY threshold already censors exactly the small, quiet, scheduled trades Cohen-Malloy-Pomorski's routine class is built from. Turkish insiders have no routine filings because routine trades are never disclosed. CMP's hypothesis is not refuted — it is inapplicable.
  • Is it stable in time? Directionally yes: net negative in both halves.

The opportunistic classifier was pre-registered (docs/stage0/OPPORTUNISTIC_CLASSIFIER.md) and frozen before it was run, because it was the signal's last plausible route to being tradeable and therefore exactly where a definition chosen after the fact would be most tempting. When the cost fix later moved the baseline, that document was appended to, never edited — a pre-registration rewritten once the answer is known is worth less than none. The frozen prior said the opportunistic subset would be "stronger gross but still not clear the spread." It was not even stronger.

Housekeeping

--help crashed with UnicodeEncodeError on a Turkish cp1254 console — the first command anyone runs. The README documented a command that does not exist. ruff and mypy are now clean (12 mypy errors, all predating this work; one of them was hiding a real defect — upsert_disclosure guarded a row for None and then dereferenced it on the next line under a type: ignore).

Changing five live INSERT statements to satisfy a type checker is exactly the edit that type-checks and then writes nothing, so they are verified rather than assumed. The one that was only reachable through a live KAP scrape — and therefore could not run at all while the WAF was blocking — now has an offline test.

250 tests pass. 4 integration failures remain and are unrelated: they assert on Ticaret Sicil and board-roster seed data that was never loaded (that scraper needs a hand-solved CAPTCHA), and on a live KAP scrape that returns nothing while the WAF is up.

Still open, and stated rather than glossed

  • Regime. The window is 2015-2018, a single regime. The backfill to 2026 is running (WAF-bound).
  • cluster_score is not yet a score. person_company_roles is empty, so the seniority term is pinned at its 0.5 default and the score is insider_count with a decimal point on it. It needs graph scrape-management, which needs KAP.
  • Kyle's lambda is uncalibrated at 1.0. At retail size the term is small enough not to change any conclusion; at institutional size it would, and the number should not be trusted there.

caganco added 2 commits July 12, 2026 03:17
… not an empty filing

Backfilling to 2015 stored 50 disclosures and produced ZERO transactions. The run
reported SUCCESS.

KAP changed the insider-filing format mid-2016. Measured disclosureType by month:

    2016-03 DUY=57    2016-04 DUY=53    2016-05 DUY=138
    2016-06 ODA=111   2016-07 ODA=50    2016-08 ODA=108

DUY-era filings are not KAP's structured form. The disclosure body just says the
explanation "ekte yer almaktadir" and points at a file the insider mailed in
(e.g. "nthol2.pdf") - free-form, often a scan, with no table to read.
parse_dkb_transactions expects the ODA form and extracts nothing from them.

Nothing said so. The ingest stored the disclosure row, wrote zero transactions, logged
disclosure_processed, and finished the run as SUCCESS. This is the same failure mode as
the list-endpoint cap: data lost quietly, with a green light on top. A backfill left to
run would have spent hours downloading unreadable PDFs and reported success the whole
way.

Two changes:

1. A DKB disclosure that yields no transactions now logs dkb_yielded_no_transactions
   with its disclosureType, and a run where most filings come back empty logs
   dkb_parse_yield_low. The whole point of a Pay Alim Satim Bildirimi is that it
   reports at least one trade; zero is a parse failure and must read as one.

2. backfill_kap_insider.py defaults to 2016-06-01 - the first month of the parseable
   format, measured, not preferred. Earlier dates cost hours and yield nothing.

Reading DUY-era filings would need OCR plus a free-form extractor. That is a separate
project, not a parameter, and it is documented as out of scope rather than left as an
implicit gap.

At ~1,100 filings/year, 2016-06 to today is ~11,000 disclosures - comfortably past the
~784 events the power gate in signals/base_rate.py requires. The reachable sample was
never the constraint; the silent failures were.
… accepted slashes

Five years of history were being thrown away by one character.

KAP writes the transaction date both ways, depending on the filing's vintage:

    14.06.2019   (dots)    - filings up to ~2020
    07/06/2023   (slashes) - filings from ~2021

_DATE_RE accepted slashes only:

    _DATE_RE = re.compile(r"\b(\d{2}/\d{2}/\d{4})\b")

So no row in a pre-2021 filing ever anchored, _extract_table_rows returned nothing, and
every filing before 2021 parsed to ZERO transactions - while the ingest logged
disclosure_processed and finished the run as SUCCESS.

The table was always there. Dumped side by side, the 2019 and 2023 layouts carry the
identical 9-column row; only the separator differs. parse_kap_date already accepted both
formats. Only this regex was turning four years of history away.

Measured against live KAP, sampling four filings per year:

    year   before   after
    2016      0%     100%
    2017      0%     100%
    2018      0%      75%
    2019      0%     100%
    2020      0%      75%
    2021+   already working, unchanged (fixture test still green)

Second bug, found while fixing the first: pdfminer sometimes emits two adjacent table
cells on one line ("174.004.552,79 174.269.552,79"). As a single token that fails
_TR_NUM_RE, and the collector silently SKIPS non-numeric tokens - so both values vanish,
every later column shifts left by two, and post_tx_share_count / post_tx_ownership_pct
are read out of the wrong cells. The row still parses, which is what makes it dangerous:
it yields plausible-looking wrong numbers. This is the source of the
implausible_ownership_pct warnings (an ownership percentage of 24.628.606,69).

Such a line is now split - but only when EVERY part is a Turkish number. Splitting
unconditionally shreds prose too: "18,45 - 18,48 TL" from the price sentence would yield
a numeric token "18,45" that the collector reads as the first table column, and
share_count comes back as a unit price. The existing fixture test caught exactly that on
the first attempt.

Also: the backfill's WAF handling slept 60s before EVERY chunk attempt, including the
first, before anything had gone wrong - ~2 hours of a 122-month run spent asleep on the
happy path. That is not protection, it is a tax on success. The escalation ladder is
kept (60s -> 10min -> 20min) but now fires only AFTER httpx.RemoteProtocolError, which
is the only thing a backoff can usefully do. The WAF is real - rapid probing tripped it
during this work - and the reactive ladder is what handles it.

Tests: 160 passing. The new suite pins the dotted date, the merged cell, and the prose
line that must NOT be split.
@caganco caganco changed the title scrapers: a DKB filing that parses to zero transactions is a failure, not an empty filing scrapers: silent parse failures were costing five years of history Jul 12, 2026
caganco added 2 commits July 12, 2026 04:11
…a confound

The floor in the previous commit (2016-06) was justified by the DUY->ODA format split:
pre-2016-06 filings looked like free-form attachments with no table. That reasoning was
wrong, and the error is worth recording because it is a textbook confound.

A 2015 backfill did produce zero transactions. But the date-separator bug in
parser._DATE_RE would have zeroed those years regardless of their format. Two causes,
one symptom - and I attributed it entirely to the one I found first.

Re-probing with the fixed parser separates them:

    year  docs with a table   parsed   type
    2013         0/3            0%     DUY   <- genuinely no table in the document
    2014         0/3            0%     DUY   <- genuinely no table
    2015         3/3          100%     DUY   <- parses fine; DUY was never the problem
    2016         1/3           33%     ODA

DUY-era 2015 filings carry the same structured table as a modern one. The format label
was a red herring. What actually ends is the table: 2013/2014 filings are a free-form PDF
the insider mailed in, often a scan, with nothing to extract. Reading those would need OCR
plus a free-form extractor - a separate project, not a parameter.

The floor moves to 2015-01-01, gaining ~1.5 years. If some early-2015 filings do turn out
to be table-less, they now surface as dkb_yielded_no_transactions rather than vanishing
silently: a measured claim with a safety net under it, not an assumption.

2015-01 -> today is ~12,500 filings, well past the ~784 events the power gate requires,
and long enough to test regime-conditionally instead of pooling a decade into one number.
… the 2015-2020 blotter

21% of stored transactions were silently wrong. Fixed indices assumed the canonical
9-column layout on every row; a missing or merged cell shifted every later column and
the row still "parsed" - share counts of 3.02 (a percentage), ownership percentages of
25,923,015 (a share count), and 4 stored transactions with share_count=0 (the totals
line of the table). The old policy for an implausible percentage was to null THAT field
and keep the rest of the row, which is precisely backwards: a row that has proven its
columns shifted cannot be trusted in any field.

Two changes.

1. Canonical rows must now prove their layout before anything is stored, using the
   arithmetic the form itself guarantees:

       start_nominal + (buy - sell) == end_nominal
       |buy - sell| == |net|
       every percentage in [0, 100]

   A shifted layout essentially cannot satisfy the first identity by accident. Rows
   that fail are logged as dkb_row_rejected with a reason and dropped - a rejected row
   is recoverable from the logs, a silently corrupt one poisons every statistic built
   on it. Zero-volume totals lines and flat round-trips (net zero) are rejected too.
   Mixed intraday rows now net to the dominant side; the old code called any row with
   sell > 0 a SELL even when buys dominated. The token collector also gets a skip
   budget: it used to wander arbitrarily far into the document stitching together
   narrative numbers, which is where the 2015-era garbage came from.

   (find_column_index / header aliases stay unused for this table: in the pdfminer
   stream the headers are far from the cells and multi-line, while position-proven-
   by-arithmetic is strictly stronger for a fixed-form table.)

2. The 2015-2020 filing is a different document, now parsed: a per-trade blotter
   (date | Alim/Satim | adet | fiyat | tutar | holdings), not the netted
   one-row-per-day form used from ~2021. pdfminer emits its cells in an unstable
   order - the 2015 fixture interleaves column-major and row-major in one table - so
   per-trade positional reconstruction is guesswork. Two anchors are order-independent
   and self-checking, and only those are used:

       TOPLAM ALIS / TOPLAM SATIS -> (buy qty, sell qty, buy amount, sell amount)
       narrative price range      -> implied avg price amount/qty must fall inside it

   Both fixtures check out to the last kurus: TEKFEN 2015, six buys netting 120.000
   shares for 526.950 TL (avg 4,39125, inside the traded 4,36-4,41); METRO 2018,
   479.593 bought at avg 1,1458 (narrated range 1,11-1,16) and 69.368 sold at exactly
   1,13. Per-filing net summaries are also the honest granularity - the modern form
   nets same-day trades into one row anyway. Post-transaction holdings are NOT
   order-independently recoverable and stay NULL rather than guessed.

   Routing is by the TOPLAM marker, which no modern fixture contains.

Both era PDFs are committed as fixtures (public KAP filings), so this never has to be
developed against the live WAF-guarded endpoint again.

Tests: 175 passing. The validator suite pins every rejection reason; the blotter suite
pins the real numbers read off the two era documents.
@caganco caganco changed the title scrapers: silent parse failures were costing five years of history scrapers: silent parse failures - five years of history, and 21% of the rest Jul 12, 2026
caganco added 24 commits July 12, 2026 04:53
…aware

The first clean backfill rejected ~65% of 2015-era filings. The reject log tallied
to three defects, all in the legacy blotter path added one commit earlier:

  18x "implied price outside narrative range"
  18x "qty/amount inconsistent"
  10x "totals block incomplete"

Root causes, in order of blame:

1. The price regex could not read two of the three period spellings. It captured
   "1.234" out of "1.234,56 TL" (parse_turkish_number then read it as one thousand
   two hundred) and read the pre-2016 dot-decimal "4.36" as 436. Either way the
   implied average landed "outside" a narrative range that had been misparsed, and
   a correct filing was rejected. _parse_price_token now treats a dot followed by
   1-2 digits as a decimal point; only exactly-3-digit dot groups are thousands.

2. One narrated range was applied to both directions. The filing template narrates
   the buy leg first, so the first range bounds the buy side and the last bounds
   the sell side; with a single range and both directions active, only the buy is
   range-checked. A sell narrated without its own range is no longer rejected
   against the buy's.

3. The TOPLAM cells' order is not stable: pdfminer emits (buy qty, sell qty,
   buy amt, sell amt) in some files and (buy qty, buy amt, sell qty, sell amt) in
   others. The parser assumed the first and rejected the second. Guessing the
   order is how columns got silently mis-read before, so it is not guessed: both
   bindings are validated against the same checks and the one that uniquely
   survives wins. Both-or-neither surviving with different substance rejects the
   filing loudly. This turns "position we hope is right" into "position the
   arithmetic proved", same principle as the canonical-row validator.

avg_price now quantizes ROUND_HALF_UP like every other Decimal in the codebase.

Tests: 186 passing, each defect pinned with a synthetic filing, including the case
that must KEEP failing (a buy avg genuinely outside its own narrated range).
…ence

A DKB filing that parses to zero transactions now keeps its PDF under
reports/parse_failures/{disclosure_id}.pdf (gitignored). The bytes were already
paid for during ingest; discarding them meant every parse failure had to be
re-fought through KAP's WAF one probe at a time just to be looked at. With the
evidence retained, remaining failure modes can be diagnosed offline in bulk and
repaired with a single re-parse pass over the quarantine directory - no new
requests at all.
…tine repair

The quarantine directory did its job on the first day. 165 unparseable PDFs
accumulated during the running backfill; diagnosing them offline (zero extra
requests against the WAF) surfaced two more real emissions of the totals block:

1. Separated labels - each label owns the pair that follows it:
       TOPLAM ALIS / qty / amt / TOPLAM SATIS / qty / amt
   The old code looked only after the SELL label, found the sell side's lonely
   zeros, and rejected the filing as "totals block incomplete" (31 times in
   month one alone).

2. Trade quantities first, totals pair after (quarantined filings 405071 and
   404639):
       [186, 133, 5, 155, 479, 0, prices...]   where 186+133+5+155 == 479
   The totals pair is findable without knowing the layout: it is the adjacent
   pair that EXACTLY equals the sum of every numeric before it. Exact Decimal
   equality makes an accidental match effectively impossible - a position is
   again accepted only because arithmetic proved it. Trade amounts are
   interleaved beyond recovery in this emission, so price_try stays NULL rather
   than guessed; share_count is the field the signal pipeline actually needs.

scripts/reparse_quarantine.py replays quarantined PDFs against the current
parser and upserts what now succeeds - no network involved. Dry-run against the
real quarantine: 81 of 165 recovered; of the rest, 36 are scanned images (OCR
territory, out of scope) and 48 are a long tail of rarer emissions that stay
quarantined with their evidence rather than being guessed at.

Tests: 189 passing, including the separated-labels case verbatim from
quarantined filing 405865.
…ree to the kurus

The dominant remaining quarantine bucket (89% of a fresh hour's failures) was a fifth
emission: the totals labels sit BEFORE the table cells entirely, so every positional
path - label-adjacent pairs, dual binding, qty-sum - had nothing to bind. The cells
carry textbook trade triples (250 x 1,35 = 337,50; 9.850 x 3,3 = 32.505), but their
positions are unusable.

What was being overlooked: the filing narrates its own volume in words. "479.593 adet
alis islemi" is data, present on every legacy filing, and it is INDEPENDENT of the
table's emission order. The fallback extracts quantities per direction from the
narrative and accepts them only when a second, disjoint source agrees exactly:

  - the table's exact qty*price==amount triples summing to the same quantity
    (price then comes out as amount_sum/qty_sum and must still clear the narrated
    range), or
  - the narrated totals appearing verbatim as table cells.

Prose lines are excluded from the table-numeric pool, so the corroboration is
genuinely cross-source - the narrative cannot confirm itself.

One contract sharpened by a test that had to change: when the quantity is doubly
corroborated but the amount contradicts the narrated price range, the old policy
rejected the whole filing. Now the quantity is kept and the price is refused (NULL) -
keep what two sources prove, never store what one source contradicts. And the
fallback is not "accept anything with a TOPLAM label": an uncorroborated filing still
dies, pinned by test.

Dry-run over the real quarantine: 903 of 1,227 files recovered (74%, from 49%),
1,044 transactions. The remainder is scanned images (OCR, out of scope) and a long
tail that stays quarantined with its evidence.

Tests: 190 passing.
A backfill was spending 83% of its wall clock waiting, and losing 12% of the data,
for the same reason.

Measured over one run: 4,200 disclosures processed, 575 lost to
httpx.RemoteProtocolError ("Server disconnected without sending a response"). Each
one cost ~33s and 555 of them were followed directly by a stall - 304 minutes, 85%
of all stall time. The run reported SUCCESS anyway.

Two defects, one cause: the code treated a WAF block as a network blip.

1. core/http._is_retryable listed RemoteProtocolError as retryable, so tenacity
   re-attempted it 5 times with exponential backoff (4+8+16+32s). But KAP's WAF does
   not lift in four seconds - retrying inline just re-pokes the block, burns up to a
   minute, and still fails. RemoteProtocolError is now NOT retried inline: it fails
   fast so the caller can handle it properly. ConnectError/ReadTimeout/429/503 - the
   failures that ARE transient - still retry as before.

2. KapInsiderScraper caught the exception, logged disclosure_error, and `continue`d -
   the disclosure was simply gone. This is the same silent-loss pattern as the list
   cap, the date regex and the column shift: data disappears, the run stays green.
   Dropped disclosures are now DEFERRED, retried once after a 90s cooldown (long
   enough for a volume-triggered throttle to lapse), and anything still missing
   downgrades the run to PARTIAL.

PARTIAL matters: the resume ledger only counts SUCCESS as done, so an incomplete
month is automatically re-run by the next pass, and because ingest is idempotent
(disclosure_exists) that re-run costs only the disclosures actually missed. A month
can no longer be silently half-ingested.

The per-disclosure body is extracted to _ingest_one so the main pass and the retry
pass share one implementation rather than duplicating it.

Expected effect: the ~300 minutes of inline WAF backoff collapse to one cooldown per
affected chunk, and the 12% loss becomes zero (or a PARTIAL flag that fixes itself).

Tests: 199 passing, including the retry-policy contract - a WAF disconnect must not
be retried inline, a genuine blip must be.
… fresh months first

Three fixes to a backfill that was sweeping the archive without advancing the frontier.

1. One 90s cooldown recovered only ~53% of the disclosures KAP's WAF dropped, so the
   month still finished PARTIAL. The resume ledger counts only SUCCESS as done, so a
   PARTIAL month gets re-run on the next pass - during which the WAF drops a fresh
   handful and the month goes PARTIAL again. Converges, but each round costs a full
   re-scan. The cooldown is now a ladder (90s, 4min, 10min): the throttle is
   volume-triggered, so waiting longer lifts it, only months that actually lost
   something pay, and a month that recovers everything reaches SUCCESS and is never
   revisited.

2. disclosure_exists ran per filing, opening a fresh DB session for each of ~200 a
   month. That is most of the cost of re-running an already-ingested month - and
   PARTIAL months get re-run by design. Now one query per chunk.

3. Chunk order: untouched months FIRST, PARTIAL stragglers LAST. A PARTIAL month is
   already ~95% ingested and only needs its WAF-dropped filings; an untouched month
   has nothing. Replaying PARTIALs in calendar order meant re-listing hundreds of
   stored filings before any new history landed - observed live: 11 chunks processed,
   ~180 of every ~200 filings skipped as already-present, and the frontier stuck at
   2017-01. Fresh data first; stragglers are cheap to collect at the end.

Tests: 199 passing.
…rs cannot vanish

Ran the pipeline end to end for the first time on real data (4,500 filings, 2015-2017,
910 clusters). It produced a convincing result and two reasons not to believe it.

1. Look-ahead guard fired: 91 transactions were dated AFTER the disclosure reporting
   them - physically impossible. The legacy blotter took max() over EVERY date in the
   document, so a reference or validity date in the prose could outrank the real trade
   date. The filing's own publication date is a hard ceiling (a trade cannot be
   disclosed before it happens), so candidates above it are now discarded, and the
   canonical path rejects such rows outright. parse_dkb_transactions takes
   published_on; both the ingest and the quarantine repair pass it.

   signals/returns.py's AssertionError was right to fire. Guards that stop the
   pipeline are worth more than results that flow through it.

2. SURVIVORSHIP_BIASED, a new verdict, and the reason this commit exists.

   The first honest numbers looked like an edge: +2.45% abnormal return at 20 days,
   t=6.07, p<0.0001, 95% CI [54.4, 62.3] excluding the null. Then: of 910 clusters,
   only 584 had any price data. The other 326 (36%) were dropped - and their tickers
   (ACSEL, ANELT, ARBUL, BISAS, BMEKS, ...) are overwhelmingly DELISTED small caps.

   yfinance does not serve dead BIST names, so the deletion is precisely of the worst
   outcomes. What remains is a survivor sample, and its mean is not an estimate of
   anything. compute_base_rate now measures that attrition and, above 10%, returns
   SURVIVORSHIP_BIASED instead of a verdict - the point estimates still print, but
   nothing may read them as evidence.

   This gate exists because +2.45% at t=6.07 was the most convincing wrong answer this
   pipeline has ever produced, and it would have been believed.

Fixing it needs a price source that covers delisted tickers. Until then the honest
output is "unusable", and the code now says so instead of leaving it to be discovered.

Tests: 199 passing.
…orship" was a bug

The base rate reported SURVIVORSHIP_BIASED because 326 of 910 clusters (36%) had no
price data, and their tickers looked exactly like delisted small caps. Most of that
was true. A large part was not - it was fetch_and_store_prices.

One request for all 245 tickers looked efficient. yfinance answers a batch containing
an unknown symbol with an error naming SEVERAL ("Quote not found for symbol: GEREL,
VRSGS.IS") and drops the innocent ones with it. ACSEL, ANELT and ATLAS each have 750+
days of history in yfinance and still ended up with zero price rows. A cluster with no
prices is silently excluded from the base rate, and the tickers a batch is most likely
to poison are the obscure ones - so the loss wore survivorship bias as a costume.

Prices are now fetched in batches of 20, and every ticker a batch fails to deliver is
retried alone. One bad symbol can no longer take its neighbours down.

Second cause, also not survivorship: two tickers were RENAMED, not delisted. KAP files
them under the old code and yfinance only serves the new one (GYHOL -> GLYHO,
AKFEN -> AKFGY, each verified individually to return a full 2015-2017 series under the
new symbol and nothing under the old). _TICKER_ALIASES maps them.

Attrition: 35.8% -> 31.0%.

The remaining 31% is the real thing, and it does not yield to engineering. MCTAS (118
clusters), MENBA (59), OZBAL (42), IZTAR (25) and others are genuinely dead, and
yfinance serves nothing for a delisted BIST name - not even the years it was trading.
For a 2015-2017 study that is a systematic hole exactly where the worst outcomes live,
so SURVIVORSHIP_BIASED still fires and the numbers below it remain unusable:

    20d  N=634  hit 60.2%  AR +2.85%  t=7.10  p<0.0001   <- NOT evidence

Closing that gap needs a price source that covers delisted tickers. Until then the
honest output is "unusable", and the pipeline says so rather than leaving it to be
discovered by whoever trusts the number.

Tests: 199 passing.
…%, and the 60-day sign flips

yfinance is survivorship-contaminated for BIST: it serves nothing for a delisted
ticker, not even the years it was actively trading. 31% of insider clusters had no
price data, and the names behind them (MCTAS, MENBA, OZBAL, IZTAR...) are exactly the
small caps that went to zero. compute_base_rate correctly refused to return a verdict.

Borsa Istanbul's official end-of-day bulletin (PP_GUNSONUFIYATHACIM) is
survivorship-clean by construction - it records what actually traded each day, so a
company that later delisted is still there for every day it lived. 2015-2026 carries
1,204,914 equity rows across 729 tickers; yfinance had 185.

    cluster attrition: 31.0% -> 1.5%
    MCTAS 743 days, OZBAL 1,525, IZTAR 1,597, SERVE 1,645 - all previously unpriceable

Corporate actions. The bulletin quotes raw prints, so a bonus issue halves the close
overnight and a naive window return reads -50%. The archive has no adjustment file, but
the bulletin restates one field itself: ONCEKI KAPANIS FIYATI, the previous close as the
exchange adjusts it. So close(t)/prev_recorded(t) is the true one-day total return, and
chaining those gives a series whose ratios are correct across any corporate action.
Verified: BOSSA 2016-06-08 reads +0.39% chained, where the raw print shows -9%.
30 of 24,967 day-pairs (0.12%) carry such an adjustment.

What the clean data says - and it is not what the biased data said:

    horizon   N     hit%    mean AR    t        verdict
    5d       934   64.2%    +1.46%    7.86     EDGE_DETECTED
    20d      934   60.4%    +2.04%    5.99     EDGE_DETECTED
    60d      934   41.0%    -2.80%   -4.58     NEGATIVE_EDGE

N clears the 784 power gate and attrition clears the 10% gate, so these verdicts are
the first this pipeline has been entitled to return.

The 60-day horizon FLIPPED SIGN: +4.18% under survivorship bias, -2.80% clean. The
dead companies' collapses were being deleted from the sample, and they were most of
what happened at 60 days. That single number is the clearest measure of how large the
bias was - and of why the gate existed.

Not a green light. Costs are still not deducted, and these are the illiquid names where
spread and impact are worst; no VBTS tradability filter is applied; and the window is
2015-2017, a single regime. The honest reading is that a short-horizon effect is
statistically present before costs, and that holding past ~20 days destroys it.

The bulletin also carries BRUT TAKAS and GECICI DURDURMA - the VBTS flags the pipeline
lacks. That is the obvious next use of this source.

Tests: 199 passing.
…sses

Half the backfill was spent asleep. Measured over 8 chunks: 124 minutes elapsed, 64 of
them (52%) inside the per-chunk WAF cooldown ladder - 7.1 minutes per month, some 14
hours across the archive - while the frontier stood still.

The ladder (90s, 4min, 10min) existed to recover a month fully before moving on. It is
unnecessary, because a blocked request is now nearly free: RemoteProtocolError fails fast
(core/http._is_retryable), so a chunk takes whatever the WAF lets through, defers the
rest, and finishes PARTIAL. The resume ledger counts only SUCCESS as done, and the ingest
is idempotent - so a later pass re-fetches exactly what was lost and skips the hundreds
already stored. Recovery belongs in a second sweep, not in a sleep inside every month.

  - cooldown ladder -> a single 90s round (it clears ~half the deferrals immediately,
    measured 76 -> 34, which keeps the PARTIAL set small)
  - the backfill now re-sweeps the PARTIAL months in up to 4 recovery passes, each
    opening with a 30-minute pause, stopping when nothing is left or a pass makes no
    progress

The pause is the part that matters. Measured: 40 minutes of silence took the block rate
from 61% to 0%, while cutting the request rate did nothing (1 RPS and 3 RPS both hit 100%
blocked once the throttle was hot). KAP's WAF meters cumulative volume, not instantaneous
rate - so the right shape is burst then rest, not a slow trickle. The rate limit is raised
to 3 RPS accordingly: when requests are getting through, 40 of them take 29s at 3 RPS
against 88s at 1 RPS, and when they are not, the rate is irrelevant.
The answer, measured end to end on survivorship-clean data:

    horizon  N=1032   gross AR   cost    net AR      t       verdict
    5d                 +0.57%    3.34%   -2.77%   -13.11   loses money
    20d                +1.76%    3.34%   -1.58%    -4.04   loses money
    60d                +2.09%    3.34%   -1.24%    -1.98   loses money

    round-trip cost: median 1.94%, p75 4.24%

The signal has genuine predictive content - gross abnormal return is significantly
positive at every horizon (20d: +1.76%, t=5.12). It is also not tradeable, because
insider clusters fire in illiquid BIST small caps and the bid-ask spread on those names
is larger than the alpha. Nothing survives the round trip.

The cost is not assumed, it is estimated per trade from that stock's own OHLC in the 30
sessions before entry (Abdi-Ranaldo 2017, RFS 30(12)) - the only estimator that works on
the delisted names the exchange bulletin carries and no quote feed does. A flat fee would
have flattered the answer; the measured median is 1.94% and the upper quartile 4.24%.

Three bugs had to be fixed first, and each one had moved the number:

1. My own cost script dropped 96% of the sample. abdi_ranaldo_spread returned None when
   a quiet window drove gamma >= 0, and I treated that as "cannot estimate" and excluded
   the trade. The survivors' gross AR came out NEGATIVE where the full sample's was
   positive - the exclusion selects on exactly the price behaviour the signal is about.
   A quiet window is not a free trade: the estimate now widens the window and is floored
   at one tick, and nothing is dropped for it.

2. The exchange bulletin changed format at 2015-12, and the older layout is a different
   file: session-level rows, DD.MM.YYYY dates, and a decimal separator that switches from
   comma to dot mid-year without touching the header. Read with the modern column map,
   2015-01..2015-11 loaded ZERO rows. Reading '54.50' as Turkish gave fifty-four thousand
   five hundred, a thousandfold error that overflowed NUMERIC(20,4) - which was lucky,
   because a subtler scale error would have gone straight into the returns. The chain now
   also rejects any one-day ratio outside [0.5, 2.0]: BIST's daily limit is +/-10%, so
   anything past that is a data fault, and chaining it compounds through the whole series.

3. get_price_and_date_after_days returned the first row after the signal date however far
   away it was. With 2015 prices missing, 1,278 outcomes were "entered" on 2015-12-01 -
   where the series happened to begin - for signals fired months earlier. Those are not
   late entries, they are fabricated ones: nobody could have bought at a price that had
   not printed yet. Entry now refuses a gap wider than 10 days.

With those fixed the sample is 1,055 clusters, 6% unpriceable, and the verdict is stable.

What this settles: the gross signal is not the question, and never was. The question is
whether it clears the cost of the names it fires in, and it does not.
…he constraint

The exchange bulletin carries BRUT TAKAS (gross settlement) and GECICI DURDURMA
(suspended). A name under gross settlement cannot be round-tripped the way a backtest
assumes; a suspended one cannot be traded at all. The pipeline had been booking entries
in both, and insider clusters fire precisely in the volatile small caps VBTS targets -
so this looked like it might matter a great deal.

Measured: 143,012 gross-settlement ticker-days across 451 names (11% of all price rows).
But of 1,055 cluster entries, only 11 (1.0%) land on a restricted day, and none on a
suspended one.

So the flags are now loaded and the gap is closed, and the honest result is that it was
never the binding constraint. The cost of crossing the spread is - a median 1.94% round
trip against a gross abnormal return of 1.76% at 20 days. Naming a limitation and then
measuring it away is worth as much as naming it and finding it fatal; what is not worth
anything is leaving it unmeasured and implied.

Tests: 199 passing.
The repository is the argument, and it was making the wrong one - opening on what the
pipeline scrapes rather than what it found. A reader had to reach docs/METHODOLOGY.md to
learn that the question has an answer.

The answer leads now: the insider-cluster signal has genuine predictive content
(20d gross abnormal return +1.76%, t=5.12, N=1,032, survivorship-clean) and is not
tradeable, because the names it fires in cost a median 1.94% to round-trip and an upper-
quartile 4.24%. A gross number is not an edge; an edge is what is left after the market
takes its cut.

The six silent data faults are listed as a table rather than buried in commit history -
each one had moved the number, and a reader deciding whether to trust the result is
entitled to see what had to be true for it to be trustworthy.

The sample-output block is replaced by the commands that reproduce the result, including
net_of_cost.py, with its real output. Claims a reader cannot re-run are decoration.
Cohen-Malloy-Pomorski's split is the signal's last plausible route to being tradeable:
pooled, it loses by a hair (gross +1.76% at 20d against a 1.94% median round trip), and
if half the trades are routine noise the opportunistic subset might clear it.

Which makes this exactly the place where a definition chosen after seeing the result
would be most tempting and least honest. The definition, the horizon, the test, the
minimum N, the falsification condition and the prior are all fixed here first.

The prior, recorded: opportunistic will be stronger gross and will still not clear the
spread. SPK II-15.1's 250,000 TRY disclosure threshold already censors the small quiet
trades CMP's routine class is built from, so there is less dilution left to strip out
than in the US sample. If the result comes out positive it comes out against the prior,
which is worth more than a confirmation.
…re are no routine filers

The pre-registered test (docs/stage0/OPPORTUNISTIC_CLASSIFIER.md, frozen before this ran).
It was the signal's last plausible route to being tradeable: pooled, the gross abnormal
return loses to the spread by a hair (+1.76% at 20d against a 1.94% median round trip), and
Cohen-Malloy-Pomorski say half of insider trades are routine noise worth stripping.

Result:

    cluster classes at 20d:  OPPORTUNISTIC=1043   ROUTINE=0   UNCLASSIFIED=12

    class          hzn     N   gross%   cost%    net%       t   verdict
    OPPORTUNISTIC  20d* 1023   +1.80    3.35    -1.55   -3.94   loses money

Not one cluster classified as routine. And that is exactly what the pre-registered prior
predicted, in the frozen document, before the number existed:

    "SPK II-15.1's 250,000 TRY disclosure threshold already censors the small quiet trades
     CMP's routine class is built from, so there is less dilution left to strip out than in
     the US sample."

Turkish insiders have no routine filings because routine trades stay under the reporting
threshold and are never disclosed at all. There is no noise to strip: the opportunistic
subset IS the sample (1,043 of 1,055), its gross return is +1.80% against the pooled
+1.76%, and it loses to the same spread.

CMP's hypothesis is not refuted - it is inapplicable. The regulation that makes this dataset
possible is the same regulation that removes the variation the test needs. That is a finding
about Turkish disclosure, not about insiders.

Stated limitation: the 36-month lookback is thin on 3.5 years of data, so a fuller backfill
would classify more insiders and might surface some routine ones. The direction is known -
stripping routine trades RAISES gross alpha - but it would have to raise +1.76% past +1.94%,
which is more than CMP's own effect size, and the disclosure threshold has already removed
most of what would do the raising.

Tests: 199 passing.
A methodology document that names a limitation it has since measured is not merely stale,
it is misleading in the direction that matters - it understates what the project knows.
All three of the gaps it listed were closed during this work, and two of them closed
against expectation.

  Transaction cost - CLOSED, and it is THE constraint. Per-trade Abdi-Ranaldo spread from
  the stock's own OHLC: median round trip 1.94%, p75 4.24%, against a gross abnormal
  return of +1.76% at 20 days. The signal does not survive being traded. This is the
  result, not a caveat on it.

  VBTS tradability - CLOSED, and NOT the constraint. 143,012 gross-settlement ticker-days
  across 451 names (11% of price rows), but only 11 of 1,055 cluster entries (1.0%) land
  on a restricted day. Named, measured, and it turned out not to matter.

  Routine/opportunistic - CLOSED, and the split does not exist here. Zero routine
  clusters. The 250,000 TRY disclosure threshold censors exactly the trades CMP's routine
  class is built from, so there is no noise to strip. The hypothesis is not refuted, it is
  inapplicable - a finding about Turkish disclosure, not about insiders.

Also recorded: the trap where the cost script dropped 96% of the sample by treating a
degenerate Abdi-Ranaldo window as "unpriceable". The survivors' gross return came out
negative where the full sample's was positive - the exclusion selects on the very price
behaviour the signal is about. That failure belongs in the methodology, not just the git
history, because the next person to write a cost model will reach for the same shortcut.

A new section 6 keeps what is genuinely still open: regime (the window is 2015-2018), the
latent single-factor cluster score, and Kyle's uncalibrated lambda.
A backfill sat silent for THREE HOURS. The process was alive, the last event was a 1200s
WAF backoff, and after waking from it the run emitted nothing at all - no chunk_start, no
error, no exit.

Every timeout in the stack should have prevented that. httpx gets KAP_TIMEOUT_S on every
request, tenacity caps at five attempts with a 60s ceiling on the wait, and the worst
legitimate path through a chunk is a few minutes. None of them fired.

Rather than guess which await never returned, the chunk now runs under asyncio.wait_for
with a 15-minute deadline. A chunk is bounded work - one list call, then two requests per
filing at a few per second - so anything past that is not slow, it is stuck. On expiry the
month is abandoned, left PARTIAL for a recovery pass, and the run moves on.

A hang that costs one month is a nuisance. A hang that costs the night is a failure, and
this one cost the night.
… the record

The previous commit added a chunk deadline and justified it with a run that "sat silent
for THREE HOURS". It had not. structlog writes timestamps in UTC, the machine clock is
UTC+3, and I compared the two directly. A healthy process looked frozen because the two
numbers were not in the same units.

    structlog : 23:54:05   (UTC)
    date      : 02:54:05   (TSS, UTC+3)

The process was working - the frontier had moved from 2018-09-14 to 2018-11-15 while I
was calling it dead - and I killed it on that diagnosis.

This is precisely the class of fault this codebase spent the day finding: two quantities
compared as if they shared a scale when they did not. The list cap, the date separator,
the column indices, the decimal comma, the survivorship gap - every one of them was some
version of the same mistake. Committing it while cataloguing it is worth recording rather
than quietly amending away.

The deadline itself stays. A chunk that cannot finish in fifteen minutes is not doing
useful work, and a run that can lose a month to a hang beats one that can lose a night.
But it now says in the source that it guards against a hang never observed, instead of
implying a bug that was really a clock.
costs.py and opportunistic.py produce the project's answer and had no unit tests. 199
tests, and none of them touched the code that decides whether the signal is tradeable.

25 added. The ones that matter:

  - A quiet window is FLOORED at one tick, never dropped. This is the trap that cost 96%
    of the sample: when Abdi-Ranaldo's gamma goes >= 0 the formula reports a spread of
    exactly zero, and treating that as "cannot estimate" excludes the trade - selecting on
    precisely the price behaviour the signal is about. The survivors' gross return came out
    negative where the full sample's was positive. Pinned so it cannot come back.

  - A bad price INSIDE the estimation window voids the estimate; one OUTSIDE it does not.
    Writing this pair caught my own test being wrong: I asserted that a zero anywhere in
    the input should void the spread, and it should not - the estimator reads only the
    trailing window, and voiding on a fault it never sees would drop the trade for nothing.
    Dropping trades on a price-data condition is exactly how the cost model went wrong the
    first time.

  - An illiquid name costs more to round-trip than the +1.76% the signal earns. The result,
    as an assertion.

  - The frozen classifier: three purchase months minimum, 60% seasonal share, repeated buys
    in one month count once, nothing after as_of is visible, one opportunistic insider makes
    the cluster opportunistic. The definition is pre-registered; these keep a later reading
    of the RESULT from quietly becoming a later reading of the DEFINITION.

Tests: 224 passing.
The result rests on one window, so it is worth knowing whether one stretch of it carries
the whole thing. Same cost test, each half separately:

    period        N     gross%   cost%    net%       t     verdict
    2015-2016    859     +2.30    3.57   -1.27    -3.04    loses money
    2017-2018    197     +0.30    2.17   -1.87    -1.57    inconclusive (N < 200)

Net abnormal return is negative in both. The second half is inconclusive rather than
confirming, and the distinction matters: N=197 is under the pre-registered minimum of 200,
so the sign agrees and the power does not. Reporting it as a second confirmation would be
claiming more than the data holds.

The gross alpha collapses from +2.30% to +0.30% across the halves. That is either a market
becoming more efficient or noise at N=197, and at that N it cannot be told apart - so it is
recorded and not interpreted.

This is explicitly NOT the regime test. The window is 2015-2018; the August 2018 crisis
leaves only 9 clusters on its far side, and the 2021-2023 negative-real-rate boom is not in
the data at all. Regime stays open, and stays in section 6 where it belongs.
… command that does not exist

Two faults found by actually running the thing, which is the only way they surface.

1. `trailingedge report --help` died with UnicodeEncodeError. A single U+2194 in a
   docstring - "Generate cross-reference brief (KAP <-> TSG)" - cannot be encoded in
   cp1254, the default console codepage on a Turkish Windows install. The first command
   anyone runs against a repository is --help. It was the one that broke.

   Help text is now ASCII, and stdout/stderr are reconfigured to UTF-8 with errors=replace
   where the runtime allows it, so the next non-ASCII character degrades to a replacement
   glyph instead of a traceback.

2. The README told readers to run `trailingedge report forensic KAPLM`. There is no such
   command. The forensic module exists and works - it produced a real PDF for SARKY on the
   first try - but the CLI exposes it as `report generate --ticker`. A README that
   documents a command that does not exist fails the reader at the exact moment they
   decide whether to trust the repository.

Both are the same category as everything else found today: something asserted, never
executed, and wrong. The fix for that class is not more care, it is running the command.
price_history.close_try is a chained, corporate-action-adjusted total-return index.
That is the correct series for returns - a bonus issue halves the print and a raw
series reads it as a 50% loss - and every return in this project is computed from it.

It is not a price. Two costs are properties of the price, not the index:

  - the tick floor is 0.01 TRY on the exchange's grid, so it depends on where the
    stock actually trades. tick_floor_pct was being handed the index.
  - ADV is price x volume, and Kyle impact scales as sqrt(1/ADV). The same index went
    into ADV.

Measured on 2018-12 bulletin data: the index sits at a median 0.98x the traded price,
but ranges 0.60x to 118x, and 32% of ticker-days are off by more than 10%. BIST
companies issue bonus shares constantly, so the two series pull apart, worst in the
serial issuers - which are small caps - which are exactly the names where the tick
floor binds and where the tradeability verdict is actually decided.

Both errors land in signals/costs.py, the module that decides the project's answer.
The factor falls on BOTH sides of 1, so this is not a conservative error that happens
to protect the conclusion: it mispriced individual trades in both directions.

  - migration 0008 keeps raw_close_try alongside the index. Nullable and NOT
    backfilled: NULL means the traded price was never stored for that row, and the
    cost model must refuse it rather than fall back to the index - falling back is
    the bug.
  - abdi_ranaldo_spread_pct / round_trip_cost take last_traded_price. The estimator
    itself is scale-free (it reads log close minus the log high-low midpoint, so a
    constant factor cancels) and correctly stays on the adjusted index; only the
    floor moves to the price.
  - net_of_cost.py and opportunistic_test.py build ADV from raw price x volume.

Prices reloaded from 2014-11 rather than 2015-01. The chain's ratios are unchanged by
an earlier anchor, so stored abnormal returns are unaffected, but the earliest KAP
clusters (first entry 2015-01-16) previously had fewer than the 22 sessions the spread
estimator needs and were dropped. That exclusion was not random: the 23 dropped
clusters averaged +8.45% at 20 days against +1.93% for the rest.

Also: KAP's relatedStocks is not always one ticker. A filing naming several stocks
became a "ticker" like `KRDMA, KRDMB, KRDMD` - a key that joins to no price row, so
those clusters silently left every result (12 distinct strings, ~131 disclosures, 14
clusters). split_related_tickers now splits it and refuses to choose: which class the
insider bought is in the filing body, not this field, and a plausible wrong ticker is
worse than a visible gap.

And a correction to an earlier claim of mine: this repo was NOT ruff- and mypy-clean.
ruff is now clean; 12 mypy errors remain, all predating this work.
…history

Two writers touch price_history and they do not mean the same thing by close_try.
load_official_prices.py stores a chained corporate-action-adjusted total-return index
with the matching raw print alongside it; data/prices.py stores yfinance's own adjusted
close and knows nothing about raw_close_try.

`prices backfill` pulls EVERY distinct KAP insider ticker from yfinance, so it would
have walked over the whole bulletin-derived series: replacing a survivorship-clean chain
(yfinance serves nothing at all for a delisted BIST ticker, not even the years it
traded) with a biased one, dropping the VBTS flags, and leaving close_try and
raw_close_try sourced from two different providers while the cost model reads both.

The upsert now refuses any row that already carries raw_close_try - i.e. any row the
bulletin wrote. yfinance keeps the job it is genuinely needed for: XU100, which is an
index and appears in no equity bulletin. Pinned by an integration test in both
directions, because this regression would raise no error, only a quietly worse number.
compute_base_rate measures the GROSS abnormal return and knows nothing about the bid-ask
spread. On this sample it correctly returns EDGE_DETECTED: +2.07% at 20 days, t = 5.36,
N = 1,079, survivorship-clean. It is a real signal.

It is also uncapturable, which is this project's entire finding - and the daily report was
printing "EDGE_DETECTED" and a 54.7% hit rate and stopping there. The tool contradicted
its own conclusion on its own screen. Anyone running it would trade a signal measured to
lose 1.35% net at 20 days.

So the net result now travels with the gross one, in both channels:
  - stdout prints the cost warning under the table
  - the JSON carries a `cost_adjusted` block, because a machine consumer reading
    `verdict: EDGE_DETECTED` never opens the README

Numbers throughout refreshed onto the corrected cost basis (raw price rather than the
total-return index, migration 0008):

  gross  20d  +2.07%  t = +5.36   N = 1,079
  net    20d  -1.35%  t = -3.31   N = 1,070
  cost        median 1.93%   mean 3.37%   p75 4.34%

  VBTS-restricted entries re-measured: 16 of 1,079 (1.5%). Still not the constraint.

  Pre-registered opportunistic test re-run on the corrected data, decision rule NOT
  re-chosen: OPPORTUNISTIC N=1,059, gross +2.08%, net -1.31%, t = -3.18. ROUTINE = 0.
  The recorded prior said the subset would be "stronger gross but still not clear the
  spread" - it was not even stronger.

docs/stage0/OPPORTUNISTIC_CLASSIFIER.md is APPENDED to, never edited. Its frozen baseline
(+1.76% / 1.94%) is left standing even though both numbers have since moved: a
pre-registration that gets rewritten once the answer is known is worth less than none.

reports/sample regenerated from a real run and re-anonymized per the repo's convention.
Its README no longer claims "nothing measured yet" - that was two findings out of date.
12 mypy errors, all predating this work, and one of them was hiding a real defect.

repository.upsert_disclosure guarded `row` for None on one line and then dereferenced it
on the next, with a `# type: ignore[return-value]` over the top. INSERT ... ON CONFLICT DO
UPDATE ... RETURNING always yields a row so it never fired, but the failure mode it was
papering over is handing the caller a None model that reads as a successfully stored
disclosure. Both invariants are now checked and raise.

The rest were typing, not behaviour:
  - pg_insert(Model.__table__) -> pg_insert(Model)  (5 call sites: prices, cluster,
    returns, and two in the graph repository)
  - Result.rowcount -> cast to CursorResult, which is what execute() actually returns
  - TsgClient's playwright handles were annotated None by inference
  - BeautifulSoup types an href as str | AttributeValueList

Changing five live INSERT statements to satisfy a type checker is exactly the kind of edit
that type-checks and then writes nothing, so they are verified rather than assumed. Four
were already covered. The fifth - insert(Person)/insert(PersonCompanyRole) in
upsert_management_roles - was only reachable through a test that scrapes KAP live, and so
could not run at all while the WAF was blocking. It now has an offline test that writes
synthetic board members and re-runs them, which also pins the idempotency the method's own
docstring claims (a NULL valid_from makes ON CONFLICT never fire, so it DELETEs first).

Remaining 4 integration failures are unrelated and pre-existing: they assert on graph and
Ticaret Sicil seed data that was never loaded (unlisted_companies is 0 - that scraper needs
a hand-solved CAPTCHA), and on a KAP board scrape that returns nothing while the WAF is up.
@caganco caganco changed the title scrapers: silent parse failures - five years of history, and 21% of the rest The signal is real. It is not tradeable. (and the six silent data faults that hid both) Jul 13, 2026
@caganco
caganco merged commit 27c84a8 into master Jul 13, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant