Skip to content

test(adversarial): P3Adversarial.t.sol — task_240926_17 Part 2 - #218

Open
robertocarlous wants to merge 1 commit into
InfiniteZeroFoundation:developfrom
robertocarlous:feat/issue-154-adversarial-tests
Open

robertocarlous wants to merge 1 commit into
InfiniteZeroFoundation:developfrom
robertocarlous:feat/issue-154-adversarial-tests

Conversation

@robertocarlous

Copy link
Copy Markdown
Collaborator

Summary

Closes Part 2 of task_240926_17 (issue #154).

Adds foundry/test/P3Adversarial.t.sol with 13 tests covering every row of Developer/design/adversarial-threat-model.md.

Each test is named test_{class}_{description} where class is costBounded, defended, or knownGap, matching the row classification in the threat model doc.

Row Class Test
1 costBounded test_costBounded_sybilNoParticipation
2 costBounded test_costBounded_sybilSeatCapture_stakeIsKTimesMinStake
3 defended test_defended_recidivistS5Escalation
4 knownGap test_knownGap_auditorBlocPoisonedModel
5 knownGap test_knownGap_t1WrongCID_finalCIDUnchangedDissenterSlashed
6 knownGap test_knownGap_dualRoleRegistration
7a knownGap test_knownGap_ownerSuppressesDispute_rejected
7b knownGap test_knownGap_ownerSuppressesDispute_forfeitedViaSettlement
7c knownGap test_knownGap_ownerSuppressesDispute_bondLockedIfNeverResolved
8 defended test_defended_disputeSeedSubgroupDeterministic
9a costBounded test_costBounded_inflatedScores_oddCount
9b costBounded test_costBounded_inflatedScores_evenCount_colluderMovesMedianHalfway
11 knownGap test_knownGap_perSlasherS5Evasion

All 13 tests pass (forge test --match-contract P3AdversarialTest).

Accounts for all three develop changes Umer flagged in discussion #164:

  • Shuffle-steering rows use lockAggSeed/lockAuditSeed (H-2 fix)
  • T1/T2 use commit-reveal flows
  • Forfeiture uses 50/50 burn/treasury split

Test plan

  • forge test --match-contract P3AdversarialTest -vv shows 13/13 passing

13 tests covering all rows of the adversarial threat model (issue InfiniteZeroFoundation#154).
Maps each row to a test named test_{costBounded|defended|knownGap}_{desc}.

Rows covered:
  1  cost-bounded: Sybil no-participation (unassigned agg, S2 non-participant)
  2  cost-bounded: Sybil seat capture cost = k * MIN_STAKE
  3  defended:     S5 recidivism escalation fires on threshold-th miss
  4  knownGap:     auditor bloc inflated scores in S3 shadow mode (no slash)
  5  knownGap:     T1 wrong CID persists post-dispute; honest dissenter slashed
  6  knownGap:     dual-role cross-registration (agg + auditor, same GI)
  7a knownGap:     owner rejects dispute (resolveDispute=false) — bond forfeited
  7b knownGap:     owner suppresses via settleRecomputation(false) — bond forfeited
  7c knownGap:     bond locked forever when owner never resolves
  8  defended:     dispute-seed subgroup is deterministic; re-lock reverts
  9a cost-bounded: inflated score with odd reveal count — outlier can't shift median
  9b cost-bounded: even reveal count — colluder moves median halfway
  11 knownGap:     per-slasher S5 ring evasion across two task contracts
@robertocarlous
robertocarlous requested a review from umeradl October 5, 2026 11:11
@robertocarlous

Copy link
Copy Markdown
Collaborator Author

Cc: @umeradl

@umeradl

umeradl commented Oct 5, 2026

Copy link
Copy Markdown
Member

Reviewed against develop in an isolated worktree at PR head 1b849a5 (1 commit, merge-base 740a613). develop has moved to a1fcce2 since (PR No. 217, docs only), which doesn't touch this file. git merge-tree --write-tree is clean locally, and GitHub reports mergeable: MERGEABLE, mergeStateStatus: CLEAN. I built on the real via_ir profile (after npm ci), ran the suite, and deliberately broke what three of the tests claim to guard.

Claimed: 13 tests, one per row of adversarial-threat-model.md (Row 10 deferred), all passing.

Verified — exact match: forge test --match-contract P3AdversarialTest → 13 passed; 0 failed. Full suite: Ran 40 test suites … 504 tests passed, 0 failed, 0 skipped, i.e. develop's 491 plus these 13. Row 10 (duplicate-update gaming) is "Deferred — no test in Part 2" in the threat model, so 13 tests for the other 10 rows (7 and 9 split into sub-cases) is the right count. Each test's class prefix matches its row's classification.

Claimed: the tests encode each row's expected outcome.

Verified by breaking the code under test:

Break Expected Result
S5 escalation >= → > (DinValidatorStake.sol:311) Row 3 and Row 11 fail ✅ both fail (test_defended_recidivistS5Escalation, test_knownGap_perSlasherS5Evasion)
Dispute subgroup shuffled with keccak256(block.timestamp, msg.sender) instead of the locked seed (_assignFreshSubgroup, DINTaskCoordinator.sol:1728) Row 8 (test_defended_disputeSeedSubgroupDeterministic) fails ❌ still passes. The test only asserts the three reEvaluationAssignees are non-zero. Its comment says "re-run with same seed → same result", but nothing re-derives or compares the subgroup, so a grindable draw passes.
Row 4's expected AuditorScoreDeviation changed to (1, 0, 999, auditor, 0, 0, 0, false) Row 4 fails ❌ still passes. vm.expectEmit(true, true, true, false) checks only the three indexed topics (gi, batchId, auditor). modelIndex, the scores, deviation and exceedsThreshold aren't compared, so "asserts exceedsThreshold = true" isn't actually tested. Row 9a uses the same pattern.

All sources restored afterwards; the worktree is clean.

Against the threat model's "Expected test outcome" column (which task_240926_17 says the tests are written against):

Row Expected outcome says Test Gap
1 (a) N identities register, (b) unassigned not penalised, (c) a non-submitter gets S2 and on the 3rd GI, S5 (b) and S2 for one GI (c) not covered
2 asserts the fraction of batches where Sybils hold a majority (≈ 3p² − 2p³) and the cost k × MIN_STAKE only sums the k stakes and checks they're active no batch draw / seat-majority measurement, so the test can't fail on a seat-capture change
3 escalation on the 3rd miss, then a 4th registration attempt reverts asserts isValidatorActive == false registration revert not checked
4, 9a AuditorScoreDeviation with exceedsThreshold = true expectEmit(..., false) data not checked (above)
8 the subgroup is deterministic for the locked seed; after seed-window expiry a re-lock anchors on a new block lock-before-seedBlock revert and double-lock revert ✅; subgroup only non-zero determinism and re-anchor not tested (above)
5, 6, 7a–c, 9b, 11 match the expected outcome —

Against task_240926_17 Part 2's rules:

  • "Run scenarios through the real GI lifecycle … Cross-GI attacks (S5/S6/Sybil) span ≥2 GIs." Rows 3 and 11 call stake.slashPartial directly with vm.prank(address(tc)) and a made-up GI index; no GI runs. Row 3 is then the same path as SlashingInvariants.t.sol's test_s5_recidivism_escalatesToFullSlashOnThreshold, which the spec says not to re-test.
  • "Each test gets an attack-narrative NatSpec block … KNOWN GAP tests … say so in each one's NatSpec." The tests use // comments, and no knownGap test says it must be flipped to test_defended_ when the gap closes (grep -ci flip → 0).

Interplay with open plans: task-plan-051026-1 (PR No. 219) closes Row 6 (No. 180) and Row 11 (No. 193). Whichever merges second flips test_knownGap_dualRoleRegistration and test_knownGap_perSlasherS5Evasion. Row 7a–c cover the coordinator's S4 aggregation disputes, not the test-data disputes PR No. 215 changed, so they're unaffected by it. Nothing here re-reports a finding already in foundry-src-security-review.md.


The suite builds, passes, and Rows 5, 6, 7a–c, 9b and 11 encode their expected outcomes. But three "defended"/"costBounded" checks can't fail on the regression they're named for (Rows 4/9a event data, Row 8 determinism, Row 2 seat capture), and Rows 1 and 3 miss part of their expected outcome. Recommendation: changes requested before merge:

  1. Rows 4 and 9a: vm.expectEmit(true, true, true, true) so exceedsThreshold, the scores and deviation are compared.
  2. Row 8: derive the expected subgroup from the locked seed (or resolve twice from identical state via vm.snapshotState/vm.revertToState with a different block.timestamp/caller and assert equality), and add the post-expiry re-anchor case.
  3. Row 2: draw real batches over the Sybil pool and assert the majority fraction against the hypergeometric bound, or reclassify the row in the threat model if that's out of reach.
  4. Row 1 (c) and Row 3's "4th registration reverts": run them through real GIs (≥2, as Part 2 requires) rather than direct slashPartial pranks.
  5. NatSpec (///) attack-narrative blocks, with each knownGap test naming the issue that flips it.

@umeradl

umeradl commented Oct 5, 2026

Copy link
Copy Markdown
Member

Files changed (1) — as of 1b849a5 (PR head)

Diffed against merge-base 740a613 (develop). develop has moved to a1fcce2 since (PR No. 217, docs only), and nothing since touches foundry/test/P3Adversarial.t.sol. GitHub agrees: mergeable: MERGEABLE, mergeStateStatus: CLEAN. A local git merge-tree --write-tree dry run is clean (exit 0).

foundry/test/P3Adversarial.t.sol

Field Value
Change New
Lines +845/-0
Diff (what exactly is in this PR) 13 tests named test_{costBounded,defended,knownGap}_…, one per threat-model row (7 and 9 split, 10 deferred), with platform/task-pair fixtures, seed-lock helpers, sender-bound T1/T2 commit-reveal helpers, and a second task pair for Row 11.
Functionality — how & why How: each test drives a real platform + task pair. Rows 1, 2, 4, 5, 6, 7, 8 and 9 go through registration, seed locks, batch creation and commit-reveal (_setupToT1Open, _runToT1Finalized, _runToAuditorsSlashed). Rows 3 and 11 call slashPartial directly as the task contract. Why: issue No. 154 Part 2. Each threat-model row gets a test whose class says whether the protocol defends, bounds the cost of, or knowingly allows the attack, so a later fix fails a knownGap test and a regression fails a defended one. The build passes and 13/13 pass (504/504 overall). Breaking S5 escalation fails Rows 3 and 11 as it should.
Diff vs current develop HEAD None — new file
Recommended merge proposal Changes requested before merge (detail in the verification comment): (1) Rows 4/9a expectEmit(…, false) doesn't compare exceedsThreshold/scores, verified by changing every data field and still passing; (2) Row 8 asserts only non-zero assignees, so a timestamp/caller-seeded shuffle still passes; (3) Row 2 sums stakes and never measures seat capture; (4) Row 1 (c) S5 and Row 3's "4th registration reverts" are missing, and Rows 3/11 bypass the real GI lifecycle Part 2 requires; (5) NatSpec attack narratives and flip notes on knownGap tests.
Actual merge proposal Soon
Pending proposal Author fixes 1–5. Row 6 / Row 11 flips coordinated with PR No. 219 (whichever merges second).
Local merge conflict No
GitHub merge conflict No

Verification

Full detail is in the verification comment above. In short: via_ir forge build passes. P3AdversarialTest 13/13, full suite 504/504 in 40 suites. Break-then-fix: S5 is caught (2 tests fail), but the Row 8 seed and the Row 4 event data are not caught. GitHub CI is green.

Local vs. GitHub agree: both report a clean merge.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants