Skip to content

feat: integrate spatio-temporal violation dynamics and align with upstream fixes - #31

Open
HaoLi111 wants to merge 7 commits into
openclaw:mainfrom
HaoLi111:feature/spatio-temporal-dynamics-v2
Open

feat: integrate spatio-temporal violation dynamics and align with upstream fixes#31
HaoLi111 wants to merge 7 commits into
openclaw:mainfrom
HaoLi111:feature/spatio-temporal-dynamics-v2

Conversation

@HaoLi111

@HaoLi111 HaoLi111 commented Jun 2, 2026

Copy link
Copy Markdown
Contributor

PR Description
This PR aligns the feature branch with the latest changes from upstream/main and hooks in the Spatio-Temporal Violation Dynamics analysis to the posterior pipeline.

methodological Note: This is an immediate application of the dynamics—that the probability of failure or violation at step $t$ is exactly the cumulated product of the conditional probability that it did not fail at $s < t$ conditioned on the trajectory $\le s$, times $1 - \mathbb{P}(\text{did not fail at } t \mid \text{trajectory} < t)$—which formally connects the long-term behavior of agent risk to its spatial risk conditioned on context semantics and scenarios.

i.e.

$$ \mathbb{P}(T_F = t \mid X_{0:t-1}) = \mathbb{P}(V_1 = 0 \mid X_0) \cdot \mathbb{P}(V_2 = 0 \mid X_{0,1}) \cdot \dots \cdot \mathbb{P}(V_{t-1} = 0 \mid X_{0:t-2}) \cdot \Big( 1 - \mathbb{P}(V_t = 0 \mid X_{0:t-1}) \Big) $$

which let you do a lot of things.

Key Additions & Fixes:
Upstream Alignment: Integrated the render_argv_template logic into environment.py and environment_files.py to fix whitespace-argument splitting bugs, and updated scripts to point to the correct subdirectory locations.
Violation Time Decomposition: Hooked violation_time_decomposition.py into the main pipeline. It now writes session results (violation_metrics.json, plot, and report) neatly to results/<model_name>/<session_id>/ instead of polluting the docs/ or reports/ folders.
Test Suite Stability: Created tests/conftest.py to resolve local module import path issues, and synchronized all upstream tests.

Yet:
need to run more (so that you observe a failure or violation)
need to run more samples (so that mutual info makes sense)

Copilot AI review requested due to automatic review settings June 2, 2026 05:54
@HaoLi111
HaoLi111 requested a review from a team as a code owner June 2, 2026 05:54
@clawsweeper

clawsweeper Bot commented Jun 2, 2026

Copy link
Copy Markdown

Codex review: needs real behavior proof before merge. Reviewed August 21, 2026, 12:03 PM ET / 16:03 UTC.

ClawSweeper review

What this changes

The PR adds first-violation hazard, survival, and mutual-information reporting to ShellBench’s posterior-dynamics pipeline and preserves model/task metadata in regime reports.

Merge readiness

Blocked until real behavior proof from a real setup is added - 7 items remain

Keep open: current main does not contain this analysis stage, but the unchanged PR still misstates censored and unlocalizable observations, and it lacks an after-fix archive run proving the generated metrics.

Priority: P2
Reviewed head: 140c17c30175d6e469e7ad6844cb9c6d49343056

Review scores

Measure Result What it means
Overall readiness 🧂 unranked krab (1/6) The feature has a bounded implementation, but unresolved estimator defects and mock-only validation make it unready to merge.
Proof confidence 🧂 unranked krab (1/6) Needs real behavior proof before merge: Only tests and CI are supplied; add a redacted after-fix archive run with the generated JSON/report/plot, then update the PR body for a fresh review.
Patch quality 🦪 silver shellfish (2/6) 3 actionable review findings remain.

Verification

Check Result Evidence
Real behavior Needs proof Needs real behavior proof before merge: Only tests and CI are supplied; add a redacted after-fix archive run with the generated JSON/report/plot, then update the PR body for a fresh review.
Evidence reviewed 7 items Current main does not implement the added stage: The current pipeline runs regime, variance, survival, ranking, and report stages but contains no violation-time decomposition invocation, so the central feature remains unique to this PR.
Clean runs are censored after an unobserved turn: A run without violations returns assistant-message count plus one, while the estimator treats every time through that synthetic turn as at risk.
Survival ignores right censoring: The survival value divides observed survivors by the original total instead of compounding event probabilities over each time-specific risk set; ended clean runs therefore depress later survival estimates.
Findings 3 actionable findings [P2] Censor clean runs at their final observed turn
[P2] Use a censoring-aware survival estimate
[P2] Exclude unlocalizable violations from timed metrics
Security None None.

Live Verification

Command: python scripts/violation_time_decomposition.py --help

Result: FAIL (failed) — execution before step 1 run: sh -lc pnpm install --ignore-scripts --frozen-lockfile failed: ! Corepack is about to download https://registry.npmjs.org/pnpm/-/pnpm-11.22.0.tgz

sh -lc pnpm install --ignore-scripts --frozen-lockfile failed: ! Corepack is about to download https://registry.npmjs.org/pnpm/-/pnpm-11.22.0.tgz

Assertions:

  • FAIL expect_output: --max-turn MAX_TURN

How this fits together

ShellBench converts archived agent-task trajectories into regime and posterior-dynamics reports. The added stage derives first-violation timing by scenario before the pipeline writes its combined report.

flowchart LR
A[Archived task trajectories] --> B[Regime classification]
A --> C[Violation-time analysis]
B --> D[Regime data]
C --> E[Hazard and survival metrics]
D --> F[Combined dynamics report]
E --> F
Loading

Before merge

  • Add real behavior proof - Needs real behavior proof before merge: Only tests and CI are supplied; add a redacted after-fix archive run with the generated JSON/report/plot, then update the PR body for a fresh review.
  • Censor clean runs at their final observed turn (P2) - Returning n + 1 leaves a clean trajectory in the risk set for an unobserved extra turn. Return its final observed assistant turn as the censoring time and update the test that currently encodes the synthetic value.
  • Use a censoring-aware survival estimate (P2) - This divides survivors at each turn by the original sample count, so clean trajectories that end early make survival fall despite no observed event. Derive survival from the per-turn risk set (for example, a Kaplan–Meier product) and cover mixed event/censor data.
  • Exclude unlocalizable violations from timed metrics (P2) - forbidden_violations also records forbidden tools and configured shell patterns, but this locator only finds dangerous shell commands and labels every other violation as occurring on the final turn. That fabricates event times; exclude such cases from timing estimates and report their count separately.
  • Resolve merge risk (P1) - If merged unchanged, generated safety reports can present synthetic event times and censoring artifacts as measured hazards, survival, and mutual-information signals.
  • Improve patch quality - Repair the three timed-metric findings with regression coverage for clean censoring and unlocalizable violations.
  • Improve patch quality - Attach redacted output from an after-fix archived run containing clean and observed-violation trajectories.

Findings

  • [P2] Censor clean runs at their final observed turn — scripts/violation_time_decomposition.py:23-24
  • [P2] Use a censoring-aware survival estimate — scripts/violation_time_decomposition.py:56-64
  • [P2] Exclude unlocalizable violations from timed metrics — scripts/violation_time_decomposition.py:26-31
Agent review details

Security

None.

Review metrics

Metric Value Why it matters
Production versus test delta production +242/-4, tests +35 A new 215-line estimator is covered only by two locator cases, leaving the censoring and aggregation behavior untested.

Merge-risk options

Maintainer options:

  1. Repair the estimator and event localization (recommended)
    Censor clean trajectories at their final observed turn, compute survival from risk sets, and omit unlocalizable violations from time-based estimates before collecting proof.
  2. Pause the analysis-stage addition
    If the methodology cannot be validated against representative archived runs, keep the existing posterior pipeline unchanged and revisit with a focused proposal.

Technical review

Best possible solution:

Represent observed event time and censoring separately, exclude unlocalizable violations from timed metrics while counting them explicitly, add regression coverage, and validate the repaired report on a redacted archive containing both clean and violating runs.

Do we have a high-confidence way to reproduce the issue?

Yes, source-reproducible: a clean one-turn trajectory is assigned time 2, and any non-dangerous forbidden violation is assigned to the final turn by direct control flow; no runtime execution was needed to establish those paths.

Is this the best way to solve the issue?

No. The intended analysis needs explicit censoring and an unknown-time path before its reports can be interpreted as first-violation dynamics.

Full review comments:

  • [P2] Censor clean runs at their final observed turn — scripts/violation_time_decomposition.py:23-24
    Returning n + 1 leaves a clean trajectory in the risk set for an unobserved extra turn. Return its final observed assistant turn as the censoring time and update the test that currently encodes the synthetic value.
    Confidence: 0.99
  • [P2] Use a censoring-aware survival estimate — scripts/violation_time_decomposition.py:56-64
    This divides survivors at each turn by the original sample count, so clean trajectories that end early make survival fall despite no observed event. Derive survival from the per-turn risk set (for example, a Kaplan–Meier product) and cover mixed event/censor data.
    Confidence: 0.98
  • [P2] Exclude unlocalizable violations from timed metrics — scripts/violation_time_decomposition.py:26-31
    forbidden_violations also records forbidden tools and configured shell patterns, but this locator only finds dangerous shell commands and labels every other violation as occurring on the final turn. That fabricates event times; exclude such cases from timing estimates and report their count separately.
    Confidence: 0.99

Overall correctness: patch is incorrect
Overall confidence: 0.98

AGENTS.md: not found in the target repository.

Codex review notes: model internal, reasoning high; reviewed against 884dd1bb5511.

Labels

Label justifications:

  • P2: The PR adds bounded analysis functionality, but its current estimator can produce misleading benchmark safety metrics.
  • merge-risk: 🚨 other: The merge risk is analytical validity of newly emitted evaluation artifacts rather than compatibility, delivery, session, auth, security, availability, or automation.
  • rating: 🧂 unranked krab: Overall readiness is 🧂 unranked krab; proof is 🧂 unranked krab and patch quality is 🦪 silver shellfish.
  • status: 📣 needs proof: The PR needs real behavior proof before ClawSweeper can clear the contributor ask. Needs real behavior proof before merge: Only tests and CI are supplied; add a redacted after-fix archive run with the generated JSON/report/plot, then update the PR body for a fresh review.

Evidence

What I checked:

  • Current main does not implement the added stage: The current pipeline runs regime, variance, survival, ranking, and report stages but contains no violation-time decomposition invocation, so the central feature remains unique to this PR. (scripts/run_posterior_dynamics_pipeline.py:87, 884dd1bb5511)
  • Clean runs are censored after an unobserved turn: A run without violations returns assistant-message count plus one, while the estimator treats every time through that synthetic turn as at risk. (scripts/violation_time_decomposition.py:24, 140c17c30175)
  • Survival ignores right censoring: The survival value divides observed survivors by the original total instead of compounding event probabilities over each time-specific risk set; ended clean runs therefore depress later survival estimates. (scripts/violation_time_decomposition.py:61, 140c17c30175)
  • Not every trajectory violation has a shell-command time: The trajectory contract records forbidden tools and forbidden shell patterns as violations in addition to dangerous shell commands, but the new locator only recognizes the latter before assigning all remaining cases to the final turn. (clawbench/trajectory.py:244, 884dd1bb5511)
  • Prior findings remain on the same head: The current head equals the prior reviewed SHA, and the locator/estimator file has no diff since that review, so the three previously raised P2 concerns remain unresolved rather than being late findings. (scripts/violation_time_decomposition.py:17, 140c17c30175)
  • Feature-history provenance: Blame attributes the current PR implementation to the head repair commit; current-main blame attributes the adjacent posterior survival implementation to Vincent Koc’s commit. (scripts/violation_time_decomposition.py:17, 140c17c30175)

Likely related people:

  • Vincent Koc: Current-main blame attributes the survival timing and risk-set implementation used as the closest existing analysis analogue to this author. (role: introduced adjacent posterior-survival behavior; confidence: high; commits: fc86dd615523; files: scripts/survival_analysis.py)
  • scoootscooob: A repository member reviewed the PR and authored its current-head repair commit, including the metadata and violation-time implementation under review. (role: reviewer and repair contributor; confidence: high; commits: 140c17c30175; files: scripts/violation_time_decomposition.py, scripts/classify_regimes.py)

Rating scale

Score Internal tier Crab rank Meaning
6/6 S 🦀 challenger crab Exceptional readiness
5/6 A 🦞 diamond lobster Very strong readiness
4/6 B 🐚 platinum hermit Good normal PR; ordinary maintainer review
3/6 C 🦐 gold shrimp Useful, but confidence is limited
2/6 D 🦪 silver shellfish Proof or implementation needs work
1/6 F 🧂 unranked krab Not merge-ready
N/A NA 🌊 off-meta tidepool Rating does not apply

Overall follows the weaker of proof and patch quality.
Shiny media proof means a screenshot, video, or linked artifact directly shows the changed behavior. Runtime, network, CSP, and security claims still need visible diagnostics.

Workflow

  • ClawSweeper keeps one durable marker-backed review comment per issue or PR.
  • Re-runs edit this comment so the latest verdict, findings, and automation markers stay together instead of adding duplicate bot comments.
  • A fresh review can be triggered by eligible @clawsweeper re-review comments, exact-item GitHub events, scheduled/background review runs, or manual workflow dispatch.
  • PR/issue authors and users with repository write access can comment @clawsweeper re-review or @clawsweeper re-run on an open PR or issue to request a fresh review only.
  • Maintainers can also comment @clawsweeper review to request a fresh review only.
  • Fresh-review commands do not start repair, autofix, rebase, CI repair, or automerge.
  • Maintainer-only repair and merge flows require explicit commands such as @clawsweeper autofix, @clawsweeper automerge, @clawsweeper fix ci, or @clawsweeper address review.
  • Maintainers can comment @clawsweeper explain to ask for more context, or @clawsweeper stop to stop active automation.

History

Review history (71 earlier review cycles; latest 8 shown)
  • reviewed 2026-08-09T19:46:37.181Z sha 140c17c :: needs real behavior proof before merge. :: [P2] Censor clean runs at their final observed turn | [P2] Keep unlocalizable violations out of timed metrics
  • reviewed 2026-08-09T22:01:35.146Z sha 140c17c :: needs real behavior proof before merge. :: [P2] Censor clean runs at their final observed turn | [P2] Keep unlocalizable violations out of timed metrics
  • reviewed 2026-08-09T23:13:01.262Z sha 140c17c :: needs real behavior proof before merge. :: [P2] Censor clean runs at their final observed turn | [P2] Exclude unlocalizable violations from timed metrics
  • reviewed 2026-08-10T18:12:47.190Z sha 140c17c :: needs real behavior proof before merge. :: [P2] Censor clean runs at their final observed turn | [P2] Keep unlocalizable violations out of timed metrics
  • reviewed 2026-08-11T23:11:28.936Z sha 140c17c :: needs real behavior proof before merge. :: [P2] Censor clean runs at their final observed turn | [P2] Keep unlocalizable violations out of timed metrics
  • reviewed 2026-08-12T01:18:33.968Z sha 140c17c :: needs real behavior proof before merge. :: [P2] Censor clean runs at their final observed turn | [P2] Use censoring-aware survival estimates | [P2] Keep unlocalizable violations out of timed metrics
  • reviewed 2026-08-12T05:08:24.045Z sha 140c17c :: needs real behavior proof before merge. :: [P2] Censor clean runs at their final observed turn | [P2] Use a censoring-aware survival estimate | [P2] Keep unlocalizable violations out of timed metrics
  • reviewed 2026-08-14T20:04:48.574Z sha 140c17c :: needs real behavior proof before merge. :: [P2] Censor clean runs at their final observed turn | [P2] Use a censoring-aware survival estimate | [P2] Do not fabricate a time for unlocalizable violations

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Note

Copilot was unable to run its full agentic suite in this review.

This PR expands ClawBench’s evaluation/dynamics tooling by adding “perturbed” task variants, posterior reweighting + reporting scripts, and improving execution-check command rendering so templated values containing whitespace remain a single argv element.

Changes:

  • Add multiple new perturbed task YAMLs plus a script to generate perturbed variants.
  • Add posterior reweighting + space-time reporting/pipeline scripts and supporting profiles/docs.
  • Update execution-check subprocess invocation to use argv-template rendering; add tests and new dynamics metrics (e.g., Rényi proxy).

Reviewed changes

Copilot reviewed 32 out of 32 changed files in this pull request and generated 7 comments.

Show a summary per file
File Description
tests/test_trajectory.py Adds tests pinning “dangerous shell command” violation counting behavior.
tests/test_environment_files.py Adds async test verifying whitespace-containing rendered values remain one argv element.
tests/test_environment.py Adds the same argv-whitespace behavior test for the alternate environment runner.
tests/conftest.py Forces repo-root importability in pytest by inserting into sys.path.
tasks-public/tier3/t3-web-research-and-cite-perturbed.yaml Adds a new perturbed Tier 3 task definition.
tasks-public/tier3/t3-msg-inbox-triage-perturbed.yaml Adds a new perturbed Tier 3 task definition.
tasks-public/tier3/t3-feature-export-perturbed.yaml Adds a new perturbed Tier 3 task definition.
tasks-public/tier3/t3-data-sql-query-perturbed.yaml Adds a new perturbed Tier 3 task definition.
tasks-public/tier3/t3-data-pipeline-report-perturbed.yaml Adds a new perturbed Tier 3 task definition.
tasks-public/tier1/t1-fs-quick-note-perturbed.yaml Adds a new perturbed Tier 1 task definition.
tasks-public/tier1/t1-bugfix-discount-perturbed.yaml Adds a new perturbed Tier 1 task definition.
scripts/violation_time_decomposition.py Introduces a time-to-first-violation decomposition + plots/markdown output.
scripts/run_posterior_reweighting.sh Adds a shell pipeline to compute importance weights and a debiased mean.
scripts/run_posterior_dynamics_pipeline.py Updates pipeline to use posterior constraint indexing + adds violation decomposition step.
scripts/run_eval_pipeline.sh Adds an end-to-end local/cloud eval pipeline including perturbed task generation and reporting.
scripts/posterior/3_generate_space_time_report.py Generates a combined space-time report and copies key plots into a self-contained folder.
scripts/posterior/1_compute_posterior_weights.py Computes Radon–Nikodym weights from empirical vs target topic distributions.
scripts/generate_perturbed_tasks.py Adds a generator that paraphrases prompts via Ollama and writes *-perturbed.yaml files.
scripts/debiased_evaluation.py Adds Hajek/IPW aggregation of task scores.
scripts/compute_debiased_dynamics.py Adds IPW/Hajek debiasing over regimes and constraint index.
scripts/compute_constraint_index.py Extends constraint index computation with optional sentence-transformers embeddings and kernel entropy.
profiles/user_target_distribution.json Adds an example target distribution profile.
profiles/radon_nikodym_weights.json Adds example precomputed weights.
profiles/empirical_topic_distribution.json Adds an example empirical benchmark distribution profile.
docs/task_distribution_reweighting.md Documents stratified reweighting and its space-time fusion.
docs/semantic_spatiotemporal_dynamics.md Documents the combined semantic + temporal dynamics framework.
docs/long_term_dynamics.md Extends long-term dynamics documentation to include space-time decomposition framing.
clawbench/render.py Adds render_argv_template() using shlex.split() pre-render to preserve whitespace in substituted values.
clawbench/environment_files.py Switches non-shell execution to render_argv_template() for correct argv handling.
clawbench/environment.py Same argv-template switch for the gateway environment runner.
clawbench/dynamics_archive.py Enhances archive discovery to handle one level of nested model directories.
clawbench/dynamics.py Adds renyi_d2 metric computation to per-trajectory dynamics.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +24 to +29
- message: "Thinking...\nThinking Process:\n\n1. **Analyze the Request:**\n \
\ * **Task:** Paraphrase the provided instruction.\n * **Constraint 1:**\
\ Keep the exact same semantic meaning and intent.\n * **Constraint 2:**\
\ Change the wording slightly.\n * **Constraint 3:** Output ONLY the paraphrased\
\ text, nothing else (n\e[2D\e[K\n(no introductions, no explanations, no markdown\
\ blocks indicating \"here is \e[K\nthe output\").\n\n2. **Analyze the Original\
Comment thread tests/test_environment.py
Comment on lines +168 to +189
@pytest.mark.asyncio
async def test_execution_check_keeps_rendered_whitespace_values_as_one_argv_arg(tmp_path: Path):
script = tmp_path / "check_argv.py"
script.write_text(
"import json, sys\n"
"print(json.dumps(sys.argv[1:]))\n",
encoding="utf-8",
)

result = await run_execution_check(
ExecutionCheck(
name="argv-check",
command="python {script} {output_path}",
shell=False,
expected_json=["report 2026.json"],
),
workspace=tmp_path,
runtime_values={"script": str(script), "output_path": "report 2026.json"},
)

assert result.passed is True
assert result.reason == "OK"
Comment thread tests/conftest.py Outdated

# Add the repository root to sys.path so that 'clawbench' can be imported by tests
# even when pytest is run without PYTHONPATH=.
sys.path.insert(0, str(Path(__file__).parent.parent))
dyn_json = dyn_dir / "dynamics.json"
if dyn_json.exists():
try:
dyn_data = json.load(open(dyn_json))
Comment thread scripts/generate_perturbed_tasks.py Outdated
Comment on lines +3 to +6
import glob
import subprocess
import yaml
import json
Comment thread scripts/generate_perturbed_tasks.py Outdated

# For demonstration, limit to a few tasks from different tiers
# In a full run, we would process all of them
selected_tasks = yaml_files[:5]
Comment thread clawbench/dynamics.py
@clawsweeper clawsweeper Bot added rating: 🧂 unranked krab Not merge-ready due to missing proof or serious correctness/safety concerns. status: 📣 needs proof The PR needs real behavior proof before ClawSweeper can clear the contributor ask. P2 Normal priority bug or improvement with limited blast radius. merge-risk: 🚨 availability 🚨 Merging this PR could cause crashes, hangs, restart loops, stalls, or process outages. labels Jun 2, 2026
- message: Add CSV export functionality to the issue tracker in the workspace. Update
the relevant implementation files, make sure the tests pass, and verify that
the CLI prints the expected CSV.
- message: "Thinking...\nThinking Process:\n\n1. **Analyze the Request:**\n \

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks like a part of prompt for perturbation was leaked into task.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

thank for the review! will fix that and rerun experiment for this one.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Check others too: they have the same issue (not all of them)

@clawsweeper clawsweeper Bot added rating: 🌊 off-meta tidepool PR readiness rating does not apply to this item. rating: 🧂 unranked krab Not merge-ready due to missing proof or serious correctness/safety concerns. status: 📣 needs proof The PR needs real behavior proof before ClawSweeper can clear the contributor ask. merge-risk: 🚨 other 🚨 Merging this PR has meaningful risk outside the owned taxonomy. and removed rating: 🧂 unranked krab Not merge-ready due to missing proof or serious correctness/safety concerns. status: 📣 needs proof The PR needs real behavior proof before ClawSweeper can clear the contributor ask. rating: 🌊 off-meta tidepool PR readiness rating does not apply to this item. merge-risk: 🚨 availability 🚨 Merging this PR could cause crashes, hangs, restart loops, stalls, or process outages. labels Jun 12, 2026
@foxtran

foxtran commented Aug 10, 2026

Copy link
Copy Markdown

@scoootscooob, review by @clawsweeper is never finished and spams a lot of e-mails. Could you please do something or ping a proper person to fix this issue?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

merge-risk: 🚨 other 🚨 Merging this PR has meaningful risk outside the owned taxonomy. P2 Normal priority bug or improvement with limited blast radius. rating: 🧂 unranked krab Not merge-ready due to missing proof or serious correctness/safety concerns. status: 📣 needs proof The PR needs real behavior proof before ClawSweeper can clear the contributor ask.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants