Skip to content

Add an experimental observer agent to trace the tracer's own tests - #12704

Draft
daniel-mohedano wants to merge 1 commit into
masterfrom
daniel.mohedano/tracing-the-tracer
Draft

daniel-mohedano wants to merge 1 commit into
masterfrom
daniel.mohedano/tracing-the-tracer

Conversation

@daniel-mohedano

Copy link
Copy Markdown
Contributor

What Does This Do

Adds an experimental way to trace this repository's own test runs with CI Visibility, while the tracer under test runs in the same JVMs.

./gradlew :dd-java-agent:observerJar builds an "observer" agent. It is the stock agent, rewritten offline into a private datadog.trace.observer namespace. You attach it to the Gradle daemon with a normal -javaagent, configure it with TRACING_OBSERVER_CONFIG_* variables or tracing.observer.config.* properties, and run tests with -PtraceTracer=true. docs/tracing_the_tracer.md explains the setup, configuration, isolation model and limitations.

Main changes:

  • Observer rewriter and runtime (buildSrc, developer-only)
    • Relocates classes, resources, service files, indexes, Byte Buddy and the field-backed context protocol, so the two tracers don't share globals or injected fields.
    • Routes System property, environment and Boolean.getBoolean reads through a private config view. The observer never reads the subject's DD/OTEL settings, config files or propagation headers.
    • Keeps host paths, config key fragments and wire names that only look like packages. Examples: the default agent and DogStatsD socket paths, .inject.datadog.attribute.enabled, DogStatsD client metric names and OTLP attribute names.
    • Fails the build if the stock layout it patches changes. Writes the jar atomically, so a running JVM keeps its old copy.
    • Only opens observer test spans for the outer, Gradle-launched JUnit/Spock engine. Nested launchers inside tests are ignored.
  • Gradle wiring
    • observerJar task. Not part of assemble, publishing or ordinary test runs.
    • -PtraceTracer=true requires an attached observer and tracks its jar and configuration as test inputs. It fails with configuration cache.
    • The modifiable-config agent recognizes the observer as an allowed extra -javaagent.
    • Spotless now also formats buildSrc Java sources.
  • Stock fixes found while doing this
    • TypeFactory resolves the transform target from the supplied bytes, not a stale cached description. Without this, a later transformer could drop an interface an earlier agent had added. This was reproduced on Mockito's DetachedThreadLocal, where it aborted a whole retransformation batch.
    • Local git detection accepts .git as a file, so runs from linked worktrees and submodules get repository, branch and commit tags again.
    • Per-test coverage loads its recording classes up front. Before, a covered defineClass hook could recurse into a StackOverflowError on the first recording, and the JVM printed java.lang.instrument ASSERTION FAILED.

Motivation

We want CI Visibility data for dd-trace-java's own test runs, collected by our own tracer, without it interfering with the tracer being tested.

Additional Notes

Compatibility

  • Nothing changes for normal builds and tests unless you pass -PtraceTracer=true and attach the observer.
  • The three stock fixes do affect all users. They are included here for now and can move to separate PRs.

Tests

  • New buildSrc tests cover the rewriter, the runtime and the config sources. The config-source tests compare the stock and relocated config providers on real artifacts. 26/26 pass.
  • TypeFactoryTransformTest passes, as do the other outline tests: 12/12.
  • A new UnknownCIInfoTest case covers a worktree-style .git file. The CI Visibility ci, coverage, git and utils tests pass: 584/584.
  • End-to-end runs against a local mock intake, on JDK 25 with JDK 17 workers. These scripts are not part of this PR.
    • Isolation checks, deliberate pass/fail/skip fixtures, logging-only, refused intake, and a representative matrix of existing suites.
    • Test outcomes and IDs matched across JUnit XML, logs and the actual requests.
    • A junit-5.3 run with coverage enabled had no assertions, correct git tags, and per-test coverage for all passing tests.
  • A manual run against a staging account showed complete sessions, git data and test results.

Limitations

  • Experimental. Verified only with this repository's Gradle 9.8 daemon, the JUnit 5 and Spock outer lifecycle, and the instrumentation test harness.
  • Test skipping, Auto Test Retries, Early Flake Detection, Test Management and Failed Test Replay have not been verified with the observer.
  • Every Gradle build opens its own session, including buildSrc, build-logic and Kotlin DSL accessor builds. Those sessions run no tests but still fetch settings. This is existing stock behavior.
  • The TypeFactory change adds one class-file parse per transformed class and has not been benchmarked.

Contributor Checklist

Build a relocated copy of the stock agent that can trace this repository's own test runs with CI Visibility while the tracer under test runs in the same JVMs.

Also fix three stock issues found while doing so: TypeFactory resolving a stale schema for the transform target, missing git data in linked worktrees, and per-test coverage recursing through covered defineClass hooks.
@daniel-mohedano daniel-mohedano added type: feature Enhancements and improvements comp: testing Testing comp: ci visibility Continuous Integration Visibility tag: experimental Experimental changes tag: ai generated Largely based on code generated by an AI or LLM labels Sep 30, 2026
@datadog-datadog-prod-us1-2

This comment has been minimized.

@cit-pr-commenter-54b7da

Copy link
Copy Markdown

CI Visibility Test Environment - sbt-scalatest

Job Status: 🟢 success

Scenario This PR (%) 7d median Δ 7d 30d median Δ 30d runs (7d/30d)
agent 56.87 56.55 $\color{red}{\blacktriangle}$ +0.32 55.43 $\color{red}{\blacktriangle}$ +1.44 98/268
agentEvpProxy 56.51 n/a n/a n/a n/a -

Baseline: median of @test.tracer_overhead on main (gitlab) over the last 7/30 days, per OSS project & scenario. Δ = this PR − baseline median; red ▲ = more overhead, green ▽ = less overhead than baseline.

Tests run to verify that CI Visibility behavior has not regressed in the current PR.

@cit-pr-commenter-54b7da

Copy link
Copy Markdown

CI Visibility Test Environment - netflix-zuul

Job Status: 🟢 success

Scenario This PR (%) 7d median Δ 7d 30d median Δ 30d runs (7d/30d)
agent 89.81 89.58 $\color{red}{\blacktriangle}$ +0.23 87.80 $\color{red}{\blacktriangle}$ +2.01 62/153
agentless 83.78 82.69 $\color{red}{\blacktriangle}$ +1.09 81.05 $\color{red}{\blacktriangle}$ +2.73 34/121
agentlessCodeCoverage 98.38 97.04 $\color{red}{\blacktriangle}$ +1.34 97.04 $\color{red}{\blacktriangle}$ +1.34 34/121
agentlessLineCoverage 115.44 113.88 $\color{red}{\blacktriangle}$ +1.56 111.62 $\color{red}{\blacktriangle}$ +3.82 33/125

Baseline: median of @test.tracer_overhead on main (gitlab) over the last 7/30 days, per OSS project & scenario. Δ = this PR − baseline median; red ▲ = more overhead, green ▽ = less overhead than baseline.

Tests run to verify that CI Visibility behavior has not regressed in the current PR.

@cit-pr-commenter-54b7da

Copy link
Copy Markdown

CI Visibility Test Environment - nebula-release-plugin

Job Status: 🟢 success

Scenario This PR (%) 7d median Δ 7d 30d median Δ 30d runs (7d/30d)
agent 39.38 37.15 $\color{red}{\blacktriangle}$ +2.23 36.42 $\color{red}{\blacktriangle}$ +2.96 62/155
agentless 37.97 36.42 $\color{red}{\blacktriangle}$ +1.55 35.70 $\color{red}{\blacktriangle}$ +2.27 34/127
agentlessCodeCoverage 46.51 45.38 $\color{red}{\blacktriangle}$ +1.13 44.48 $\color{red}{\blacktriangle}$ +2.03 34/127
agentlessLineCoverage 60.49 56.55 $\color{red}{\blacktriangle}$ +3.94 56.55 $\color{red}{\blacktriangle}$ +3.94 34/128

Baseline: median of @test.tracer_overhead on main (gitlab) over the last 7/30 days, per OSS project & scenario. Δ = this PR − baseline median; red ▲ = more overhead, green ▽ = less overhead than baseline.

Tests run to verify that CI Visibility behavior has not regressed in the current PR.

@cit-pr-commenter-54b7da

Copy link
Copy Markdown

CI Visibility Test Environment - heliboard

Job Status: 🟢 success

Scenario This PR (%) 7d median Δ 7d 30d median Δ 30d runs (7d/30d)
agent 10.25 9.92 $\color{red}{\blacktriangle}$ +0.33 9.92 $\color{red}{\blacktriangle}$ +0.33 33/121

Baseline: median of @test.tracer_overhead on main (gitlab) over the last 7/30 days, per OSS project & scenario. Δ = this PR − baseline median; red ▲ = more overhead, green ▽ = less overhead than baseline.

Tests run to verify that CI Visibility behavior has not regressed in the current PR.

@cit-pr-commenter-54b7da

Copy link
Copy Markdown

CI Visibility Test Environment - reactive-streams-jvm

Job Status: 🟢 success

Scenario This PR (%) 7d median Δ 7d 30d median Δ 30d runs (7d/30d)
agent 21.18 21.65 $\color{green}{\blacktriangledown}$ -0.47 21.65 $\color{green}{\blacktriangledown}$ -0.47 63/160
agentless 18.15 19.20 $\color{green}{\blacktriangledown}$ -1.05 19.20 $\color{green}{\blacktriangledown}$ -1.05 33/122
agentlessCodeCoverage 18.72 19.99 $\color{green}{\blacktriangledown}$ -1.27 19.99 $\color{green}{\blacktriangledown}$ -1.27 33/122
agentlessLineCoverage 26.05 26.98 $\color{green}{\blacktriangledown}$ -0.93 26.98 $\color{green}{\blacktriangledown}$ -0.93 22/70

Baseline: median of @test.tracer_overhead on main (gitlab) over the last 7/30 days, per OSS project & scenario. Δ = this PR − baseline median; red ▲ = more overhead, green ▽ = less overhead than baseline.

Tests run to verify that CI Visibility behavior has not regressed in the current PR.

@dd-octo-sts

dd-octo-sts Bot commented Sep 30, 2026

Copy link
Copy Markdown
Contributor

🟢 Java Benchmark SLOs — All performance SLOs passed

Suite Status
Startup 🟢 pass

SLO thresholds are defined here based on automatically generated metrics. A warning is raised when results are within 5% of the threshold.

PR vs. master results
Scenario Candidate master Δ (95% CI of mean)
startup:insecure-bank:iast:Agent 13.99 s 14.02 s [-0.9%; +0.5%] (no difference)
startup:insecure-bank:tracing:Agent 12.92 s 13.02 s [-1.5%; -0.1%] (maybe better)
startup:petclinic:appsec:Agent 17.51 s 17.46 s [-0.6%; +1.2%] (no difference)
startup:petclinic:iast:Agent 17.56 s 17.67 s [-1.5%; +0.2%] (no difference)
startup:petclinic:profiling:Agent 17.45 s 17.25 s [+0.0%; +2.3%] (maybe worse)
startup:petclinic:sca:Agent 17.60 s 17.58 s [-0.7%; +1.0%] (no difference)
startup:petclinic:tracing:Agent 16.71 s 16.77 s [-1.3%; +0.6%] (no difference)

Commit: a2dbd3da · CI Pipeline · Benchmarking Platform UI


Load and DaCapo benchmarks can be triggered manually in the GitLab pipeline. Results will appear in the Benchmarking Platform UI after completion.

@cit-pr-commenter-54b7da

Copy link
Copy Markdown

CI Visibility Test Environment - pass4s

Job Status: 🟢 success

Scenario This PR (%) 7d median Δ 7d 30d median Δ 30d runs (7d/30d)
agent 9.22 9.54 $\color{green}{\blacktriangledown}$ -0.32 9.54 $\color{green}{\blacktriangledown}$ -0.32 62/137
agentless 9.14 9.54 $\color{green}{\blacktriangledown}$ -0.40 9.35 $\color{green}{\blacktriangledown}$ -0.21 34/109
agentlessCodeCoverage 17.69 17.73 $\color{green}{\blacktriangledown}$ -0.04 16.69 $\color{red}{\blacktriangle}$ +1.00 33/103

Baseline: median of @test.tracer_overhead on main (gitlab) over the last 7/30 days, per OSS project & scenario. Δ = this PR − baseline median; red ▲ = more overhead, green ▽ = less overhead than baseline.

Tests run to verify that CI Visibility behavior has not regressed in the current PR.

@cit-pr-commenter-54b7da

Copy link
Copy Markdown

CI Visibility Test Environment - sonar-kotlin

Job Status: 🟢 success

Scenario This PR (%) 7d median Δ 7d 30d median Δ 30d runs (7d/30d)
agent 15.76 13.40 $\color{red}{\blacktriangle}$ +2.36 13.13 $\color{red}{\blacktriangle}$ +2.63 62/159
agentless 14.19 12.37 $\color{red}{\blacktriangle}$ +1.82 11.88 $\color{red}{\blacktriangle}$ +2.31 34/122
agentlessCodeCoverage 17.30 15.11 $\color{red}{\blacktriangle}$ +2.19 15.11 $\color{red}{\blacktriangle}$ +2.19 34/122
agentlessLineCoverage 18.15 17.73 $\color{red}{\blacktriangle}$ +0.42 17.38 $\color{red}{\blacktriangle}$ +0.77 33/122

Baseline: median of @test.tracer_overhead on main (gitlab) over the last 7/30 days, per OSS project & scenario. Δ = this PR − baseline median; red ▲ = more overhead, green ▽ = less overhead than baseline.

Tests run to verify that CI Visibility behavior has not regressed in the current PR.

@cit-pr-commenter-54b7da

Copy link
Copy Markdown

CI Visibility Test Environment - jolokia

Job Status: 🟢 success

Scenario This PR (%) 7d median Δ 7d 30d median Δ 30d runs (7d/30d)
agent 94.04 93.23 $\color{red}{\blacktriangle}$ +0.81 93.23 $\color{red}{\blacktriangle}$ +0.81 62/158
agentless 88.18 89.58 $\color{green}{\blacktriangledown}$ -1.40 89.58 $\color{green}{\blacktriangledown}$ -1.40 34/130
agentlessCodeCoverage 98.08 99.00 $\color{green}{\blacktriangledown}$ -0.92 99.00 $\color{green}{\blacktriangledown}$ -0.92 34/130
agentlessLineCoverage 99.04 101.00 $\color{green}{\blacktriangledown}$ -1.96 101.00 $\color{green}{\blacktriangledown}$ -1.96 33/134

Baseline: median of @test.tracer_overhead on main (gitlab) over the last 7/30 days, per OSS project & scenario. Δ = this PR − baseline median; red ▲ = more overhead, green ▽ = less overhead than baseline.

Tests run to verify that CI Visibility behavior has not regressed in the current PR.

@cit-pr-commenter-54b7da

Copy link
Copy Markdown

CI Visibility Test Environment - okhttp

Job Status: 🟢 success

Scenario This PR (%) 7d median Δ 7d 30d median Δ 30d runs (7d/30d)
agent 18.79 20.39 $\color{green}{\blacktriangledown}$ -1.60 19.99 $\color{green}{\blacktriangledown}$ -1.20 62/148
agentless 18.15 19.99 $\color{green}{\blacktriangledown}$ -1.84 19.59 $\color{green}{\blacktriangledown}$ -1.44 35/118
agentlessCodeCoverage 21.90 23.45 $\color{green}{\blacktriangledown}$ -1.55 22.99 $\color{green}{\blacktriangledown}$ -1.09 35/118
agentlessLineCoverage 38.24 40.25 $\color{green}{\blacktriangledown}$ -2.01 39.45 $\color{green}{\blacktriangledown}$ -1.21 35/119

Baseline: median of @test.tracer_overhead on main (gitlab) over the last 7/30 days, per OSS project & scenario. Δ = this PR − baseline median; red ▲ = more overhead, green ▽ = less overhead than baseline.

Tests run to verify that CI Visibility behavior has not regressed in the current PR.

@cit-pr-commenter-54b7da

Copy link
Copy Markdown

CI Visibility Test Environment - spring_boot

Job Status: 🟢 success

Scenario This PR (%) 7d median Δ 7d 30d median Δ 30d runs (7d/30d)
agent 16.19 16.36 $\color{green}{\blacktriangledown}$ -0.17 16.36 $\color{green}{\blacktriangledown}$ -0.17 62/146
agentless 9.91 9.92 $\color{green}{\blacktriangledown}$ -0.01 9.92 $\color{green}{\blacktriangledown}$ -0.01 34/116
agentlessCodeCoverage 13.43 13.67 $\color{green}{\blacktriangledown}$ -0.24 13.67 $\color{green}{\blacktriangledown}$ -0.24 34/116
agentlessLineCoverage 22.98 22.54 $\color{red}{\blacktriangle}$ +0.44 22.54 $\color{red}{\blacktriangle}$ +0.44 33/117

Baseline: median of @test.tracer_overhead on main (gitlab) over the last 7/30 days, per OSS project & scenario. Δ = this PR − baseline median; red ▲ = more overhead, green ▽ = less overhead than baseline.

Tests run to verify that CI Visibility behavior has not regressed in the current PR.

@cit-pr-commenter-54b7da

Copy link
Copy Markdown

CI Visibility Test Environment - sonar-java

Job Status: 🟢 success

Scenario This PR (%) 7d median Δ 7d 30d median Δ 30d runs (7d/30d)
agent -11.64 13.40 $\color{green}{\blacktriangledown}$ -25.04 14.23 $\color{green}{\blacktriangledown}$ -25.87 62/159
agentless -17.54 11.19 $\color{green}{\blacktriangledown}$ -28.73 11.19 $\color{green}{\blacktriangledown}$ -28.73 36/131
agentlessCodeCoverage 79.03 87.80 $\color{green}{\blacktriangledown}$ -8.77 77.88 $\color{red}{\blacktriangle}$ +1.15 36/131
agentlessLineCoverage 109.00 123.36 $\color{green}{\blacktriangledown}$ -14.36 116.18 $\color{green}{\blacktriangledown}$ -7.18 35/132

Baseline: median of @test.tracer_overhead on main (gitlab) over the last 7/30 days, per OSS project & scenario. Δ = this PR − baseline median; red ▲ = more overhead, green ▽ = less overhead than baseline.

Tests run to verify that CI Visibility behavior has not regressed in the current PR.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

comp: ci visibility Continuous Integration Visibility comp: testing Testing tag: ai generated Largely based on code generated by an AI or LLM tag: experimental Experimental changes type: feature Enhancements and improvements

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant