Abstained

test_service_serves_warm_config

test_R13_production_scope_race.py · failed in 50% of captured runs

The proven race is between operations in application code (/Users/umangsingh/Chronotrace/benchmark/cases/R13_production_scope_race/app_service.py), not test code. Repairing the test would hide a real product defect. Re-run with --allow-production-repair to proceed deliberately.

Where the two runs diverge

The same test, twice, on a shared time axis. The outlined pair is the ordering that differs.

The same test twice on a shared time axis. Each operation is anchored at the moment it started. In the failing run cache_get#0 starts before cache_fill#0, and cache_fill never ran at all.
Passing run3 operations in 1.11 ms
cache_fill
cache_get
assert
Failing run2 operations in 0.26 ms
cache_get
assert
cache_fill — never ran
0 ms0.56 ms1.11 ms

Outlined: cache_get#0 starts before cache_fill#0 in the failing run, and after it when the test passes.

cache_fill never started in the failing run. The run flushed its spans normally, so that absence is evidence: the operation had not happened by the time the assertion read the state.

Evidence

Suspicion comes from comparing runs. The decision comes from forcing the ordering and seeing what happens.

OrderingSuspiciousnessWhen forcedVerdict
cache_get#0 → cache_fill#0the assertion depends on this1.00fails 100%reproduces the failure every time it is forced

Policy gate

Every check the proposed patch had to pass before it was allowed to run.

No patch was proposed, so there was nothing for the policy gate to review.

Verification

What was established, and at which strength. A weaker check is never presented as proof.

Verification did not run: no patch reached it.

← All incidents