test_service_serves_warm_config
test_R13_production_scope_race.py · failed in 50% of captured runs
The proven race is between operations in application code (/Users/umangsingh/Chronotrace/benchmark/cases/R13_production_scope_race/app_service.py), not test code. Repairing the test would hide a real product defect. Re-run with --allow-production-repair to proceed deliberately.
Where the two runs diverge
The same test, twice, on a shared time axis. The outlined pair is the ordering that differs.
Outlined: cache_get#0 starts before cache_fill#0 in the failing run, and after it when the test passes.
cache_fill never started in the failing run. The run flushed its spans normally, so that absence is evidence: the operation had not happened by the time the assertion read the state.
Evidence
Suspicion comes from comparing runs. The decision comes from forcing the ordering and seeing what happens.
| Ordering | Suspiciousness | When forced | Verdict |
|---|---|---|---|
| cache_get#0 → cache_fill#0the assertion depends on this | 1.00 | fails 100% | reproduces the failure every time it is forced |
Policy gate
Every check the proposed patch had to pass before it was allowed to run.
No patch was proposed, so there was nothing for the policy gate to review.
Verification
What was established, and at which strength. A weaker check is never presented as proof.
Verification did not run: no patch reached it.