FixedReproduced on demandcause established — the ordering is now harmless

test_request_uses_refreshed_token

test_R04_fixture_dirty_read.py · failed in 10% of captured runs

Forcing use_token to start before refresh_token reproduced the failure on every attempt, so that ordering is a sufficient condition for the failure.

The two operations in `proven_inversion` name the same `resource`, and one has `access: 'write'` while the other has `access: 'read'`. The reader observed state the writer had not yet published.

Where the two runs diverge

The same test, twice, on a shared time axis. The outlined pair is the ordering that differs.

The same test twice on a shared time axis. Each operation is anchored at the moment it started. In the failing run use_token#0 starts before refresh_token#0, and refresh_token never ran at all.
Passing run3 operations in 0.60 ms
refresh_token
use_token
assert
Failing run2 operations in 0.43 ms
use_token
assert
refresh_token — never ran
0 ms0.30 ms0.60 ms

Outlined: use_token#0 starts before refresh_token#0 in the failing run, and after it when the test passes.

refresh_token never started in the failing run. The run flushed its spans normally, so that absence is evidence: the operation had not happened by the time the assertion read the state.

Evidence

Suspicion comes from comparing runs. The decision comes from forcing the ordering and seeing what happens.

OrderingSuspiciousnessWhen forcedVerdict
use_token#0 → refresh_token#0the assertion depends on this1.00fails 100%reproduces the failure every time it is forced

One scheduling constraint is enough to reproduce this failure.

Policy gate

Every check the proposed patch had to pass before it was allowed to run.

All 15 checks passed. A patch runs only when every one of them does.

Refused shortcuts

  • no sleep calls introduced
  • no aliased sleep imports introduced
  • no timeout marker added or inflated
  • no retry or flaky decorator added
  • no retry loop wrapped around the assertion
  • no assertion removed or weakened
  • no exception handler swallowing the failure
  • no test skipped, xfailed or renamed out of collection

Required substance

  • a real synchronization primitive was added
  • the primitive is reachable from both the signal and the wait site
  • the patch is not a no-op

Structural safety

  • the proposed wait edge creates no wait-for cycle
  • no production-scope file modified without opt-in
  • no third-party or vendored file modified
  • the patched module parses

Verification

What was established, and at which strength. A weaker check is never presented as proof.

  • Reproduced the exact interleaving on demand

    before the fix it failed every time under the forced ordering; after the fix it passed every time under the identical ordering (confirmed)

  • Adversarial schedules

    not attempted for this incident, and so not claimed

  • Residual flake check

    20 of 20 ordinary runs stable, against 11 failures in the same number of runs before the fix

Measured overhead
+0.212 msno fixed delay introduced
Isolation
one process per run
Reproduction seed
random_seed_base 1729random_seed_sweep 1729..1733pythonhashseed 0python_version 3.12.13forced_order use_token#0 -> refresh_token#0
Regression guard (executed and confirmed)
benchmark/cases/R04_fixture_dirty_read/test_R04_fixture_dirty_read.py::test_chronotrace_regression_ef5d95de

Proposed change

Nothing is merged automatically. This is a diff for a human to review.

--- a/benchmark/cases/R04_fixture_dirty_read/test_R04_fixture_dirty_read.py
+++ b/benchmark/cases/R04_fixture_dirty_read/test_R04_fixture_dirty_read.py
@@ -7,6 +7,17 @@
from benchmark.support import io_latency
from chronotrace.capture.instrument import assertion, operation
+from chronotrace.schedule.harness import ScheduleHarness, force_order
+
+_chronotrace_gate_session_token = asyncio.Event()
+
+
+@pytest.fixture(autouse=True)
+def _chronotrace_reset_session_token():
+ """Provide a fresh synchronization gate for each test."""
+ global _chronotrace_gate_session_token
+ _chronotrace_gate_session_token = asyncio.Event()
+ yield
class Session:
@@ -20,11 +31,13 @@
async def refresh_token(session: Session) -> None:
"""Install a fresh token on the session."""
session.token = "tok-2"
+ _chronotrace_gate_session_token.set()
@operation("use_token", resource="session.token", access="read")
async def use_token(session: Session) -> str | None:
"""Read the token the request will be signed with."""
+ await _chronotrace_gate_session_token.wait()
return session.token
@@ -48,3 +61,19 @@
with assertion("session.token"):
assert token == "tok-2"
await task
+
+
+@pytest.mark.asyncio
+async def test_chronotrace_regression_ef5d95de(session) -> None:
+ """Reproduce the interleaving that used to fail, deterministically.
+
+ Generated by ChronoTrace for incident ef5d95de. Before the repair this
+ race appeared in roughly 10% of runs; this guard forces the
+ exact ordering that caused it, so a regression fails here on every run
+ rather than once in a while.
+ """
+ forced_order = ['use_token#0', 'refresh_token#0']
+ harness = ScheduleHarness(forced_order, timeout_s=5.0)
+ with force_order(harness):
+ await test_request_uses_refreshed_token(session)
+ assert harness.reached == forced_order
← All incidents