test_sample_is_above_threshold
test_N01_random_seed_flake.py · failed in 20% of captured runs
Passing and failing runs executed operations in the same order, and this test draws on unseeded random number generation. The nondeterminism is in the data, not the schedule.
Where the two runs diverge
The same test, twice, on a shared time axis. The outlined pair is the ordering that differs.
Evidence
Suspicion comes from comparing runs. The decision comes from forcing the ordering and seeing what happens.
No ordering differed between the passing and failing runs.
Policy gate
Every check the proposed patch had to pass before it was allowed to run.
No patch was proposed, so there was nothing for the policy gate to review.
Verification
What was established, and at which strength. A weaker check is never presented as proof.
Verification did not run: no patch reached it.