ENH: reproducible Monte Carlo via per-simulation-index seeding - #1054
ENH: reproducible Monte Carlo via per-simulation-index seeding#1054thc1006 wants to merge 24 commits into
Conversation
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## develop #1054 +/- ##
===========================================
+ Coverage 82.18% 83.91% +1.73%
===========================================
Files 122 128 +6
Lines 16355 16765 +410
===========================================
+ Hits 13441 14068 +627
+ Misses 2914 2697 -217 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
|
The seeding logic is unit-tested in The lines codecov still shows uncovered are all in the parallel path: |
0d37ed6 to
761c092
Compare
phmbressan
left a comment
There was a problem hiding this comment.
The implementation is very clear and throughout, nice work.
The explanation on the concepts behind per index seeding (both in the issue and PR description) were rather helpful. I agree having reproducible results was an issue with the parallel per worker seeding.
Regarding the decisions on parameter naming, I agree with most of the decisions taken here. Moreover, the rng attribute is well docstringed, so it shouldn't be a matter of confusion to the user.
@MateusStano could you give your two cents on the changes here before we proceed with a merge?
Addresses review feedback on RocketPy-Team#1054. Parallel workers claimed the next index with an unlocked keep_simulating() + increment(), so near the end of a run two workers could both pass the count < n check and then claim sim_idx == n; the per-index child_seeds lookup turned that into an IndexError (before, it only wrote one extra record). Move the claim into a _claim_next_index helper that holds the shared mutex across the check and the increment, so each index is handed out once and the counter never overshoots. A deterministic unit test (a barrier plus a widened check-to-increment window) over-claims and fails if the lock is dropped. __root_seed_sequence returned the caller's SeedSequence, and spawn() advances its child counter, so passing the same object to simulate() twice produced different children. Copy it from its full state instead, which leaves the caller untouched and keeps repeated calls reproducible. Also drop Generator/BitGenerator from the accepted types: a stateful generator is not a seed, and reducing it to its underlying SeedSequence ignores how far it has been consumed. random_seed now takes an int, a sequence of ints, or a SeedSequence. Signed-off-by: thc1006 <84045975+thc1006@users.noreply.github.com>
Addresses review feedback on RocketPy-Team#1054. Parallel workers claimed the next index with an unlocked keep_simulating() + increment(), so near the end of a run two workers could both pass the count < n check and then claim sim_idx == n; the per-index child_seeds lookup turned that into an IndexError (before, it only wrote one extra record). Move the claim into a _claim_next_index helper that holds the shared mutex across the check and the increment, so each index is handed out once and the counter never overshoots. A deterministic unit test (a barrier plus a widened check-to-increment window) over-claims and fails if the lock is dropped. __root_seed_sequence returned the caller's SeedSequence, and spawn() advances its child counter, so passing the same object to simulate() twice produced different children. Copy it from its full state instead, which leaves the caller untouched and keeps repeated calls reproducible. Also drop Generator/BitGenerator from the accepted types: a stateful generator is not a seed, and reducing it to its underlying SeedSequence ignores how far it has been consumed. random_seed now takes an int, a sequence of ints, or a SeedSequence. Signed-off-by: thc1006 <84045975+thc1006@users.noreply.github.com>
9c020b6 to
3e22729
Compare
Addresses review feedback on RocketPy-Team#1054. Parallel workers claimed the next index with an unlocked keep_simulating() + increment(), so near the end of a run two workers could both pass the count < n check and then claim sim_idx == n; the per-index child_seeds lookup turned that into an IndexError (before, it only wrote one extra record). Move the claim into a _claim_next_index helper that holds the shared mutex across the check and the increment, so each index is handed out once and the counter never overshoots. A deterministic unit test (a barrier plus a widened check-to-increment window) over-claims and fails if the lock is dropped. __root_seed_sequence returned the caller's SeedSequence, and spawn() advances its child counter, so passing the same object to simulate() twice produced different children. Copy it from its full state instead, which leaves the caller untouched and keeps repeated calls reproducible. Also drop Generator/BitGenerator from the accepted types: a stateful generator is not a seed, and reducing it to its underlying SeedSequence ignores how far it has been consumed. random_seed now takes an int, a sequence of ints, or a SeedSequence. Signed-off-by: thc1006 <84045975+thc1006@users.noreply.github.com>
3e22729 to
6bf8bb6
Compare
|
@MateusStano friendly ping when you have a moment. Both points from your last pass are addressed: the parallel index claim now holds the shared mutex across the check-and-increment (with a deterministic test that goes red if the lock is removed), and a supplied |
6bf8bb6 to
c529d0a
Compare
Addresses review feedback on RocketPy-Team#1054. Parallel workers claimed the next index with an unlocked keep_simulating() + increment(), so near the end of a run two workers could both pass the count < n check and then claim sim_idx == n; the per-index child_seeds lookup turned that into an IndexError (before, it only wrote one extra record). Move the claim into a _claim_next_index helper that holds the shared mutex across the check and the increment, so each index is handed out once and the counter never overshoots. A deterministic unit test (a barrier plus a widened check-to-increment window) over-claims and fails if the lock is dropped. __root_seed_sequence returned the caller's SeedSequence, and spawn() advances its child counter, so passing the same object to simulate() twice produced different children. Copy it from its full state instead, which leaves the caller untouched and keeps repeated calls reproducible. Also drop Generator/BitGenerator from the accepted types: a stateful generator is not a seed, and reducing it to its underlying SeedSequence ignores how far it has been consumed. random_seed now takes an int, a sequence of ints, or a SeedSequence. Signed-off-by: thc1006 <84045975+thc1006@users.noreply.github.com>
|
Heads up that this changed enough since the last look to be worth a fresh pass rather than merging on the earlier approval. @MateusStano @phmbressan when you have a moment. What is new since the review:
Both earlier concerns are still handled: the parallel claim holds the mutex across the check and the increment, and a supplied |
c529d0a to
e5ba940
Compare
|
@MateusStano @phmbressan a follow-up pass turned up a few more things worth fixing, so I have pushed them and would appreciate another look when you have time. Since your reviews:
I also marked the earlier threads resolved. The race and the SeedSequence copy are both fixed in the current code, and the dangling-files question checked out: the run writes only under A few larger items from the same review are better as their own issues, so I opened #1075 (append continuation), #1076 (a full parallel test under spawn and forkserver) and #1077 (a seed for |
5a9c119 to
da2ba5c
Compare
|
A quick note on the red CI here, so it isn't mistaken for a regression from this change: the failing jobs crash in Re-running usually clears it. Happy to help look at the flaky animation tests on their own if that would be useful. Filed #1078 to track the flaky animation tests. |
Two things create_object samples were left out of the per-simulation reseed. Air brakes were never in the reseed loop at all, so they drew from wherever the generator had been left rather than from the simulation index. Every seeding test passed because no fixture had an air brake, which is exactly how it stayed hidden. The collections are now declared in one place and walked from there, and a test scans create_object's source so a collection added later cannot quietly miss the reseed. CP and thrust eccentricity were validated once, at add time, against the generator as it stood then. Reseeding replaced the generator but not those values. Keep the specs as given and reapply them after each reseed. Signed-off-by: thc1006 <84045975+thc1006@users.noreply.github.com>
In the worker: sim_idx and inputs_json are bound before the try. A failure in the index claim used to raise UnboundLocalError inside the error handler, so nothing was written and nothing was printed and the run ended with no record of what went wrong. The handler now writes a JSON line either way, since _read_log_file parses that file with json.loads. Reporting a failure is best effort and must never replace the failure it is reporting. Setting the shared event can raise on its own once the manager has gone, and so can taking the mutex or writing the file. All of it is guarded, the mutex is released if it was taken, and the original exception is what leaves the worker. The worker re-raises so its exit code says it died. In the parent: join() returns None however a child ended, so the shared event was the only signal a run had. A worker can leave without setting it: SystemExit, os._exit, a segfault in a native extension, a target that will not unpickle under spawn, or its own error handler failing. Check the exit codes too. Workers are started inside the try, so a start() that fails part way through the fleet does not leave the running ones with nobody to reap them. After the run, check that every index this run claimed left exactly one input row and one output row. Neither file shows this on its own: the rows look well formed, and reading them back keyed by index hides a duplicate behind the row that overwrote it. A row cut off mid-write is reported as the index that went missing rather than failing to parse, which is what actually happened to it. A run stopped with Ctrl-C is exempt, since both run paths catch it, keep what they have and return, and being short is the point rather than a fault. Signed-off-by: thc1006 <84045975+thc1006@users.noreply.github.com>
The existing tests cover the seed arithmetic everywhere and the real loop under fork. Neither reaches multiprocess.Process, __sim_producer, the manager proxies or pickling the stochastic object graph anywhere but fork, and spawn is what Windows and macOS run, and forkserver is Python 3.14's POSIX default. Serial, two workers and four workers are compared per index under each available start method. Object identity is stripped before comparing: a Function's signature hash and its serialised source encode the object rather than the value drawn for it, and a child that re-imported the module cannot agree with the parent about those. Six fields differ across the boundary on a real run and all six are these. The fixtures are built so the properties can actually fail. The shared stochastic environment has zero wind at every altitude, and zero times any factor is zero, so a compounding baseline cannot show up in it; this one sets a wind that is actually blowing. A bare StochasticAirBrakes gives every parameter a standard deviation of zero, so it gets one that varies. The assertions check the eccentricities and the air brake are among the compared fields, or stripping identity could quietly empty the comparison. Also covers the parent-side checks: a run missing an input row, a run missing an output row, a row cut off mid-write, appending onto an earlier run, and a run stopped with Ctrl-C. Signed-off-by: thc1006 <84045975+thc1006@users.noreply.github.com>
Assigning to montecarlo._MonteCarlo__evaluate_flight_inputs and friends trips pylint's invalid-name, which exits 16 and fails the Linters job even though the score is 10.00. monkeypatch.setattr takes the name as a string, so the check does not fire, and it puts the original back afterwards instead of leaving the instance patched for whatever runs next. Signed-off-by: thc1006 <84045975+thc1006@users.noreply.github.com>
A review of the seeding change turned up failure paths where a run that went wrong could still be reported as a success. Six of them, all in the machinery around the simulations rather than in the seeding itself. The parent waited for every worker with an unbounded join, in the order they were started. One worker stuck in a native call held it there while another had already set the error event, so neither the error nor the cleanup after it was ever reached, and Ctrl-C hung on the same join a second time. The wait is bounded and gives up as soon as the event is set. Shutdown signals the whole fleet before waiting on any of it, then falls back to kill, so a worker that ignores the first signal does not keep the others, the manager and the open files alive behind it. The completeness check accepted a corrupt file. Rows it could not parse were skipped, rows carrying no index or an index outside the run were ignored, and JSON true or 1.0 passed for the index 1 because both compare equal to it. Every row now has to be an object with a plain non-negative int index, the two files have to agree on the exact set, and an interrupted run is allowed to be short but not to be corrupt. Both run paths cleared the current payload after the call that can be interrupted rather than before it. In the serial path Ctrl-C on the first lap reached the handler with it unbound, so the interrupt surfaced as an UnboundLocalError, and between laps it still held the row that had just been written. In the worker the same ordering meant a claim that failed on a later lap reported the simulation that had just succeeded. The normal write path also released the mutex in finally whether or not acquire had returned. The error record kept either the inputs or the traceback, never both, so every failure after sampling left no traceback in the file the run points the user at. n_workers was validated after the logs were opened "w+", so asking for a worker count the run cannot use destroyed the previous results on the way to raising. All argument checking happens before any file is touched, and number_of_simulations is checked too. Signed-off-by: thc1006 <84045975+thc1006@users.noreply.github.com>
Three problems from review, and they compound. The check judged the whole file, so a duplicate, a torn row or a pair that disagrees anywhere in it failed the run. The documented way to resume a Monte Carlo is to interrupt one and carry on with append=True, and an interrupted run is exactly what leaves that damage behind, so a file could be damaged once and never resumed again. It now judges only the indices this run claimed. Damage below _initial_sim_idx is an earlier run's and is warned about rather than raised on, while a run that wrote the whole file is still held to all of it. The parent stopped waiting the moment a worker reported an error and terminated the fleet immediately, giving a worker part way through a write no chance to finish it. Measured: 0.0 ms on that path against the 5000 ms the interrupt path already gave. So the shutdown produced the torn rows the check then reported. Both paths give the same window now. Not by reusing _bring_the_fleet_down: that sets the error event, which on a run that finished cleanly is what the crash check reads next, and wiring it in made every successful parallel run report itself as failed. A torn row also makes the two files disagree, and the cross-file check ran first, so the message named the symptom. The damage check runs first now. The real parallel path was only exercised by a test marked slow, and pull-request CI skips those, so the path this work exists to support gated nothing. There is a small version of it now that is not marked slow and uses the platform's default start method, so each CI job gates the one it actually runs: spawn on Windows and macOS, forkserver on Python 3.14's POSIX default. It caught the _bring_the_fleet_down mistake above within seconds of being written. Signed-off-by: thc1006 <84045975+thc1006@users.noreply.github.com>
The resume point came from a line count rather than from the indices on disk, and the completeness check trusted it. A blank line makes the two disagree: two rows plus one blank load as three simulations, so the next run starts at index 2, index 1 is never written, and a check scoped to the new range reports success. Measured at 3 for one blank line and 4 for two, with the file holding only 0 and 1 either way. Appending now reads both logs first and refuses unless what it finds is the run it is being asked to continue: every row readable, no index twice, the inputs and outputs holding the same set, and the indices forming exactly the range below the resume point. Nothing is opened for writing until that passes, so a checkpoint that cannot be resumed is left as it was found. A run that was not interrupted is then held to the whole range rather than to its own share of it. A file numbered from 1 is named rather than reported as an off-by-one. Serial runs used to be numbered that way, and appending onto one would rewrite its last index instead of continuing, so the answer is to re-baseline. This replaces the tolerance added in the previous commit for damage an earlier run left behind. That belonged at the wrong end: a torn row holds an index nobody can recover, so the resume point cannot be trusted either, and the preflight refuses before any simulation is spent rather than after. Two other things that could destroy a previous run: multiprocess is an optional extra, and it was imported inside the parallel path, which runs after both logs have been opened "w+" and emptied. An install without rocketpy[monte-carlo] lost its previous results on the way to the ImportError. It is imported with the other argument checks now. The rejection message for a Generator advised rng.bit_generator.seed_seq, which NumPy only grew in 1.25 while this package declared numpy>=1.13, so the advice raised AttributeError on versions it claimed to support. It now says to pass the seed the generator was built from, mentions seed_seq as a 1.25 option, and lists integer sequences among the accepted inputs. The floor moves to 1.17, which default_rng has needed all along. Also fixes the custom sampler fixture, which built a Generator in reset_seed and dropped it while sample() drew from the process-global np.random, so nothing in it answered to a seed and the 128-bit path went untested. And states in _nominal that construction-time snapshot semantics apply to every stochastic model rather than only to the environment, with a test to hold it there. Signed-off-by: thc1006 <84045975+thc1006@users.noreply.github.com>
`dict_generator` walks the whole instance, so `parachutes` and `air_brakes` were drawn from as ordinary lists. Since this branch started seeding list choices from the model's own generator, that draw moved every later one, and `StochasticRocket.dict_generator` discards it a few lines further down. Attaching a main and a drogue changed the sampled mass under a fixed seed, which is the property this branch exists to establish. One component does not show it: `integers(1)` has a single outcome and NumPy returns it without consuming any state, so the test covers 0, 1 and 2. Signed-off-by: thc1006 <84045975+thc1006@users.noreply.github.com>
A worker killed outright sets no error event. If it died holding the shared lock, its siblings never return either, so `any(is_alive())` stayed true and the unbounded wait never reached the exit-code check below it. The wait now also ends on a non-zero exit code. Only the unbounded one: the shutdown grace period is bounded already and must not be cut short. `type(value) in (int, np.integer)` is False for every NumPy integer, because `type(np.int64(3))` is `np.int64`. It was written that way to keep `True` out, which `isinstance` lets through, so both are now checked explicitly. `number_of_simulations` is the total to reach when appending, not a batch to add. Below the checkpoint it ran nothing and returned success, leaving a file with more simulations than the caller asked for. Signed-off-by: thc1006 <84045975+thc1006@users.noreply.github.com>
4334438 to
069558d
Compare
The unmarked test took the platform default and its docstring claimed that gated spawn on macOS and forkserver on 3.14. Neither is true: multiprocess hard-codes fork on every POSIX platform, macOS and 3.14 included, with a `#FIXME: spawn` still beside the darwin branch. So the shipped parallel path was gated on fork everywhere except Windows, and the failure message named the stdlib start method rather than the one that made the workers. The thorough test already covers all three through the real path and takes 26 s against 89 s for the rest of the directory, so it is unmarked now and the default-taking one is deleted rather than corrected. Its assertions were a subset. The start-method list is asked of multiprocess as well. That import is at module level behind a try/except because it happens while tests are collected, where importorskip would take the module down instead of skipping it. Signed-off-by: thc1006 <84045975+thc1006@users.noreply.github.com>
SeedSequence keeps a sequence entropy by reference, and the capture stored that
reference, so a caller who passed a list and later edited it changed the child
seeds of a run that had already read the seed. The docstring called it an
immutable snapshot, which it was only for an int.
entropy = [1, 2, 3]
mc.simulate(2, random_seed=entropy)
entropy[0] = 999999 # moved every index of that run
Deep-copied on the way in now, with spawn_key made a tuple while it is there.
Signed-off-by: thc1006 <84045975+thc1006@users.noreply.github.com>
|
Two commits on top of your rebase, @Gui-FernandesBR. Both are defects in this branch's own code that I found while going back over it, so they belong here rather than in a follow-up.
if sys.platform == 'darwin':
_default_context = DefaultContext(_concrete_contexts['fork']) # FIXME: spawn
else:
_default_context = DefaultContext(_concrete_contexts['fork'])There is no version check in that file, so 3.14 does not change it either. The shipped parallel path was therefore gated on The thorough test already covers all three through the real path, and it costs 26 s against 89 s for the rest of that directory, so it is unmarked now and the default-taking one is deleted rather than corrected. Its assertions were a strict subset.
An int seed was never exposed to this, which is why it went unnoticed. Deep-copied on the way in now. Each fix has a test that fails when the fix is removed, and the mutation only takes the new test with it. Local: ruff clean, On the rebase itself: I had rebased locally onto the same @MateusStano, your review is from 10 July and the branch has been through a rebase and a fair amount of rework since. Whenever you have time, a fresh look at the current head would be welcome. |
Each worker got the full grace to itself, and twice over, once after terminate and once after kill. Eight stubborn workers could therefore hold the parent for sixteen grace periods rather than two, which is a 5 s promise turning into 80 s. The deadline is shared now, so the wait costs the same whatever the fleet size. Every worker is still joined, so exit codes are still reaped. Signed-off-by: thc1006 <84045975+thc1006@users.noreply.github.com>
7d0c5da to
f00c205
Compare
|
Third and last of the ones I found going back over this branch,
With The test asserts on the time each worker was granted rather than on how long the call took. My first version measured wall clock with 43% headroom, which would have been a flake waiting for a loaded runner; the granted times are exact and the same on any machine. Removing the fix fails it at fleet size 6 and leaves fleet size 1 passing, which is the point. Verification on this headEverything the workflows run, plus the slow suite: The slow failure is Each of the three fixes has a mutation that fails a named test, and each mutation takes only its own test with it. |
…meet The assertion checked one fleet against the grace itself, so its slack had to cover the machine's timer granularity. On Windows that is about 15 ms against a 200 ms grace, and the job failed at 0.213s under a 0.21s bound. Comparing a fleet of six against a fleet of one carries the same granularity on both sides, so it cancels. Six against twice one leaves roughly half the bound spare on the Windows numbers, and the per-worker grace it replaced would grant six times as much. Signed-off-by: thc1006 <84045975+thc1006@users.noreply.github.com>
|
Windows 3.10 was red on a test I wrote two commits ago, and it failed for exactly the reason I claimed it could not. Fixed in I said the granted-time assertion was "exact and the same on any machine", which replaced a wall-clock one I had called a flake waiting to happen. It is not machine-independent. Checking a single fleet against the grace means the slack has to absorb the machine's timer granularity, and on Windows that is roughly 15 ms against a 200 ms grace: Three milliseconds. The second worker was granted a sliver instead of zero, because The property I actually want is that the wait does not grow with the fleet, so it now measures a fleet of one and a fleet of six and compares them. Both carry the same granularity, so it cancels: On the Windows numbers from that failed job the crowd would come in near 0.43s against the same 0.8s bound. Removing the fix grants six times as much and fails it. Local on this head: ruff clean, Worth saying plainly: this is the second timing assertion I have written for this one behaviour, and the first was wrong in the way I had just finished warning about. The 3 ms is a fair thing to have missed on Linux, the confident sentence in my last comment was not. |
|
A note on the one line this touches in The change here is What I had not checked is whether 1.17 is enough for the Python this package declares. It is not, quite. So a resolver on 3.10 that lands between 1.17 and 1.21.3 builds from source rather than failing outright, which is slow where it works and unpleasant where it does not. I have not raised it here. Moving a floor past what this branch needs is a packaging decision rather than part of a seeding change, and there is a larger one next to it that has nothing to do with this PR: Filed both together as #1107, since they are one edit to one file. Happy to fold the NumPy half into this PR instead if you would rather it travelled with the line that is already changing. |
The refusal message suggested re-running "or renumber the file down by one", while the comment two lines above it said the fix is to re-baseline rather than retry. The comment was right. Renumbering lines the indices up and leaves the seeds behind. Those rows came from the old sequential scheme, so a renumbered file would carry rows 0..n-1 that this release's per-index derivation would never have produced for those indices, and appending onto it would join two different seedings without saying so. Signed-off-by: thc1006 <84045975+thc1006@users.noreply.github.com>
|
Three things from this round, one of which was mine and wrong. The renumber advice, Renumbering lines the indices up and leaves the seeds behind. Those rows came from the old sequential scheme, so a renumbered file carries rows 0..n-1 that this release's per-index derivation would never have produced for those indices, and appending onto it joins two different seedings without saying so. The message now says that, and a test fails if the old advice comes back. Hole recovery. I went looking for the same kind of over-promise there and did not find one. The docstring already says what it does:
Leaving that as it is. The #1102 interaction. I said earlier that the two conflict in one line of imports. That was true when #1102 was The current heads happen to merge with no textual conflict at all, which is the part worth not trusting, so I built the merge and ran it rather than reading it: Both sides survive and the call sites line up. Whichever lands second should still be rebased and rerun rather than merged on that evidence, but there is nothing to resolve by hand today. Local on this head: ruff clean, |
|
Moved the scope statement to the top of the description. It was already there, under "What this PR does and does not promise", but at line 48 of 218, and three rounds of review have now asked me not to claim things it explicitly declines to claim. Written but not read is the same as not written, and that is on me rather than on anyone reading it. It now opens with what is and is not guaranteed, and names #1090 and #1091 with the measurement rather than leaving them to be found further down. While tidying that I noticed something worse in the same area. I have been pointing at the Saying "that one is pre-existing" without leaving anything behind for someone to pick up is not much better than not saying it. On why this branch narrows that rather than closing it, since the two look alike and the difference decides the scope.
Nothing samplable is left uncovered, so the regression this branch introduced is closed. The wider rewrite, generating only from the declared names, is the right fix for #1109 and I would rather send it separately than add it to a change this size. |
A failure after sampling wrote the inputs to .errors.txt with no traceback,
while a worker writes {index, ...inputs, error: traceback}. The file the run
tells the user to read named which inputs failed and not why.
Both paths now build the row through one helper. Three further points on that
path:
- inputs_json is cleared once the pair is on disk, so a failure inside
print_update_status() reports itself rather than reporting an already
committed row as one that never finished.
- sim_idx is bound before the loop, so a failure on the first iteration has an
index to record.
- the re-raise is bare, so the handler's own line does not join the traceback.
KeyboardInterrupt keeps its own handler: an interrupt is not a failure with a
traceback worth recording, and it still logs the inputs that did not finish.
The append docstring said only that results are appended. It now says
number_of_simulations is the target total rather than a number to add, that a
lower value is refused, and that the root seed is not stored in the files.
Signed-off-by: thc1006 <84045975+thc1006@users.noreply.github.com>
|
Pushed A failure after sampling wrote the inputs to Three smaller things on the same path, since they only show up once the record is worth reading:
On the frame point, my first test for it asserted the failing frame survives, which is true of both spellings, so it passed with the fix removed. The assertion is now that the handler frame appears once: Three tests, each pinned by reverting the line it covers and watching it fail. Also rewrote the One claim from the review I could not stand behind: that Local: ruff clean, pylint 10.00/10 exit 0, 2159 passed and 45 skipped across unit and integration. |
Scope, up front, because three reviews have now asked me not to claim things this does not claim.
This makes the sampled inputs reproducible: for a given root seed, simulation index
idraws the same stochastic parameters whether the run is serial or parallel and however many workers it uses. That is what.inputs.txtrecords and what the tests compare.It does not make the flown inputs or the trajectory reproducible, and it does not close #1053. Two built-in random sources sit outside the seed tree this builds, and both are filed rather than folded in:
MonteCarlodraws the flight dictionary three times, so the row written to.inputs.txtis not the row that was flown. Measured on the shared fixture: logged inclination 85.60, flown 84.46.Parachutetakes its pressure noise from the process-globalnp.random, which can move the deployment time and the descent with it.A third, #1109, is in the code this touches but predates it:
dict_generatorwalks the whole instance, so a valid tupleinitial_solutionis read as a distribution and its last element called. This PR narrows the blast radius of that rather than closing it, and the reasoning is under "New behavior" below.Pull request type
Current behavior
MonteCarlo.simulate()seeds the stochastic models per worker in parallel mode (from a fresh, unseedednp.random.SeedSequence().spawn(n_workers)) and once at construction in serial mode. So the sampled inputs depend on the execution mode and the number of workers, and parallel runs are not reproducible run to run. This is #1053.New behavior
Adds a keyword-only
random_seedtosimulate(). From that root, simulation indexiis seeded from its own child of the root seed, derived before simulationiruns, so indeximaps to the same seed no matter which worker runs it. The sampled inputs come out identical across serial, parallel(2) and parallel(N), and reproducible from the seed.A few specifics:
iis built by extending the root'sspawn_key, which is exactly howSeedSequence.spawnderives it, sochild(i)is bit-identical toroot.spawn(number_of_simulations)[i]. Nothing pre-spawns a full list: a worker reconstructs any index from a small root state (entropy, spawn_key, pool_size, counter) that travels with the pickled instance, so nothing O(N) is sent to each process.int, not aSeedSequence. An int is the seed typenumpy.random.default_rngand the stdlibrandom.Randomboth accept (aSeedSequenceraisesTypeErrorinrandom.Randomsince Python 3.11), so a custom sampler whosereset_seeddocuments an int keeps working. All fouruint32words are combined by value, so the seed is byte-order independent and keeps the full 128-bit pool rather than collapsing to 32 bits.StochasticModel.dict_generatordrew list attributes with the stdlibrandom.choice(an unseeded global instance), sorandom_seeddid not govern them. It now draws the index from the model's own seeded generator, which also avoidsnumpy.random.choicecoercing a heterogeneous list (Function, paths, arrays) to a single dtype.random_seedis a seed, not a live RNG: it takes an int, a numpy integer, a sequence of ints, or aSeedSequence, withNone= fresh entropy so existing behavior is unchanged unless you pass a seed. A suppliedSeedSequenceis copied from its full state before use, so it is never mutated and repeated calls with the same object reproduce the same run. This is informed by SPEC 7 and NumPy's parallel idiom, but keeps immutable seed-snapshot semantics rather than SPEC 7's statefulrng: aGenerator/BitGeneratoris not accepted, because reducing it to its underlyingSeedSequencewould ignore how far it has been consumed. Passrng.bit_generator.seed_seqto seed from an existing generator.Relation to #1071
#1071 targets the same issue. This PR takes the two ideas it got right, deriving each index's seed on demand instead of pre-spawning a list, and handing the samplers a plain int, and combines them with the parallel-claim lock below, the full 128-bit width (a single 32-bit word collides near 2**16 streams), the list-sampling fix, and cross-platform tests. Happy to reconcile the two however the maintainers prefer.
Notes from review
keep_simulating()+increment(). Near the end of a run two workers could both passcount < nand then both claim an index, running past the requested count. The claim now holds the shared mutex across the check and the increment, so each index is handed out once.SeedSequencewas returned as-is, andspawn()advances its child counter, so passing the same object twice was not reproducible. It is now copied from its full state, andGenerator/BitGeneratorare no longer accepted (see above).Follow-up review round
A closer pass after the first reviews turned up four more fixes, all pushed here:
StochasticRocket._set_stochasticgave the same seed to the rocket body and to every surface, motor, rail button and parachute, so components that sample the same distribution (a main and a drogue parachute, for instance) drew identicalcd_sandlagquantiles. Each component now gets its own child of the run's seed, in a fixed order, so they stay independent and reproducible.dict_generatorandStochasticRocket._randomize_positionsampled list-valued attributes (component positions included) with the stdlibrandom.choice, whichrandom_seeddid not govern. Both now draw the index through the model's seeded generator, via a shared_random_choicehelper.simulate()set up (and, forappend=False, truncated) the output files before the seed was validated, so passing a rejected seed destroyed a previous run's results. The seed is validated first now.RandomState, and moved it torocketpy.toolsso the stochastic models can share it.Known limitations and follow-ups
The 128-bit int does not fit the legacy
numpy.random.RandomState, which caps seeds at 2**32 - 1. A custom sampler built on the moderndefault_rng(or the stdlibrandom.Random) takes it fine; one built onRandomStatewould need to reduce it. SinceRandomStateis the discouraged legacy path this felt like the right trade for keeping the full 128-bit decorrelation, but I am happy to revisit if you would rather cap the width.Larger items from the review are better handled on their own, so they are filed separately rather than growing this PR:
append=Truedoes not continue the same seeded run #1075:append=Truedoes not yet continue the same seeded run (the root is not persisted, and the next index comes from a row count).simulate_convergence()cannot reproduce a study from a seed #1077:simulate_convergence()has no seed, so a convergence study is not reproducible.Failure safety and log integrity
A later review round found several paths where a run that went wrong could still
be reported as a success. Those are fixed here too, with tests:
join(), in start order.One worker stuck in a native call held it there while another had already set
the error event, so neither the error nor the cleanup after it was reached,
and Ctrl-C hung on the same join a second time. The wait is bounded now and
returns as soon as the event is set. Shutdown signals the whole fleet before
waiting on any of it, with a kill fallback.
rows with no index or an index outside the run were ignored, and JSON
trueor
1.0passed for index1because both compare equal to it. Every row nowhas to be an object with a plain non-negative
intindex, the two files haveto agree on the exact set, and an interrupted run may be short but not corrupt.
than before it. Serial Ctrl-C on the first lap surfaced as an
UnboundLocalErrorover the interrupt, and between laps the handler stillheld the row that had just been written. In the worker, a claim that failed on
a later lap reported the simulation that had just succeeded.
finallywhether or notacquire()had returned, so a manager that died during acquire raised asecond error over the first.
n_workerswas validated after the logs were opened"w+", so asking for aworker count the run cannot use destroyed the previous results on the way to
raising.
Tests
Seed handling is unit tested in
tests/unit/simulation/test_monte_carlo_determinism.py:accepted seed types, the
SeedSequencecopy preserving the full.state, theO(1) child equal to
spawnbit-for-bit including a root whose counter hasadvanced and indices past 2**32, the 128-bit width, and the parallel index claim.
tests/integration/simulation/test_monte_carlo_determinism.pyruns the realparallel path, no stub, under fork, spawn and forkserver, comparing serial
against parallel(2) and parallel(4) per index. The fixtures are built so the
properties can fail: the shared stochastic environment has zero wind at every
altitude and zero times any factor is zero, so a compounding baseline cannot
show up in it, and a bare
StochasticAirBrakesgives every parameter a standarddeviation of zero. The assertions check the eccentricities and the air brake are
among the compared fields, or stripping object identity could quietly empty the
comparison.
tests/unit/simulation/test_monte_carlo_log_integrity.pycovers what the run isallowed to call a success and how the fleet comes down when it is not.
The thorough version of that test is marked
slowand pull-request CI skipsthose, so it gated nothing. I said earlier that the weekly run would still
cover it; that was wrong.
Scheduled Teststriggers onscheduleand onpushtomaster, with nopull_request, so it only ever runs the defaultbranch and never sees a test that is still on a PR. There is a small version of it now that is not
marked, and because it takes the platform default each job ends up gating the
start method it actually runs: spawn on Windows and macOS, forkserver on Python
3.14's POSIX default, fork below that. It costs about six seconds.
That turned out to be worth more than the argument for it. While fixing the
shutdown timing below I reused a helper that sets the error event, and every
successful parallel run started reporting itself as failed. The new gate caught
it within seconds of being written, on a path the slow-marked version would not
have run in CI at all.
Later review round
Three more from a closer look at the head, all fixed here:
pair that disagrees anywhere in it failed the run. The documented way to
resume is to interrupt a run and carry on with
append=True, and aninterrupted run is exactly what leaves that damage, so a file could be damaged
once and never resumed. It judges only the indices this run claimed now, and
reports rather than raises on what an earlier run left.
worker part way through a write was cut off. Measured at 0.0 ms against the
5000 ms the interrupt path already gave. The shutdown was producing the torn
rows the check then reported. Both paths give the same window now.
first, so the error named the symptom rather than the cause. Reordered.
Checkpoint validation
The resume point came from a line count rather than from the indices on disk,
and the completeness check trusted it. A blank line makes the two disagree:
So the next run starts at 2, index 1 is never written, and a check scoped to
the new range reports success. Appending now reads both logs first and refuses
unless what it finds is the run it is being asked to continue: every row
readable, no index twice, both files holding the same set, and the indices
forming exactly the range below the resume point. Nothing is opened for writing
until that passes. A run that was not interrupted is then held to the whole
range rather than to its own share of it.
A file numbered from 1 is named rather than reported as an off-by-one, since
serial runs used to be numbered that way and the answer is to re-baseline.
A file with a hole in it is refused rather than repaired. Filling holes needs
workers to claim from a plan instead of counting on from the end, which is
#1075. Until then, refusing loudly beats resuming in the wrong place quietly.
The test that used to empty both logs and then assert a four-simulation result
holding only indices 2 and 3 was a success is gone. That was the shape of the
bug rather than a guard against it.
Two other ways a previous run could be lost:
multiprocessis an optional extra and was imported inside the parallelpath, which runs after both logs have been emptied. An install without
rocketpy[monte-carlo]lost its results on the way to theImportError.Generatorrejection advisedrng.bit_generator.seed_seq, which NumPygrew in 1.25 while this package declared
numpy>=1.13, so the advice raisedAttributeErroron versions it claimed to support. The floor moves to 1.17,which
default_rnghas needed all along.Also: the custom sampler fixture built a
Generatorinreset_seedanddropped it while
sample()drew from the process-globalnp.random, sonothing in it answered to a seed and the 128-bit path went untested.
Breaking change
The exact numbers a run produces change (per-index seeding, the env/rocket/flight
split and the per-component split within a rocket, the 128-bit int seeds, the
serial index now counting from 0 to match parallel, and list-valued attributes
and positions now sampled through the seeded generator), so external code that
pinned exact Monte Carlo samples would need to re-baseline. The in-repo Monte
Carlo tests do not pin exact values (
test_monte_carlo_simulatechecks apogeeand impact velocity within a tolerance and still passes), and
random_seedisopt-in.
Partially addresses #1053. The per-index seeding for serial and parallel runs is
here;
append=Truecontinuing the same seeded stream is #1075, and the runtimerandom sources are #1090 and #1091. I would rather leave #1053 open until those
land than close it on a guarantee that only covers the sampled inputs.
What this PR does and does not promise
The guarantee here is over the sampled inputs: for a given root seed,
simulation index
idraws the same stochastic parameters whether the run isserial or parallel and however many workers it uses. That is what
.inputs.txtrecords and what the tests compare.
It is deliberately not a guarantee about the whole trajectory yet, because two
built-in random sources sit outside the seed tree this PR builds. I found both
while going back over this change and filed them rather than growing the PR
further:
Parachutetakes its pressure noise from the process-globalnp.random.Flightadds that noise to the pressure it hands the trigger, soit can move the deployment time and the descent with it. The shared fixtures
already use non-zero noise, so this is on the ordinary path, not an opt-in.
MonteCarlocalls_randomize_rail_length,_randomize_inclinationand
_randomize_heading, and each one draws the whole flight dictionaryagain. The
Flightgets rail length from the first draw, inclination from thesecond and heading from the third, while the row written to
.inputs.txtholds the third. Measured on the shared fixture, the logged inclination is
85.60 while the flight used 84.46.
Until those land, two runs agreeing on
.inputs.txtdoes not prove they flewthe same thing. Once they do, the promise can be restated in terms of results.