From tracer positions to a final void catalog, then everything you can do with
the catalog without running the finder again. Every block below is
copy-pasteable. IDENTIFY_VOIDS.md documents every argument
of identify_voids.
Contents: install · a first run · loading your own data · the pyVIDE run · reading your catalog back · exactly reproducing a VIDE run · members and your own columns · post-processing · merging without re-running · weights without re-running · a density ceiling · the output files · other features in one line each · before you plot anything · where to find more details
git clone https://github.com/nicosmo/pyVIDE.git
cd pyVIDE
pip install -e .This pulls in numpy and scipy, which is the whole required dependency list. No
compiler needed. pyVIDE is pure python, so pip is optional: you can also
download the repository and add its folder to your path
(sys.path.insert(0, "/path/to/pyVIDE")), or start python from the repository
root, and import pyvide works.
python3 example_run.pyThis is the quickest way to see whether pyVIDE works on your system, and it walks through the whole pipeline: finding, loading, filtering, memberships, re-running with weights, rebuilding at another merging threshold, and field mode. It looks for data in three places, in this order: the block at the top of the file, where you can put your own catalog; the public reference catalog from the pyVIDE data record, if you have downloaded it into the folder named in that block; and, failing both, a synthetic box of tracers it generates itself, which then takes a few minutes to run through the pipeline. It prints which of the three it used.
For developers: pip install pytest and pytest tests runs the whole suite,
including the frozen end-to-end regression fixture that re-runs the pipeline on
a stored 100 000-tracer box and compares every catalog column against the
stored reference. pytest tests -m "not slow" skips those.
pyVIDE takes numpy arrays, so any reader that gets your positions into an
(N, 3) array will do. Three common cases:
import numpy as np
import pyvide
# a text catalog: one row per halo, columns mass, x, y, z, vx, vy, vz
table = np.loadtxt("halos.txt")
masses = table[:, 0] # (N,)
positions = table[:, 1:4] # (N, 3), Mpc/h, required
velocities = table[:, 4:7] # (N, 3), km/s, only needed for doRSD
# an HDF5 snapshot (Gadget-style layout shown; adapt the group names)
import h5py
with h5py.File("snapshot.hdf5") as f:
positions = f["PartType1/Coordinates"][:]
velocities = f["PartType1/Velocities"][:]
# a numpy file: pass the path and pyVIDE records where the data came from
# positions = "halos.npz:pos"pyVIDE has no snapshot readers of its own: the two or three lines above are all
a reader would do. The run checks the positions you pass and reports three
things: how many lie outside the box, how many lie so far outside that a single
wrap cannot bring them back into the box, and whether the coordinates span less
than a tenth of the box on some axis, which usually means wrong units, or a
sub-volume passed with the full boxLen. It reports and carries on, leaving
your array untouched.
What then happens to a coordinate outside the box depends on the mode.
'pyvide' wraps the periodic axes with a modulo, so any excess comes back, and
stops with an error if a non-periodic axis carries anything outside,
because such an axis is never wrapped. The two VIDE modes wrap the periodic
axes once by a box length, exactly as VIDE does, and drop whatever a single
wrap leaves outside; on a walled axis nothing is wrapped, and a tracer outside
the box is dropped. Both modes treat the walls themselves the same way from
there on.
catalog = pyvide.identify_voids(
saveDir="voids", # path to output folder, created if missing
saveName="run1", # prefix of every output file
boxLen=640.0, # scalar for a cubic box, (Lx, Ly, Lz) otherwise
positions=positions, # input tracer positions
velocities=velocities, # optional input velocities; stored, and used only if doRSD is not False
otherProperties={ # optional extra per-tracer columns, see section 7
"mass": masses,
},
VIDE_mode="pyvide", # required: 'pyvide' for a new catalog, see the table below
numDivisions=2, # sub-boxes per axis
numThreads=4, # sub-boxes tessellated in parallel; bit-identical to 1
saveIntermediate=True, # cache the tessellation, so weights and the density ceiling can change later (sections 10 and 11)
)
print(catalog)The run prints its progress and ends with a summary block: total void count,
box dimensions, periodic axes, applied merging threshold, hierarchy depth,
range of void radii, mode employed, and the time spent per stage.
print(catalog.summary()) prints it again for an old run, without the timings,
which live in the _info.txt file.
The one decision you cannot skip is VIDE_mode, because it changes the
arithmetic:
VIDE_mode |
what it does | use it when |
|---|---|---|
'VIDE_halos' |
VIDE's --halos chain: positions pass through a %e text file, then float32 |
reproducing or comparing with a VIDE run on a halo or galaxy catalog |
'VIDE_direct' |
the float32 chain without the text step | reproducing a VIDE run on matter particles |
'pyvide' |
float64 end to end, no tracer dropped at the box edge, buffer sized from the catalog, periodic images with a single sub-box |
the default choice, unless you are reproducing a VIDE run |
Which mode you want follows from what the catalog is for. If it has to match
VIDE's, whether you are comparing against an existing VIDE run or producing the
catalog that run would have produced, use the mode it used. If it does not, use
'pyvide': the same finder in double precision, with the buffer sized from
your catalog; IDENTIFY_VOIDS.md lists everything the mode
changes. In the VIDE modes buffer stays at VIDE's 0.1 whatever the catalog,
which can be too thin for a sparse sample: when tracers end up as neighbours of
the guard shell rather than of each other, the run says so and tells you to
raise it. PERFORMANCE.md explains how to size it. To
reproduce an existing VIDE run setting by setting,
FOR_VIDE_USERS.md has the template.
pyvide.explain('VIDE_mode') prints the long version, and explain() does the
same for any other argument.
None of the arguments below appear in the call above. Each has a default that suits most runs, but each is worth knowing before you hit the case it covers.
duplicateMode decides what happens when two tracers land on the same
coordinate. The default 'error' stops the run and names them. 'merge'
instead collapses each coincident group into one tracer carrying the group's
summed weight, so the surviving position keeps the density the whole group had,
and the tracer file still records every merged-away tracer and the survivor it
belongs to (the run then uses the weighted path even if you passed no weights).
A VIDE mode's float32 cast can create coincidences your float64 input did not
have. WEIGHTS_CEILING_AND_DUPLICATES.md has the details.
mergingThreshold (default outputs
(default 'full') decides how much is written: 'voids' keeps the catalog
alone, and that costs you sections 7, 9 and 10, since member lookups,
rebuilding at another threshold and re-running with weights all read the files
it skips. The rebuilds of sections 9 and 10 take outputs too, and there it is
often the right choice. FILES.md says how large each file gets. A
second run under the same saveName refuses to overwrite the first unless you
pass overwrite=True, and even then pyVIDE deletes nothing: a run that would
leave files of the old run behind refuses to overwrite them and instead names
them in order to avoid mixing files from different runs.
cat = pyvide.load_catalog("voids", "run1") # saveDir, saveName; reads a finished run from disk
cat.numVoids # rows in this catalog, one per void
cat.radius, cat.macrocenter, cat.numPart # columns are numpy arrays
cat.columns() # every column name
cat.parameters # how the run was made (a dict)
cat.statistics # what it found (a dict)
radius = cat.radius # a column as an attribute ...
radius = cat.column("radius") # ... or by name, same array
coreDens = cat.coreDens # core densities, in units of the mean density
big = cat.radius > 20.0 # plain numpy masks
print(cat.macrocenter[big])The columns are the catalog's own arrays; add .copy() if you intend to modify
one. Column meanings and units are described in
CATALOG_COLUMNS.md. One rule to learn now: every column
is indexed by row, while voidID, parentID, members(), children() and
descendants() speak in void IDs. The two agree only in a fully periodic run
at the default settings, because a wall or a cut removes rows without
renumbering ids, and cat.rows_of(ids) converts between them.
To reproduce an existing VIDE run rather than make a new catalog, the call of section 4 changes in one place, the mode, and takes the settings VIDE used:
vide_run = pyvide.identify_voids(
saveDir="voids",
saveName="vide_run",
boxLen=640.0,
positions=positions,
velocities=velocities,
otherProperties={"mass": masses},
VIDE_mode="VIDE_halos", # 'VIDE_halos' for a run prepared with --halos, 'VIDE_direct' without
mergingThreshold=1e-9, # the VIDE run's value
numDivisions=2, # VIDE's numZobovDivisions
buffer=0.1, # VIDE's default buffer value
numThreads=4,
)That is the whole recipe for the catalog.
FOR_VIDE_USERS.md has the full template, including the
three settings that matter only when a VIDE run used them (minRadius,
maxCentralDen, doRSD='VIDE') and the volumeMethod='exact' switch for
bit-comparing per-tracer volumes. VIDE also applied its rMin cut at the mean
tracer separation. Pass minRadius=-1 to include it, as the template in
FOR_VIDE_USERS.md does.
pyVIDE keeps every density basin as a catalog row. VIDE's text catalogs leave out one kind of basin, a zone holding only a single tracer, so a VIDE catalog of the same tracers is potentially a few rows shorter. The pyVIDE run says so in its summary, and the number VIDE would print is stored:
cat.statistics["numVideVoidDescRows"] # the length VIDE's catalog hasTo get exactly VIDE's rows, in VIDE's order, do:
vide_like = pyvide.filter_catalog(cat, "untrimmed_all")filter_catalog computes nothing. It only drops the single-tracer zones, sorts
by density contrast the way jozov2 does, and applies the cuts of the variant
you name. 'untrimmed_all' adds no cuts of its own, so it is the like-for-like
list. VIDE's default catalog, the centers_all_* file with no filename prefix,
is the "trimmed" variant instead: top-level voids only, with the
central-density cut applied at VIDE's merging threshold. VIDE fed one number
into both roles. pyVIDE keeps them apart, so you name the cut yourself:
vide_default = pyvide.filter_catalog(cat, "trimmed_all", maxCentralDen=1e-9)At "untrimmed_dencut_all" gives the
same rows there, but they differ once merging is on.
FOR_VIDE_USERS.md has the full table of VIDE file names to
pyVIDE calls.
The catalog is one of several files a run writes. The tracer file holds one row per kept tracer, and the membership file connects voids to their tracers:
cat = pyvide.load_catalog(
"voids", "run1",
loadTracers=True, # positions, Voronoi volumes, zone ids, weights, the columns you added
loadMembership=True, # the void -> zone -> tracer lists (optional: members() reads them when first asked)
)
pos = cat.tracers["positions"] # (N, 3), the kept tracers, in input order
vol = cat.tracers["volume"] # each tracer's Voronoi cell volume
zone = cat.tracers["zoneID"] # which zone/void each tracer belongs to
mass = cat.tracers["mass"] # the column you passed as otherProperties
voidID = int(cat.voidID[np.argmax(cat.radius)]) # ID of the void with the largest radius
inVoid = pyvide.members(cat, voidID) # rows of that void's tracers in cat.tracers (positions, volume, ...)
print(inVoid.size, "member tracers")
print(pos[inVoid], vol[inVoid], mass[inVoid]) # their positions, volumes, massesAbout otherProperties: it is a dict of name -> array with any number of
entries, each (N,) or (N, k), so for example a magnetic-field vector would
be one (N, 3) column. The finder never reads them and they are stored only
for the tracers the run kept, row-aligned with positions, volume and
zoneID, so mass[members(cat, id)] gives the masses of one void's tracers.
Names must be valid python identifiers, must not collide with the existing
columns pyVIDE writes itself, and must not start with merged_, which the
coincident-tracer bookkeeping uses. Each entry may also be a path,
"halos.npz:temperature".
members takes a void ID, not a row number, and the indices it returns are a
third numbering: rows of the tracer file, not of your input array. The box cut
drops tracers, so indexing your own array with them silently gives the wrong
objects. The kept tracers are those left inside the box after the wrap
described in section 3. cat.tracers["inputIndex"] maps every kept row back to
your input array, and cat.statistics["numDroppedTracers"] counts the rest.
None of this re-runs the finder: these are selections over rows that already
exist. The VIDE variants exist to reproduce VIDE's catalogs. For a new
analysis, start from the full catalog instead and make your own cuts. The
catalog pyVIDE writes keeps every void, including the single-tracer zones VIDE
never wrote at all; 'untrimmed_all' is as close as VIDE got. Your own boolean
mask through select says what you mean more directly than a variant name
does.
# VIDE's eight catalog variants, as filters (two shown; see the paragraph
# below, and section 6 for the two common ones)
pyvide.filter_catalog(cat, "trimmed_nodencut_all", minRadius=5.0)
pyvide.filter_catalog(cat, "trimmed_central", minRadius=5.0, maxCentralDen=1e-9)
# your own cuts: a boolean mask over the rows gives a smaller catalog object
usable = cat.select(cat.shapeReliable) # voids with a meaningful shape tensor
large = cat.select(cat.radius > 20.0) # a radius cut, in your units
inner = cat.select(~cat.boundaryFlag) # ~ negates the mask: voids NOT bordering a wall
# the hierarchy (only interesting once merging is on, see section 9)
cat.children(voidID) # direct sub-voids, as IDs
cat.descendants(voidID) # every void inside this one, as IDs
rows = cat.rows_of(cat.descendants(voidID), missing="mask")
# rows of those IDs; -1 where a filtered catalog no longer has the void
# (missing="drop" would leave them out; the default raises an error)The variant names are untrimmed, untrimmed_dencut, trimmed_nodencut and
trimmed, each with _all or _central. Every one of them first drops the
"wrong" voids (densCon above radius < minRadius, which does nothing at the default 0.0) and, on a
box with walls, the ones bordering a wall. On top of that, dencut removes
voids whose centralDen exceeds maxCentralDen, and trimmed keeps only
top-level voids. all versus central is VIDE's survey-region split, which
makes no difference for a simulation box.
The default merging threshold of
merged = pyvide.rebuild_catalog("voids", "run1", mergingThreshold=0.2)This re-runs the watershed step over the stored graph, recomputes every
property and the hierarchy, and writes a full file set named run1_mt0.2:
without a newSaveName, rebuild_catalog adds the merging threshold it
applied to the run's name, so that a rebuild never clashes with the run it
read. Pass newSaveName to choose a name yourself. Its output is column for
column identical to a from-scratch run at that threshold (checked by
tests/test_rebuild.py), and those from-scratch runs are the ones verified
against VIDE.
merged.statistics["maxTreeLevel"] # how deeply nested the catalog is
merged.isTopLevel.sum() # voids contained in no other void
merged.numZones.max() # the most zones any one void merged
top = int(merged.voidID[np.argmax(merged.radius)]) # the largest void, as an ID
merged.children(top) # its direct sub-voids (IDs)
merged.descendants(top) # everything inside it, at any depth
merged.treeLevel # per void: how many voids contain it (0 = top level)
merged.numDescendants # per void: size of the subtree below itparentID gives the ID of the smallest void containing a given one, or -1 if
nothing contains it. treeLevel counts how many voids contain it, so a
top-level void is 0, and large catalogs can reach ten or more. Which set you
use depends on the question: isTopLevel alone covers space without overlaps,
and the full set of rows gives you voids on every scale.
Two things about the value 0.0. It means "merge everything the topology
allows", so the top level collapses to a single void spanning the whole box,
and the useful output is the hierarchy below it: the complete merger tree, with
no threshold chosen by hand. That is what makes it worth having, since any
other threshold imposes a scale whose effect depends on the tracer population
and its bias. It is also the slowest to rebuild, because every tracer then
belongs to every void above it.
A sweep is one rebuild per threshold, and a rebuild takes a fraction of the
time a typical run costs. Each one writes its own run1_mt<threshold> files,
so the thresholds never collide with each other and the original run1 is
never touched. overwrite=True is only needed when you run the same sweep
again, over file sets a previous pass already wrote:
for threshold in (0.1, 0.2, 0.5, 1.0):
pyvide.rebuild_catalog("voids", "run1", mergingThreshold=threshold,
overwrite=True, verbose=False)On a sweep it is usually worth passing outputs='voids'. The tracer file that
a threshold rebuild writes is a byte for byte copy of the one it originally
read, so a ten-threshold sweep otherwise leaves ten copies of the same file on
disk. The catalogs themselves are unchanged by this setting:
for threshold in (0.1, 0.2, 0.5, 1.0):
pyvide.rebuild_catalog("voids", "run1", mergingThreshold=threshold,
newSaveName="run1_lean_mt{0:g}".format(threshold),
outputs="voids", verbose=False)A lean rebuild cannot be rebuilt from this output in turn, since it did not
write the needed zone graph. Every threshold in the loop reads run1 anyway,
so nothing is lost. One thing to know before switching a sweep to 'voids': if
a previous pass wrote full file sets under the exact same names, the lean pass
refuses, because replacing only the catalog would leave the older membership
and tracer files in the same folder. Delete those first, or use other names.
Two conditions for this. The original void identification must have been run
with outputs='full' (the default), because a rebuild reads its tracer,
membership and merge-event files. A rebuild can only change the threshold: the
stored zone graph came out of the original run's densities, so changing the
weights needs rerun_watershed, which is the next section.
A weight scales how much density a tracer contributes: give one tracer twice
the weight of another and it counts for twice as much density. Only the density
side of the run changes. Each Voronoi cell is rescaled to V * mean(w) / w and
the watershed runs on that, so the zones, merging and the densities that follow
from them (coreDens, densCon, leakDens, prob) are weighted, while
everything geometric stays geometric: voidVol, radius, macrocenter,
centralDen, maxRadius and the shape tensor.
Both sides are kept, so nothing is lost: the weighted volumes are reported
alongside as voidVolWeighted, zoneVolWeighted and radiusWeighted, and the
unweighted density at each core survives as coreDensGeometric. densCon,
leakDens and prob have no geometric counterpart. In a weighted run they are
weighted quantities and nothing else.
Weights change the density field the watershed floods, and with it the zones
and the zone graph, which is one stage more than a rebuild has to redo.
rerun_watershed starts from the cached Delaunay graph and Voronoi volumes and
redoes only the zones, the watershed and the properties, so the tessellation,
which takes most of the runtime, is reused. It requires saveIntermediate=True
on the original run (section 4).
weighted = pyvide.rerun_watershed(
"voids", "run1",
newSaveName="run1_mass", # required with weights; a rerun never writes over run1 on its own
weights=cat.tracers["mass"], # one weight per kept tracer, same order
)
weighted.voidVol, weighted.voidVolWeighted # the same cells, unweighted and weighted
both = pyvide.rerun_watershed( # new weights and a new threshold in one pass
"voids", "run1", newSaveName="run1_mass_mt0.2",
weights=cat.tracers["mass"], mergingThreshold=0.2,
)
lean = pyvide.rerun_watershed( # catalog only, for a sweep over weights
"voids", "run1", newSaveName="run1_mass_lean",
weights="halos.npz:mass", outputs="voids",
)A rerun that only changes the threshold is written to run1_mt<threshold>,
similar to a rebuild. In contrast, a rerun that uses weights has no natural
name, since a weight vector cannot be summed up in a suffix, so it asks you for
a new name. Giving run1 itself as the name is refused unless you also pass
overwrite=True, and then it only works when nothing of the old run would be
left behind.
weights can take a path in the same forms as identify_voids does, and the
file it read is recorded in the new run's inputSources. Whatever you pass,
the run writes one sentence about it into parameters['weightsSource']:
inherited from run1, read from a file, passed as an array, or switched off.
That sentence matters most with outputs='voids': the weight vector itself
lives in the tracer file, which a lean rerun does not write, so the sentence
and, for a file, the recorded path are what remain.
Weights can also be given to identify_voids directly. They must be strictly
positive, one per tracer, in the order you passed the positions. Equal weights
reproduce the unweighted catalog bit for bit, so the weighted path is a
generalisation rather than a separate code route. Unequal ones are a different
definition of a void, not a correction to the old one: the core of a region can
move to another tracer, because the minimum of w / V need not be the minimum
of 1 / V. Typical use cases could be halo or galaxy mass, luminosity, and a
selection or completeness weight.
WEIGHTS_CEILING_AND_DUPLICATES.md has the definition in full.
Normally every tracer belongs to some void, so all voids together fill the box.
maxCellDensity changes this: a cell denser than the chosen ceiling (in units
of the mean density) cannot belong to any void, so denser structures are left
out (e.g., the walls between voids) and voids cannot merge across them. A void
whose lowest point is already above the ceiling disappears from the catalog
entirely. Like weights, the ceiling enters after the tessellation, so a rerun
can add, change or remove it without a new tessellation:
capped = pyvide.rerun_watershed(
"voids", "run1", newSaveName="run1_ceiling1.0",
maxCellDensity=1.0, # cells denser than 1.0x the mean density belong to no void
)
capped.statistics["excludedVolumeFraction"] # share of the box left out of every voidA ceiling can also be given to identify_voids directly and removed again
later:
dense = pyvide.identify_voids(..., saveName="run2", maxCellDensity=1.0,
saveIntermediate=True)
plain = pyvide.rerun_watershed("voids", "run2", newSaveName="run2_plain",
maxCellDensity=None)A rerun that leaves out maxCellDensity keeps the ceiling of the run it reads,
and rebuild_catalog cannot change the ceiling at all.
WEIGHTS_CEILING_AND_DUPLICATES.md explains how it interacts with
merging.
| file | contents | needed by |
|---|---|---|
_voids.npz |
the catalog, plus parameters and statistics | everything |
_tracers.npz |
kept tracers: positions, volumes, zone ids, weights, your own added columns | members, rerun_watershed, rebuild_catalog |
_void_details.npz |
void→zone and zone→tracer lists, child lists | members, children, descendants, rebuild_catalog |
_merge_events.npz |
the zone graph with its linking densities | rebuild_catalog |
_tessellation.npz |
the Delaunay graph and volumes; only with saveIntermediate=True |
rerun_watershed |
_zones.npz |
the pre-merging basins; only with zones=True |
optional |
_info.txt, _logs.txt |
parameters, statistics, timings, environment; the run log | you |
outputs='full' (the default) writes the four npz files and the two text
files. outputs='voids' writes only the catalog and the two text files, which
is what you might want in a loop over thousands of runs where the catalog is
all that is ever read. rebuild_catalog and rerun_watershed can use this as
well, and on a sweep it saves more than it does on a fresh run, because the
files a rebuild skips are almost exclusively copies of the ones it just read.
They also take saveHDF5. Every key of every file is listed in
FILES.md.
Every created file carries two stamps: the run that wrote it, and the
tessellation it descends from. A rebuild gets a fresh run stamp but keeps the
tessellation stamp, and its _info.txt names the run it came from as
parentSaveName. If files from a different run exist under the same name,
loading refuses with a message rather than joining them silently, and that now
holds between two rebuilds of the same run as well. The first line of every
_logs.txt states the run's stamps and its parent information.
# non-cubic box, per-axis periodicity (voids touching a wall get boundaryFlag)
catalog = pyvide.identify_voids(..., boxLen=(1000.0, 1000.0, 500.0),
periodicBox="xy") # or periodicBox=(True, True, False)
# coincident tracers merged into one tracer with the summed weight, instead of refused
catalog = pyvide.identify_voids(..., duplicateMode="merge")
# redshift-space distortions (standard convention; doRSD='VIDE' reproduces VIDE's)
catalog = pyvide.identify_voids(..., doRSD=True, omegaM=0.31, redshift=0.5,
velocities=velocities)
# voids in a density field on a grid instead of in tracers
catalog = pyvide.identify_voids_from_field("voids", "grid1", 640.0, field)
# additionally one HDF5 file with every table; the npz and txt files are still written
catalog = pyvide.identify_voids(..., saveHDF5=True)
# a KD-tree that respects the box's periodicity, for stacking and profiles;
# non-periodic axes stay walls, and results index the positions you passed
tree = pyvide.build_tree(positions, 640.0)
# the tracers within twice the first void's radius of its centre, as indices
# into positions; a sphere query never returns the same tracer twice
inside = tree.query_ball_point(cat.macrocenter[0], 2.0 * cat.radius[0])
# what does this argument do?
pyvide.explain("buffer")The README explains each of these in plain language and lists which features VIDE has that pyVIDE deliberately does not have.
Two columns are easy to misread, and both are VIDE's definitions:
np.mean(cat.ellipticity)is NaN whenever the catalog holds a one-member void, which at the default threshold it usually does. Usecat.ellipticity[cat.shapeReliable]orcat.select(cat.shapeReliable).centralDenis exactly zero for most voids, because the sphere it counts in is a sixty-fourth of the void and only members count. Those zeros are upper limits, not measurements.cat.numCentralis the raw count behind the column, so cut on it before you average or plot.
The README's "Before you plot anything" table adds voidID versus row number,
and CATALOG_COLUMNS.md has the rest. Read them once.
- IDENTIFY_VOIDS.md: every argument of the main function.
- CATALOG_COLUMNS.md: every column of a void catalog.
- FILES.md: every file and key a run can write.
- FOR_VIDE_USERS.md: how the two codes line up.
- PERFORMANCE.md: sizing a run, and what it costs.
- WEIGHTS_CEILING_AND_DUPLICATES.md: weights, the density ceiling, and coincident tracers.
- ALGORITHM.md: the pipeline, stage by stage.
- FIDELITY.md: what "reproduces VIDE" means and how it was verified.