Skip to content

Latest commit

 

History

History
613 lines (507 loc) · 30.8 KB

File metadata and controls

613 lines (507 loc) · 30.8 KB

Getting started

From tracer positions to a final void catalog, then everything you can do with the catalog without running the finder again. Every block below is copy-pasteable. IDENTIFY_VOIDS.md documents every argument of identify_voids.

Contents: install · a first run · loading your own data · the pyVIDE run · reading your catalog back · exactly reproducing a VIDE run · members and your own columns · post-processing · merging without re-running · weights without re-running · a density ceiling · the output files · other features in one line each · before you plot anything · where to find more details

1. Install

git clone https://github.com/nicosmo/pyVIDE.git
cd pyVIDE
pip install -e .

This pulls in numpy and scipy, which is the whole required dependency list. No compiler needed. pyVIDE is pure python, so pip is optional: you can also download the repository and add its folder to your path (sys.path.insert(0, "/path/to/pyVIDE")), or start python from the repository root, and import pyvide works.

2. A first run, no data needed

python3 example_run.py

This is the quickest way to see whether pyVIDE works on your system, and it walks through the whole pipeline: finding, loading, filtering, memberships, re-running with weights, rebuilding at another merging threshold, and field mode. It looks for data in three places, in this order: the block at the top of the file, where you can put your own catalog; the public reference catalog from the pyVIDE data record, if you have downloaded it into the folder named in that block; and, failing both, a synthetic box of tracers it generates itself, which then takes a few minutes to run through the pipeline. It prints which of the three it used.

For developers: pip install pytest and pytest tests runs the whole suite, including the frozen end-to-end regression fixture that re-runs the pipeline on a stored 100 000-tracer box and compares every catalog column against the stored reference. pytest tests -m "not slow" skips those.

3. Loading your own data

pyVIDE takes numpy arrays, so any reader that gets your positions into an (N, 3) array will do. Three common cases:

import numpy as np
import pyvide

# a text catalog: one row per halo, columns mass, x, y, z, vx, vy, vz
table = np.loadtxt("halos.txt")
masses     = table[:, 0]          # (N,)
positions  = table[:, 1:4]        # (N, 3), Mpc/h, required
velocities = table[:, 4:7]        # (N, 3), km/s, only needed for doRSD

# an HDF5 snapshot (Gadget-style layout shown; adapt the group names)
import h5py
with h5py.File("snapshot.hdf5") as f:
    positions  = f["PartType1/Coordinates"][:]
    velocities = f["PartType1/Velocities"][:]

# a numpy file: pass the path and pyVIDE records where the data came from
# positions = "halos.npz:pos"

pyVIDE has no snapshot readers of its own: the two or three lines above are all a reader would do. The run checks the positions you pass and reports three things: how many lie outside the box, how many lie so far outside that a single wrap cannot bring them back into the box, and whether the coordinates span less than a tenth of the box on some axis, which usually means wrong units, or a sub-volume passed with the full boxLen. It reports and carries on, leaving your array untouched.

What then happens to a coordinate outside the box depends on the mode. 'pyvide' wraps the periodic axes with a modulo, so any excess comes back, and stops with an error if a non-periodic axis carries anything outside, because such an axis is never wrapped. The two VIDE modes wrap the periodic axes once by a box length, exactly as VIDE does, and drop whatever a single wrap leaves outside; on a walled axis nothing is wrapped, and a tracer outside the box is dropped. Both modes treat the walls themselves the same way from there on.

4. The pyVIDE run

catalog = pyvide.identify_voids(
    saveDir="voids",              # path to output folder, created if missing
    saveName="run1",              # prefix of every output file
    boxLen=640.0,                 # scalar for a cubic box, (Lx, Ly, Lz) otherwise
    positions=positions,          # input tracer positions
    velocities=velocities,        # optional input velocities; stored, and used only if doRSD is not False
    otherProperties={             # optional extra per-tracer columns, see section 7
        "mass": masses,
    },
    VIDE_mode="pyvide",           # required: 'pyvide' for a new catalog, see the table below
    numDivisions=2,               # sub-boxes per axis
    numThreads=4,                 # sub-boxes tessellated in parallel; bit-identical to 1
    saveIntermediate=True,        # cache the tessellation, so weights and the density ceiling can change later (sections 10 and 11)
)
print(catalog)

The run prints its progress and ends with a summary block: total void count, box dimensions, periodic axes, applied merging threshold, hierarchy depth, range of void radii, mode employed, and the time spent per stage. print(catalog.summary()) prints it again for an old run, without the timings, which live in the _info.txt file.

The one decision you cannot skip is VIDE_mode, because it changes the arithmetic:

VIDE_mode what it does use it when
'VIDE_halos' VIDE's --halos chain: positions pass through a %e text file, then float32 reproducing or comparing with a VIDE run on a halo or galaxy catalog
'VIDE_direct' the float32 chain without the text step reproducing a VIDE run on matter particles
'pyvide' float64 end to end, no tracer dropped at the box edge, buffer sized from the catalog, periodic images with a single sub-box the default choice, unless you are reproducing a VIDE run

Which mode you want follows from what the catalog is for. If it has to match VIDE's, whether you are comparing against an existing VIDE run or producing the catalog that run would have produced, use the mode it used. If it does not, use 'pyvide': the same finder in double precision, with the buffer sized from your catalog; IDENTIFY_VOIDS.md lists everything the mode changes. In the VIDE modes buffer stays at VIDE's 0.1 whatever the catalog, which can be too thin for a sparse sample: when tracers end up as neighbours of the guard shell rather than of each other, the run says so and tells you to raise it. PERFORMANCE.md explains how to size it. To reproduce an existing VIDE run setting by setting, FOR_VIDE_USERS.md has the template. pyvide.explain('VIDE_mode') prints the long version, and explain() does the same for any other argument.

None of the arguments below appear in the call above. Each has a default that suits most runs, but each is worth knowing before you hit the case it covers.

duplicateMode decides what happens when two tracers land on the same coordinate. The default 'error' stops the run and names them. 'merge' instead collapses each coincident group into one tracer carrying the group's summed weight, so the surviving position keeps the density the whole group had, and the tracer file still records every merged-away tracer and the survivor it belongs to (the run then uses the weighted path even if you passed no weights). A VIDE mode's float32 cast can create coincidences your float64 input did not have. WEIGHTS_CEILING_AND_DUPLICATES.md has the details.

mergingThreshold (default $10^{-9}$, no merging, as in VIDE) can be changed after the run without re-tessellating, which is what section 9 does. outputs (default 'full') decides how much is written: 'voids' keeps the catalog alone, and that costs you sections 7, 9 and 10, since member lookups, rebuilding at another threshold and re-running with weights all read the files it skips. The rebuilds of sections 9 and 10 take outputs too, and there it is often the right choice. FILES.md says how large each file gets. A second run under the same saveName refuses to overwrite the first unless you pass overwrite=True, and even then pyVIDE deletes nothing: a run that would leave files of the old run behind refuses to overwrite them and instead names them in order to avoid mixing files from different runs.

5. Reading your catalog back

cat = pyvide.load_catalog("voids", "run1")     # saveDir, saveName; reads a finished run from disk

cat.numVoids                                   # rows in this catalog, one per void
cat.radius, cat.macrocenter, cat.numPart       # columns are numpy arrays
cat.columns()                                  # every column name
cat.parameters                                 # how the run was made (a dict)
cat.statistics                                 # what it found (a dict)

radius   = cat.radius                          # a column as an attribute ...
radius   = cat.column("radius")                # ... or by name, same array
coreDens = cat.coreDens                        # core densities, in units of the mean density

big = cat.radius > 20.0                        # plain numpy masks
print(cat.macrocenter[big])

The columns are the catalog's own arrays; add .copy() if you intend to modify one. Column meanings and units are described in CATALOG_COLUMNS.md. One rule to learn now: every column is indexed by row, while voidID, parentID, members(), children() and descendants() speak in void IDs. The two agree only in a fully periodic run at the default settings, because a wall or a cut removes rows without renumbering ids, and cat.rows_of(ids) converts between them.

6. Exactly reproducing a VIDE run

To reproduce an existing VIDE run rather than make a new catalog, the call of section 4 changes in one place, the mode, and takes the settings VIDE used:

vide_run = pyvide.identify_voids(
    saveDir="voids",
    saveName="vide_run",
    boxLen=640.0,
    positions=positions,
    velocities=velocities,
    otherProperties={"mass": masses},
    VIDE_mode="VIDE_halos",       # 'VIDE_halos' for a run prepared with --halos, 'VIDE_direct' without
    mergingThreshold=1e-9,        # the VIDE run's value
    numDivisions=2,               # VIDE's numZobovDivisions
    buffer=0.1,                   # VIDE's default buffer value
    numThreads=4,
)

That is the whole recipe for the catalog. FOR_VIDE_USERS.md has the full template, including the three settings that matter only when a VIDE run used them (minRadius, maxCentralDen, doRSD='VIDE') and the volumeMethod='exact' switch for bit-comparing per-tracer volumes. VIDE also applied its rMin cut at the mean tracer separation. Pass minRadius=-1 to include it, as the template in FOR_VIDE_USERS.md does.

pyVIDE keeps every density basin as a catalog row. VIDE's text catalogs leave out one kind of basin, a zone holding only a single tracer, so a VIDE catalog of the same tracers is potentially a few rows shorter. The pyVIDE run says so in its summary, and the number VIDE would print is stored:

cat.statistics["numVideVoidDescRows"]          # the length VIDE's catalog has

To get exactly VIDE's rows, in VIDE's order, do:

vide_like = pyvide.filter_catalog(cat, "untrimmed_all")

filter_catalog computes nothing. It only drops the single-tracer zones, sorts by density contrast the way jozov2 does, and applies the cuts of the variant you name. 'untrimmed_all' adds no cuts of its own, so it is the like-for-like list. VIDE's default catalog, the centers_all_* file with no filename prefix, is the "trimmed" variant instead: top-level voids only, with the central-density cut applied at VIDE's merging threshold. VIDE fed one number into both roles. pyVIDE keeps them apart, so you name the cut yourself:

vide_default = pyvide.filter_catalog(cat, "trimmed_all", maxCentralDen=1e-9)

At $10^{-9}$ every void is top-level, so "untrimmed_dencut_all" gives the same rows there, but they differ once merging is on. FOR_VIDE_USERS.md has the full table of VIDE file names to pyVIDE calls.

7. Member tracers, and your own columns

The catalog is one of several files a run writes. The tracer file holds one row per kept tracer, and the membership file connects voids to their tracers:

cat = pyvide.load_catalog(
    "voids", "run1",
    loadTracers=True,       # positions, Voronoi volumes, zone ids, weights, the columns you added
    loadMembership=True,    # the void -> zone -> tracer lists (optional: members() reads them when first asked)
)

pos  = cat.tracers["positions"]     # (N, 3), the kept tracers, in input order
vol  = cat.tracers["volume"]        # each tracer's Voronoi cell volume
zone = cat.tracers["zoneID"]        # which zone/void each tracer belongs to
mass = cat.tracers["mass"]          # the column you passed as otherProperties

voidID = int(cat.voidID[np.argmax(cat.radius)])   # ID of the void with the largest radius
inVoid = pyvide.members(cat, voidID)              # rows of that void's tracers in cat.tracers (positions, volume, ...)
print(inVoid.size, "member tracers")
print(pos[inVoid], vol[inVoid], mass[inVoid])     # their positions, volumes, masses

About otherProperties: it is a dict of name -> array with any number of entries, each (N,) or (N, k), so for example a magnetic-field vector would be one (N, 3) column. The finder never reads them and they are stored only for the tracers the run kept, row-aligned with positions, volume and zoneID, so mass[members(cat, id)] gives the masses of one void's tracers. Names must be valid python identifiers, must not collide with the existing columns pyVIDE writes itself, and must not start with merged_, which the coincident-tracer bookkeeping uses. Each entry may also be a path, "halos.npz:temperature".

members takes a void ID, not a row number, and the indices it returns are a third numbering: rows of the tracer file, not of your input array. The box cut drops tracers, so indexing your own array with them silently gives the wrong objects. The kept tracers are those left inside the box after the wrap described in section 3. cat.tracers["inputIndex"] maps every kept row back to your input array, and cat.statistics["numDroppedTracers"] counts the rest.

8. Post-processing: filters and selections

None of this re-runs the finder: these are selections over rows that already exist. The VIDE variants exist to reproduce VIDE's catalogs. For a new analysis, start from the full catalog instead and make your own cuts. The catalog pyVIDE writes keeps every void, including the single-tracer zones VIDE never wrote at all; 'untrimmed_all' is as close as VIDE got. Your own boolean mask through select says what you mean more directly than a variant name does.

# VIDE's eight catalog variants, as filters (two shown; see the paragraph
# below, and section 6 for the two common ones)
pyvide.filter_catalog(cat, "trimmed_nodencut_all", minRadius=5.0)
pyvide.filter_catalog(cat, "trimmed_central", minRadius=5.0, maxCentralDen=1e-9)

# your own cuts: a boolean mask over the rows gives a smaller catalog object
usable = cat.select(cat.shapeReliable)           # voids with a meaningful shape tensor
large  = cat.select(cat.radius > 20.0)           # a radius cut, in your units
inner  = cat.select(~cat.boundaryFlag)           # ~ negates the mask: voids NOT bordering a wall

# the hierarchy (only interesting once merging is on, see section 9)
cat.children(voidID)                             # direct sub-voids, as IDs
cat.descendants(voidID)                          # every void inside this one, as IDs
rows = cat.rows_of(cat.descendants(voidID), missing="mask")
# rows of those IDs; -1 where a filtered catalog no longer has the void
# (missing="drop" would leave them out; the default raises an error)

The variant names are untrimmed, untrimmed_dencut, trimmed_nodencut and trimmed, each with _all or _central. Every one of them first drops the "wrong" voids (densCon above $10^{4}$ or a non-finite volume), the too-small ones (radius < minRadius, which does nothing at the default 0.0) and, on a box with walls, the ones bordering a wall. On top of that, dencut removes voids whose centralDen exceeds maxCentralDen, and trimmed keeps only top-level voids. all versus central is VIDE's survey-region split, which makes no difference for a simulation box.

9. Different merging thresholds and the void hierarchy

The default merging threshold of $10^{-9}$ is small enough that no basins merge in practice, so every zone is its own isolated void. Raise it and neighbouring basins join across their lowest ridge, and the catalog gains a hierarchy: small voids inside larger ones. Changing the threshold needs no new tessellation, because the run stored the zone graph:

merged = pyvide.rebuild_catalog("voids", "run1", mergingThreshold=0.2)

This re-runs the watershed step over the stored graph, recomputes every property and the hierarchy, and writes a full file set named run1_mt0.2: without a newSaveName, rebuild_catalog adds the merging threshold it applied to the run's name, so that a rebuild never clashes with the run it read. Pass newSaveName to choose a name yourself. Its output is column for column identical to a from-scratch run at that threshold (checked by tests/test_rebuild.py), and those from-scratch runs are the ones verified against VIDE.

merged.statistics["maxTreeLevel"]                 # how deeply nested the catalog is
merged.isTopLevel.sum()                           # voids contained in no other void
merged.numZones.max()                             # the most zones any one void merged

top = int(merged.voidID[np.argmax(merged.radius)])  # the largest void, as an ID
merged.children(top)                              # its direct sub-voids (IDs)
merged.descendants(top)                           # everything inside it, at any depth
merged.treeLevel                                  # per void: how many voids contain it (0 = top level)
merged.numDescendants                             # per void: size of the subtree below it

parentID gives the ID of the smallest void containing a given one, or -1 if nothing contains it. treeLevel counts how many voids contain it, so a top-level void is 0, and large catalogs can reach ten or more. Which set you use depends on the question: isTopLevel alone covers space without overlaps, and the full set of rows gives you voids on every scale.

Two things about the value 0.0. It means "merge everything the topology allows", so the top level collapses to a single void spanning the whole box, and the useful output is the hierarchy below it: the complete merger tree, with no threshold chosen by hand. That is what makes it worth having, since any other threshold imposes a scale whose effect depends on the tracer population and its bias. It is also the slowest to rebuild, because every tracer then belongs to every void above it.

A sweep is one rebuild per threshold, and a rebuild takes a fraction of the time a typical run costs. Each one writes its own run1_mt<threshold> files, so the thresholds never collide with each other and the original run1 is never touched. overwrite=True is only needed when you run the same sweep again, over file sets a previous pass already wrote:

for threshold in (0.1, 0.2, 0.5, 1.0):
    pyvide.rebuild_catalog("voids", "run1", mergingThreshold=threshold,
                           overwrite=True, verbose=False)

On a sweep it is usually worth passing outputs='voids'. The tracer file that a threshold rebuild writes is a byte for byte copy of the one it originally read, so a ten-threshold sweep otherwise leaves ten copies of the same file on disk. The catalogs themselves are unchanged by this setting:

for threshold in (0.1, 0.2, 0.5, 1.0):
    pyvide.rebuild_catalog("voids", "run1", mergingThreshold=threshold,
                           newSaveName="run1_lean_mt{0:g}".format(threshold),
                           outputs="voids", verbose=False)

A lean rebuild cannot be rebuilt from this output in turn, since it did not write the needed zone graph. Every threshold in the loop reads run1 anyway, so nothing is lost. One thing to know before switching a sweep to 'voids': if a previous pass wrote full file sets under the exact same names, the lean pass refuses, because replacing only the catalog would leave the older membership and tracer files in the same folder. Delete those first, or use other names.

Two conditions for this. The original void identification must have been run with outputs='full' (the default), because a rebuild reads its tracer, membership and merge-event files. A rebuild can only change the threshold: the stored zone graph came out of the original run's densities, so changing the weights needs rerun_watershed, which is the next section.

10. Weights without re-tessellating

A weight scales how much density a tracer contributes: give one tracer twice the weight of another and it counts for twice as much density. Only the density side of the run changes. Each Voronoi cell is rescaled to V * mean(w) / w and the watershed runs on that, so the zones, merging and the densities that follow from them (coreDens, densCon, leakDens, prob) are weighted, while everything geometric stays geometric: voidVol, radius, macrocenter, centralDen, maxRadius and the shape tensor.

Both sides are kept, so nothing is lost: the weighted volumes are reported alongside as voidVolWeighted, zoneVolWeighted and radiusWeighted, and the unweighted density at each core survives as coreDensGeometric. densCon, leakDens and prob have no geometric counterpart. In a weighted run they are weighted quantities and nothing else.

Weights change the density field the watershed floods, and with it the zones and the zone graph, which is one stage more than a rebuild has to redo. rerun_watershed starts from the cached Delaunay graph and Voronoi volumes and redoes only the zones, the watershed and the properties, so the tessellation, which takes most of the runtime, is reused. It requires saveIntermediate=True on the original run (section 4).

weighted = pyvide.rerun_watershed(
    "voids", "run1",
    newSaveName="run1_mass",              # required with weights; a rerun never writes over run1 on its own
    weights=cat.tracers["mass"],          # one weight per kept tracer, same order
)

weighted.voidVol, weighted.voidVolWeighted   # the same cells, unweighted and weighted

both = pyvide.rerun_watershed(               # new weights and a new threshold in one pass
    "voids", "run1", newSaveName="run1_mass_mt0.2",
    weights=cat.tracers["mass"], mergingThreshold=0.2,
)

lean = pyvide.rerun_watershed(               # catalog only, for a sweep over weights
    "voids", "run1", newSaveName="run1_mass_lean",
    weights="halos.npz:mass", outputs="voids",
)

A rerun that only changes the threshold is written to run1_mt<threshold>, similar to a rebuild. In contrast, a rerun that uses weights has no natural name, since a weight vector cannot be summed up in a suffix, so it asks you for a new name. Giving run1 itself as the name is refused unless you also pass overwrite=True, and then it only works when nothing of the old run would be left behind.

weights can take a path in the same forms as identify_voids does, and the file it read is recorded in the new run's inputSources. Whatever you pass, the run writes one sentence about it into parameters['weightsSource']: inherited from run1, read from a file, passed as an array, or switched off. That sentence matters most with outputs='voids': the weight vector itself lives in the tracer file, which a lean rerun does not write, so the sentence and, for a file, the recorded path are what remain.

Weights can also be given to identify_voids directly. They must be strictly positive, one per tracer, in the order you passed the positions. Equal weights reproduce the unweighted catalog bit for bit, so the weighted path is a generalisation rather than a separate code route. Unequal ones are a different definition of a void, not a correction to the old one: the core of a region can move to another tracer, because the minimum of w / V need not be the minimum of 1 / V. Typical use cases could be halo or galaxy mass, luminosity, and a selection or completeness weight. WEIGHTS_CEILING_AND_DUPLICATES.md has the definition in full.

11. A density ceiling

Normally every tracer belongs to some void, so all voids together fill the box. maxCellDensity changes this: a cell denser than the chosen ceiling (in units of the mean density) cannot belong to any void, so denser structures are left out (e.g., the walls between voids) and voids cannot merge across them. A void whose lowest point is already above the ceiling disappears from the catalog entirely. Like weights, the ceiling enters after the tessellation, so a rerun can add, change or remove it without a new tessellation:

capped = pyvide.rerun_watershed(
    "voids", "run1", newSaveName="run1_ceiling1.0",
    maxCellDensity=1.0,               # cells denser than 1.0x the mean density belong to no void
)
capped.statistics["excludedVolumeFraction"]   # share of the box left out of every void

A ceiling can also be given to identify_voids directly and removed again later:

dense = pyvide.identify_voids(..., saveName="run2", maxCellDensity=1.0,
                              saveIntermediate=True)
plain = pyvide.rerun_watershed("voids", "run2", newSaveName="run2_plain",
                               maxCellDensity=None)

A rerun that leaves out maxCellDensity keeps the ceiling of the run it reads, and rebuild_catalog cannot change the ceiling at all. WEIGHTS_CEILING_AND_DUPLICATES.md explains how it interacts with merging.

12. What the run writes

file contents needed by
_voids.npz the catalog, plus parameters and statistics everything
_tracers.npz kept tracers: positions, volumes, zone ids, weights, your own added columns members, rerun_watershed, rebuild_catalog
_void_details.npz void→zone and zone→tracer lists, child lists members, children, descendants, rebuild_catalog
_merge_events.npz the zone graph with its linking densities rebuild_catalog
_tessellation.npz the Delaunay graph and volumes; only with saveIntermediate=True rerun_watershed
_zones.npz the pre-merging basins; only with zones=True optional
_info.txt, _logs.txt parameters, statistics, timings, environment; the run log you

outputs='full' (the default) writes the four npz files and the two text files. outputs='voids' writes only the catalog and the two text files, which is what you might want in a loop over thousands of runs where the catalog is all that is ever read. rebuild_catalog and rerun_watershed can use this as well, and on a sweep it saves more than it does on a fresh run, because the files a rebuild skips are almost exclusively copies of the ones it just read. They also take saveHDF5. Every key of every file is listed in FILES.md.

Every created file carries two stamps: the run that wrote it, and the tessellation it descends from. A rebuild gets a fresh run stamp but keeps the tessellation stamp, and its _info.txt names the run it came from as parentSaveName. If files from a different run exist under the same name, loading refuses with a message rather than joining them silently, and that now holds between two rebuilds of the same run as well. The first line of every _logs.txt states the run's stamps and its parent information.

13. The other features, in one line each

# non-cubic box, per-axis periodicity (voids touching a wall get boundaryFlag)
catalog = pyvide.identify_voids(..., boxLen=(1000.0, 1000.0, 500.0),
                                periodicBox="xy")   # or periodicBox=(True, True, False)

# coincident tracers merged into one tracer with the summed weight, instead of refused
catalog = pyvide.identify_voids(..., duplicateMode="merge")

# redshift-space distortions (standard convention; doRSD='VIDE' reproduces VIDE's)
catalog = pyvide.identify_voids(..., doRSD=True, omegaM=0.31, redshift=0.5,
                                velocities=velocities)

# voids in a density field on a grid instead of in tracers
catalog = pyvide.identify_voids_from_field("voids", "grid1", 640.0, field)

# additionally one HDF5 file with every table; the npz and txt files are still written
catalog = pyvide.identify_voids(..., saveHDF5=True)

# a KD-tree that respects the box's periodicity, for stacking and profiles;
# non-periodic axes stay walls, and results index the positions you passed
tree = pyvide.build_tree(positions, 640.0)
# the tracers within twice the first void's radius of its centre, as indices
# into positions; a sphere query never returns the same tracer twice
inside = tree.query_ball_point(cat.macrocenter[0], 2.0 * cat.radius[0])

# what does this argument do?
pyvide.explain("buffer")

The README explains each of these in plain language and lists which features VIDE has that pyVIDE deliberately does not have.

14. Before you plot anything

Two columns are easy to misread, and both are VIDE's definitions:

  • np.mean(cat.ellipticity) is NaN whenever the catalog holds a one-member void, which at the default threshold it usually does. Use cat.ellipticity[cat.shapeReliable] or cat.select(cat.shapeReliable).
  • centralDen is exactly zero for most voids, because the sphere it counts in is a sixty-fourth of the void and only members count. Those zeros are upper limits, not measurements. cat.numCentral is the raw count behind the column, so cut on it before you average or plot.

The README's "Before you plot anything" table adds voidID versus row number, and CATALOG_COLUMNS.md has the rest. Read them once.

15. Where to go next