lab/ is a locally built Sagan engine and a suite of checks that pin down what
it actually does, so that the claims this project makes about Sagan's behaviour
can be verified by execution instead of by reading C.
Last verified against Sagan main at 3b9b0fa, RSigma 0.22.0 and the rule
corpus at a1cf3b3, on 2026-09-17: 9,331 rules judged by the corpus differential, no
disagreement, 9,314 of them exercised; the correlation differentials adding 921
after boundaries and 10 xbits state machines, also without disagreement; and
the model differential agreeing with the engine on 4,308 rules. That line is the lab's support statement: it says what was
measured and when, and it is meant to be updated by whoever runs the suite
next.
It is not part of the test suite and CI never runs it. A full pass starts
several hundred Sagan processes and takes the better part of an hour, which does
not belong on every push. The repository's own tests/ cover the conversion
against a Python model of the engine and against the real RSigma, in seconds,
and that is what runs on every change.
It lives here all the same, for two reasons. The claims it produced are all over the converter's comments, the design notes and the upstream bug reports, and a claim nobody else can re-measure is a claim on trust. And it imports the converter and the probe generator from this same checkout, so keeping the two together is what stops them drifting apart.
| Prerequisites | Nix (the harness runs the engine inside nix-shell), a C toolchain, Python 3.10+ |
| One-off build | lab/build/build-sagan.sh, a few minutes, clones Sagan main |
| Full check suite | lab/run-all.sh, the better part of an hour |
| One corpus differential | 30 to 60 minutes, and it needs rsigma on PATH |
| Optional | vector on PATH to judge the shipped pipeline with the rules |
Two of the three engine builds need patches to Sagan that this repository does not carry, so part of the suite cannot run from a fresh clone. What runs and what does not is listed below.
Reading Sagan's C is not the same as knowing what Sagan does. Building the engine and running it found three converter defects that the source reading had missed, and all three made the converted rules noisier than the originals:
| Defect | Scope | Nature |
|---|---|---|
after: count N emitted gte: N |
970 rules | alerted one event too early |
after: track by_string grouped on the username |
5 rules | grouping too fine |
both emitted a bare disjunction |
latent | would fire on one address alone |
flexbits ignored its direction |
9 correlations | keyed on an address where Sagan keys on the user |
flexbits knew 6 of 14 direction tokens |
latent | the direction would be taken for the bit name |
after honoured keys the parser ignores |
4 correlations | grouped more finely, so fired less often |
facility, level, tag accepted as aliases |
latent | converts a rule Sagan cannot load |
syslog_priority unknown |
latent | a real selector refused as unknown |
PCRE A and x dropped as inert |
latent | both change what the pattern matches |
json_*_strstr treated as *_contains |
latent | emits |contains where Sagan compares whole values |
json_base64_decode spellings accepted |
latent | converts a rule Sagan cannot load |
blacklist: all without a parsed position |
latent | fires where Sagan cannot alert at all |
The first is the one worth remembering. The project's own design document had
asserted "alert from N+1" since the beginning, while the code emitted gte: N.
Documentation and implementation had disagreed for the entire life of the
project, and only running both engines side by side exposed it.
bin/ sagan-upstream pristine upstream, hardened as Nix builds it
sagan-patched upstream plus one local patch (not shipped)
sagan-sane every patch; use this to measure what a rule means
(generated, not committed: see build/)
build/ build-sagan.sh build the binaries from source
config/ sagan.yaml the harness configuration; @LAB@ is substituted
with this checkout's path on every run
make-country-mmdb.py
builds the tiny fixed GeoIP database (see below)
country.mmdb generated by the script above, not committed
rules/ classification, reference, protocol map, rulebases
(including lab-normalize.rulebase, see below), and
the fixed denylist and Zeek intel feeds: lab-* for
the checks, diff-* for the corpus differential
checks/ check_*.py the verification suite, one file per family
differential/
engine_differential.py
detection semantics, corpus-wide, Sagan vs rsigma
after_differential.py
the after correlations, at their N/N+1 boundary
xbits_differential.py
the xbits state machine, setter then tester
model_differential.py
the Python model of Sagan against the real one
attribution.py which probe an alert belongs to, shared by the above
vector_pipeline.py
runs the shipped VRL transforms, for the profiles
whose fields a syslog line alone cannot carry
work/ scratch: rules, events and logs of the last run (wiped each time)
harness.py the measuring instrument
run-all.sh run every check
From a fresh clone, once:
python lab/config/make-country-mmdb.py # needs mmdb-writer and netaddr
lab/build/build-sagan.sh # clones and builds Sagan mainThen:
lab/run-all.sh # every check
lab/run-all.sh correlation # one family, by filename substringEach check prints its own pass/fail lines, plus a [n/total] name ok (Ns elapsed) line as it starts and finishes, so a long run says where it is rather
than sitting silent. The script exits non-zero if any file failed.
LAB_PYTHON overrides the interpreter.
Expect the full suite to take the better part of an hour: every single case
starts a real Sagan process inside a nix-shell, and there are over two hundred
of them.
Concurrent runs are serialised for you. Two at once used to destroy each
other, since each wipes work/ and the shared state in /dev/shm, and the
result was never an error but a plausible wrong answer: a load failure naming a
keyword the check never used, or a file where nothing fires. This page warned
against it and it still happened three times, twice costing a wrong conclusion,
because a warning binds only whoever reads it. run_sagan now takes an
exclusive lock, so a second caller waits and says so. If a run is killed mid
-flight the lock is released with the process; a stale one is reported by name
after fifteen minutes.
Sagan must be built for this machine's Nix store. If bin/ is stale or
missing, run build/build-sagan.sh. The binaries are build output and are not
committed: they are only valid for the store of the machine that produced them.
The engine needs three patches before anything about it can be measured
reliably: two for its after correlation path, and one for an address parser
that aborts the process on an ordinary log line. They are not in this
repository: they fix defects that are not public, and publishing the patches
here would amount to disclosing them, which is not a converter's call to make.
Whoever holds them points SAGAN2SIGMA_LAB_PATCHES at the directory containing
them and gets the two extra builds.
Without them, a fresh clone can run:
| Needs the patched builds? | |
|---|---|
| 11 of the 16 check files | no, they run on the plain upstream build |
check_correlation, check_track_keys, check_flexbits, check_engine_limits, check_silence |
yes, each skips itself and says so |
the three differentials in differential/ |
yes, each refuses to run and says so |
A skip is printed, counted and named. It is never reported as a pass, so a short run cannot be mistaken for a full one.
That is the honest limit of what this directory offers a contributor today: the
engine behaviours the checks pin down are reproducible, the corpus-wide
comparison against the engine is not. The comparison that does run everywhere
is the one in tests/differential/, against a Python model of Sagan and the
real RSigma, which CI executes on every change.
from harness import Report, event, rule, sagan
report = Report("what this file establishes")
fired = sagan(
[rule(1, 'program: sshd; content:"Hello"')],
[event("Hello world"), event("goodbye")],
)
report.check("matches exactly", "1" in fired.get("Hello world", set()), True)
raise SystemExit(0 if report.done() else 1)sagan() returns message text → set of SIDs that alerted, so give every probe
a distinct message. alerts() additionally returns the source and destination
address the engine resolved, which is what makes Parse_IP observable.
loads() answers whether Sagan accepts a rule at all, which settles questions
an alert count cannot: after: track by_string is rejected at load, and that
is only possible if the key contributes nothing to the parser.
Use binary=PATCHED for anything involving after, and guard the file with
skip_unless_available(report, PATCHED) so that it skips rather than fails
where that build is absent; see the caveats below. Pass at_time= to run under
faketime and timezone= to set TZ, which is what makes alert_time
observable. Give at_time an epoch ("@1787063400") when the timezone is the
thing under test: a wall-clock string is read in the current zone, so it reads
the same in every zone and tests nothing.
Two habits matter more than the code. State what the file establishes, in its
docstring, in terms a reader who has not read Sagan's C would use. And give
every assertion a control: a measurement that passes because both sides did
nothing is the failure mode this whole directory exists to avoid, which is why
run-all.sh prints skips separately and why the differentials count how many
rules actually fired.
check_matching.py
: Case sensitivity in content, meta_content and program, and the fact that
nocase binds only to the content it follows. Negation. |41| decoded as
hex in content but treated as regex alternation in pcre. Program
alternatives and globs.
check_positional.py
: A zero-valued distance or a lone within is inert, which is what lets 245
corpus rules convert. A non-zero distance is an absolute offset from the
start of the message, not a gap from the previous match, which is why the
converter refuses it rather than emitting an ordered regex. Plus Parse_IP
position and delimiter handling, and the fallback to the syslog sender when
the requested position holds no address.
check_correlation.py
: after: count N alerts from the N+1th event. by_string is honoured by
threshold and inert under after. threshold caps alert volume for both
limit and suppress without changing detection. xbits set/isset/isnotset
and the ordering of ip_pair.
check_track_keys.py
: Which after: track keys the engine honours. The parser uses strcmp, so
by_user, byusername, by_tag and by_hostname are inert and contribute
nothing to the counter key, while a rule whose only key is inert is
rejected at load. by_srcport and by_dstport are real. Observable because
count 1 makes "do these two events share a bucket?" visible.
check_meta_content.py
: How the helper is split from the values. The first comma separates wherever
it sits, including inside the quotes; Between_Quotes drops every quote it
meets, so ""%sagan%" yields the helper %sagan% and the 72 Cisco rules
search for %ASA rather than "%ASA; a value keeps the stray closing quote
a rule leaves on it. A single space after the separating comma does not reach
the search, which is the shape 156 corpus options use.
check_enrichment_feeds.py
: blacklist and zeek-intel against local fixtures, since both ship disabled
and expect feeds that are not on this machine. Which address each direction
tests, that both fires when either end is listed, that an unrecognised
direction leaves the flag clear so the rule fires on its other conditions,
and that all scans only the addresses Parse_IP found. Plus an engine
defect: mask bits leak between denylist lines, so a shorter prefix following
a longer one is silently narrowed. bluedot is absent on purpose, being a
network query to a closed service.
check_alert_time.py
: Which clock the window is read against. Run under faketime, an event
stamped Sunday 03:00 fires a Tuesday-afternoon window when the machine
believes it is Tuesday afternoon, and stays silent on a window matching its
own stamp: the wall clock decides, not the event. localtime, so one fixed
instant is inside a 1400-1500 window in UTC and outside it in Tokyo. Plus
inclusive bounds, tm_wday numbering, the midnight-crossing branches, and
the fact that hours without days loads and can never fire.
check_event_id.py
: The fallback window when no json_map binds event_id: nine characters, not
the ten the strlcpy size argument suggests, and the searched string is
" <id>: " with both spaces, so an ID at offset 0 never matches. Plus the
structured path, which compares the decoded value whole.
check_json_ops.py
: json_meta_content is an OR over values, each compared whole, with
json_meta_contains switching to a substring search and json_meta_strstr
loading but doing nothing. json_pcre is unanchored, and on a key the event
does not carry it matches unconditionally, unlike its two siblings.
check_pcre_flags.py
: What each PCRE flag letter does. The engine's switch handles i s m x A E G
and has no default case, so any other letter is ignored rather than rejected,
which is why the converter must not refuse U or H. A (anchored) and x
(extended) do change matching, so dropping them silently is not safe.
check_envelope.py
: Which envelope selectors exist. Only the syslog_ forms are keywords: a bare
facility, level or tag aborts the ruleset. syslog_priority is real and
matches a field distinct from syslog_level, with | alternation and exact
comparison.
check_flexbits.py
: The fourteen direction tokens Flexbit_Type() accepts, and that anything else
is rejected at load. That the bit name is the third argument, proved with a
bit literally named by_src. Set/isset over time and the keying of by_src
and none. Plus an engine defect: address directions compare the printable
address buffer, so a correlation whose address comes from parse_src_ip
depends on the message text following it.
check_normalize.py
: liblognorm normalization overrides parse_src_ip and parse_dst_ip when
it resolves an address, and positional parsing only fills in what
normalization left unset, including one address out of two. This is the claim
D_NORMALIZE_PRECEDENCE rests on for 88 corpus rules, and the converter
reproduces the fallback half deliberately. Uses its own rulebase; see the
caveat below.
check_enrichment.py
: json_map: "message" redirects a content search to that key, and without it
content reads the raw body. json_content is an exact whole-value match,
not a substring search. A pass rule alerts before it short-circuits.
country_code requires a resolved country for both is and isnot, so an
address the database cannot place fires neither.
checks/ pins one behaviour at a time. differential/ asks the other question:
does the converted corpus behave like the original, rule by rule.
lab/differential/engine_differential.py --rules <corpus> # detection
lab/differential/after_differential.py --rules <corpus> # correlation boundary
lab/differential/model_differential.py --rules <corpus> # the model, not the conversionBoth batch heavily, which is what makes the corpus tractable: one Sagan run carries hundreds of rules and thousands of events, and one rsigma run does the same, so thousands of rules cost dozens of invocations.
Both take --profile, defaulting to rsigma-syslog:
lab/differential/engine_differential.py --rules <corpus> --profile vector-enriched
lab/differential/after_differential.py --rules <corpus> --profile vector-enrichedUnder vector-enriched the converter emits rules that plain syslog refuses,
754 of them on c6fddfd. The fields those rules match on do not come from a
rendering of the pipeline: they come from the pipeline. vector_pipeline.py
builds a stdin-to-console Vector configuration from pipeline_transforms(), the
same list and order sagan2sigma --emit-vector-config writes, and the syslog
line handed to Sagan is handed to that. Whatever comes out is what rsigma is
asked about, so the rules and the transforms are judged as the one deliverable
the documentation says they are. Requires vector on PATH; --events model
falls back to the probe generator's rendering, and --events pipeline refuses
to run without it.
Three of the optional transforms run here, and their data is local and fixed:
GeoIP from config/country.mmdb, the denylist from
config/rules/diff-blacklist.txt and Zeek intel from
config/rules/diff-zeek-intel.dat. The Sagan side is pointed at the same two
feed files, so both engines consult one list and a disagreement about a listed
address means something. Since a rule reading a denylist flag never says which
address has to be on the list, the probe is given the five the feed holds,
ahead of the rule's own literals, so the position it parses is listed whichever
one it declares.
Bluedot does not run and cannot: it is Quadrant's closed threat-intel service,
queried over the network, and the converted rule matches flags an operator's own
feeds produce instead. That is a substitution rather than a translation, which
the conversion states as D_BLUEDOT_SUBSTITUTION, so those 134 rules are
excluded as a declared divergence rather than counted as a missing field: even a
reachable service would not make the two sides comparable.
What remains under enrichment (<field>) is a field no running transform
produces. The list is derived from the profile and from which transforms are on,
so it is empty for rsigma-syslog and shrinks by itself when data is added.
alert_time needs a second flag and a second run. --at-time puts the engine
under faketime and stamps the pipeline's events with the same instant, which
is what lets the window arithmetic be compared even though the two sides read
different clocks by design (D_ALERT_TIME_EVENT_CLOCK). The corpus declares
two windows and no more, 0700-1800 on weekdays for 27 rules and 1800-0800
for 3, so two runs cover every one of them:
lab/differential/engine_differential.py --rules <corpus> --profile vector-enriched \
--at-time "2026-08-18 14:30:00" # a weekday afternoon
lab/differential/engine_differential.py --rules <corpus> --profile vector-enriched \
--at-time "2026-08-18 23:00:00" # the same weekday, at nightThe two windows are disjoint, and each run lists under silent the rules the
other one judges, so a reader can check that the pair covers the family rather
than take it on trust.
For the correlation differential the same pipeline is what makes a rule
grouping on parse_src_ip judgeable at all: its probe is given five addresses
to be parsed, one set per case so that a neighbour's event lands in a different
group on both sides, and each side then derives its own group key from the same
line. Measured against sagan-rules@a1cf3b3, 921 correlations are judged
under the enriched profile against 748 under syslog, with no disagreement in
either run, at N events or at N+1. The enriched run sets 34 cases aside as no trigger and 12 as over-count; the syslog run, 5 and 7. What the two profiles
part company over is named in the skip counters: 120 rules need the enriched
pipeline and 65 want an address normalize alone cannot resolve, which is the
whole of the gap.
Those figures replace 454 and 373, measured on an earlier corpus and, more to the point, before the group-by work. Nothing about the corpus accounts for the difference: it gained a handful of rules over the same period. The harness learned to drive correlations it used to set aside.
Both also carry a flag that reintroduces a defect the project has actually
shipped, and both must report it: --case-policy relaxed drops |cased and the
case-flipped probe has to disagree, --reintroduce-off-by-one restores
gte: N and every correlation has to fire a step early. A differential that
cannot fail is worth nothing, and this one was wrong three times before it was
right.
Read the counters, not just the verdict. no trigger means the generated event
never satisfied the rule, and over-count that another rule in the batch fed it
extra matching events: in both cases the two engines agree, but the boundary was
not tested, so those rules are covered by the run and not judged by it.
model_differential.py asks a different question from the other three, and the
only one whose answer is unambiguous. They compare the converted rule against
something; this compares the two readings of Sagan, the engine and
tests/differential/sagan_reference.py, on the same probes. A divergence
therefore cannot be an arbitration between two opinions: the engine is Sagan, so
the model is wrong. And the model is what CI runs on every push, which is the
point of running this in the lab: what it corrects strengthens the light net
without the heavy one ever leaving here.
On a1cf3b3: 4,308 rules, 28,792 probes, no divergence. It took one defect to
get there. Two paths of a document can clip to the same stored key, a nested one
and the object holding it, and the engine's table keeps both in the order it
walked them while src/json-content.c stops at the first key that matches. The
shallower entry decides and the deeper one is unreachable; the model kept the
deeper one, so five confluent.rules rules that are dead in the engine read as
alive. They are the rules upstream rewrote to the clipped form deliberately, and
sagan2sigma's own upstream detector already called them dead, so the converter
and the model had disagreed about them for as long as both existed with nothing
to arbitrate.
The other 23 divergences of that first run were defects in the tool rather than
in the model, and all three had been solved next door in engine_differential.py
before: attribution has to key on the program as well as the message, since
base and wrong_program differ only by it; every candidate has to be re-judged
alone, since Sagan's pass action silences the rules that follow it in a batch;
and a rule carrying json_map: "program" has its program replaced by the body's
value, logged that way, so those rules are attributed by message alone. Reading
the neighbour first would have saved all three.
xbits_differential.py asks the same boundary question of the other state
machine, one setter event then one tester, and runs each correlation twice: once
with the bit primed and once without, since a rule that fires either way has not
been judged at all. It needs no corpus-wide batching, the family being small.
On a1cf3b3: 14 cases, 10 judged, no disagreement in either state. The four it
does not judge are named individually, sids 5009793, 5003985, 5014047 and
5003390, each because the rule that sets the bit does not match its own probe.
That is a generated event the tool could not build, reported rather than counted
as agreement.
exercised says the same thing for the detection run: how many rules made
both sides fire on a probe that satisfies every positive condition, the base
one or, for a rule matching either shape, its plain-text twin. Two silent evaluators agree about
nothing, and on a JSON-bodied rule that is the easy failure, because the engine
searches the serialised document and a literal carrying quotes may be
impossible to place in one. A rule counted as silent was run and not decided.
A rule carrying json_map and no other JSON keyword is probed twice, once
as a document and once as a plain line, because the engine matches it either
way: measured here, program: sshd; json_map: "src_ip", ".ip"; content:"needle"
fires on a plain syslog line exactly as the same rule without the binding does.
41 corpus rules are in that state. The document arm is dropped under a profile
whose pipeline keeps no raw body, where the conversion covers the plain half and
says so (D_JSON_BODY_ARM_LOST); probing it against a document would measure
that profile's blind spot rather than the conversion. Before this, the probe's
shape was read off the rule, so a plain-text probe was handed to rsigma under
the JSON envelope names and a conversion that could match no plain line agreed
with a rendering no pipeline produces.
Both differentials re-judge every disagreement with its rule alone, and report
what did not survive as batch-only. A batch puts hundreds of rules and
thousands of events in front of both engines at once, so a neighbour's event
that also satisfies a rule's detection joins its counter, and a pass rule
ahead of it silences it outright. Either way the verdict belongs to the batch
rather than to the rule, and eleven extraHop rules and one AWS brute-force rule
were reported that way before the pass existed.
The after correlation path cannot be measured on a plain build. Two
defects sit in it, and each is enough to make a measurement wrong rather than
merely fail: one aborts the process on a hardened build, and the other lets a
rule that declares no correlation behave as though it did, alerting only from
the second event and silently costing a detection. Both are fixed by the two
local patches described above. Measure detection semantics against
bin/sagan-sane, which carries both, and never against bin/sagan-patched
alone: that build stops the crash while leaving a rule silently correlated.
The defects themselves are not described here on purpose. They are unfixed in a security product, and a repository that converts its rules is not the place to publish them.
Sagan's correlation state outlives the process. xbits, flexbits,
threshold and after counters live in /dev/shm/sagan-*.shared. A harness
that does not wipe them inherits the previous run's counters, and a rule already
over threshold alerts on the very first event. This produced a confidently wrong
conclusion before it was found: the results looked like memory corruption, and
patching the overflow did not fix them. harness.py wipes them on every run.
Fifty threads make the result non-deterministic. Sagan's default thread
count means the same input can yield different alerts run to run. The harness
forces -t 1 -b 1.
Attribution is by message text, so whitespace in a probe matters. sagan()
keys its result on the message Sagan logged, and the reader used to .strip()
that line. A probe beginning with a space then hashed to a different key than
the alert, so a rule that had fired read as "did not fire" — a false negative
with no error anywhere. It cost an event_id result that looked like a genuine
engine behaviour. The reader now removes only the single separator space Sagan
writes after Message:. When a check fails in a way that surprises you, dump
work/log/alert.log before believing it.
skip_networks defaults to skipping 8.8.8.8. The lab config sets it empty.
Left at the default, the GeoIP "country in the list" case silently becomes a
skip and the truth table reads backwards.
The bundled normalization.rulebase normalizes nothing here. Every rule in
it is written rule=: <pattern>, and the space after the colon is part of the
pattern, so liblognorm expects the message to begin with a space. Its own
documented sample message comes back as unparsed-data. The failure is silent:
normalization simply never fires, every address falls back to positional
parsing, and a precedence check written against it passes while measuring
nothing. check_normalize.py therefore supplies its own rulebase through
config_with_rulebase(). If you write a check involving normalize, assert
first that normalization resolves an address at all, and run the engine with
-d normalize when it does not: that prints liblognorm's parse of every
message and is the only quick way to see the difference between "no match" and
"no rule".
Five of those files are Sagan's own and are redistributed here unchanged, under
the same licence this project carries: classification.config,
reference.config, protocol.map, json-input.map and
normalization.rulebase. The engine refuses to start without them, so a lab
that did not carry them would not run at all.
The rest are fixtures written for this directory: the two lab-* and two
diff-* feeds below, lab-normalize.rulebase, and test.rules.
Four fixtures stand in for feeds that cannot be redistributed, and each is fixed so the expected answers are known without a live source. Every address in them comes from a documentation range, so nothing here is third-party threat intelligence.
config/rules/lab-blacklist.txt and config/rules/lab-zeek-intel.dat belong to
check_enrichment_feeds.py. The denylist one carries a /24 before a host
address on purpose: Sagan never resets its mask buffer between lines, so a
shorter prefix following a longer one is silently narrowed, and that quirk is
what the last section of that check pins.
config/rules/diff-blacklist.txt and config/rules/diff-zeek-intel.dat belong
to the corpus differential and are deliberately separate from the two above, so
that a change to one measurement cannot move another. The denylist holds a
network and one host outside it, in that order for the same reason; the Zeek
intel file holds five hosts, one per position a rule can declare.
Both sides read the same two files: Sagan as its feeds, and the Vector side
through tools/build_denylist_mmdb.py, the script the documentation tells an
operator to use. The differential therefore exercises the lookup the shipped
configuration declares rather than a stand-in for it.
config/country.mmdb is generated, tiny and fixed, so the expected results are
known without depending on a live feed:
| Network | Country |
|---|---|
| 5.5.5.0/24 | RU |
| 8.8.8.0/24 | US |
| 203.0.113.0/24 | FR |
| anything else | unknown |
$HOME_COUNTRY in config/sagan.yaml is US,CA, so 5.5.5.5 is outside the
home country, 8.8.8.8 is inside, and 198.51.100.7 cannot be placed at all.
It is generated rather than committed, by config/make-country-mmdb.py, which
needs mmdb-writer and netaddr:
python lab/config/make-country-mmdb.pyA binary blob in a repository is something a reader can neither check nor rebuild, so the script is both the recipe and the record of what the database contains. No third-party data is redistributed here: the three networks are documentation ranges and the two feeds below are fixtures written for this purpose.
Sagan's --file reads its traditional pipe format, one event per line:
host|facility|priority|level|tag|date|time|program|message
event() builds these. The parser is src/input-pipe.c; the field order is not
documented anywhere else.
The non-address half of normalize: it also sets the username, the ports and
the protocol, which feed track by_username and a rule's default ports, and
only the addresses are pinned here. Reproducing the rest would mean shipping
liblognorm rulebases per log format, which is data rather than behaviour, and
the converter says as much where it depends on it
(username-extraction.vrl is a starting point, not a port).
Everything else the checks once listed as missing is now covered:
check_event_id.py, check_json_ops.py and check_alert_time.py were written
after this section first said they were not. alert_time needed the engine to
be run under faketime, since it compares its window against the wall clock at
processing time rather than the event's own stamp; that divergence from the
converter's event-time model is recorded as D_ALERT_TIME_EVENT_CLOCK.