Convert Sagan rules into Sigma rules.
Sagan's rule corpus is large, actively maintained and covers a lot of ground that SigmaHQ does not, particularly network appliances and Unix daemons. The Sagan engine itself has seen little movement in years. This tool moves the rules to a format other engines can run, notably RSigma.
On the upstream corpus (10,025 active rules across 343 files) it converts
86.9% into 9,528 Sigma documents, with zero parse failures, zero documents
rejected by pySigma, and zero rules the RSigma engine refuses to load. With
--profile vector-enriched, which ships the transforms needed to recreate the
fields Sagan derived from raw text, the rate rises to 94.4%, and to 97.1%
once --sagan-yaml supplies the site's own variables, $HOME_COUNTRY among
them, and its GeoIP transform resolves the country of the address each
country_code rule names.
Everything it does not convert is reported with a stable code and the reasoning behind it, so the gap in your coverage is explicit rather than silent.
The upstream corpus, already converted, is committed so you can take the Sigma
rules without installing anything: converted/ is the default
rsigma-syslog profile, and converted-vector-enriched/
is the vector-enriched profile, which recovers more of the corpus but needs the
Vector pipeline it ships to run. A scheduled job keeps both in step with the
upstream corpus, reconverting the whole set on each change, and
converted/VERSIONS.md records which snapshot they were
built from, for both profiles, and how old it is.
Not published to PyPI yet, so install from source:
git clone https://github.com/NRGLine4Sec/sagan2sigma.git
cd sagan2sigma
pip install .Or in one line, without keeping the checkout:
pip install "git+https://github.com/NRGLine4Sec/sagan2sigma.git"Either way you get a sagan2sigma command on your PATH. Python 3.10 or newer.
Once the project is released, pip install sagan2sigma will work as well; see
RELEASING.md.
git clone --depth 1 https://github.com/quadrantsec/sagan-rules.git
sagan2sigma sagan-rules --output convertedYou get three things:
| Path | What it is |
|---|---|
converted/rules/*.yml |
Sigma rules, one file per Sagan source file |
converted/CONVERSION-REPORT.md |
every refusal and every semantic loss, grouped by product family |
converted/conversion-report.json |
the same data, untruncated, for CI |
converted/vector/ |
with vector-enriched, a runnable Vector pipeline carrying the VRL transforms those rules depend on |
Useful flags:
# recover the correlations Sagan grouped on src_ip or username, and get a
# runnable Vector pipeline that recreates those fields
sagan2sigma sagan-rules -o converted --profile vector-enriched
# target a pipeline where Vector parses syslog into JSON, without enrichment
sagan2sigma sagan-rules -o converted --profile vector-json
# resolve $USERS and friends from your own configuration
sagan2sigma sagan-rules -o converted --sagan-yaml /etc/sagan/sagan.yaml
# trade exact case fidelity for recall
sagan2sigma sagan-rules -o converted --case-policy relaxed
# fail the build if conversion regresses
sagan2sigma sagan-rules -o converted --min-rate 80 --fail-on-validationsagan2sigma --help lists the rest.
A companion command answers a question a migration always raises: which of the converted rules already have an equivalent in SigmaHQ, so you can deploy SigmaHQ for those and keep the converted rules only where they add coverage. It answers it by running, not by comparing text: every rule from both sets is turned into events that satisfy it, and the RSigma engine decides which rules each event fires. When it reports that SigmaHQ covers a converted rule, a test event that fires both is attached.
pip install "sagan2sigma[overlap]" # pulls in the two extra dependencies
sagan2sigma-overlap \
--converted converted/rules \
--sigmahq /path/to/sigmahq \
--output overlap --cache .overlap-cacheIt needs the rsigma binary on your PATH, and writes OVERLAP-REPORT.md (the
actionable list) and overlap-report.json (every verdict, with its witness
event). The method, the taxonomy and the results are in
docs/SIGMAHQ-OVERLAP.md.
A second, separate command answers a softer question the behavioural one cannot: which converted rules look like they detect the same thing as a SigmaHQ rule, even when they can never fire the same event because one matches raw text and the other a structured field. It compares the distinctive terms rules search for and their ATT&CK techniques, and produces review candidates, never verdicts:
sagan2sigma-conceptual \
--converted converted/rules \
--sigmahq /path/to/sigmahq \
--output conceptualIt needs no engine and no extra dependency. It is a triage aid, not grounds for
retiring a rule, and the two analyses are almost disjoint by design; see
docs/CONCEPTUAL-OVERLAP.md.
The converter refuses rather than approximates. A missing rule is recoverable; a rule that looks right and matches the wrong thing is not. It will not convert:
- effective positional matching, a non-zero
offset,depthordistance, which pins a pattern to a byte position Sigma string modifiers cannot express. A zero-valued positional is a no-op in the Sagan engine and is converted. - Bluedot threat-intelligence lookups, which query a closed Quadrant
service. GeoIP
country_code,blacklistdenylists andzeek-intelfeeds are no longer refused outright: under--profile vector-enrichedthey convert to a match on a field the bundled transforms derive from a database or public feed (DShield, CriticalPathSecurity), refused only when that profile is not in use. - negative correlations (
xbits isnotset), which Sigma cannot express. - group-by keys that only liblognorm produced, since its rulebases are
per-format data files with no algorithm to reproduce. The regex-extracted
ones are recovered by
--profile vector-enriched, which takes this category from 313 rules down to 14.
Each of these carries a stable code in the report, with the reasoning attached.
This is a 0.6.0 release and the rules it emits are marked status: experimental for a reason.
What is verified. Every emitted document is parsed by pySigma, the reference implementation, so the output is valid Sigma rather than YAML that resembles it. The full upstream corpus converts with zero parse failures and zero rejected documents, and the conversion is deterministic: two runs are byte-identical.
Beyond shape, behaviour is checked too. A differential harness runs every corpus rule it can judge, 4,303 of them, through two independent evaluators: a reference implementation of Sagan semantics written from the engine C source, and the real rsigma engine evaluating the converted rule. Tens of thousands of event evaluations, no disagreements. This is what caught the field-naming defect that silently broke a quarter of the corpus before release.
That harness has a limit it states itself: the Sagan side is a reference
implementation, so a misreading it shares with the converter survives both. It
did. after: count N alerted an event early across 970 correlations while the
harness reported perfect agreement, because the same belief shaped the model and
the code.
So behind it sits a second line of checking, which runs a locally built Sagan
instead of modelling one. It judges the rules no model can decide too, pcre
and effective positional constructs among them. Measured against sagan-rules
at a1cf3b3 under the enriched profile: 9,331 rules judged, no
disagreement, and 9,314 of them made both engines fire on a probe satisfying
every positive condition, which is what says the agreement was not two silences.
The seventeen that did not are undecidable rather than unmeasured, and the run
names each: twelve carry byte windows a single probe cannot satisfy at once,
four a literal no serialised document can hold, and one an event_id the engine
cannot resolve from a document.
It separately walks every after correlation it can drive to its threshold,
checking that each stays silent at N events and alerts at N+1: 921 of them under
the same profile, again with no disagreement, and 10 xbits state machines
judged both primed and unprimed. All three were shown able to fail before being
believed: with the old gte: N reinstated, the correlation check flags every
rule.
That exercise corrected a dozen behaviours the source reading had missed, most
of them making the converted rules noisier than the originals. It also found
limits in the engine that no converter can reproduce, which are recorded in
docs/DESIGN-DECISIONS.md rather than imitated, and defects in the upstream
rules themselves, eighteen of which are now fixed upstream.
The instrument that produced all of it is in lab/, documented in
docs/LAB.md. CI never runs it: it needs a compiled Sagan and
the better part of an hour. It is here so that the claims it produced can be
re-measured by someone else rather than taken on trust.
The bundled VRL transforms are executed against a real Vector binary in CI, and
their address extraction is checked case by case against the branches of
Sagan's own Parse_IP().
What is not verified. The differential harness covers detection semantics
only. Correlation rules, pcre and effective (non-zero) positional constructs
are outside what the reference evaluator can judge, and are skipped rather than
approximated. No test replays real production traffic. Treat the first deployment as a tuning
exercise, not a migration, and read the conversion report before trusting any
of it.
If you run the output through another Sigma engine, or against traffic the harness does not model, reports of divergence are the most useful contribution this project can receive.
What is out of scope. This project does not re-check that the upstream Sagan
rules are valid Sagan: quadrantsec/sagan-rules already validates them in its own
CI, so the corpus is trusted as input and the effort goes into converting it
faithfully. See docs/DESIGN-DECISIONS.md under "We do not re-validate the
upstream Sagan rules".
The rule corpus this tool reads is the work of
Quadrant Information Security and the
contributors to sagan-rules,
maintained continuously for well over a decade. This project only translates
it; the detection engineering is theirs.
docs/SIGMAHQ-OVERLAP.mdto see which converted rules SigmaHQ already covers, and how that is established by testingdocs/CONCEPTUAL-OVERLAP.mdfor the separate, lexical review-candidate analysis that covers the rules testing cannot reachdocs/OVERLAP-INVENTORY.mdfor the merged, confidence-tiered list of overlapping rules, pinned to a commit of each corpus (a point-in-time snapshot; regenerate withsagan2sigma-inventory)converted/for the pre-converted rules and how they are kept currentdocs/PIPELINE.mdto get the output running under RSigmadocs/MAPPING.mdfor the keyword-by-keyword mappingdocs/DESIGN-DECISIONS.mdfor the traps this converter avoids, and whydocs/ARCHITECTURE.mdto work on the codedocs/LAB.mdto measure the engine yourself, which is how every claim about Sagan's behaviour in this repository was establishedCONTRIBUTING.mdto add a keyword handler
GPL-2.0-only, matching quadrantsec/sagan-rules.
The rules this tool produces are derivative works of the Sagan corpus and inherit its licence. If you redistribute converted rules, they are GPL-2.0 too. Running them in your own SOC is unaffected.