How Detection Explorer normalizes "where the data comes from" across 8 detection rule repositories with completely different schemas.
This is a technical reference for engineers who need to:
- Understand how a rule's
platforms/data_sources/event_typesvalues get computed from its source YAML/TOML. - Add a new rule repository to the system (Issue 6 work).
- Edit a mapping when the canonical vocabulary changes or a vendor schema updates.
- Add a new canonical platform / data source / event type.
Status: Phase 1 of 3 (Issue 2 from
TODO.md). The new fields coexist with the legacy ones (log_sources,platform,event_category,data_source_normalized) until Phase 3 cuts the legacy fields out entirely.
These principles were formalized after the first review pass against
real sigma/lolrmm data. They apply to event_types specifically, but
the "no inference" rule applies to the other two dimensions too.
If a rule's logsource is a broad channel (e.g. windows/security,
okta, m365), we do NOT try to guess which event type(s) the rule
might be looking at. Windows Security log contains authentication
events (4624, 4625), process events (4688), and audit events (4728,
4732) all intermixed — a rule could be filtering for any of them. We
tag the rule with a single coarse event_type (audit_event) rather
than listing all plausible sub-types.
Don't do this:
windows/security:
event_types: [authentication, audit_event] # inference — wrongDo this:
windows/security:
event_types: [audit_event] # coarse but accurateThe future work is a per-EventID dictionary that refines the
classification by parsing the rule's EventID selection. That's a
separate undertaking (see TODO items) and not something we fudge with
heuristics in the meantime.
When Sigma gives us a specific category, we preserve it 1:1 as its
own canonical event_type. We do NOT collapse related categories into
a coarser parent.
Don't do this:
windows/file_delete:
event_types: [file_event] # lossy generalizationDo this:
windows/file_delete:
event_types: [file_delete] # preserve the specific kindDetection engineers filter by these specific activities — a rule about
LSASS access (process_access) is meaningfully different from one
about process creation. Collapsing them makes the taxonomy less useful
for real analysts.
Okta System Log, Microsoft 365 Unified Audit, GitHub Audit Log, GCP Audit, Entra ID Audit, and AWS CloudTrail are all API-call event streams at the architectural level. Per Okta's own docs, every System Log entry represents an API call (reference). Same is true for the others — they're all REST-API event records.
We classify all of them as api_call for consistency. Tagging some
as api_call (AWS) and others as audit_event (Okta) would be
inconsistent — the feeds are architecturally identical.
Sign-in logs are the exception — Entra ID signinlogs and Azure AD
sign-in events are exclusively authentication activity, so they get
authentication instead of the generic api_call.
Every detection rule answers three independent questions about its
telemetry source. We model each as a separate multi-value field on the
Detection row.
| Field | Question | Examples |
|---|---|---|
platforms |
WHERE does the telemetry live? | windows, aws, okta, email |
data_sources |
WHAT product/integration produces it? | sysmon, aws_cloudtrail, crowdstrike_fdr |
event_types |
WHICH activity is being detected? | process_creation, network_connection, authentication |
All three are list[str] so a rule can span multiple sources. Elastic's
cross-platform Node.js rule, for example, lists six index patterns and
naturally resolves to:
platforms = ["windows", "linux", "macos"]
data_sources = ["elastic_defend", "sysmon", "windows_security_event_log",
"crowdstrike_fdr", "sentinelone", "auditd"]
event_types = ["process_creation"]When the vendor data doesn't supply enough info to determine a value,
the list contains ["unknown"] — never silently empty. Users can
filter for "rules with unknown platform" to spot vendor coverage gaps.
backend/app/services/taxonomy/
├── __init__.py # Public API: resolve_for_repo, PLATFORMS, ...
├── canonical.py # SINGLE SOURCE OF TRUTH for valid values
├── resolver.py # Dispatcher: routes a parsed rule to its vendor
├── _loader.py # YAML loading + canonical validation
├── vendors/ # Per-vendor resolver functions
│ ├── sigma.py
│ ├── elastic.py
│ ├── elastic_hunting.py
│ ├── elastic_protections.py
│ ├── splunk.py
│ ├── sublime.py
│ ├── lolrmm.py
│ └── sentinel.py
└── mappings/ # Per-vendor YAML mapping rules (USER-EDITABLE)
├── sigma.yaml
├── elastic.yaml
├── elastic_hunting.yaml
├── elastic_protections.yaml
├── splunk.yaml
├── sublime.yaml
├── lolrmm.yaml
└── sentinel.yaml
Separation of concerns:
canonical.pydefines what values are legal. Three frozensets:PLATFORMS,DATA_SOURCES,EVENT_TYPES. If you want to introduce a new value (e.g. addproofpointas a platform), edit this file first.vendors/<name>.pyimplements the logic — how to walk the vendor-specific parsed rule and look up entries in the YAML.mappings/<name>.yamlis the data — the actual mapping rules for that vendor. Reviewable by humans, easy to edit.
The flow for a single rule:
parsed rule (vendor schema)
│
▼
resolve_for_repo("sigma", parsed) ← public entry point
│
▼
vendors/sigma.resolve(parsed) ← vendor-specific logic
│
├─ Reads parsed.log_source.product / service / category
│
▼
load_mapping("sigma") ← reads mappings/sigma.yaml
│
▼
Look up keys most-specific to least:
"windows/powershell/process_creation"
"windows/powershell" ← match!
"windows/process_creation"
"windows"
│
▼
Return canonical sets:
{platforms: ["windows"],
data_sources: ["sysmon", "windows_security_event_log"],
event_types: ["process_creation"]}
The resolver dispatcher (resolver.py) does final defensive validation:
ensures three lists, applies [UNKNOWN] fallback if any list is empty,
sorts and deduplicates.
Each vendor has its own YAML structure tailored to its rule schema. The structures are documented inline in each YAML file. Quick reference:
by_key:
windows/process_creation: # most specific match wins
platforms: [windows]
data_sources: [sysmon, windows_security_event_log]
event_types: [process_creation]
windows: # broader fallback
platforms: [windows]
data_sources: [windows_security_event_log]The resolver tries the most specific key first
(product/service/category) and falls back to less specific ones if no
match. First match wins; missing dimensions can be filled in from a
broader fallback.
index_patterns:
"logs-aws.cloudtrail*": # wildcard suffix matched by prefix
platforms: [aws]
data_sources: [aws_cloudtrail]
event_types: [api_call]
"logs-endpoint.events.process-*":
platforms: [windows, linux, macos] # multi-platform
data_sources: [elastic_defend]
event_types: [process_creation]
integrations: # fallback when no index matches
okta:
platforms: [okta]
data_sources: [okta_system_log]A single rule can list many indices — the resolver unions all matching
mappings. So a rule with index: ["logs-aws.cloudtrail*", "logs-azure.signinlogs*"]
gets platforms: [aws, azure_ad] and data_sources: [aws_cloudtrail, entra_id_signin].
os_to_platforms:
windows: windows
linux: linux
eql_category_to_event_types:
process: process_creation # `process where ...` → process_creation
network: network_connection
file: file_event
always_includes: # every rule includes these
data_sources: [elastic_defend]Elastic Protections rules are agent-resident — there's no index. The
data source is always elastic_defend. Platforms come from os_list,
event type comes from parsing the EQL query head.
data_source_labels:
"sysmon eventid 10": # exact match preferred
platforms: [windows]
data_sources: [sysmon]
event_types: [process_creation]
"sysmon": # falls through to substring match
platforms: [windows]
data_sources: [sysmon]
event_types: [process_creation]Splunk's data_source field carries free-form labels like
"Sysmon EventID 10" or "ASL AWS CloudTrail". The resolver does
exact match first, then substring match.
connectors:
awssecurityhub:
platforms: [aws]
data_sources: [aws_security_hub]
event_types: [audit_event]
data_types: # cross-cutting overrides
securityalert: # SecurityAlert table from any connector
data_sources: [siem_alert]
event_types: [audit_event]Sentinel rules carry requiredDataConnectors with connectorId +
dataTypes. We map by both. data_types overrides apply on top of
connectors matches.
always_includes:
platforms: [email]
data_sources: [email_message_metadata]
event_types: [email_message]Every Sublime rule is email — no per-rule extraction needed.
If you're adding a new vendor (e.g. Google SecOps, Panther) in Issue 6, follow this checklist:
Sample 5-10 rules from the new repo. Identify which canonical
platforms, data_sources, and event_types they map to. If
the repo introduces a brand-new platform (e.g. Google SecOps content
about Chrome Enterprise), add it to canonical.py first.
In backend/app/services/taxonomy/canonical.py, add the new value(s)
to PLATFORMS, DATA_SOURCES, or EVENT_TYPES with a comment
explaining what the value covers.
Create backend/app/services/taxonomy/mappings/<repo_name>.yaml. Use
whichever structure matches the vendor's rule schema:
- Keyed by some categorical field → use a
by_keymap (Sigma/LOLRMM style) - Index/path pattern matching → use
index_patterns(Elastic style) - Connector/integration ID → use a
connectorsmap (Sentinel style) - Free-form labels → use
data_source_labelswith substring matching (Splunk style) - Implicit / always the same → use
always_includes
When in doubt, copy the closest existing vendor YAML and adapt.
Create backend/app/services/taxonomy/vendors/<repo_name>.py. Pattern:
from typing import TYPE_CHECKING
from app.services.taxonomy._loader import load_mapping
if TYPE_CHECKING:
from app.parsers.base import ParsedRule
_MAPPING = load_mapping("<repo_name>") # loads <repo_name>.yaml
def resolve(parsed: "ParsedRule") -> dict:
"""Resolve canonical taxonomy values for a parsed <vendor> rule."""
platforms: set[str] = set()
data_sources: set[str] = set()
event_types: set[str] = set()
# ... walk parsed rule fields, look up entries in _MAPPING,
# ... union into the three sets
return {
"platforms": platforms,
"data_sources": data_sources,
"event_types": event_types,
}The function MUST return three lists/sets. The resolver dispatcher
applies [UNKNOWN] fallback if any are empty, so you don't need to
handle that yourself.
Add an import to vendors/__init__.py and an entry to
_VENDOR_RESOLVERS in resolver.py.
In backend/tests/test_services/test_taxonomy.py, add at least:
- One happy-path test for a representative rule.
- One test for the "unknown product" / empty parsed-rule case.
cd backend && python -m pytest tests/test_services/test_taxonomy.py -v
If the YAML loader logs Mapping <repo>.yaml references non-canonical ... warnings on import, fix the typos in the YAML (or add the missing
canonical value).
Trigger a manual sync so the new resolver populates the taxonomy columns for the new repo's rules:
POST /api/scheduler/trigger {"repository": "<repo_name>"}
The mapping YAML files are designed to be edited by humans.
backend/app/services/taxonomy/_loader.py validates references at
import time — if you accidentally write data_sources: [crowdstrik_fdr]
(typo), the worker will log a WARNING at startup but won't crash.
Workflow:
- Open the relevant
mappings/<repo>.yaml. - Find the entry to change. Add/remove canonical values from the
platforms/data_sources/event_typeslists. - If you're using a value that doesn't exist yet, add it to
canonical.pyfirst. - Run the test suite locally to catch obvious problems.
- Commit, push. The worker auto-redeploys; the next sync picks up the new mappings.
Tip: during Phase 1, the new fields coexist with the legacy ones. To compare what the new resolver would produce vs. what the legacy code produced, query the database for both — both columns are populated.
If a vendor introduces a new product or platform that doesn't fit any existing canonical value:
- Add the value to
canonical.pywith a short comment. - Add a mapping entry in the relevant
mappings/<vendor>.yaml. - (Future, after Phase 2) Add a display name + color in
frontend/src/constants/taxonomy.ts. - Re-run a sync to backfill the column for affected rules.
If the value is genuinely orthogonal (a real new dimension, not a new value within an existing dimension), reach out for design discussion before adding a 4th field — the 3-field model is intentional and adding a 4th one is a significant change.
attack_typesfor Sublime rules (phishing, BEC, etc.) — that's a threat-classification dimension, not a telemetry-source one. It belongs in a separate field (or intags). Out of scope for Issue 2.- MITRE ATT&CK tactics/techniques — already a separate first-class dimension on the Detection model. Not part of taxonomy.
- Severity, status, rule type — already separate fields.
- Detection methods (e.g. Sublime's "Header analysis", "NLU") — too vendor-specific to normalize meaningfully.
- Phase 1 (Issue 2 first ship): Add new fields. Both legacy and new fields are populated by every ingest. Nothing user-visible changes yet.
- Phase 2: Update API and frontend FilterPanel to read the new fields. Legacy fields are still in the DB but unused.
- Phase 3: Drop legacy fields entirely. Rename
taxonomy_*columns toplatforms/data_sources/event_types.
The phasing is to allow incremental rollout and testing without ever breaking the live site.