https://projects.knownelement.com/issues/825
6.3 KiB
Wazuh Alert Tuning — KNEL SOC
Date: 2026-09-06 · Ticket: Redmine #825 · Change: GLPI CR 23
Applied on: tsys-siem (Wazuh manager 4.14.1), /var/ossec/etc/rules/local_rules.xml
Discourse: https://community.turnsys.com/t/soc-security-operations-lane-wazuh-tuning-823-follow-through-soc-team/340
Rollback: cp /var/ossec/etc/rules/local_rules.xml.bak-20260906-soc /var/ossec/etc/rules/local_rules.xml && systemctl restart wazuh-manager
Why
The manager emails every alert of level >= 7 to coo@turnsys.com (max 12/hour). September 2026 through 09-06: 664k alerts this month (518k from tsys-cloudron-new web 400s alone, ingested via journald = every Cloudron app's container stdout), and a permanent L7+ email storm: scanner-storm composites (31151 @ L10), CVE backlog refires (23505 High @ L10, 23506 Critical @ L13, 23504 Medium @ L7 — the whole backlog re-fires every scan), SCA failures (19007 @ L7), ssh/PAM fail amplifiers (2502 @ L10, 5551 @ L10), dpkg events (2902/2904 @ L7), generic syscheck churn (550 @ L7).
Design rule going forward: email (L7+) is reserved for genuine incidents. Everything else lives in dashboards and the SOC weekly loop (triage notes on #825, OpenVAS/GLPI tickets for vulnerabilities, hardening backlog for SCA).
Mechanism note (learned the hard way)
Wazuh 4.14 rule overwrite="yes" in local_rules.xml REPLACES the original rule
definition — it is not a merge. An override with only a level would match every
event (observed: rule 2502 flagging kernel i2c lines at 1000/min). Therefore every
override below carries the FULL original rule body with only level/frequency/timeframe
changed. Also: nested <group> elements are invalid in rule files — single root group
only. Level-0 rules do not index but still count as "fired" for frequency composites
(verified: 31151 keeps firing over the L0 31101 parent).
Changes
| Rule | Meaning | Was | Now | Rationale |
|---|---|---|---|---|
| 31101 | web 400 per-line | L5 | L0 | 492k/mo bot noise; raw logs stay in journald |
| 31120 / 31122 | web 500 per-line | L5 | L3 | real app errors, dashboard-only |
| 31151 | web 400 storm composite | L10 freq 14/90s | L6 freq 40/240s | sustained scans only, no email |
| 5501 / 5502 | PAM session open/close | L3 | L0 | session accounting |
| 23502 / 23507 | CVE solved / cleared | L3 | L0 | lifecycle info |
| 23503 | CVE low | L5 | L3 | |
| 23504 | CVE medium/untriaged | L7 | L4 | was emailing; backlog refire |
| 23505 | CVE high | L10 | L6 | vulns = weekly OpenVAS/GLPI loop, not email |
| 23506 | CVE critical | L13 | L6 | same — backlog refire storm made email useless |
| 19007 | SCA check failed | L7 | L5 | hardening backlog |
| 2902 / 2904 | dpkg installed / half-configured | L7 | L3 | ops inventory, not incidents |
| 550 | syscheck checksum changed (generic) | L7 | L5 | churn; high-signal subcases below |
| 2502 | "missed password >once" | L10 | L5 | dup of 5503/5760 |
| 5551 | PAM multiple failed logins | L10 | L6 | 2.3k/mo on routine churn |
| 31510 | CMS brute force | L8 freq 8/30s | L6 freq 20/120s | bot storms |
| 40704 | systemd unit failed | L5 | L3 | 145k/mo churn; root-cause dig tracked on #825 |
Kept at L7+ deliberately: 5712 (sshd brute, L10), 5763 (sshd max-auth), 31152/31153 (401/403 storms, L10), 31162 (suspicious URI, L10), plus all auth failures at L5 (5710/5760/5503) — cheap and real now that SSH is tailnet-only (#769).
New detections (the #823 pattern)
The July 2026 ezeda.org compromise (Redmine #823) persisted via wp-config.php backdoor + 41 webshells in uploads. New rules + Cloudron agent syscheck watchers catch a repeat:
| Rule | Level | Fires when |
|---|---|---|
| 100500 | L12 (email) | any wp-config.php under Cloudron app data changes (syscheck) |
| 100501 | L10 (email) | a .php file appears/changes inside any WP uploads/ dir (syscheck) |
| 100510 | L8 | web request reaches .php inside an uploads/ path (access-log) |
Agent 001 (tsys-cloudron-new) ossec.conf additions: syscheck
/home/yellowtent/appsdata/*/data/wp-config.php + */data/wp-content/uploads
(frequency 12h); quarantine dirs excluded (#823 evidence holds known malware).
Verification (2026-09-06, clean window after deploy)
| Metric | Before (per hour) | After (per hour) |
|---|---|---|
| Total alerts | ~19,000+ | ~2,500 |
| L5 web-400 flood | ~25,000 | 0 |
| 2502 misfire flood | ~60,000 at peak | 0 (7 per 10 min = genuine) |
| L7+ (email to coo@) | constant (12/h cap hit) | 0 |
| L10 scanner storms (31151) | continuous | L6, sustained scans only |
Remaining L3 volume = systemd churn (40704) + apparmor denials (52002) + web 500s + sshd successes: visible for triage, below every notification path.
Week-loop plan (adjustment cadence)
- Daily: triage new L6+ in the Wazuh dashboard; note dispositions on #825.
- Watch 5710/5760/5503 — should stay near-zero with tailnet-only SSH; a spike = someone knocking from inside the tailnet.
- 100500/100501/100510 must stay at zero; any fire = possible #823 recurrence → incident.
- Investigate 40704 root cause on tsys-siem (which unit churns) as an ops item.
- Re-check email volume at coo@ weekly; tune further only on evidence.
- Vulnerability backlog: OpenVAS scans (GLPI-fed targets) → tickets per the triage standard; Wazuh vuln alerts feed the same weekly review.
Sweep additions
2026-09-06 (CR 24): rule 100520 (L3) — UCS machine-account SSH probe suppression:
tsys-ucs-02$ → tsys-ucs-01 every 20 min is a UCS domain artifact (verified in
auth.log), not an intrusion. Narrow match: if_sid 5710 + srcip 100.68.10.17 + pcre2
Invalid user tsys-ucs-02\$. All other invalid-user attempts stay L5. Gotchas that
bit during apply: srcuser is not a valid rule element in 4.14 (use regex on the log),
OSMatch $ anchors and does not escape reliably — use <regex type="pcre2">.
Known remaining churn (L3, dashboard-only, tracked): 40704 systemd failures
dominated by nut-driver@tripplite.service (pfv-tsys5) and dell-fan-control.service
(pfv-tsys9) arriving via remote syslog — infra-ops crash-loops, not security.