Purple teaming: turning red-team findings into detections
The most uncomfortable line in a red-team debrief is "and your SOC didn't see any of it." It's also the most valuable — if you do something with it. Purple teaming is how you turn "you missed this" into "we now detect this," deliberately and quickly.
Buying detection isn't detecting
Organisations invest heavily in EDR, SIEM and SOAR and reasonably assume they're covered. But tooling deployed is not the same as techniques detected. The only way to know what you actually catch is to run real attacker techniques and watch what fires. That's the purple-team exercise: red and blue in the same room, running techniques on purpose and measuring the response.
How it runs
- Plan. Agree the techniques to test — mapped to MITRE ATT&CK — and what success looks like.
- Execute together. Run each technique in a controlled way while the defenders watch their tooling in real time.
- Measure. Record what was prevented, what alerted, and what passed silently. Map coverage to the ATT&CK matrix so the gaps are visible.
- Tune and re-run. Build or improve detections for the gaps, then run the technique again to prove the detection works.
Why it's efficient
A pure red team measures your defences at a point in time. Purple teaming improves them during the engagement. Instead of a report that says "we evaded detection," you leave with a coverage matrix, a set of new and tuned detection rules, and proof that the gaps you cared about are now closed. It's the fastest route we know from a red-team finding to a measurable blue-team improvement.
Detections decay — re-run them
A detection that fired perfectly last quarter may be silently broken today. A tooling upgrade changed a log format; a noisy rule got disabled during an incident and never re-enabled; a cloud migration moved the telemetry somewhere the SIEM no longer ingests. None of this announces itself. That’s why the highest-value form of purple teaming isn’t a one-off — it’s a regular cadence that re-validates your most important detections still work against the techniques you care about. Detection coverage is a moving target, and the only way to know where it is today is to test it today.
From exercise to engineering discipline
The lasting value of purple teaming isn’t the report — it’s the habit. Detection engineering stops being a reaction to the last red team and becomes a repeatable practice: every exercise leaves behind tested rules, a clearer picture of coverage, and a team that knows how to build the next detection themselves. The exercise is the seed; the discipline is the point.
Mapping coverage to ATT&CK, and reading the matrix honestly
A purple-team exercise generates a natural map: each technique the red team executes is a cell in the MITRE ATT&CK matrix, and each one either fired an alert, produced logs you could have alerted on, or passed in silence. Laying those outcomes over the matrix gives you a coverage heat map that executives love and that quietly lies if you let it.
The dishonesty creeps in through three habits. First, counting a technique as "covered" because one of its procedures was detected, when the technique has a dozen variations and you tested a single, convenient one. Second, conflating log availability with detection: having the telemetry is not the same as having a rule that fires on it. Third, scoring by technique count rather than by risk, so a wall of green hides the fact that the handful of techniques most likely to appear in a real intrusion against your estate are the amber ones.
Read the matrix with those traps in mind. Grade each cell on a scale that distinguishes no telemetry, telemetry but no detection, detection with gaps and robust detection, and weight the picture towards the techniques relevant to your sector and architecture rather than the whole encyclopaedia.
Putting it into practice
Record coverage at the procedure level, not just the technique level, and annotate each green cell with the exact variation you tested. A cell you tested one way is a hypothesis about the others, not a result.
Detection-as-code and atomic, technique-level testing
Findings that live in a spreadsheet decay the moment the exercise ends. Findings that live in version control become durable engineering. Treat each detection as code: a rule definition, the logic that backs it, a reference to the technique it addresses, and a test that proves it still fires. When a rule changes, it moves through review, a pull request and a pipeline, exactly as application code does.
Pair that with atomic testing. Rather than only re-running the full red-team scenario, break each technique into the smallest action that should trigger a detection and fire it in isolation, on a schedule. Atomic tests catch the silent regressions that end-to-end exercises miss, and they turn "did we fix it?" from an argument into a passing or failing check.
- Store rule logic, enrichment and test cases together, so a change to any one is visible alongside the others.
- Tag every rule with its ATT&CK technique and the exercise that produced it, so coverage reporting generates itself.
- Run atomic tests continuously, not once a quarter, and alert on a detection that stops firing as loudly as on one that never existed.
Finding telemetry gaps across EDR, SIEM and cloud or identity logs
Most silent failures are not rule failures; they are visibility failures. The endpoint agent sees process execution and never sees the identity provider issuing a token. The SIEM holds months of firewall data and none of the API audit logs from the cloud control plane. An attacker moving from a stolen session to a cloud role to a mailbox rule crosses three telemetry domains, and a gap at any boundary is a corridor they can walk unobserved.
Hunt the gaps deliberately. For each technique, ask which log source should witness it and confirm that source is actually ingested, parsed and retained — not merely available in principle. The boundaries between EDR, network, SIEM and cloud or identity telemetry are where coverage maps are most often optimistic.
Putting it into practice
Build a source-to-technique inventory before the next exercise: for every technique in scope, name the log that evidences it and the system that collects it. An empty cell there is a blind spot you found on paper rather than in an incident.
Tuning without going blind, and the metrics that prove it
Every noisy rule is a candidate for suppression, and every suppression is a small act of self-blinding. The discipline is to cut false positives by making logic more precise — narrowing on context, enriching with asset or identity data, excluding known-good behaviour by exact characteristic — rather than by broad exclusions that an attacker can hide inside. A filter that drops a whole process name is a gift to anyone who can masquerade as it.
Judge the result with metrics that resist gaming. Mean time to detect tells you whether tuning is making you faster or merely quieter. Detection coverage, graded honestly as above, tells you how much of the relevant matrix genuinely fires. Alert fidelity — the share of alerts that turn out to be true positives — tells you whether analysts can still trust the queue. Watch them together: fidelity that climbs while coverage falls means you tuned by going blind.
Finally, fix the cadence. Purple teaming earns its value only as a repeated rhythm — a standing schedule where new techniques are added, prior findings are re-tested atomically, and the coverage map is re-graded against what has changed in the estate. One exercise is a snapshot; a cadence is a capability.
What does your tooling actually catch?
A purple-team exercise measures your detection coverage against real techniques — and closes the gaps while we're there.
Plan a purple-team exercise