Putting Detection Coverage to the Test: A Windows Security Case Study
A case study of 144 Sigma analytics using Windows Security log sources shows why ATT&CK mappings alone do not reveal the depth or quality of …
By Antonia Feffer • September 29, 2026
ATT&CK heatmaps are a useful way to see where detections are mapped, but a green box can only tell us so much. It doesn’t necessarily tell us how much of a technique we can detect, how easily an adversary could evade those detections, or how well the underlying signals distinguish malicious from benign activity. That challenge motivated our work on the Coverage Calculator, which looks beneath the ATT&CK mapping to evaluate two complementary dimensions:
Implementation Coverage - which behaviorally distinct ways of performing a technique can be detected
Detection Quality - based on the robustness and precision of the signals providing that coverage.
We explain the methodology and the work behind the Coverage Calculator in Beyond the Heatmap: Measuring What Detection Coverage Really Means. But building a methodology is one thing. We also wanted to see what it could tell us when applied to real detections.
For that, we turned to Sigma, an open-source, community-driven detection format that has become an important part of the detection engineering ecosystem. Its repository provides a large body of portable detection content contributed and maintained by practitioners across the security community, making it a useful corpus for this kind of analysis.
Now to clarify, our intent is not to grade Sigma as a project. Open-source content inherently reflects different contributors, environments, assumptions, telemetry sources, ATT&CK versions, and points in time. Instead, we use the repo to explore a more focused question:
How accurately do ATT&CK mappings communicate the actual detection coverage provided by analytics?
To explore that question, we evaluated 144 Sigma analytics using Windows Security log sources. Windows Security events are particularly interesting because they are widely used in enterprise detection engineering and have applicability across a substantial portion of ATT&CK. But they also come with an important constraint: Windows Security auditing was designed primarily to support security and identity auditing, not to provide a complete behavioral record of everything occurring on a Windows system. As a result, the telemetry can be extremely useful while still lacking the execution context needed to distinguish among different implementations of some ATT&CK techniques, and sometimes the detection logic isn’t the limiting factor - the sensor is.
The 144 analytics we evaluated carried mappings representing 84 ATT&CK techniques. At first glance, that sounds like substantial breadth. The Coverage Calculator, however, told a different story.
Based on our findings, only 26 of those 84 techniques had at least one implementation - or method of executing a given technique - for which the associated analytics provided meaningful detection coverage. Across the set of implementations represented by those techniques, the analytics collectively provided approximately 26.8% implementation coverage. Most strikingly, 64 of the 84 mapped techniques had no implementation-level coverage at all under the calculator’s behavioral criteria.
That does not mean the rules are useless or that Windows Security telemetry lacks defensive value. A better conclusion is that an ATT&CK tag attached to an analytic and a behavioral detection of an ATT&CK implementation are not necessarily the same thing and shouldn’t be treated as such. That gap is exactly what the Coverage Calculator is designed to expose.
The gap between mapped ATT&CK coverage and measured implementation coverage did not have a single explanation. Looking across the zero- and low-coverage results, we found several recurring patterns, some related to ATT&CK mappings, some to the analytics themselves, and others to the limits of the underlying telemetry. Importantly, these findings do not suggest that an analytic receiving a low implementation-coverage score is necessarily a bad analytic – a narrowly targeted rule may be perfectly useful for its intended purpose. The issue arises when that narrow detection is interpreted as evidence of broader ATT&CK coverage than its logic and telemetry can actually support.
Some of the clearest cases were simply mapping problems. As ATT&CK evolves, techniques change, sub-techniques are introduced, and definitions are refined. An analytic written against an earlier version can remain operationally useful while its ATT&CK metadata slowly drifts out of alignment, which leads to some of the misalignment that we observed. We also encountered cases where the analytic appeared to detect behavior more consistent with a different technique than the one attached to the rule. In either situation, the detection itself may still work exactly as intended, but the statement being made about what that detection covers is no longer valid. That makes ATT&CK mapping maintenance more than documentation housekeeping. If mappings are used to calculate defensive coverage, stale or inaccurate metadata eventually becomes stale or inaccurate coverage beliefs.
Another recurring pattern involved analytics built around highly specific artifacts, like a filename, executable, service name, command string, or other indicator associated with a tool known to participate in a particular attack behavior. Those detections can be valuable, but the problem comes when detecting the identity of a tool is treated as equivalent to detecting the behavior represented by the ATT&CK technique. An analytic that says “I saw Tool X” may catch a known implementation very effectively, but if the adversary can rename the tool, alter a command, substitute another utility, or accomplish the same objective through a different implementation, the ATT&CK technique has not disappeared but the observable has.
This is the Pyramid of Pain showing up in coverage measurement. Atomic and attacker-controlled indicators can provide useful detections, but they often represent a narrow slice of the behavioral space. Implementation coverage helps make that limitation visible.
Closely related were mappings based on inferred attacker intent. The reasoning may look something like: “An adversary using this utility could use it to perform Technique X, therefore detecting the utility provides coverage of Technique X.” That is a reasonable lead for an analyst to start with, however, it is a much weaker basis for claiming detection coverage. The distinction is between observing evidence that the behavior is occurring and observing something that could be used to perform the behavior. The Coverage Calculator deliberately sets a higher bar for the former.
A useful question when reviewing an ATT&CK mapping becomes “what in the detection logic demonstrates the mapped behavior”? If answering that question requires several sentences beginning with “Well, the attacker might…”, then the mapping may be describing possibility rather than observable coverage.
The opposite problem also appeared. Instead of being specific, some analytics relied on telemetry that was too generic to establish meaningful behavioral coverage. An Event ID, for example, shows that a class of system activity occurred, but if the analytic doesn’t do more than just identify that event without examining fields that distinguish the malicious behavior of interest, the event itself may not provide enough context to demonstrate an ATT&CK implementation.
This is an important distinction because collecting telemetry is not the same as detecting behavior. An organization can have excellent visibility into an event source and still need additional analytic logic to turn that visibility into a useful detection. The Coverage Calculator considers whether an analytic references the right sensor and what observable fields and conditions it uses.
These patterns point towards the broader issue that incentivizing ATT&CK coverage can work against meaningful coverage measurement. An analytic may plausibly support several ATT&CK techniques. A tool it identifies could be used for Technique A, might facilitate Technique B, and has been observed during Technique C. Tagging the analytic with all three creates an impressive-looking coverage map. Each additional mapping turns another check box green. But the number of green boxes has increased without the detection logic changing at all.
That is what we mean by checkbox chasing: treating the breadth of ATT&CK mappings as the objective rather than asking how much observable defensive capability sits behind each mapping. The result can be contradictory, as a detection program appears to gain coverage as more ATT&CK tags are added, even though it has gained no new telemetry, no new analytic logic, and no ability to detect an additional adversary behavior. Implementation coverage changes the incentive; instead of asking “How many techniques can we reasonably associate with this analytic?”, it asks “What behavior does this analytic observe, and which implementations does that evidence allow us to detect?” That may produce fewer green boxes, but the boxes that remain green mean considerably more.
None of this means ATT&CK mappings are unimportant, it’s the opposite. Mappings become more valuable when we expect them to describe a defensible relationship between detection logic and adversary behavior. The gap we observed was not exclusively a Sigma problem, a mapping problem, an analytic problem, or a telemetry problem – it was often a relationship problem between those components. A technique tag tells us where an analytic says it belongs, but meaningful coverage requires us to follow the evidence underneath it.
Like any automated analysis, the Coverage Calculator depends on good inputs. An analytic must contain a valid ATT&CK mapping because the calculator evaluates detection logic against the implementation catalog for that technique; it does not infer missing mappings. Likewise, robustness and precision scoring depend on recognizable telemetry field names that can be resolved against the project’s sensor mappings and scoring dictionary. These boundaries are intentional: the Coverage Calculator evaluates the relationship between an analytic’s stated ATT&CK mapping, detection logic, and telemetry, rather than guessing when those relationships are missing or ambiguous.
For transparency and reproducibility, the detailed spreadsheet results from this Windows Security case study are provided alongside this analysis, including technique-level implementation coverage and analytic-level robustness and precision scores.
Use coverage analysis to prioritize engineering work. Measuring coverage is to identify where we are blind, where perceived coverage is shallow, which analytics depend on fragile signals, where telemetry improvements will unlock better detection. Visit our Detection Coverage Calculator to improve your detection!
Maintain ATT&CK mappings like detection logic. ATT&CK metadata should be reviewed as techniques evolve. A stale mapping can create the appearance of coverage long after the relationship between the analytic and the technique has changed.
Favor durable behavior over disposable artifacts. Specific indicators still have value, particularly for threat hunting and rapid response. But analytics intended to provide durable ATT&CK coverage should move toward system interactions and behavioral signals that adversaries cannot change cheaply.
Know when telemetry is the bottleneck. If the fields required to distinguish an implementation simply are not present, rewriting the query for the tenth time will not solve the problem. Sometimes the appropriate engineering decision is to add or change telemetry.
Measure depth and quality, not just breadth. Ten green ATT&CK boxes supported by narrow or fragile detections may provide less defensive value than five techniques covered across several meaningful implementations with robust, precise signals.
© 2026 The MITRE Corporation. Approved for Public Release. ALL RIGHTS RESERVED. Document number
26-0334.
A case study of 144 Sigma analytics using Windows Security log sources shows why ATT&CK mappings alone do not reveal the depth or quality of …
Summiting the Pyramid introduces implementation coverage and detection quality to help defenders measure the depth and effectiveness of detection …
Attack Flow v4 helps defenders turn incident evidence into connected flows, add the context needed for action, and communicate decisions across …