Applying Security Brutalism, A Playbook for Leadership and Practitioners
For Security and IT Leadership
The goal of a security program is survivability, not the elimination of risk. Eliminating risk is not achievable, and treating it as the goal produces programs that spend heavily on low-consequence problems while critical systems stay exposed. Survivability means reducing how badly a failure hurts and reducing how long the organization stays in that failed state, and three questions should drive every decision that follows from it. Does this reduce susceptibility, the realistic attack paths into systems that carry real consequence? Does it limit damage, the blast radius once compromise happens? Does it reduce recovery time, how fast the team detects, contains, and restores? A control, tool, policy, or process that cannot answer yes to at least one of these, backed by evidence, is consuming resources and adding complexity without improving the organization's security posture, and complexity brings a real cost of its own. Every additional tool, integration, and process that cannot justify its contribution to survivability adds attack surface, consumes staff time, and introduces new ways for the system to fail, so removing what does not meet that bar is itself a security improvement.
Now, none of this can be evaluated without a consequence map, the ranked list of systems ordered by what failure actually costs. Building that map is covered in Building the Consequence Map, and turning it into an active exercise is covered in The Second Job. Everything below assumes that ranking already exists, and treats the systems at the top of it as the ones every question that follows is really aimed at.
Evaluating Your Current Program
A consequence map turns into a real evaluation once you use it to test the program you already have. Most programs get built around compliance frameworks, vendor capabilities, or historical precedent rather than around consequence, so this test often produces a different picture than an audit would. For each system on the map, ask your team whether they can describe the realistic attack paths that lead to its compromise. An answer they cannot give points to a visibility gap that needs fixing before anything else here is worth addressing.
The next question is what a realistic attacker can do once inside a consequential system, including what data becomes reachable, what actions become possible, and where the attacker can move from that position. Without clear answers, the blast radius stays undefined, and undefined usually means larger than assumed. How long detection takes is just as important, counting from the event itself to the moment a person is looking at the right data and understands what happened, not from the moment an alert theoretically fires. A simple answer measured in weeks means detection is not working yet.
Restoration time closes out the list. Ask how long it takes to fully restore a consequential system from a known-good state, and whether that number came from a test run in the last ninety days under realistic conditions rather than a runbook estimate. A plan that has never run under pressure tends to fail in ways nobody predicted, once an actual incident forces the question. A program can score well on a compliance audit and carry a large tooling budget while still failing all four of these tests, because the score and the budget describe what was purchased, not what the purchase accomplished.
And speaking of investing, when a team proposes a new tool, control, or initiative, apply the same three questions from above, asked specifically against the systems highest on the consequence map. A proposal that cannot answer yes to at least one, regardless of how it gets framed, should not move forward on security grounds, though compliance requirements can still justify it under a separate evaluation. The cost side deserves the same scrutiny since every new tool becomes new attack surface, a new integration that can be misconfigured, and a new dependency running through the supply chain. Security tooling itself has served as the entry point in several major supply chain compromises, so the burden of proof for adding anything should stay high, and removing tools that cannot meet that burden counts as a security decision in its own right.
A team working this way should produce a specific set of things without months of preparation. A ranked list of the organization's most consequential systems, each paired with the top two or three realistic attack paths against it, comes first. Current detection coverage follows, showing which of those paths would produce a signal today and which would not. Recovery status for each consequential system is just as important, including the last tested restoration and the measured time to restore, alongside a list of standing access grants and long-lived credentials touching those systems and when each was last reviewed. A team unable to produce these on short notice is running a program organized around something other than survivability, and reorienting it becomes the first task.
Compliance and Security
Compliance requirements are real, and meeting them is not optional. Regulatory exposure, insurance requirements, and contractual obligations are legitimate constraints on any organization. The problem shows up when compliance work gets treated as equivalent to security work, since the two solve different problems. Compliance satisfies external requirements from auditors, frameworks, regulators, and customers, while security reduces how badly a failure hurts and how long the organization stays in that state. These two goals overlap sometimes, and that overlap is efficient when it happens, but when they diverge they need separate funding and separate management.
A useful test for any control asks whether it reduces how bad is the damage or how long recovery takes, or whether it exists mainly to satisfy an auditor. Both are legitimate reasons to fund something. The mistake is funding one under the name of the other, or letting compliance work absorb resources that survivability needs.
Which brings us to the metrics we are often asked to show. Most reported metrics describe the existence of controls rather than their effectiveness. Tool coverage percentages, vulnerability counts, and compliance scores describe what got purchased and deployed, not whether any of it works under pressure. A better set follows from the same three questions used throughout this playbook. Time to detect a high-consequence compromise, measured from the event to the moment a person understands what happened, shows whether detection actually functions, while time to contain and time to restore, measured from real incidents and test exercises rather than runbook estimates, show whether recovery capability is real rather than assumed. Blast radius per system, tested, reveals the true scope a compromise can reach, and alert signal quality, the share of alerts tied to real activity worth attention, reveals whether detection produces something useful or just noise. Honeytoken activation, canary credentials placed where only an attacker would find them, gives near-certain evidence of an active intrusion the moment one trips. Producing these numbers takes testing and operational discipline that most programs skip, which is exactly why they tell you more than the metrics already sitting on most dashboards.
A program run this way looks different from one built around audits and dashboards, since every decision traces back to the same three questions instead of a growing list of separate requirements. That single thread is what makes the whole approach usable at the leadership level, not just legible to the security team applying it.