Nobody Wanted to Be the One to Call It
Key Intel / TL;DR
  • Every incident response plan begins at declaration, and nothing in it covers the hours before somebody declares.
  • The delay is rarely ignorance. It is competent people waiting for more certainty because being wrong in public is costly.
  • Name one person per shift who can declare alone, and make declaring cheap enough that being wrong costs them nothing.
  • Write the threshold as observable conditions, so the call does not depend on somebody feeling confident.
  • Track time from first signal to declaration and review it, because it is the only number in the whole process nobody measures.

The first hour of a serious incident usually contains four people who each privately believe something is wrong, and no incident. Somebody in the network team has noticed traffic that does not fit. Somebody on the help desk has taken a third call about the same odd behavior. An analyst has an alert they cannot dismiss and cannot explain. A manager has heard two of those things secondhand and is waiting for a third.

Every one of them is competent. Every one of them is doing what their role trained them to do, which is gather a bit more before escalating. And the clock that matters is running the entire time.

The Plan Starts After the Hard Part

Pull out your incident response plan and read the first page. Ours covers the planning steps in detail, and like every other one I have read, it opens at the moment an incident has been declared. Roles activate, the bridge opens, and notifications go out on a schedule. Everything downstream of that first sentence is well specified, rehearsed in tabletops, and generally executed competently by people who know their jobs.

The declaration itself gets a line. Sometimes it gets a title, usually something like the incident commander or the duty manager, and no description of how that person decides or what they are supposed to do when they are not sure.

So the plan is a detailed set of instructions for a machine that somebody has to switch on, and the switch is undocumented. I have reviewed a lot of these and the pattern is consistent enough that I now read the plan backward, starting at the declaration criteria, because if that section is thin the rest of the document describes a response that will begin ninety minutes late.

Why Competent People Wait

This is not an intelligence problem or a training problem, and treating it as either one is why the usual fixes do not work. More training on the escalation path does nothing to the thing that is actually stopping people.

Declaring an incident is a public act with asymmetric consequences. If you declare and you are right, you did your job and the outcome is the outcome. If you declare and you are wrong, you woke up executives at two in the morning, pulled a dozen people off delivery work, possibly notified a customer, and everyone now has a story about the time you panicked. The cost of being wrong is concentrated on one person and it is social, immediate, and remembered.

So people do the rational thing under that incentive, which is gather more evidence. Each individual decision to wait fifteen minutes for one more data point is defensible on its own. Four of them in sequence is an hour, and an hour is the difference between containing something at one host and containing it at forty.

There is a physiological layer under this too. The people making these calls are frequently tired, frequently mid-task, and operating with the specific kind of uncertainty that pushes a stressed brain toward inaction, because doing nothing feels reversible and doing something does not. Under load, the option that requires no justification wins. Your staff are running standard equipment under load, and any process that depends on somebody overriding that equipment at three in the morning is badly designed.

Add the converged dimension, which is where this gets worse. The physical security officer who sees a door forced at an odd hour, the network analyst who sees the beacon, and the fraud team who sees the transaction pattern are frequently three different reporting lines with three different escalation paths, and none of them is looking at the other two signals. Each one individually falls below anybody’s threshold. Together they are obvious, and nobody is standing where they can be seen together.

Make the Call Cheap

The fix is to change what declaring costs the person who does it. A better threshold does nothing while the price of using it stays where it is.

If a declaration wakes twelve people and cancels a release, nobody will make one on partial information, and partial information is all anybody ever has in the first hour. So separate the act of declaring from the size of the response it triggers.

A declaration should start something small: a bridge opens, three named people join, someone starts a timeline document, and the investigation gets an owner. That costs one hour of three people’s time. If it amounts to nothing, it cost one hour of three people’s time, and the person who called it did exactly the right thing. Escalation beyond that tier is a second decision, made with more information, by more people.

Once being wrong costs an hour instead of a reputation, people declare early, and early is the only thing in this entire process that reliably reduces damage. Everything else you can buy operates on whatever is left after that first hour.

Five Things to Fix Before the Next One

Name the person, per shift

Not a role and not a team. A named individual for every hour of the week, including the ones your organization pretends do not exist, who has explicit authority to declare alone and without consulting anybody. If two people have to agree, you have built a system where each can wait for the other, which is the failure this whole article is about.

Write the threshold as conditions, not as judgment

“Use your judgment” transfers the risk onto whoever is on shift at three in the morning. Observable conditions do not. Write down the specific things that mean declare now regardless of how confident you feel, such as credentials confirmed in use from an unexpected location, any encryption of a shared file store, any authentication anomaly on an administrative account, or two unrelated teams reporting anomalies inside the same hour. That last one is your converged tripwire and almost nobody has it.

Publish that declaring is free

Say out loud, from someone senior, that a declaration that finds nothing is still a correct decision and will be treated that way. Then the first time it happens, do exactly that, visibly. Everyone is watching what occurs the first time somebody is wrong, and that single event sets the behavior for the next two years. This is the same mechanism as blameless reporting, applied to the one decision that gates everything else.

Give it a tiered first response

Define the smallest possible response tier and make it genuinely small. Three people, one hour, one timeline document, one owner. Nobody needs approval to trigger it and nothing about it is disruptive enough to make somebody hesitate.

Measure time from first signal to declaration

Every organization measures time to containment and time to recovery. Almost nobody measures the gap between the first thing somebody noticed and the moment it became an incident, because it is uncomfortable to reconstruct and it is usually the largest number in the timeline. Pull it from your last three incidents. If that number surprises you, it is because it has never been anybody’s metric.

What This Looks Like When It Works

An analyst sees something at 2:40 in the morning, cannot explain it, and declares at 2:44 without calling anybody first. Three people join a bridge. By 3:30 they have either found the thing or established that it was a misconfigured backup job, and in the second case the analyst gets told they made the right call and everyone goes back to bed.

The organizations that recover well are the ones where the person who noticed first did not have to be brave about it. Tooling and plan length matter far less than that one condition.

If you want an outside read on your declaration criteria and who actually holds the authority at three in the morning, contact Grab The Axe. We will ask to see the plan, and then we will ask what happened the last time somebody was wrong. You can also take our free Human Attack Surface Score.

Jeff Welch is CEO of Grab The Axe.

Distribute Intel
Jeff Welch
Chief Executive Officer
Jeff Welch
Architect of the 'Cognitive Firewall.'

A PhD candidate in Health Psychology and former Corrections Officer, Jeff founded GTA to dismantle passive security models. He focuses on the 'Human Zero-Day', mitigating executive burnout and decision fatigue before they become security breaches.

View Author Page →