---
title: "Your Decisions Are the Bottleneck, Not Your AI"
summary: "AI can do security work in seconds, but the work still waits in queues and SLAs built for human speed. Compare each SLA with how fast AI can do the same step, automate the steps that need no judgment, and set clear rules for when containment blocks and when it goes to a person."
canonical: "https://cisoexpert.com/blog/your-decisions-are-the-bottleneck"
publishedAt: 2026-09-27
---

# Your Decisions Are the Bottleneck, Not Your AI

# Your Decisions Are the Bottleneck, Not Your AI

**Example scenario:** An executive asks for a security exception. It's a simple request, and the risk review takes a few minutes once someone opens it. Getting someone to open it takes the rest of the day. By then you have an upset VIP and a ticket nobody wants to own.

The same week, a detection fires on a production server: a process is doing something it shouldn't. The SOAR (Security Orchestration, Automation and Response) tooling has already pulled the host details, the login history and the vulnerability data, and the block is one click away. It sits in a queue because blocking a process on a production system needs approval, and approval runs on human time. If the wait runs long enough (minutes to hours), a suspicious process becomes a security incident.

I see this pattern across security, from incident response to GRC. **AI has cut the time to do the work to seconds, but our SLAs, approval gates and queues were built for human speed. Until we redesign when and how we apply human judgment, the operating model is the constraint, not the technology.**

My measurement is simple. Take every SLA your teams already work to, and next to it write how fast AI or automation can do the same step. The gap between those two numbers is where your program loses time, and in my experience most of that gap is spent waiting, not thinking.

## SLAs built for slow work

Most security SLAs were set when the work itself took hours. Investigate the alert, write the ticket, schedule the change, run the containment. A queue that added a few hours was a rounding error next to work that took a day.

That arithmetic has flipped. Here's how it looks for three SLAs almost every team has, two easy and one hard:

| Step                                | What the SLA measures                                           | What AI or automation can do now                                                  | Where the time actually goes                      |
| ----------------------------------- | --------------------------------------------------------------- | --------------------------------------------------------------------------------- | ------------------------------------------------- |
| **Assign an incoming alert**        | Alert to a named owner                                          | Tag and route to the on-call list the moment the alert lands                      | Waiting for someone to pick it up                 |
| **Ticket a critical vulnerability** | Detection to a ticket that can still meet the patching deadline | Match the finding to the asset, owner and policy deadline, then open the ticket   | Triage queues, then arguments over who owns it    |
| **Contain a threat**                | Detection to block or isolation                                 | Pull host, identity and exposure context from several systems and stage the block | Waiting for approval, and gathering facts by hand |

In most environments I've seen, the first two don't need a human in the loop for the routine case. The third sometimes does. The real work is deciding when, and where you draw those lines depends on your business.

The industry numbers show the same thing. IBM's [2025 Cost of a Data Breach Report](https://www.ibm.com/reports/data-breach) put the global average time to identify and contain a breach at 241 days, the lowest in nine years. Organizations using security AI and automation extensively shortened that lifecycle by about 80 days and cut breach costs by about $1.9M compared with those that didn't. My read: the tools help, but they only move the number as far as the operating model around them lets them.

## Easy case: ticket routing

Time to assign is the cleanest example. Whatever number you picked, whether 15 minutes for criticals or 24 hours for lows, the SLA exists because it was sized to how fast a person could read the alert and decide who gets it. That decision is a classification on a handful of factors: severity, the asset, and which team is on call.

Auto-tag the alert to the on-call list and the assignment SLA shrinks to the time it takes the alert to arrive. The analyst's time goes into the investigation, where their judgment counts.

If your assignment metric still depends on a person reading a queue, the SLA is measuring staffing, not security.

The catch is that "easy" doesn't mean static. A routing rule written once is right on the day you write it and wrong soon after. Teams reorganize, on-call rotations change, and a system that was a low-priority test box last quarter now runs month-end close. Good routing reads the business context when the alert arrives, not when the rule was written: who owns the asset today, how critical it is this week, who is actually on call. Automating the routing is the easy part. Keeping it current as the business changes is the work.

## Middle case: vulnerability ticketing

Ticketing a critical vulnerability looks like judgment, but for the routine case it usually isn't. If your patching policy already sets the deadline, the judgment was made when the policy was written. The only question is whether the ticket gets created, routed and dated in time for the fix to meet it.

Published deadlines show how little slack is left. [PCI DSS v4.0.1](https://www.pcisecuritystandards.org/document_library/) requirement 6.3.3 requires critical security patches to be installed within one month of release. CISA's [BOD 26-04](https://www.cisa.gov/news-events/directives/bod-26-04-prioritizing-security-updates-based-risk), issued in June 2026 for US federal civilian agencies, replaced the flat 15- and 30-day windows of the now-[revoked BOD 19-02](https://www.cisa.gov/news-events/directives/bod-19-02-vulnerability-remediation-requirements-internet-accessible-systems-revoked) and the known-exploited deadlines of BOD 22-01. In their place are risk tiers: 3 days, 14 days, 60 days, or fix at the next upgrade.

Two things about that directive matter even if you're not a federal agency. First, it sorts vulnerabilities by facts about the asset and the exploit: public exposure, known exploitation, whether exploitation can be automated, and technical impact. No person reads each finding. Second, CISA warns that threat actors' use of AI "may further narrow the time defenders have to react between patch release and possible exploitation."

When the deadline is three days, a ticket that spends two of them in a triage queue has already failed. Let automation do the matching and ticketing. The human handles the exception: the asset that can't be patched in time and needs a compensating control.

Setting that automation up once is easy. Keeping it right is not. The matching depends on knowing who owns each asset, and ownership drifts: people leave, teams merge, servers move between business units. When it drifts, tickets land with someone who no longer owns the system, bounce back, and burn the days the deadline gave you. Treat the ownership data as part of the automation, not something next to it. Every ticket that comes back as "not mine" is a signal that a mapping is wrong, and it should fix the mapping, not just re-route the ticket. Assets with no clear owner should go to a named default owner and get flagged for review, not sit in a queue.

## Hard case: containment

Containment is where "just automate it" goes wrong. Blocking a process on a production server can stop an attack. It can also take down the system that pays everyone's salary.

Before I let a block run automatically, I ask two questions. Is there a blocking policy that covers this action? Do we have the facts? For a process block on a server, these are some of the facts I want. It's an illustration, not a complete list:

| Fact                                       | Why it matters to the block decision                                  |
| ------------------------------------------ | --------------------------------------------------------------------- |
| Where the system is hosted                 | Cloud, on-premises or third party changes who can act and what breaks |
| What runs on it                            | A build server and a payments database are not the same risk          |
| What we know about the host                | Criticality, owner, recent changes, prior alerts                      |
| Who has logged in                          | A known admin at 2 PM reads differently from a new account at 2 AM    |
| What that identity can reach               | How far an attacker could get if the identity is compromised          |
| Whether they've run this command before    | The line between routine and new                                      |
| Whether the host has a known vulnerability | A plausible way in raises confidence that the activity is malicious   |
| Whether it's externally facing             | Exposure shortens the time you can afford to wait                     |
| What we know about the user                | Role, recent access changes, prior incidents                          |

These facts live in different places: the asset inventory, the endpoint tooling, the identity provider, the vulnerability scanner, HR systems. Pulling them by hand means logging into five consoles and hoping they agree. An agent with read access can gather them in seconds. This is the SLA table's gap at its widest, and it's why containment is where getting the operating model right pays off most.

## Block or escalate?

A caveat before the rule: most programs can't answer every question in that table today. The asset inventory is stale, nobody baselined which commands are normal, the identity data doesn't say what an account can actually reach. That's normal, and it isn't a reason to wait. It decides which way the rule leans. Where the facts are missing, the decision goes to a person. Every gap you close moves a class of decisions from "ask someone" to "act".

Here's where I land. **The block runs automatically when the asset is externally facing, the activity is new and the host is vulnerable.** With that combination, waiting is costly and the evidence is strong. Stop it, notify the owner, log everything, and have a human review it afterwards.

The decision goes to a person when we're blind or the business impact is high. If the tooling can't tell you what runs on the host or who owns it, you don't know what the block will break. A confident automated action based on missing facts is how a security control turns into an outage.

The hard middle is activity that's simply new. It could be an attacker, or it could be a legitimate job nobody has run before. This is the old problem with anomaly and outlier detection: there are always outliers, depending on how small you draw the circle.

"First time this user ran this command" is a fact, not a verdict. On its own, it should lower confidence and send the decision to a person. Combined with an externally facing asset and a vulnerable host, it should tip the decision to block.

In an orchestration layer, that logic can be this plain:

```
facts = gather(hosting, runtime, host_intel, logins, identity_access,
               command_history, vulns, exposure, user_context)

if not policy.covers(action):
    alert(human, reason="no blocking policy")
elif facts.missing("runtime", "owner") or impact.business_critical:
    alert(on_call_senior, sla="15m", context=facts)
elif exposure.external and activity.first_seen and vulns.present:
    block(); notify(owner); log(facts, "auto-blocked")
else:
    alert(analyst, sla="1h", context=facts)
```

That's not production code, and the SLA values are examples, not recommendations. What matters is that the person who gets the alert gets it with the facts already gathered. Their fifteen minutes go on judgment, not on collecting data or initial triage work.

## Sort by risk and reversibility

Containment is one type of decision. A security program has dozens, and in the environments I've seen, most of them share one approval path. Blocking a known-bad IP from a threat feed waits in the same queue as isolating a production server.

Two questions sort most of them: how bad is it if this is wrong, and how easily can you undo it? Where each action lands is a judgment call for your environment. The grid below is how I'd sort them, not a standard.

| | **Reversible** | **Hard to reverse** |
|---|---|---|
| **Low risk** | Run it; review afterwards. Threat-feed blocks, alert enrichment, auto-assignment. | Run it with logging; audit on a schedule. Sinkholing a known-bad domain, quarantining a known-bad file hash. |
| **High risk** | Human approval on a short SLA, facts pre-gathered. Endpoint isolation with rollback, firewall rules that expire. | Human approval, always. Disabling a privileged account, segmenting a production network, anything that destroys evidence. |

The goal isn't fewer humans. It's to stop putting a reversible, low-risk action through the same wait as an irreversible one.

## Scoped overrides with an expiry

Nothing executes without approval: that's the right default to start with. It's the wrong place to stay. If you never move off it, your AI tooling is permanently limited by who happens to be on shift.

The way forward is scoped overrides. When a category of decisions is well understood, open the gate for that category only. Give each override a scope (what it covers), an expiry (when it goes back to needing approval) and an audit trail (what happened while it was open).

No permanent bypasses and no toggles that stay on. Every extension of autonomous authority is deliberate, bounded and reversible, the same principle as break-glass access. You extend trust in steps rather than giving up control.

An override doesn't have to be all or nothing. It can open on a trigger. A process on a critical server goes to a person by default, but if a second high-severity alert fires on the same host, the evidence has changed and the block runs. Or the approval has a clock: if nobody responds inside the SLA, the action runs rather than waiting. That last one is a decision you should make on purpose, ahead of time. When nobody answers, does the system hold or act? Either answer can be right for your business. What I see most often is that teams have never chosen, which means they've chosen to hold, and the attacker gets the wait.

## Log every automated decision

Some rules for logging AI decisions already exist. The EU AI Act requires high-risk AI systems to [record events automatically](https://artificialintelligenceact.eu/article/12/) and to be designed for [effective human oversight](https://artificialintelligenceact.eu/article/14/). If your AI tooling is classed as high-risk, that means a high level of monitoring: every decision logged, and a person able to see and step in. Whether it is depends on how and where you use it, and that's a question for your legal team, not a blog post.

Even if yours isn't, I wouldn't wait to be told. The same [IBM 2025 report](https://www.ibm.com/reports/data-breach) found that 63% of breached organizations either had no AI governance policy or were still developing one. The AI Act articles already show what an auditor will ask for: what the system recorded, and where a human could step in.

For every automated action, log what facts the agent had, which policy it applied, what it did or recommended, and whether a human approved, overrode or reviewed it afterwards. That record is also how you earn the next scoped override. You can't argue for widening autonomy without evidence of how the current scope behaved.

## From making calls to designing them

For twenty years, the security leader who thrived was the one who made the right call under pressure. That skill still matters, but it doesn't scale to decisions arriving at machine speed.

The leader who thrives now designs the system that makes the calls correctly. That means the SLA that measures the right thing, the fact checklist behind a block, the rule for when a human steps in, and the log that proves it worked. You still make decisions, and they're bigger ones: decisions about how decisions get made.

That design isn't a project you finish. Every case in this piece has the same failure mode: the rules are right when you write them and drift after. Routing goes stale when the org chart changes. Ticketing breaks when ownership moves. The block rule gets looser or tighter than you meant as your data gets better or worse. So run your decision rules the way you already run a firewall rulebase: an owner for each rule, a scheduled review, and a record of why it exists.

It also changes what you measure. Time to assign and time to contain still matter, but they don't tell you whether the design is working. These do:

- **Wait versus work.** For each SLA, how much of the elapsed time was a queue and how much was someone thinking. The goal is to shrink the queue, not the thinking.
- **Human reversals.** How often a person overrides or undoes an automated action. Near zero on a new scope may mean nobody is looking. Rising means the rule or the data has drifted.
- **Decisions sent to people for missing facts.** Each one points to a gap in your data. That count is your data roadmap.
- **Override age.** How many scoped overrides are past their expiry and still open. That's the number that tells you whether "bounded" is true.

None of this needs new tooling to start. It needs someone accountable for how decisions are made, not just for making them. In most programs I've seen, nobody holds that job today.

I'm confident about the direction and less certain where the lines should sit for your environment. That depends on your policies, your tooling and how well your data answers the questions in the checklist. If you've drawn them differently, tell me. I'd rather be corrected than confident.

## Your next moves

**Good (this week):** List your team's existing SLAs, including assignment, vulnerability ticketing, containment and exception approval. Next to each, write how fast automation or AI could do the same step, and mark which part of today's time is waiting and which is judgment.

**Better (this quarter):** Take the two easy cases off human time. Auto-assign alerts to the on-call list, and auto-ticket critical vulnerabilities with the owner and the deadline from your patching policy. Track your own before and after. There's no universal target here: the right number is the one that works for your business, and your own data makes the case for the next step.

**Best (this half):** Write the blocking policy for one containment action. Map each fact on the checklist to the system that holds it, and put the block-versus-alert rule in your orchestration layer, starting as a scoped override with an expiry. Log every decision it makes, and review the log before you widen the scope.

---

*David O'Neil is a CISO and builder who designs operating models for AI-augmented security teams. He writes about the gap between what AI can do and what organizations are ready to let it do at cisoexpert.com.*
