A Practical Framework for Prioritizing Engineering Incidents Under Pressure

A field guide for SRE incident execution in high-pressure environments
Ricardo Acosta
Grow from Software Engineer to Engineering Manager with practical, actionable guidance
Get in touch

When people think about major production incidents, they usually picture engineers frantically debugging systems, checking logs, and executing commands.

While all of that certainly happens, I discovered after becoming an Engineering Manager that my role during an incident is very different.

I'm not usually the person fixing the technical problem.

I'm the person making sure the people who can fix it have everything they need to succeed.

That wasn't immediately obvious to me when I first became a manager. Coming from an engineering background, my instinct was to jump into the technical details. Over time, after participating in dozens of incidents across a global cloud environment, I realized that the biggest contribution I could make wasn't writing commands—it was creating order while everyone else focused on solving the technical problem.

One particular incident reinforced that lesson more than any framework or management book ever could.

The page nobody likes to receive

My phone buzzed.

The message simply read:

"Join the war room."

There was no incident summary.

No customer name.

No explanation.

Just a Zoom link.

Even today, after participating in many escalations, I still feel that rush of adrenaline every time I receive a page like that.

You never know what you're walking into.

Sometimes it's a false alarm that is resolved before you even join the call.

Sometimes it's a critical outage affecting one of your biggest customers.

No amount of experience completely removes that feeling of uncertainty.

And honestly, that's one of the things I enjoy about this job. Every incident is different, and no framework can completely prepare you for the unknown.

I clicked the link.

Around forty people were already on the call.

As usual, I introduced myself, explained which team I represented, and immediately started listening to the incident host as they summarized the situation.

At the exact same time, I opened my team's Slack channel and typed:

"I just got pulled into a war room. Has anyone received critical tickets in the last thirty minutes?"

Within seconds, I was receiving information from two completely different directions.

The war room was describing the customer impact, the affected region, and the estimated blast radius.

Meanwhile, one of my engineers replied that several servers had suddenly become unreachable and that he was already investigating.

Without realizing it, I had already started building situational awareness.

Understanding the real impact

The first thing I did was ask that engineer to join the war room.

He had information the rest of us didn't.

Once he joined, we compared what he was seeing with the list of affected servers that had just been shared during the incident.

That's when the atmosphere in the call changed.

There were sixteen unreachable servers.

On paper, sixteen servers doesn't necessarily sound catastrophic.

In our environment, it represented something much larger.

Those servers supported roughly one hundred database clusters for one of our largest customers.

At that moment, we didn't know whether those clusters had successfully failed over or whether some of them were now operating without redundancy.

We knew performance degradation was likely.

We also knew that if additional failures occurred before redundancy was restored, the impact could become much more severe.

Nobody panicked.

But everyone understood the seriousness of the situation.

Following the evidence

My engineer's investigation suggested that the affected hosts were connected to only a couple of network switches.

That immediately narrowed our search.

Rather than continuing to investigate every possible component, we focused on the infrastructure those servers had in common.

We searched our change management records.

About thirty minutes before the incident started, a network configuration change had been executed.

At that point we didn't have proof.

But we finally had a strong hypothesis.

We invited the Network team into the war room, explained the symptoms we were observing, shared the affected infrastructure, and pointed them toward the recent change.

After that...

The call became much quieter.

The Network engineers spent the next ten to fifteen minutes validating the configuration.

Those minutes felt much longer.

Everyone understood that if our hypothesis was correct, recovery could be straightforward.

If it wasn't, we were back to square one.

Eventually they confirmed the issue.

An incorrect configuration had been applied to one of the network switches.

The team rolled back the change.

Within minutes, the unreachable servers began recovering.

Less than forty-five minutes after I had joined the war room, services were returning to normal.

While engineers investigated, my work looked very different

One misconception about engineering management is that managers spend incidents directing every technical decision.

That wasn't my role.

While the engineers investigated the network issue, I focused on everything surrounding the investigation.

I asked my engineer to send customer notifications for the affected clusters so end users were aware that an incident was underway.

Because this was one of our largest customers, I also contacted their management directly through our established Slack channel.

I shared what we knew, what we didn't know yet, and promised regular updates as new information became available.

At the same time, I kept my own leadership informed.

My manager needed updates.

His manager needed updates.

Her manager needed updates.

Whether the root cause belonged to our team or not, I knew this incident would eventually reach executive conversations. My responsibility was to make sure those conversations were based on accurate information rather than speculation.

Just as importantly, I wanted my engineers to stay focused.

I didn't want them stopping every few minutes to answer questions from customers or executives.

Their job was solving the technical problem.

My job was handling the communication around it.

Looking back, I realized I was following the same pattern

After participating in enough incidents, I noticed I was approaching each one in a remarkably similar way.

Eventually, I gave that mental model a name: F.O.C.U.S.

F — Filter the signal

Before trying to solve anything, understand what is already happening.

Who is already investigating?

What information already exists?

Where is the real signal, and where is the noise?

O — Organize the response

Create a shared understanding across teams.

The faster everyone is working from the same picture, the fewer assumptions people make.

C — Coordinate ownership

Put the right specialists on the right problems.

Leadership isn't about solving every technical issue yourself.

It's about making sure the people best equipped to solve it are connected and have the context they need.

U — Unblock engineers

Remove obstacles.

Whether that's bringing another team into the conversation, finding additional information, or removing unnecessary interruptions, your role is to keep technical work moving.

S — Shield execution

Protect engineers from communication overload.

Customers deserve timely updates.

Leadership deserves visibility.

But engineers deserve uninterrupted time to investigate.

Keeping those responsibilities separate allows everyone to do their best work.

The incident ended, but the work didn't

Once the rollback restored connectivity, the mood in the war room noticeably changed.

The Network team acknowledged that the incorrect configuration had caused the incident and accepted ownership of the corrective and preventive actions to reduce the chances of it happening again.

For me, that was enough.

I had the preliminary root cause analysis I needed.

I could confidently brief my leadership, provide meaningful updates to the customer, and follow up on the remaining actions after the incident.

After the larger war room ended, only my engineer, my manager, and I stayed online for a few more minutes.

We sent the final customer notification confirming that services had been restored.

I also informed the customer's leadership that recovery was complete.

There were very few questions at that point.

Like us, they were busy validating that everything was healthy again from their side.

Final thoughts

That morning reinforced something I wish I had understood earlier in my management career.

Engineering managers don't create value during incidents by becoming the best debugger in the room.

They create value by reducing uncertainty.

By connecting the right people.

By maintaining clear communication.

By protecting engineers' focus.

And by helping dozens of people move in the same direction while the technical specialists solve the problem.

The F.O.C.U.S. framework wasn't born from a single incident.

It emerged from many incidents like this one.

Each one reinforced the same lesson:

When systems become chaotic, leadership isn't about having every answer.

It's about creating the conditions that allow the right answers to emerge.

Ready to find the right
mentor for your goals?

Find out if MentorCruise is a good fit for you – fast, free, and no pressure.

Tell us about your goals

See how mentorship compares to other options

Preview your first month