Every engineering team I have worked with has one. The person whose name comes up in every incident channel. The one who knows why the payments service got built the way it did in 2019. The one who gets paged at 3am because nobody else knows what to do when the queue backs up.
Leadership loves this person. They call them indispensable. They put them on stage at the all-hands.
They are also the biggest risk in your architecture, and you built them yourself.

How you built the hero
Nobody sets out to do this. It happens through a hundred small, sensible decisions.
Someone fixes a nasty production bug at midnight. You thank them. Next time something breaks, you route it to them, because they handled it well last time. They handle it again. Their reputation grows. Now the escalation path has their name hard-coded into it.
Then the ratchet kicks in. Every incident they resolve makes them the fastest person to resolve the next one. Every hour they spend firefighting is an hour they do not spend writing the runbook to let someone else do it. The gap between them and the rest of the team widens, and the widening gap justifies routing more work to them.
You did not decide to overload one person. You made a series of locally optimal choices, and they added up to it.
The org calls it excellence
Here is where it gets ugly. The organisation looks at this pattern and reads it as a virtue.
Your best engineer is working the most hours. They are visible in every crisis. They are the one management name-checks when things go well. Performance review season rolls around and they get the top rating, because the review process measures output, and their output is enormous.
So the incentive is clear to everyone watching. Want to be recognised here? Be the one who never says no.
I have spent years asking people about their bosses. My own research found 99.5% of respondents said they have had one or more types of bad boss. Not "a bad boss once." Types, plural. The most common story is not the shouty tyrant... it is the boss who took someone's willingness and quietly spent it until there was none left.
What you have is key-person risk
Strip away the language of heroism and look at the system.
You have a critical path running through one human being. They have no redundancy, no documented failover, and a hard operational limit of roughly one of them. If they take a two-week holiday, your incident response degrades. If they quit, you lose institutional knowledge existing nowhere else.
If you designed a service like this, you would fail your own architecture review.

Google's SRE team wrote this down years ago and it still gets ignored. In their chapter on eliminating toil, they define toil as work which is "manual, repetitive, automatable, tactical, devoid of enduring value, and that scales linearly as a service grows." They cap it at 50% of an SRE's time and run at about 33% in practice.
Their reasoning is not soft. They write: too much toil "leads to burnout, boredom, and discontent" and will "motivate the team's best engineers to start looking elsewhere for a more rewarding job."
Read it again. The people you lose to toil are not the weakest ones. They are the strongest ones, because they are the ones with somewhere else to go.
The numbers are not comforting
The 2025 Stack Overflow Developer Survey found only 24.5% of developers describe themselves as happy at work. 28.4% say they are not happy. Nearly half sit in the middle, described as complacent, which is a polite word for checked out.
The same survey asked what matters most for job satisfaction. Top of the list was not money. It was autonomy and trust to manage your own tasks. Money came second.
Your hero has neither. Their day is dictated by whatever broke. They have no control over their calendar, because their calendar belongs to the incident queue.
DORA's 2024 research puts a finer point on it. They found "unstable organisational priorities cause meaningful decreases in productivity and substantial increases in burnout." The part worth worrying about comes next... the effect "is highly resistant to mitigation and persists even in environments with strong leaders and high-quality documentation."
You cannot manage your way out of chaos with a good attitude. Being a nice boss does not undo a broken system.
Resilience is not the fix
The standard response to all this is to send people on a resilience course. Offer mindfulness. Put fruit in the kitchen.
This is the wrong layer of the stack.
If a database falls over under load, you do not tell the database to be more resilient. You look at the load. You add capacity, shed traffic, or fix the query. Nobody would accept "the database needs better coping strategies" as an incident postmortem.
Yet we accept exactly this when the overloaded component is a person.
Five fixes worth making
None of these are difficult. They are unpopular, because they slow you down this quarter to protect you next year.
1. Measure the concentration
Pull your last 50 incidents. Count who resolved them. If one name is on more than a quarter of them, you have a design flaw, not a star.
Do the same for code review approvals and for deployments. The pattern usually shows up in all three.
2. Make the expert the teacher, not the fixer
When the alert fires, the expert does not touch the keyboard. Someone else drives while they narrate. It is slower the first three times and faster forever after.
This is the single highest-leverage change on this list, and the one most teams refuse to make, because it costs minutes during an incident.
3. Give toil a budget and enforce it
Pick a number. 30% is a reasonable starting point. Track how much of each engineer's week goes to reactive work. When someone breaches the cap two sprints running, treat it as an escalation, the same as breaching an error budget.
A cap you do not enforce is not a cap. It is a wish.
4. Rotate on-call properly
Not a rota where one person is technically on-call and one person gets called every time. If your escalation path has a name in it, replace the name with a role. Then make sure everyone in the role has run through a real incident with support.
5. Take the holiday seriously
Send your key person on a two-week holiday with their laptop switched off. Whatever breaks while they are gone is your list of undocumented dependencies. It is not a disaster, it is a free audit.

The uncomfortable bit
Fixing this means your best engineer gets less visible. Their incident count drops. Their heroic saves stop happening, because there is nothing left to save you from.
If your recognition system still rewards firefighting, you have quietly punished them for fixing the problem.
So change what you measure. Reward the person who made themselves unnecessary. Reward the runbook, the automation, the boring Tuesday where nothing broke. Reward the engineer who taught three people to do the thing only they knew how to do.
It is a harder story to tell at the all-hands than a midnight rescue. Tell it anyway.
I write more about the leadership side of this over at Step It Up HR, where the same pattern shows up in every industry, not only ours.
Where to start
Open your incident history right now. Count the names.
If one person is carrying more than their share, you already know what you have. The question is whether you are going to call it excellence for another year, or call it what it is and fix it.