Incident playbook
One problem. Every affected customer.
The volume of an incident is driven less by the fault than by the silence after it. Almost everything that goes wrong in incident handling is a communication failure rather than a technical one.
Notice the cluster early
Three similar reports within an hour is a pattern worth stopping for, even when the wording is completely different. The instinct is to answer them individually because each one looks small; that instinct costs several hours and produces three inconsistent explanations.
Count reported and affected separately
For most incidents the customers who contacted you are a fraction of the customers affected. Communicating only with the reporters means the larger group finds out from your silence, which is the worst possible version.
Say something before you know the cause
Teams delay the first update until they have an answer, which is precisely backwards. What is broken, who it affects, what you are doing, and when you will next update — that can be sent within fifteen minutes and prevents most of the second wave.
- What is broken, in the customer's terms
- Who is affected
- What you are doing about it
- When you will update next — and then actually do
The next-update time is a promise
Committing to an update in an hour and not sending one does more damage than the original fault. If you cannot commit to an hour, commit to three. Incident updates that stop arriving are how a technical problem becomes a trust problem.
Publish somewhere public
Many affected customers never contact you. They check for a status page, find nothing, and form a view. Publishing early costs very little and is read by far more people than your replies are.
Tell everyone it is fixed
This is the most commonly skipped step in all of customer service. The fault is resolved, the engineers go home, and nobody closes the loop with the people who reported it. It should be one action, and it should not depend on anybody having the energy for it afterwards.
Follow up individually where it is owed
Some customers need more than a broadcast — a call, a credit, an apology from the person who owns the account. Identify those while the incident is live rather than afterwards, because afterwards nobody does.
The third time is not a coincidence
When the same underlying fault causes several incidents, that is a root cause, and it stays open after each incident closes. Quantify what it has cost — cases, customers, hours, account value — because that total is the only argument that reliably gets it prioritised.
Take better care of every customer.
Give your team the context, knowledge and AI they need to resolve problems properly — and know who needs attention before they ask.
Keep the mailbox you already use · The AI is never metered · [email protected]