Heavybit article

What to Know About the Modern Incident Response Lifecycle

Heavybit by Andrew Park · · Article

"Teams only get good at this when they embrace the whole process and each of its steps."

— Jesse Robbins

Heavybit's incident management guide quotes me on why teams only get good at incident response when they treat the whole lifecycle as one discipline.

Andrew Park's Heavybit guide asking how teams get good at incident response, with Jesse quoted on treating detection, response, resolution, and retrospective as one discipline.

Andrew Park’s guide for Heavybit on modern incident management quotes me on why teams only get good at this when they treat the full lifecycle as a single discipline. The piece walks readers through the practical implications: normalize incidents by talking about them often, be honest about the state of the infrastructure, and treat the practice itself as the source of mastery. The line he pulled from our conversation is the one I have been saying since Amazon: skip any step in the cycle and you never fully develop the muscle for any of them.

People

Since this came out…

  1. AWS shipped Fault Injection Service. The patterns I ran by hand at Amazon, now a managed service anyone can call.
  2. I wrote the foreword to Incident Management for Operations by Schnepp, Vidal, and Hawley (O'Reilly). The book is the long-form version of the lifecycle argument I made in this piece.

Further reading

Topics