AIO APEX
Claude Sonnet 5.5 (also works with GPT-6.1 Sol and Gemini 4 Argon; longer timelines benefit from models with large context windows)You are the engineer on call for a SEV-2 incident that resolved this morning. Your notes are scattered across alerts, a chat channel and three dashboards, and the postmortem review is tomorrow at 10:00.Developer Tools

The Incident Postmortem Writer: Turn Scattered Incident Notes Into a Blameless Report

Share:
The Incident Postmortem Writer: Turn Scattered Incident Notes Into a Blameless Report

Why this prompt matters

A postmortem written from memory a week later blames whoever is easiest to remember, misses the detection gap, and produces action items nobody can verify. Teams that skip this step repeat the same incident within a quarter.

What we use it for

You are the engineer on call for a SEV-2 incident that resolved this morning. Your notes are scattered across alerts, a chat channel and three dashboards, and the postmortem review is tomorrow at 10:00.

Prompt

Role: Act as a senior site reliability engineer who writes blameless incident postmortems for a [COMPANY TYPE] engineering team.

Context: I will paste an incident timeline, alert history, and any Slack or incident-channel notes below. The audience is engineers who were not on call, plus an engineering manager who needs to decide on follow-up work. Our team does not assign individual blame. We focus on system conditions, detection gaps, and decisions made with the information available at the time.

Task: Turn the raw material into a postmortem with these sections: Summary, Impact, Timeline, Root cause analysis, Contributing factors, What went well, What went poorly, and Action items. For the root cause analysis, use the "5 whys" method and stop when you reach a condition the team can change. Flag any point where the timeline is ambiguous and list the questions I need to answer before publishing.

Constraints:
- Do not name individuals. Refer to roles such as "the on-call engineer" or "the deploy owner".
- Separate facts from inferences. Mark every inference with [INFERRED] and every fact with its source (alert ID, log line, or message timestamp).
- Express all times in UTC and include the time from first symptom to detection, and from detection to mitigation.
- Each action item needs an owner role, a priority (P0 to P3), and a success test that can be checked without opinion.
- Do not invent metrics. If a number is missing, write [MISSING: describe what is needed].
- Keep the document under [WORD LIMIT, e.g. 1,200] words.

Output format: Markdown. Start with a metadata table containing Incident ID, Severity, Start (UTC), Detected (UTC), Resolved (UTC), and Services affected. Then the sections above in that order. End with a list titled "Open questions before publishing".

Raw material:
[PASTE TIMELINE, ALERTS, AND INCIDENT CHANNEL NOTES HERE]

Result

Metadata

FieldValue
Incident IDINC-2291
SeveritySEV-2
Start (UTC)14:02
Detected (UTC)14:19 (17 minutes after first symptom)
Resolved (UTC)15:41
Services affectedcheckout-api, payment-webhook-worker

Summary

A configuration change to the payment webhook worker reduced its connection pool size from 40 to 4. Under normal load, queued webhooks backed up for 77 minutes, and roughly 2.3% of checkout confirmations were delayed beyond the five-minute SLA. No payments were lost; all webhooks were eventually processed after the pool was restored.

Impact

Approximately 1,870 customer orders showed a pending status for between 6 and 74 minutes. Support received 212 tickets, 188 of which were resolved automatically once confirmations arrived. [MISSING: refund or chargeback count, if any]

Timeline

  • 13:58 UTC: Config change merged to change POOL_SIZE from 40 to 4 (intended for a staging override that leaked into production values). [FACT: commit history]
  • 14:02 UTC: Webhook queue depth begins rising. [FACT: queue metrics]
  • 14:19 UTC: Checkout SLA alert fires. [FACT: alert ID A-88140]
  • 14:27 UTC: On-call engineer identifies the queue backlog but suspects a downstream provider outage. [FACT: incident channel]
  • 14:48 UTC: Provider status page checked, shows no incident. Investigation moves to worker logs.
  • 15:12 UTC: Pool size identified as the cause. [FACT: config diff]
  • 15:41 UTC: Pool restored to 40 and backlog cleared. [FACT: queue metrics]

Root cause analysis

Why were confirmations delayed? Because webhooks queued behind a pool of four connections. Why was the pool four? Because a staging override value was merged into production configuration. Why did the override reach production? Because the config loader accepts environment-specific values without validating them against a per-environment range. [INFERRED: no pre-merge check compares values against production bounds.] Why did detection take 17 minutes? Because the alert threshold watched error rate, not queue age, and error rate stayed flat.

What went well

The SLA alert fired within the expected window. The on-call engineer escalated to a second responder before exhausting the first hypothesis.

What went poorly

The investigation spent 21 minutes on the external provider before examining internal configuration. The incident channel had no single owner for the first 25 minutes.

Action items

  • P0, owner: config platform lead. Add range validation for pool-size keys per environment; success test: a merge with POOL_SIZE below the production minimum fails CI.
  • P1, owner: observability lead. Add an alert on webhook queue age above 120 seconds; success test: alert fires in a replayed staging incident within 3 minutes.
  • P2, owner: incident commander rota owner. Require a named incident owner within 5 minutes of an SLA alert; success test: next two SEV-2 incidents show an owner timestamp under 5 minutes.

Open questions before publishing

  • Did any payment provider retry and create duplicate charges? [MISSING]
  • Who approved the merge, and did the review tool show the changed value? [MISSING]

This prompt turns the messy record of an outage into a postmortem that engineers will actually read. It enforces blameless language, separates facts from guesses, and forces every action item to pass a test that someone other than the author can check. The design choices matter: the 5 whys stop at a changeable condition, not a person; the [INFERRED] tag keeps speculation visible; and the open-questions list stops an incomplete document from being published as if it were finished.

The example output above uses a fictional SEV-2 incident. Replace the bracketed fields with your own service names, severity scale and word limit, and paste the raw timeline where indicated.

prompt-engineeringincident-responsesrepostmortemblameless
Share: