AIO APEX
Claude Opus 4.7 or GPT-5 (a structured-reasoning task; any strong general-purpose model works, no need for a specialized coding model)An on-call engineer wrapping up a 8-hour shift with three alerts, one resolved incident with a lingering root cause, and a flaky Redis layer, who has ten minutes before the next engineer logs on and needs to hand off clearly instead of firing off a rushed Slack message that leaves out half the context.Developer Tools

The On-Call Handoff Writer: Turn a Chaotic Shift Change Into a Clean Incident Handoff

Share:
The On-Call Handoff Writer: Turn a Chaotic Shift Change Into a Clean Incident Handoff

Why this prompt matters

Vague handoffs are a well-documented driver of extended incident resolution times: the next on-call engineer re-investigates problems that were already diagnosed, misses that a 'resolved' incident still has an open root cause, or doesn't learn about a degrading system until it pages again at 3am. A five-minute handoff write-up costs far less than the hour a team loses re-discovering context that already existed in someone's head.

What we use it for

An on-call engineer wrapping up a 8-hour shift with three alerts, one resolved incident with a lingering root cause, and a flaky Redis layer, who has ten minutes before the next engineer logs on and needs to hand off clearly instead of firing off a rushed Slack message that leaves out half the context.

Prompt

Act as a senior site reliability engineer who writes clear, actionable on-call handoffs that the next engineer can read in under five minutes.

CONTEXT:
- Shift window: [START TIME/DATE] to [END TIME/DATE]
- Alerts and incidents during the shift: [LIST EACH ALERT/INCIDENT WITH ROUGH TIME AND SYSTEM AFFECTED, e.g. "02:14 — payment-api p99 latency breached 2s threshold for 8 minutes"]
- Actions taken and outcomes: [WHAT YOU DID FOR EACH ITEM AND WHETHER IT RESOLVED, MITIGATED, OR IS STILL OPEN]
- Current system state: [ANYTHING DEGRADED, FLAKY, OR BEING WATCHED EVEN IF IT DIDN'T PAGE]
- Scheduled or in-flight changes: [DEPLOYS, MIGRATIONS, OR MAINTENANCE THAT STARTED OR IS PLANNED DURING THE NEXT SHIFT]

TASK:
Convert these shift notes into a handoff document the next on-call engineer can act on immediately, without needing to ask follow-up questions or dig through dashboards to understand current state.

CONSTRAINTS:
- Do not write vague status language like "keep an eye on things" or "should be fine." Every open item needs a specific next action and, if known, an owner.
- Separate items that need immediate attention from items that are informational only — do not bury a live risk among resolved items.
- If an incident was resolved, state the resolution and whether a follow-up (postmortem, ticket, permanent fix) is still outstanding — "resolved" and "fully closed" are not the same thing.
- If you don't have enough information to assess severity or next steps for an item, say so explicitly rather than guessing.

OUTPUT FORMAT:
1. One-line TL;DR at the top — the single most important thing the next engineer needs to know before anything else
2. Still-Open Issues — table with columns: Issue | Current State | Next Action | Owner (if known)
3. Resolved During Shift — what broke, what fixed it, and whether a follow-up ticket/postmortem is still outstanding
4. Watch List for Next Shift — things that didn't page but are trending toward a problem
5. In-Flight Changes — anything deploying or migrating that could affect the next shift

Result

Shift: Oct 9, 22:00 – Oct 10, 06:00 UTC

TL;DR

Payment API latency spike is resolved but the root cause (a slow downstream fraud-check call) is not fixed — it will likely recur under similar load. Redis cache layer is flapping and needs attention before it causes a real outage.

Still-Open Issues

IssueCurrent StateNext ActionOwner
Redis cache-layer connection drops (3x this shift, ~90s each)Self-recovers each time, no customer impact yet, but frequency doubled vs. last weekCheck connection pool sizing and recent Redis client version bump; escalate to infra if it recurs before 10:00Unassigned — flagging for infra team
Fraud-check service p99 latency at 1.4s (up from 400ms baseline)Caused the payment-api latency spike at 02:14; mitigated by raising payment-api's timeout, not fixed at the sourceFile a ticket with the fraud-check owners; current mitigation is a band-aid that will mask further degradationPayments team (ticket not yet filed)

Resolved During Shift

02:14 — payment-api p99 latency breach (8 min). Root cause: fraud-check service slowdown. Fix applied: raised payment-api's downstream timeout from 1s to 2s, which stopped the cascading retries. This is a mitigation, not a fix — follow-up ticket still outstanding (see Still-Open Issues above).

Watch List for Next Shift

Redis connection pool: if drops hit 4+ occurrences or exceed 90 seconds, treat as a page-worthy incident rather than a watch item. Fraud-check latency: if it crosses 2s p99, the current timeout mitigation will stop working and payment-api will start erroring instead of just slowing down.

In-Flight Changes

None scheduled for the next shift. A database index migration on the orders table was originally planned for 08:00 but has been pushed to tomorrow — confirm with the DB team before assuming it's still off.

Every on-call rotation has the same failure mode: the outgoing engineer knows exactly what happened during their shift, and none of that knowledge survives the handoff. The incoming engineer gets a Slack message that says something happened with payments, inherits a dashboard full of unexplained spikes, and spends the first twenty minutes of their shift reconstructing context that already existed in someone else's head twelve hours earlier. This prompt exists to stop that specific, recurring loss.

Why the Prompt Forces a Severity Split

The structure deliberately separates «Still-Open Issues» from «Resolved During Shift» from «Watch List,» rather than producing one flat list of things that happened. That split exists because the single most dangerous handoff failure isn't a missing fact — it's a live risk getting buried in the middle of a wall of text next to five routine, already-resolved items. A reader skimming a long handoff under time pressure will treat everything in a flat list as roughly equal priority, which is exactly backwards. Forcing the open issues into their own table with an explicit next action and owner means the next engineer can triage severity in seconds instead of reading every line to figure out what still matters.

Why «Resolved» and «Fully Closed» Are Treated as Different States

The constraint requiring a distinction between an incident being resolved and being fully closed targets a specific, common mistake: applying a mitigation, watching the symptom disappear, and then writing it up as done. In the example output, raising a timeout stopped the payment-api latency spike, but the actual cause — a slow fraud-check service — is untouched and will resurface under the same load conditions. A handoff that just says «fixed» hides that the same incident is likely to recur on the next shift. Forcing an explicit note about outstanding follow-up work means the next engineer isn't caught off guard by a repeat of something they were told was handled.

Why the Watch List Exists as Its Own Section

Not everything worth mentioning triggered a page. A connection pool that's flapping more than usual, a latency metric trending upward but still under threshold, a disk filling up slowly — these are exactly the signals that get lost in a verbal handoff because they didn't rise to the level of «something happened,» even though they're often the leading indicator of what will page on the next shift. Giving this category its own dedicated space in the output format means these signals get written down instead of staying in the outgoing engineer's head until they forget them by the time they're off shift.

Use this at the end of any on-call shift, incident rotation, or support handoff where the cost of losing context is measured in minutes of re-investigation, or worse, a repeat incident nobody saw coming because nobody wrote it down.

Share: