Build an outage timeline from dated evidence, separate observations from recollections, calculate impact honestly and turn the record into assigned improvements.

Direct answer: Reconstruct the outage from timestamped evidence, recording the time, observation or action, source and level of certainty for each entry. Separate the start of user impact from detection, mitigation and verified recovery, and keep possible causes labelled as hypotheses. Use the timeline to choose a specific improvement with an owner and completion test. If the incident may involve compromised accounts, lost data or legal reporting duties, preserve evidence and follow the appropriate incident process before doing ordinary editorial cleanup.

A chronological story can sound convincing while giving the wrong explanation. The last change before an outage is a useful lead, not proof of cause. A colleague remembering that the system failed “around nine” is not the same evidence as a failed transaction recorded at 09:12.

Your aim is to make the response understandable enough to improve it. This gives a small-team technology retrospective something firmer than recollections and frustration, without turning a short disruption into an unnecessarily elaborate investigation.

Applies to: documenting a business-software outage in a UK business context after immediate response needs are being handled. It is not a forensic investigation procedure or a substitute for qualified local advice. Legal, regulatory and contractual requirements vary by country, sector and agreement.

Build the evidence-timed incident record

The evidence-timed incident record is an editorial method for keeping every time-linked claim attached to its evidence. It separates what happened, what someone thought was happening and what you now believe caused it.

Before editing a summary, preserve the relevant original records in an authorised location. These might include error messages, service-status updates, support tickets, application logs and the incident conversation. Restrict access because logs and screenshots can contain personal information, private URLs or credentials.

Do not paste unfiltered incident material into an unapproved AI tool. A timeline can be drafted manually in a spreadsheet. If an approved assistant helps order redacted material, verify every time and attribution against the original and leave unavailable facts blank.

Use these fields:

FieldPurpose
Date, time and zoneEstablish the sequence without ambiguous local times
EventState the observed result or action plainly
EvidenceLink or reference the original record
CertaintyDistinguish recorded time, estimate and unresolved conflict
SignificanceExplain whether it changed impact, diagnosis or response

My recommendation is to publish a useful internal timeline with clearly marked unknowns rather than delay it until someone can tell a perfectly complete story. For a serious incident requiring forensic preservation or external reporting, let the responsible specialists determine what can be circulated and when.

Separate four different clocks

Keep user impact, detection, response and recovery distinct. A monitoring alert can arrive after users are already affected. A mitigation can restore part of the service while another essential task still fails. The incident team may continue checking after the service appears restored.

Google's published postmortem example explicitly separates outage, mitigation and incident-management milestones and states the time zone. Its fictional service details are not a benchmark to copy; the useful practice is to avoid treating those milestones as interchangeable. Google SRE example postmortem.

Choose one display time zone for the timeline and preserve the original timestamp alongside any conversion. If a source lacks a date or zone, mark that gap rather than guessing. “London time” without the date can be ambiguous around seasonal clock changes.

Be careful with screenshot times and messages forwarded later. A message's send time may show when someone reported an error, not when it first occurred. A vendor status update may describe a broader incident than the one affecting your account. Record its scope rather than making your local timeline conform to it.

Write observations before explanations

A good event entry says, “The accounts team could not export an invoice; the application returned the recorded error.” A weaker entry says, “The vendor's update broke exports”, unless the supporting evidence actually establishes that link.

Use separate entries for an attempted fix and the result of checking it. “Restart completed” does not establish recovery. State which task was retried, with which safe input, and what the result showed. Keep failed attempts because they explain elapsed time and may prevent the same unproductive sequence next time.

When accounts disagree, preserve both claims and their sources. Ask a narrow follow-up: was the time observed on a clock, inferred from a meeting or copied from a log? Do not resolve uncertainty by choosing the more senior person's recollection.

Avoid silently correcting original logs. If a system clock is known to differ, record the evidence for the adjustment and retain the original value. If you cannot establish the difference, show the ordering uncertainty rather than forcing every entry into precise minute-by-minute sequence.

Calculate duration without adding parallel work

Consider this entirely illustrative outage record. All times use UTC on the same fictional incident date; no real outage or product failure is being reported.

TimeRecorded eventWhat it establishes
09:12First retained failed exportEarliest confirmed failure in the available evidence
09:18Team sees the incident messageDetection by the responding team
09:27An initial workaround fails its checkAttempted mitigation, not recovery
09:43A manual export route worksPartial mitigation for the urgent task
10:07Normal exports pass the agreed checksConfirmed operational restoration for those checks

The elapsed interval between the first retained failure and verified restoration is 55 minutes: 48 minutes from 09:12 to 10:00, plus seven minutes to 10:07.

The detection gap is 09:18 - 09:12 = six minutes. The first working mitigation arrives 31 minutes after the first retained failure. Neither calculation establishes when the underlying fault began; the evidence might start late.

Suppose three responders spend 38, 29 and 17 active minutes respectively. Their total response effort is 84 person-minutes. Because work overlaps, you must not call this an 84-minute outage. If four other staff are unable to perform a particular task for all 55 minutes, that is 4 × 55 = 220 person-minutes of task unavailability, not necessarily 220 minutes of paid time wasted: they may have done other work.

Do not add response effort and blocked time unless the populations and activities are distinct. Record preparation and review costs separately too. An illustrative 25-minute timeline draft and a 15-minute review by two people adds 25 + 30 = 55 person-minutes of follow-up effort. None of these figures automatically represents cash lost.

Use the record to choose a repairable weakness

Look for evidence that a specific part of the response can improve. A six-minute detection gap might justify a clearer reporting route. A failed workaround might need a verified instruction. A late recovery check might reveal that nobody knew who could authorise a return to normal work.

Do not use the timeline to assign blame based solely on who performed the last visible action. Google's guidance on blameless postmortems focuses on contributing conditions and preventive action rather than punishment. That is an established operational practice, not a framework invented for this article. Google SRE guidance on postmortem culture.

Choose an action with an observable finish. “Improve communication” is not enough. “Publish the fallback export procedure, assign its owner and have a second person complete a dummy export using it” provides evidence of completion.

Keep root-cause analysis separate where the timeline cannot settle it. A supplier investigation may continue after your local process is repaired. Mark the causal question open and update the record when supported evidence arrives, rather than presenting a guess as the final explanation.

Produce the first record by the next working day

  1. Preserve authorised evidence while it is still available, without changing logs or broadly sharing sensitive records. Follow any security or legal preservation requirements first.
  2. Within the next working day, draft the time-ordered entries and distinguish evidence from recollection. Keep unknowns explicit.
  3. Ask the people involved to check the events attributed to them and the scope of the impact. Resolve factual discrepancies, not disagreements about blame.
  4. Choose the most useful improvement, name its owner and set a review date. Verify its completion through a safe exercise or other observable evidence.

Stop ordinary cleanup and seek qualified help if evidence suggests compromised access, data loss or a reporting obligation. A tidy timeline is not more important than protecting the investigation or meeting the relevant duties.

Frequently asked questions

Is the timeline the same as the root cause analysis?

No. The timeline records the sequence and evidence, while root cause analysis tries to explain the contributing conditions and mechanisms. The sequence can identify useful leads, but an event happening first does not prove it caused the next one. Keep causal hypotheses separate until the evidence supports them. For a vendor-managed service, you may never obtain every internal detail. You can still improve your own detection, fallback and communication processes using the facts available. Do not delay a practical local improvement solely because the supplier's final explanation is pending.

What if our team did not record anything during the outage?

Start with the evidence that still exists, such as support tickets, error notifications, message timestamps and the first confirmed successful task. Ask participants for recollections and label them as estimates, including the basis for the timing. Do not manufacture minute-level precision from a general memory. Record the evidence gap as a finding and choose a simple way to capture key events next time. For a minor incident, a short record with honest uncertainty may be sufficient. If the incident is serious, preserve what remains and involve the appropriate specialists before reconstructing more extensively.

Should we include every message from the incident chat?

No. Keep the original conversation where retention and access rules permit, but select timeline entries that explain impact, decisions, actions and verification. A complete message dump can obscure the sequence and expose unnecessary personal or confidential information. Use references so authorised reviewers can inspect context when needed. Include a disagreement or mistaken assumption if it affected the response, without turning the summary into a transcript of frustration. Where legal or forensic preservation applies, do not delete the original material simply because it is not useful in the concise operational timeline.

Can AI turn the chat into a timeline for us?

It can help organise approved material, but a person must verify every event, time, source and attribution. A model may turn an estimate into an exact time, merge separate attempts or mistake a proposed action for one that actually happened. Use a permitted service and minimise sensitive input before uploading anything. Ask it to preserve uncertainty rather than complete missing details. If redaction and checking take longer than writing the short timeline manually, skip the assistance. The record's value comes from its evidence, not from how quickly a plausible narrative appears.

How should we handle a vendor status page that disagrees with our evidence?

Record both accounts with their scope and timestamps. The vendor may be describing a region, feature or customer group that differs from yours, or may publish an update after the event it describes. Your local failed and successful checks establish something specific about your workflow, not the entire service. Ask the supplier to clarify material discrepancies and retain its response. Do not rewrite observed times solely to match the public status page. If the disagreement affects a contractual claim or reporting duty, use the appropriate professional advice and supporting records.

How much detail should we share outside the response team?

Share what the audience needs to understand the impact, current status and relevant next actions, while protecting sensitive technical and personal information. An internal operational timeline may contain details unsuitable for a client update or public post. Prepare a separate summary where necessary and have the authorised owner review it. Do not imply a confirmed cause or complete recovery beyond the evidence. If the event involves legal, regulatory or contractual notification requirements, those rules and the relevant jurisdiction govern the communication; an informal summary is not a substitute for the required process.

Sources and verification

  • Google SRE: example postmortem, checked 11 September 2026 for its explicit timeline milestones and time-zone practice. Its fictional incident figures are not reused.
  • Google SRE: postmortem culture, checked 11 September 2026 for the established blameless approach and action-oriented purpose.
  • The parent was read in the supplied website source; its public route could not be retrieved. This article's outage, times and effort figures are illustrative rather than observed results.
Twokq Tech

This article is practical guidance. Apply it in proportion to your tools, evidence, risks, and responsibilities.