Measure blocked work, workaround effort and recovery time without double-counting colleagues, then judge whether a software fix is worth its ongoing cost.

Direct answer: Record the people actually affected by each interruption, their blocked minutes, and the additional work needed to recover. Exclude colleagues who continued normally and avoid counting the same person's waiting and troubleshooting minutes twice. Compare the avoidable portion of that workload with the setup and continuing effort of the proposed fix, keeping released time separate from cash savings.

A team's most frustrating application is not necessarily its most expensive problem. An irritating delay may be easy to work around, while an infrequent failure can consume a morning reconciling records. Counting incidents alone cannot distinguish them.

My recommendation is a short, transparent incident diary before buying replacement software. Escalate security concerns, data loss or a consequential service failure immediately; you do not need a completed cost model to respond to those risks.

Measure interrupted work, not general annoyance

Use The interruption workload account, an editorial method for separating an event's duration from its effect on people and from the work needed to put things right. It is a lightweight decision aid, not a financial reporting standard.

This is narrower than the broader assessment in how to know whether your tech stack is working. You are measuring one recurring problem well enough to decide what evidence or intervention it warrants.

Choose a specific symptom, such as “saving an enquiry fails and the record must be recreated”. Avoid categories such as “the system is rubbish”. The first can be timed and investigated; the second combines many possible complaints.

Google's site reliability guidance distinguishes observing externally visible behaviour from measuring a system's internal state. That distinction is useful here: you need evidence about the work that failed, not just an alarming chart. This article's workload calculation is an editorial approach, not Google's monitoring framework. See Monitoring Distributed Systems.

Agree the purpose with the team before collecting information. The diary should explain a system problem, not secretly rank staff. Record task-level facts without copying client information, message contents or credentials into a new log.

Build one event record people can complete quickly

For each occurrence, capture the time, symptom, attempted task, affected people or roles, and what restored normal work. Use an event identifier so four reports of the same outage are not mistaken for four independent failures.

Add separate fields for blocked time, workaround effort and recovery work. Blocked time means the person could not continue useful planned work. Workaround effort means extra steps taken to complete the task another way. Recovery work means restoring or reconciling the intended result afterwards.

These categories describe time; they do not automatically add together. If somebody spends a ten-minute outage troubleshooting for six minutes and waiting for four, the total is ten, not sixteen. If they then spend five additional minutes repairing a record, the total becomes fifteen.

Do not demand second-by-second accuracy. A clearly labelled estimate is better than an invented precise measurement. Mark the difference between a timestamped event, a contemporaneous estimate and a recollection supplied days later.

Keep a note of useful work completed during the disruption. Somebody who switched to an already-planned task was inconvenienced, but those minutes were not necessarily lost. Any additional switching or recovery effort needs its own defensible observation rather than a universal multiplier.

Count exposure before multiplying duration

A service being unavailable for ten minutes does not mean a six-person team lost sixty minutes. Some people may not use it at that time, and others may have continued through a functioning alternative.

For each event, list the affected roles and their relevant minutes. If several people experienced identical consequences, grouping them is reasonable. If the experiences differ, keep separate entries rather than forcing an average.

Treat an unresolved task as a pending outcome. For example, “customer response delayed until tomorrow” is not automatically a full day's labour loss. Record the delay and check whether it caused additional work, a missed commitment or a documented commercial consequence.

That distinction also helps prioritise. A short failure at an approval deadline can matter more than a longer interruption during flexible internal work. Keep the consequence visible beside the time account rather than pretending one hourly rate captures every risk.

If the team cannot agree whether incidents share a cause, group them by observed symptom first. A consistent diary can still support investigation without an unsupported claim about the underlying fault.

Work through a six-person team's example

Consider a six-person studio recording eleven interruptions over one working week. All figures below are illustrative assumptions, not observations or a benchmark. Assume the same four colleagues are genuinely blocked for seven minutes per interruption; the other two continue unaffected.

Blocked workload is:

11 interruptions × 4 people × 7 minutes = 308 person-minutes.

After each interruption, one of those four people spends an additional five minutes reconciling records. Because this happens after the blocked interval, it can be added:

11 × 5 = 55 person-minutes of recovery.

The team spends a further 25 person-minutes documenting the pattern and preparing a support request. Total recorded effort is 308 + 55 + 25 = 388 minutes, or 6 hours 28 minutes.

It would be wrong to calculate eleven events across all six people: 11 × 6 × 7 = 462 minutes. That would add 154 minutes of supposed blockage experienced by nobody in the example.

Suppose the team assigns an illustrative internal capacity value of £30 per person-hour. The account is 388 ÷ 60 × £30 = £194. That is the value assigned to occupied capacity, not proof that £194 left the bank account or can be recovered as revenue.

Keep the operational description attached to the number. Without it, a plausible-looking total can travel into a budget discussion stripped of its assumptions.

Compare only the workload a fix can actually remove

Suppose a proposed configuration change is expected, but not yet demonstrated, to prevent eight of the eleven weekly events. Each prevented event would release 4 × 7 + 5 = 33 person-minutes in this example. Eight events would therefore release 264 minutes.

Assume the change takes 180 person-minutes to implement and adds a ten-minute weekly check. The expected recurring release is 264 − 10 = 254 minutes per week. Setup would be recovered in 180 ÷ 254 = about 0.71 comparable weeks if the assumptions hold.

Do not present that as a promised result. Run the reversible change, repeat the diary and check whether the same symptom actually falls. If the quieter week had much less work, compare interruption opportunities as well as raw counts.

The 25-minute support preparation from the original example may be a one-off effort. It should not be credited as a recurring weekly saving without evidence. Equally, continuing supervision or manual checks belong in the proposed fix's ongoing cost.

For a paid replacement, add its actual quoted price, migration, training and parallel running. A large theoretical capacity benefit can disappear if the replacement moves work elsewhere instead of removing it.

Keep cash, capacity and consequences separate

Use separate headings in the decision note. Cash costs include verified additional charges or paid overtime attributable to the failure. Capacity is staff time that could have supported other useful work. Consequences include missed commitments, unreliable records or customer impact that you can substantiate.

Do not count the same consequence twice. If paid overtime is how the team recovered work, explain whether the associated minutes are already included in the workload account. Summing every available label can make one incident look like several independent losses.

Avoid guessing how much revenue a saved hour will produce. For a capacity-constrained team with confirmed demand, that time may be commercially useful. For another team, it may improve response time, reduce pressure or create space for maintenance. Those are legitimate outcomes without fictitious cash savings.

Document uncertainty in ordinary language. “Four events timed, seven estimated by participants” tells a decision-maker more than an unexplained confidence score. If the purchase decision changes under plausible alternative estimates, collect better evidence before committing.

Run a bounded investigation this fortnight

  1. Spend twenty minutes agreeing one symptom, the diary fields and who will collate reports. Confirm that sensitive information stays out of the log.
  2. Record occurrences for a normal working week. Investigate urgent risk immediately rather than waiting for the diary to finish.
  3. Reserve thirty minutes to remove duplicates, separate overlapping time and calculate a range where estimates are uncertain.
  4. During the following week, trial one authorised, reversible remedy with a saved baseline and a clear rollback owner. Measure both interruptions and new maintenance effort.
  5. Keep the change only if the task improves without unacceptable new risks or burdens. If the evidence cannot distinguish competing explanations, send the symptom record to qualified support instead of repeatedly changing unrelated settings.

Frequently asked questions

Should I include the time it takes to regain concentration?

Include additional recovery effort when you can describe and estimate it credibly, but do not attach a standard concentration penalty to every event. Someone may resume immediately, while another person may need to reconstruct a complicated decision. Ask what work had to be repeated and record that separately from the outage itself. If the estimate comes from memory, label it accordingly. For a decision sensitive to this amount, observe a small number of future events more carefully with the team's knowledge. The exception is a supplied, relevant measurement for your actual task; use it with its method and limitations stated.

What if people stop reporting because the problem is so common?

Simplify the diary and agree a bounded reporting period. A problem that has become routine can disappear from support records even while consuming effort. Ask for one event identifier, the blocked task and a short estimate rather than an essay each time. A designated person can consolidate reports of the same incident. Make clear when reporting will end and what decision it informs. If logging every event would itself be burdensome, sample a defined period and label the limitation. Do not extrapolate a particularly bad hour into a whole year's loss without evidence that its workload and conditions are representative.

Can I use salary divided by working hours as the hourly cost?

You can use a clearly defined internal rate as a capacity valuation, but it is not automatically an avoidable cash cost. Salary continues to be paid when a software problem disappears. State what the rate includes and use a consistent basis across the options being compared. If finance already has an approved costing convention, follow it instead of inventing another one. Actual paid overtime or an external recovery invoice belongs in the cash account with supporting records. The calculation becomes misleading when a capacity estimate is described as guaranteed savings or assumed new revenue without a corresponding operational change.

Should a rare but serious failure outrank frequent small delays?

It can. Time totals cannot adequately represent data loss, compromised access, missed statutory obligations or a failure at a critical customer deadline. Escalate the particular risk through the appropriate organisational process and obtain qualified advice where needed. Continue to document ordinary interruption effort, but do not make a serious incident wait for a favourable arithmetic result. For less consequential rare failures, compare the recovery process and actual impact rather than guessing a dramatic worst case. The important distinction is between an evidenced consequence requiring action and an unsupported possibility used to justify whichever purchase somebody already prefers.

How do I compare weeks with different workloads?

Record the number of relevant opportunities for the failure as well as the incident count. If the symptom occurs when saving enquiries, compare failures alongside attempted saves, not only total working hours. Also retain total affected workload because a smaller failure rate can still create more effort during a busier week. Use the same event definition and recording method before and after a change. If those conditions differ, state the limitation rather than declaring the fix successful. A short observation period may justify further testing, but it rarely supports a confident annual savings forecast across seasonal or changing work patterns.

Do we need monitoring software to do this properly?

Usually not for an initial small-team decision. A short shared diary can establish who was affected, what failed and how much additional work followed. Automated telemetry becomes useful when it answers a defined question that people cannot reasonably observe, such as the timing of an intermittent technical error. It does not automatically explain the human consequence. Before introducing monitoring, agree the information collected, access and retention, and check the applicable organisational and legal requirements. Do not install intrusive employee tracking to estimate a software annoyance. If deeper diagnostics are necessary, involve authorised technical support and collect the narrow evidence required.

Sources and verification

  • Google SRE: Monitoring Distributed Systems, checked on 11 September 2026 for the distinction between externally visible behaviour and internal monitoring. The workload account and numerical example are editorial methods and illustrative assumptions, not Google benchmarks.
  • The assigned parent guide was read in the supplied site source. Its public route could not be retrieved during verification; the supplied canonical path is retained. No product prices or salary benchmarks are asserted.
Twokq Tech

This article is practical guidance. Apply it in proportion to your tools, evidence, risks, and responsibilities.