Mikko Kotila

Topics: Research and development

Research and development

Adverse Events Are Invitations for Process Improvement

This article is part of Short essays covering what I learned from running software projects for three decades.

All kinds of projects suffer from repeated occurrence of various adverse events and from organizations responding to such events reactively. Basically, something happens, the adverse event, and then something else happens as a reaction to that, the response. The adverse event is treated as a problem, and the response is treated as a solution, but they are actually both part of the problem.

Simply put, a purely reactive incident → response stance is a high-probability long-term loser in technology projects for a simple reason: it fights glaring symptoms in systems where the real damage accumulates quietly, upstream, and combinatorially.

1) Adverse events rarely arise from single causes; they’re alignments of latent conditions.

A more systematic approach distinguishes active failures from latent conditions (design gaps, staffing, incentives, tooling, coupling, etc.,) and argues that serious incidents usually require both. If you only react to the active failure, you leave the latent conditions in place — so the system remains primed to generate the next “surprising” failure.

2) In complex, tightly coupled systems, “after-the-fact fixes” systematically miss the real dynamics.

In systems with high interactive complexity and tight coupling, small issues can combine and cascade in unpredictable ways — making “we’ll fix it when it breaks” structurally unreliable as a control strategy. To mitigate this, successful projects essentially evolve into complexity management engines.

3) Reactive mode creates a reinforcing loop: firefighting consumes the very capacity needed to prevent fires.

Literature makes this clear: toil, often celebrated in software projects, tends to be interrupt-driven and reactive and is often work with no enduring value. When such work dominates, actual reliability work gets crowded out — so incidents keep coming, and toil grows. There is an aggrevating hidden force at play here: firefighting produces heroes, which creates cultural incentives to tolerate the conditions that produce fires. This is a dangerous cultural pattern.

4) Complex systems “drift” into failure — so waiting for a clear alarm is waiting too long.

Global risk increases gradually through locally reasonable decisions made under pressure (schedule, cost, workarounds, etc.,), exhausting available margins. A reactive posture detects problems late, when options are costlier and side-effects larger. Drift is particularly insidious because each individual step is defensible. No single decision looks reckless — the danger lives in the aggregate trajectory, which nobody in a reactive organization is tracking.

5) In software, late discovery and late correction is economically untenable.

Defect-fix cost rises sharply the later issues are found. A reactive strategy is, by design, a late-discovery strategy — so its cost curve tends to steepen over time. Fixing a bug in design costs $1. Fixing it in QA costs $10. Fixing it in Production during an outage costs $100 (plus reputation cost). Reactive organizations are permanently operating on the most expensive part of the curve.

Conclusion

Reactive postures aren’t just slower — they’re systematically mismatched to the failure dynamics of complex systems. They address the proximate trigger while leaving the generative structure intact. This is why adverse events really are an invitation for process improvement.

You might also enjoy* The Three Fundamental Principles of Software Projects.*