When a slope collapses, a drain bursts or a railway track becomes unsafe, the public reaction is usually the same: surprise, outrage and a quick hunt for someone to blame. But if you have ever managed quality properly, with metrics and consequences, what you feel is something else: ‘It was only a matter of time.’ The circle of quality applies to everything we do, not just the industrial environment.

Because these episodes do not ‘occur’ on the day of the storm. They are created over months or years with small decisions: maintenance that is postponed, inspections that are reduced, incidents that are recorded… and remain dormant in an endless queue. Then the heavy rain comes, the ground becomes saturated, the load does its job… and the system shows the full bill, with interest.
It’s not bad luck. It’s not an isolated incident. It’s a pattern.
Problems with critical infrastructure—railways, roads, drainage systems, embankments, pavements—are often the predictable result of an explosive combination:
- Maintenance debt (what is not fixed today, but continues to deteriorate).
- Weak or bureaucratic risk management (lots of procedures, little redress).
- More extreme weather (new ranges of rainfall, heat and variability that break with previous assumptions).
And there is a silent accelerator that makes everything worse: administrative latency. When an incident reported by technicians can take up to 18 months to resolve, you are not ‘managing maintenance’. You are accumulating risk.
The uncomfortable truth about quality
There is an uncomfortable truth about quality: serious failures almost never arise from a single major error. They arise from a chain of small compromises.
The mechanism repeats itself with depressing precision:
1) Latent faults
Minor deterioration that does not seem urgent: clogged drains, overgrown gutters, cracks, subsidence, erosion on a slope, contaminated ballast, joints that no longer seal as they should. Nothing ‘catastrophic’… yet.
2) Incidents detected but not closed
Technicians report. Specialised staff notify. An order is opened. It is added to a backlog. And this is where the real problem begins: the system is designed to process, not to repair.
If the normal circuit takes months, incidents pile up. And the backlog is not an inventory: it is a time bomb.
3) Trigger conditon
An episode of heavy rainfall, saturated soil, flooding, accumulated vibration, thermal changes. What was previously ‘tolerable’ is no longer so because the safety margin had already been exhausted.
4) Missing barriers
In a mature system, there are barriers: risk-based inspections, extraordinary controls after events, sensors at critical points, preventive shutdowns based on criteria, rapid response with SLAs based on severity.
In an immature system, there is an Excel spreadsheet, a procedure, and a lethal phrase: ‘it’s in progress’.
5) Crises, emergencies and costly expenditure
Reactions are slow. Emergencies are paid for. Improvisation takes place. Services are cut. The reputational cost is assumed. And the cycle begins again.
In terms of quality, this is not a one-off failure. It is a broken PDCA cycle: planning (P) and partial execution (D) are followed by weak checking (C) and little or late improvement (A). Without real ‘A’, the system learns to repeat mistakes.

Recent storms have left a clear picture: floods, rising waters, landslides and power cuts. The weather is not an excuse; it is an operational scenario. And today that scenario is harsher and more frequent.
But climate alone does not explain why one system holds up and another collapses. What makes the difference is accumulated vulnerability: drains with no real capacity, slopes with deferred maintenance, inspections that never happen, findings that are never acted upon.
There is one piece of information that, without exaggeration, explains everything: the cycle time between ‘detected’ and ‘corrected’. If that time is measured in quarters or years, the system has already decided that reliability is secondary. And when reliability is secondary, security ends up paying the price.
As one journalist said—and it’s a surgical phrase—: ‘Maintenance work is not inaugurated.’ And that is the crux of the problem: what does not make for a good photo does not compete well against what does. But physics, water and gravity do not vote; they just charge.
What would happen with a private company?
Here comes the painful part, because in a private company (industrial, logistics, aerospace, energy) this would be managed differently. Not out of moral virtue, but because of incentives and operational discipline.
1) SLA by criticality
In a company with a culture of quality, a ‘critical’ incident cannot wait months. It is classified by severity and assigned an SLA: hours, days or a few weeks. Anything that affects security or service continuity is given top priority.
2) KPIs that matter
You don’t argue with opinions, you argue with numbers:
- Backlog by severity (critical/major/minor)
- Average closure time (and percentiles, not just average)
- Percentage of inspections completed on time
- Percentage of findings closed
- Recurrences (if it happens again, the system did not learn)
If these KPIs do not exist or are not published internally, it means that no one is managing: they are merely surviving.
3) Root cause and true CAPA
When there is a serious incident, it is investigated. But not to point the finger at the last person who touched the piece, but to find system failures: missing barriers, prioritisation decisions, lack of resources, faulty process design. And a CAPA (corrective and preventive actions) is created with subsequent verification.
If the incident is repeated, it is a management failure, not ‘bad luck’.
4) Change of design if the environment changes
If the climate changes, the plan changes. Period. Frequencies, thresholds, materials, drainage, stabilisation, post-event protocols. In a serious company, the new climate is not up for debate: it is incorporated into the risk assessment.
In the public sphere, climate change too often remains a mere statement of intent, while plans continue to be designed for the “normality” of decades ago.
The solution is not a slogan. It is a change of model: moving from ‘infrastructure as a construction project’ to infrastructure as a product.

¿How the public infrastructure should be managed?
A serious product has requirements (security, availability, resilience), a control plan, traceability, auditing, and continuous improvement. Critical infrastructure should be managed in the same way, with four pillars:
1) Criticality-based maintenance (not calendar-based)
Not every section, embankment or drainage system poses the same risk. Priorities must be set based on exposure, history, consequences and vulnerability. This allows for better spending and action where it reduces real risk.
2) Reducing latency: From processing to repair
If there is backlog, there is accumulated risk. Cycle time must be addressed by:
- triage by severity
- Binding SLAs
- protected maintenance budgets (non-cannibalizable)
- contracting and permits designed to execute, not to prolong
3) Quality control with evidence
It is not enough to simply ‘plan inspections’. It is necessary to ensure:
- infrastructure resources
- operqative windows
- evidence (photographs, measurements, records)
- closure of findings and verification
If control is only carried out ‘on paper’, then control does not exist.
4) Climate risk built into the heart of the system
Heavier rainfall requires properly sized drainage systems, extraordinary post-event inspections, instrumentation at critical points, and preventive operating protocols. Hotter temperatures require checking for expansion, fatigue, materials, and fire risk.
The control plan must be dynamic, not an annual ritual.
In light of the above
We can continue to act as if these episodes were inevitable surprises, or accept a simple truth: critical infrastructure is deteriorating every day, and extreme weather only makes visible what was already wrong.
If incidents take months or years to resolve, if the backlog grows, if controls are only partially complied with and findings are not verified, we are not managing quality. We are accumulating risk until nature decides when to make us pay.
And here is the phrase that should be displayed in every investment committee:
Quality does not prevent the weather; it prevents the weather from breaking your system.

About the author: Xavier Conesa
Xavier Conesa is the Chief Growth Officer at Kapture.io, a company specialising in the digitisation of data. in the digitisation of data through its QMS.

