A production line suddenly comes to a halt. Within seconds, several alarms appear: temperature too high, insufficient flow, unusual vibrations, product out of tolerance.
The failure is detected almost immediately. Yet its cause remains unknown.
Operators check the parameters. Technicians review alarm histories, inspect different components, and try to reach a colleague who knows the equipment well. Meanwhile, production remains interrupted.
In many plants, the challenge is therefore no longer simply detecting an anomaly. It is understanding quickly why it occurred so it can be resolved and prevented from happening again.
Repairing a part may sometimes take only a few minutes. Identifying the right part to repair can take much longer. Every minute spent on diagnosis increases the cost of unplanned downtime.
Summary
Generative AI, supported by the strengths of large language models, can accelerate industrial failure diagnosis by identifying the correct source of a problem, even while being mistaken about why it occurred. A fast diagnosis is only useful if it is reliable and understandable to users. It must be possible to verify that the AI’s reasoning reflects the actual operation of the machine and that its answers remain consistent over time. Data quality, process knowledge, and human validation remain essential
What Is the True Cost of Production Downtime?
The cost of production downtime is not limited to the value of the units that were not produced.
When a line stops, several categories of expenses and losses can accumulate at the same time:
- Lost production capacity
- Idle employees and equipment
- Urgent mobilization of maintenance teams
- Overtime
- Expedited purchasing of parts
- Scrap or out-of-tolerance products
- Cleaning, recalibration, and restart operations
- Delivery delays
- Contractual penalties
- Damage to customer relationships
Depending on the process involved, there may also be consequences related to safety, the environment, or regulatory compliance.
The scale of the problem is considerable. According to Siemens’ The True Cost of Downtime 2024 report, Fortune Global 500 companies worldwide may lose nearly US$1.4 trillion each year because of unplanned downtime. This amount would represent approximately 11% of their revenue.
An international survey published by ABB indicates that average losses associated with unplanned downtime approach US$170,000 per hour. Nearly half of respondents report equipment-related interruptions at least once a month.
These averages naturally vary by industry, site, and equipment. One hour of downtime on a small production cell does not have the same consequences as one hour of downtime in an automotive plant, a steel mill, or a continuous process facility.
In every case, the clock keeps running until production has returned to a stable state.
A Failure Includes Several Distinct Phases
To understand where losses occur, it is useful to break an industrial shutdown into several stages.
-
Detection: A control system, sensor, or operator notices an anomaly, such as a product defect, unusual vibration, or rising temperature. An alarm is triggered and the equipment slows down or stops. In modern facilities, this stage can be almost instantaneous.
-
Diagnosis: The team tries to determine what happened and why. It analyzes alarms, reviews data, speaks with operators, and inspects the components that may be involved. The duration of this stage is often difficult to predict.
-
Intervention: Once the probable cause has been identified, the team performs the required adjustment, cleaning, repair, or replacement.
- Restart:The equipment is restarted and tested. The team must confirm that the parameters are stable and that product quality meets requirements. Companies generally track the total duration of a shutdown. However, they less often measure the proportion of time spent on each of these stages. This distinction is important. Once the cause is known, the repair may be relatively simple. The longest and most uncertain part often occurs before the intervention, while technicians are still trying to determine which component to inspect.
Diagnostic Time: An Often-Underestimated Cost
During an unplanned shutdown, the maintenance team does not always receive a clear message indicating that a filter is clogged or a pump has failed.
Instead, it receives a series of clues:
- Rising temperature
- Decreasing flow
- Increased motor power consumption
- Emerging vibration
- Declining product quality
- Several alarms triggered within seconds of one another
Diagnosis involves reconstructing the sequence of events and determining which event caused the others.
This investigation may be slowed by several factors.
Data may be spread across the control system, maintenance software, local files, and paper reports. Timestamps are not always synchronized. Previous interventions may be poorly documented. The terminology used in reports may not match the names assigned to sensors or components.
The team may also need to proceed through trial and error by resetting a parameter, replacing a sensor, checking a motor, or dismantling a subsystem.
Each of these checks extends the shutdown.
Diagnostic time is therefore a major opportunity for reducing losses. Even when it is impossible to prevent every failure, an organization can reduce its impact by identifying the true cause more quickly.
Why More Sensors Do Not Always Mean a Faster Restart
Modern plants generate an increasing amount of data. Equipment can continuously measure pressure, flow, temperature, vibration, electricity consumption, speed, position, and product quality.
These data significantly improve the ability to detect anomalies. However, they do not automatically provide a diagnosis. An alarm primarily answers the question:
Is there a problem?
A diagnosis must answer a more complex question:
Why is it happening?
Consider an increase in temperature. It could be caused by a faulty sensor, insufficient lubrication, a malfunctioning pump, a clogged filter, excessive friction, or an electrical failure.
The sensor that triggers the alarm identifies the visible symptom. It is not necessarily responsible for the failure.
Several variables may also change at nearly the same time because they are affected by a common cause. A decrease in cooling flow may result in higher temperatures, dimensional variation, and declining product quality.
The challenge is therefore not simply to collect more data. It is necessary to:
- Put the data in the correct sequence
- Distinguish causes from consequences
- Understand relationships between components
- Compare the incident with previous situations
- Eliminate hypotheses that are incompatible with the machine’s actual operation
Without this process knowledge, an abundance of alarms may even make technicians’ work more difficult.
The Hidden Costs of an Incorrect Diagnosis
A slow diagnosis prolongs downtime. An incorrect diagnosis can create even more costly consequences.
For example, the team may replace a sensor that was functioning properly, modify a parameter that was not involved, or inspect a subsystem unrelated to the failure.
These interventions mobilize personnel, consume parts, and delay the inspection of the actual problem.
In some cases, the equipment restarts temporarily even though the root cause has not been corrected. The same failure then occurs again a few hours or days later. The organization experiences another shutdown, another investigation, and another intervention.
Inaccurate diagnoses can also lead to:
- Increased spare-parts inventory
- Unnecessary preventive replacements
- More frequent maintenance than required
- Loss of confidence in data and tools
- Increasingly lengthy troubleshooting procedures
- Greater dependence on a small number of experts
Even when the team eventually identifies the correct component, the reasoning behind the diagnosis is not always documented. The next technician who encounters the same issue must then repeat the investigation.
Reducing downtime costs therefore also depends on the ability to retain and reuse what was learned during previous incidents.
When Diagnosis Depends on a Few Experts
In many organizations, experienced technicians possess highly detailed knowledge of the equipment.
They know which alarms truly matter, which sounds are unusual, which components to check first, and which combinations of signals correspond to a known failure.
Some of this knowledge is documented in procedures. Another portion is based on experience accumulated over many years.
This dependence becomes problematic when the expert:
- Works on another shift
- Is located at another plant
- Is absent when the incident occurs
- Leaves the organization
- Retires
The shortage of skilled workers intensifies this challenge. A less experienced team may have access to the same alarms and manuals without having the reasoning required to connect the clues quickly.
Formalizing operational knowledge therefore becomes an investment in business continuity. It makes failure sequences, component relationships, priority checks, and known exceptions explicit.
This documentation can then support training, procedure improvement, and the development of diagnostic support tools.
Three Ways to Reduce Diagnostic Time
No single technology can eliminate unplanned downtime on its own. However, companies can act on three complementary areas.
1. Make Maintenance Information More Accessible
The data required for diagnosis often already exist, but they are scattered.
A technician may need to consult the following separately:
- Control-system alarms
- Sensor trends
- Work orders
- Replacement histories
- Intervention reports
- Technical manuals
- Internal procedures
Connecting these sources reduces the time spent searching for information and makes it easier to compare an incident with previous cases.
Centralization does not necessarily require adding sensors to every piece of equipment. In many cases, the data already available in automation and maintenance systems provide a sufficient starting point.
2. Structure Process Knowledge
Data must be interpreted according to how the equipment actually operates.
It is therefore useful to document:
- The function of each component
- Dependencies between subsystems
- Normal operating ranges
- Physically possible sequences
- Known failure scenarios
- Checks performed by experts
This work helps distinguish simple correlation from a cause-and-effect relationship. It also helps preserve the knowledge of experienced employees.
3. Use Artificial Intelligence as an Investigation Support Tool
Artificial intelligence can accelerate certain tasks that currently require significant time.
A system can, for example:
- Reconstruct the chronology of an incident
- Summarize the events preceding a shutdown
- Retrieve similar cases
- Compare multiple information sources
- Suggest possible causes
- Recommend an order of inspection
- Present results in accessible language
Large language models can make it easier to interact with this information. For example, a technician could ask which events preceded an increase in temperature or which previous interventions were associated with a specific alarm.
However, AI must remain a support tool. A clear and detailed answer is not necessarily correct. The recommendation must be based on reliable data, reflect the machine’s actual configuration, and be verifiable by a human.
How Can You Determine Whether Your Plant Can Reduce Diagnostic Time?
Before investing in a new solution, an organization can begin by considering a few simple questions:
- Which equipment causes the most costly shutdowns?
- How long does it generally take to detect, diagnose, repair, and restart?
- What proportion of downtime is spent searching for the cause?
- Do certain failures occur repeatedly?
- Are the actual causes documented after incidents?
- Are the required data accessible and synchronized?
- Does diagnosis depend heavily on a few individuals?
- Do technicians sometimes replace parts without certainty?
- Are there enough historical cases to compare incidents?
- Can teams explain why a particular cause was selected?
These questions help identify the equipment and scenarios for which better use of data could generate value quickly.
It is generally preferable to begin with a limited scope: one critical machine, relatively frequent failures, and data that are already available. The organization can then measure its current diagnostic time, structure the necessary knowledge, and evaluate the resulting gains.
Reducing Downtime Begins with Better Understanding Its Causes
Unplanned downtime is among the most costly operational risks in manufacturing. Its impact does not depend only on the severity of the failure. It also depends on how quickly the organization can understand what happened and make the right decision.
Sensors and alarms make it possible to recognize quickly that a problem exists. They are not always sufficient to identify its source.
By making data more accessible, structuring technicians’ expertise, and using artificial intelligence as a diagnostic support tool, manufacturers can reduce the time spent investigating and accelerate the return to production.
But how can a truly reliable diagnostic assistant be developed? What data should it receive? How can it be prevented from confusing a cause with a symptom, and how can organizations ensure that its explanation reflects the actual operation of the machine?
Our reference guide outlines the possibilities offered by artificial intelligence and large language models for diagnosing industrial failures. It also presents the main pitfalls to avoid, the conditions for success, and the steps required to evaluate this approach in a production environment.
Download the reference guide to learn how AI can support the diagnosis of unplanned downtime while keeping human expertise at the centre of decision-making.