Building a Predictive Maintenance Loop in Maximo: From Failure History to Total Cost of Ownership
Predictive maintenance programs stall when the loop between failure history, asset risk, scheduling priority, executed work, and total cost of ownership is left open. This article walks that loop end to end in Maximo terms, with the configuration obligations and organizational realities that…
Building a Predictive Maintenance Loop in Maximo: From Failure History to Total Cost of Ownership
Predictive maintenance programs fail in a predictable way. Someone buys sensors, connects them to a historian, produces a stream of condition alerts, and then discovers that nobody scheduled anything in response. Eighteen months later the sensors are still running, the alerts are still firing, and the maintenance strategy in Maximo looks exactly as it did before the program started.
The technology was never the hard part. The hard part is the loop: the closed path that takes failure history, converts it into risk, converts risk into scheduling priority, converts priority into executed work, and then feeds the resulting failure and repair data back into the model that started it. IBM Community practitioners have been writing about this loop explicitly, and the most useful framing currently circulating breaks it into roughly ten concrete steps from failure history through risk, priority, and total cost of ownership. This article walks that loop end to end in Maximo terms, with the configuration obligations and the organizational realities that determine whether it holds.
Whether you are running Reliability Centered Maintenance on a critical asset class, standing up a condition monitoring program on rotating equipment, or trying to make an existing predictive maintenance investment pay for itself, the loop is the same. What differs is where in the loop your organization happens to be weakest.
Step One: Establish Failure History That Is Actually Usable
Every predictive maintenance program claims to be data-driven. Very few have failure history that will support the claim. The gap between those two statements is where most programs quietly stall.
Usable failure history requires three things in Maximo. First, failure codes that discriminate between meaningful failure modes rather than collapsing everything into a single "equipment failure" code that tells you nothing about mechanism. Second, work orders that are actually closed with the failure data populated, rather than closed with the default code because the technician wanted the record off their queue. Third, a problem and cause structure that maps to how your reliability engineers think about the asset, not to how someone happened to configure the code tables in 2014.
The practical starting point is an audit of your failure code usage distribution. Pull a year of corrective work orders and count records per failure code. In most organizations the result is a long tail of codes with a single use and a handful of codes absorbing seventy percent or more of the volume. That distribution tells you two things: which codes are doing real work, and which parts of your code structure nobody understands well enough to use.
Then look at the unpopulated rate. If a substantial share of corrective work orders close with no failure code at all, your predictive model is training on a subset that may or may not be representative. The fix here is not a reporting mandate. It is making the code selection fast and obvious in the technician's workflow, and giving supervisors visibility into closure quality so the data issue surfaces while the job is still fresh.
There is a temptation to skip this step and go straight to analytics. Resist it. A predictive model built on ambiguous failure codes will produce confident-looking output that sends maintenance effort to the wrong assets, and the credibility cost of a wrong prediction is much higher than the cost of a month spent cleaning up code usage.
One more consideration: distinguish between failure history and condition history. Failure history tells you what broke and how. Condition history tells you what the asset looked like as it approached failure. A robust predictive loop needs both, which in Maximo terms means meter readings, condition monitoring results, and inspection findings recorded against the asset in a way that can be joined to the eventual failure event. If your inspection findings live in a spreadsheet, that join does not exist.
Step Two: Convert Failure History Into Asset Risk
Raw failure history is an operational record. Risk is a decision construct. The conversion between them is where reliability engineering earns its keep, and in Maximo it is usually expressed through criticality analysis and failure mode effects analysis.
The standard approach assigns each asset a criticality informed by consequence and likelihood. Consequence covers safety, environmental impact, production loss, and cost. Likelihood draws on the failure history you just validated. The result is a ranked asset population that tells you where to spend predictive effort, because predictive maintenance is expensive and cannot be applied uniformly.
The common configuration mistake is treating criticality as a one-time classification exercise rather than a living attribute. An asset that was critical when the plant ran at full throughput may be far less critical after a product line change, and an asset that was routine may have become a single point of failure after a redundancy was removed. Build the review cadence into your reliability process and record criticality changes as governed events, because a predictive program driving off stale criticality will systematically misallocate sensors and analysis hours.
The more sophisticated layer is failure mode specific. Criticality ranks assets. Failure mode analysis ranks the specific ways an asset can fail, and that is the level at which predictive technique selection actually happens. Vibration analysis detects developing bearing faults and imbalance and misalignment. Oil analysis catches wear metals and contamination. Thermography finds electrical and thermal anomalies. Ultrasonic detection catches leaks and early-stage lubrication problems. Each technique has a detection envelope, and applying a technique outside that envelope produces either false confidence or alert fatigue.
Maximo's condition monitoring and APM capabilities support this structure through asset and location hierarchies, meter and characteristic definitions, and the association of condition data to specific measurement points. Getting the measurement point model right is a substantial configuration effort, but it is the difference between condition data that is analytically useful and condition data that is a stream of numbers with no failure mode association. Budget real time for it, and involve the reliability engineers who will consume the output in the design, not just in the review.
There is also a data quality dimension that predicts program success better than almost any other factor. Where failure history, criticality, and failure mode associations are all maintained, predictive programs reach benefit quickly. Where any one of the three is neglected, the program tends to stall in a pilot state. Fixing all three before scaling is the highest-leverage investment available.
Step Three: Turn Risk Into Scheduling Priority
Risk that does not change the schedule is academic. This step is where most predictive maintenance programs break, and it breaks for organizational rather than technical reasons.
The mechanism is straightforward. Condition monitoring produces an indication that an asset is degrading. That indication needs to enter the same planning and scheduling pipeline as every other work demand, with a priority that reflects the risk. In Maximo, that usually means work order generation from a condition or inspection finding, with priority driven by the asset criticality and the severity of the indication, routed into the planner's queue rather than directly to a technician.
The failure mode to watch for is the alert queue that exists outside the work management system. If condition alerts accumulate in an APM dashboard, an email inbox, or a monitoring center's own tooling and never become Maximo work orders, then the loop is open and the program cannot demonstrate value. Practitioners who have made this work consistently report the same design choice: condition indications create work orders in Maximo, and Maximo is the only queue anyone works from.
Priority assignment deserves explicit design. A simple high, medium, low scheme will not survive contact with a real asset population. What works better is a priority derived from criticality and degradation rate together, because two assets with identical severity can warrant very different urgency depending on how fast the condition is deteriorating and how much consequence the failure carries. Build the priority logic into your condition to work order automation and document it, so that when a planner disagrees with a priority they can see the reasoning rather than overriding an opaque number.
The planner's role is the human judgment layer in this step. Automation should produce a well-prioritized, well-specified work demand. The planner decides when it fits in the schedule, whether it can be bundled with planned work already going to that asset, and whether the window available at the site permits the corrective action. This is genuine value-adding work, but only if the planner receives condition-derived demands that are specific enough to schedule. A work order that says "vibration high on pump" is not schedulable. One that says "accelerating bearing defect indication on drive-end bearing, trending to failure threshold, recommend inspection and probable bearing replacement within two weeks" is.
That specificity requirement has an organizational consequence. Someone has to write the analysis, and that someone needs both reliability expertise and enough time to do the interpretation. Programs that assume the monitoring system will generate schedulable work orders automatically usually discover that the interpretation step is the real bottleneck, and staffing for it is what separates a functioning loop from an alert factory.
Step Four: Execute the Work and Preserve the Feedback Path
Predictive maintenance only pays when the recommended work actually happens, and the work actually happening is where many programs discover that their authority model does not support the strategy.
Consider the case where reliability analysis recommends replacing a bearing that is trending toward failure. The asset is running. Production wants it to keep running. The window is tight. The planner has to trade off the cost of an unplanned failure against the cost of a planned intervention that interrupts output. Organizations that have not established the authority for condition-driven work to claim schedule priority will systematically defer these jobs, and the program will revert to reactive behavior with extra sensors attached.
Resolving this is a governance task. Decide in advance which criticality levels carry the authority to force a planned outage window, who makes that call when production objects, and what evidence the reliability function needs to present. The existence of that decision path, and the fact that everyone knows it exists, is what makes condition-driven scheduling credible.
Then the harder half: preserving the feedback path. When the bearing job executes, the technician's findings need to flow back into the record. Was the defect present? Was it as severe as predicted? What was the actual condition? Did the repair resolve it? What did the job actually cost in labor, parts, and downtime?
This is the step that most organizations lose, and it is the step that determines whether the loop is genuinely closed. Without feedback, the model never learns. Thresholds stay where they were set at initial configuration. Technique selection never improves. And the program cannot report benefit with any rigor, because it has no record of what was prevented versus what merely happened to not fail.
In Maximo terms, the feedback path means the corrective work order carries the failure code that was eventually confirmed, the actual repair scope that was executed, and the labor and material that it consumed. That data is what ties a predicted failure to a cost, and that tie is what makes the next section possible.
Make the feedback obligation explicit in the workflow. A condition-driven corrective work order should not be closeable without a disposition: confirmed as predicted, partially confirmed, not confirmed, or no fault found. That single field turns a stream of completed jobs into a training set for your program's improvement.
Step Five: Close the Loop With Total Cost of Ownership
Total cost of ownership is where predictive maintenance stops being an engineering activity and becomes a financial argument, and it is the argument that determines whether your program survives its next budget cycle.
The calculation is conceptually simple. For a given asset or asset class, compare the cost of the predictive program, including sensors, monitoring infrastructure, analysis labor, and planned intervention downtime, against the avoided cost of unplanned failures, secondary damage, emergency labor premiums, and lost production. The difficulty is entirely in the data, which is why the earlier steps are prerequisites rather than preliminaries.
Unplanned failure cost comes from the corrective work orders that actually occurred before the program started, which is why step one matters so much. Avoided failures come from condition-driven interventions where the feedback record confirms a real defect was caught. Planned intervention cost comes from the work orders generated by the program itself. Monitoring and analysis cost comes from the program's own operating budget, which is often tracked somewhere other than Maximo and needs to be joined in deliberately.
Where organizations get into trouble is claiming benefit without the feedback record. A program that reports "prevented twelve failures this year" without being able to show, per event, the condition indication, the planned work order, the confirmed defect, and the counterfactual failure cost is reporting a belief rather than a measurement. Finance will eventually ask for the underlying records, and a program that cannot produce them loses credibility at exactly the moment it needs budget renewal.
There is a second, more strategic use of TCO in Maximo terms: life cycle strategy optimization. Once you can associate failure modes with their detection techniques and their actual costs, you can reason about whether a given technique is worth its expense on a given asset class. Some assets justify continuous vibration monitoring. Others are better served by operator rounds with a handheld route. Others justify run-to-failure with good spares strategy. The TCO framework is what lets you make those calls on evidence, and it is the difference between a predictive program that scales intelligently and one that spreads uniformly because nobody can defend a differentiated allocation.
Finally, feed the TCO findings back into the maintenance strategy itself. If a failure mode is being detected reliably and cheaply, the strategy for that asset should reflect a condition-based maintenance task rather than a time-based one. That strategy change in Maximo is the tangible output of the whole loop, and it is the artifact that demonstrates the program has actually changed how the organization maintains assets rather than merely observing them.
Practical Implications
For reliability engineers, the loop described here is largely your existing practice formalized into a data architecture. The discipline it demands is refusing to let the interpretation step become informal. Analysis that lives in an engineer's notebook cannot be audited, cannot be improved, and cannot be handed to a successor, and it makes the TCO argument impossible to defend.
For Maximo administrators, the configuration surface implied by a functioning loop is broader than most upgrade projects assume: failure code structures that discriminate, condition monitoring measurement points associated with failure modes, automation that converts condition indications into prioritized work orders, and a disposition field that enforces feedback. Each of these is modest individually. Together they are the infrastructure that makes predictive maintenance operational rather than aspirational.
For planners, the loop's value depends heavily on the specificity of the demands it produces. A condition-derived work order that arrives with a clear scope, a recommended window, and a stated consequence is genuinely easier to schedule than the reactive demand it replaces. Push back on vague condition work orders, because accepting them trains the analysis function to stop doing the interpretation work that makes the loop worthwhile.
For maintenance and operations leadership, the governance decisions are yours. Which criticality levels can force a maintenance window, who adjudicates when production objects, and how the program's operating cost is tracked against its avoided failures. Without those decisions, the loop exists on paper but stalls at the first schedule conflict.
For finance, the ask is patience with a specific structure. The first year of a predictive program usually produces more data infrastructure than measurable savings. The savings become legible in years two and three, once failure history, condition data, planned interventions, and cost are all recorded against the same assets in the same system. Fund the loop, not the sensors.
Bottom Line
A predictive maintenance loop in Maximo is a chain of five conversions: failure history into risk, risk into priority, priority into executed work, executed work into confirmed feedback, and confirmed feedback into total cost of ownership. Every link is necessary. A program with excellent condition monitoring and no feedback path cannot demonstrate value. A program with good analytics and no schedule authority cannot execute. A program with strong execution and no cost model cannot survive a budget review.
The encouraging part is that the constraint is organizational rather than technological. The Maximo capabilities needed here are mature and available today. What separates programs that work from programs that stall is whether someone owns the whole loop rather than a segment of it, and whether the organization has decided in advance who has the authority to act when condition data contradicts a production plan.
Start where your loop is weakest. If your failure codes are not discriminating, fix that first. If condition data never becomes a work order, fix the automation before adding sensors. If nobody records what the technician actually found, make disposition mandatory before you claim a single avoided failure. Work the weak link, close the loop, and the measured benefit will follow.