The hardest part of an enterprise reliability strategy isn't choosing between reliability-centered maintenance, condition-based monitoring, or standard preventive maintenance. It's that all three coexist — at every site, across hundreds of asset types — and most organizations never build a clear, defensible way to decide which method goes where.
So what fills that vacuum? Every site does its own thing. One region runs everything on calendar PMs because that's what the last director set up. Another bought vibration sensors for pumps that fail once every seven years. A third has a spreadsheet someone's calling "RCM" that's really just a list of PM tasks with fancier column headers. When corporate tries to roll up reliability numbers, nothing reconciles, and nobody can explain why identical air handlers are maintained three different ways across three buildings.
This is about the system that fixes that — not a single method, but the decision layer sitting above the methods. How you assign RCM, CBM, or PM per asset class, what governance gates keep the roll-out honest, and how you sequence sites and build competency so the whole thing doesn't collapse when the original champion moves on.
Why method selection breaks at multiple sites (but not at one)
At a single facility, method selection is almost informal. A good reliability engineer knows the equipment, knows which pumps are cranky, and makes reasonable calls. It works because one person holds all the context.
Multiply that across 15, 40, or 80 sites and the informal approach becomes the actual problem. Context doesn't travel. What one engineer "just knows" about a chiller isn't written anywhere, so the next site reinvents it — usually worse. And because there's no shared logic, you can't audit whether decisions were reasonable. You just end up with a pile of maintenance programs that grew like weeds.
A pattern worth naming: the failure usually isn't wrong method selection at the asset level. It's inconsistent selection across identical assets. Two sites with the same rooftop units, same climate, same run hours — one on quarterly PMs, one on condition monitoring — and neither can explain why. That inconsistency kills your ability to benchmark, standardize spares, or negotiate service contracts at scale.
There's a second, quieter breakdown. Method decisions made at launch almost never get revisited. An asset gets tagged "PM" in year one and stays PM for a decade even after it moves into a critical process, even after failure data clearly justifies condition monitoring. Nobody owns re-evaluation, so the program calcifies.
The decision matrix: assigning RCM, CBM, or PM per asset class
The core artifact you need is a decision matrix that makes method selection repeatable instead of a judgment call every time. It doesn't eliminate engineering judgment — it constrains it so decisions stay consistent and explainable.
Eliminate downtime with proactive maintenance.
Openfixit helps you plan, track, and complete maintenance efficiently—maximizing asset reliability.
- Centralized asset management
- Automated maintenance scheduling
- Inventory and parts tracking
No credit card required
-
Consequence of failure — safety, compliance, production or occupancy impact, and the cost of the failure event itself.
-
Failure pattern — does the asset degrade in a detectable way (bearings, insulation, heat), or does it fail randomly with no useful warning?
-
Detection economics — can you actually sense the degradation cheaply enough to justify it?
Here's how that plays out across common facility asset classes:
| Asset class | Consequence if it fails | Failure pattern | Recommended primary method | Why |
|---|---|---|---|---|
| Central chillers / large AHUs | High (occupancy + energy) | Detectable degradation | CBM + targeted RCM | Vibration/thermal/oil analysis pays off; failure is expensive |
| Critical pumps (fire, process) | High | Detectable | CBM | Early warning is cheap relative to downtime |
| Standard exhaust/supply fans | Medium | Partly detectable | PM with condition checks | Sensors rarely justify the cost at volume |
| Electrical switchgear / transformers | Very high, low frequency | Slow degradation | RCM-driven inspection + thermography | Rare but catastrophic; needs structured analysis |
| Lighting, small motors, dampers | Low | Random/wear | PM or run-to-failure | Monitoring costs more than the failure |
| Elevators / life-safety systems | Very high, regulated | Mixed | PM (compliance-driven) + RCM overlay | Method partly dictated by code, not just risk |
| Emergency generators | Very high | Detectable under test | CBM + rigorous PM | Must work on demand; testing is the detection method |
The mistake that comes up constantly: teams push everything toward CBM because it sounds advanced. But condition monitoring on a $400 exhaust fan that fails predictably every four years is a net loss. Method selection is an economics problem as much as a reliability one. Getting PM intervals right on assets that stay on PM matters just as much — and there's real money in optimizing preventive maintenance intervals to balance labor and parts costs rather than defaulting to whatever the OEM manual says.
One more note on RCM: full RCM analysis is slow and expensive. Reserve it for the small slice of assets where failure consequences are severe and failure modes aren't obvious. For everything else, a streamlined version is enough to assign a defensible method without a six-week workshop per asset class.
Governance gates: keeping the roll-out from drifting
A decision matrix on paper does nothing if sites can quietly ignore it. What holds it together is a set of governance gates — checkpoints where a method assignment has to pass review before it goes live.
-
Gate 1 — Asset criticality sign-off. Before any method is assigned, the criticality rating has to be reviewed. Most errors originate here — a mis-rated asset gets the wrong method automatically. This gate depends on clean asset data, which is why getting your asset hierarchies and naming conventions standardized has to happen before method assignment, not after.
-
Gate 2 — Method assignment review. A reliability lead confirms the matrix was applied correctly and that any deviation from the default has a written justification. "We chose CBM for these fans because they're in a cleanroom" is fine. No reason at all is not.
-
Gate 3 — Instrumentation readiness (for CBM only). Before an asset moves to condition-based monitoring, someone confirms that sensors, baselines, and alarm thresholds actually exist and are validated. Assigning CBM without working detection is one of the most common quiet failures — the asset is "on CBM" but nobody's actually watching anything.
-
Gate 4 — Program change control. Any later change to an asset's method has to go through the same review, so the program can evolve without silently fragmenting.
Require one person to own criticality sign-offs so it doesn't become a checkbox exercise.
Governance gates aren't bureaucracy for its own sake. Reliability programs fail slowly and invisibly. A skipped gate doesn't cause a problem this week — it causes a mystery failure eighteen months later that nobody can trace to a decision. Gates create the paper trail that makes the program auditable and defensible when finance asks why you're spending on monitoring.
Staged roll-out: sequencing sites so you don't overwhelm the org
Rolling out a new reliability method to every site simultaneously is how these programs die. The people running maintenance still have day jobs. Ask 40 sites to reclassify assets, install sensors, learn new workflows, and change how they close work orders all at once, and you'll get shallow compliance everywhere and real adoption nowhere.
Stage it instead. A sequence that tends to work:
-
Pilot cohort (2–4 sites). Pick sites with a mix of strong and average maintenance maturity — not just your best performers, or you'll build something that only works for A-teams. Prove the matrix, the gates, and the workflows here first.
-
Reference build-out. Turn pilot learnings into templates
pre-filled criticality guides per asset class, standard CBM thresholds, a concise "how we decide" document. The goal is that site ten doesn't have to think as hard as site one.
-
Regional waves (5–8 sites per wave). Roll out by region, not by asset count. Regional consistency makes spares, contracts, and benchmarking far easier than scattered adoption.
-
Backfill and re-evaluation. Return to earlier sites to check drift and re-rate assets whose criticality changed. This is the step most organizations skip and later regret.
A staged CBM roll-out especially benefits from proving value small before committing capital at scale. The discipline of a tight pilot — clear scope, minimal sensors, real go/no-go criteria — matters, and it's worth borrowing that structure from a proper predictive maintenance pilot designed to prove value before you roll out across dozens of buildings.
Roll-out readiness checklist — a site shouldn't enter a wave until:
-
[ ] Asset register is clean and consistently named
-
[ ] Criticality ratings reviewed and signed off
-
[ ] A named site owner for the reliability program exists
-
[ ] Baseline PM compliance is above a reasonable floor (a site drowning in backlog can't absorb a new program)
-
[ ] Technicians have completed the relevant competency tier (see below)
-
[ ] CBM assets, if any, have validated sensors and thresholds
That backlog point deserves emphasis. Piling a reliability transformation onto an already-overwhelmed team just buries them further. Stabilize first, transform second.
Competency and training milestones: the part everyone underfunds
Method selection and governance mean nothing if the technicians executing the work can't distinguish a bearing signature from noise, or don't understand why an asset's method changed. The biggest predictor of whether a CBM program survives isn't the sensors — it's whether the people reading the data actually know what they're looking at.
Tie training to milestones rather than a single onboarding session. A tiered structure works because it matches skill level to responsibility:
Tier 1 — Foundation (all technicians). Understand the three methods, why assets are assigned differently, and how to close work orders with the data quality the program depends on. Sloppy CBM inspection records corrupt your entire detection layer.
Tier 2 — Condition monitoring basics (CBM sites). How to take readings correctly, recognize obvious out-of-range conditions, and escalate appropriately. Not full analysis — reliable data capture and first-line judgment.
Tier 3 — Analysis and diagnosis (reliability leads). Interpreting trends, distinguishing real degradation from sensor drift, adjusting thresholds, and feeding findings back into method decisions.
Tier 4 — Program governance (site and regional owners). Running the gates, applying the decision matrix, managing change control, and owning re-evaluation cycles.
The trap to avoid: training everyone to Tier 1 and stopping because it's cheap and fast. You end up with a program that generates data nobody's qualified to act on. Six months in, alarms get ignored because "they're always false," and you're back to run-to-failure with expensive sensors attached.
Milestones should gate participation. A site doesn't operate CBM assets until Tier 2 coverage exists. A region doesn't self-govern until owners reach Tier 4. Competency becomes a prerequisite, not an afterthought.
When this approach makes sense — and when it doesn't
When it makes sense: You're running enough sites that inconsistency is measurably costing you — duplicated spares, incomparable reliability numbers, service contracts you can't standardize. You have critical assets where failure consequences justify structured analysis. And you have leadership willing to fund the governance layer, not just the sensors.
When it's overkill: A handful of sites with a stable, experienced team and mostly low-consequence assets don't need this level of formality. The informal approach genuinely works at small scale, and forcing enterprise governance onto it just adds friction without payoff.
Who should not start here: Anyone whose asset data is a mess. If your register is incomplete, inconsistently named, or full of duplicates, method selection is essentially impossible — you'll be assigning methods to assets that don't reflect reality. Clean the foundation first. Every gate described above assumes you can actually trust your asset list.
A realistic scenario
A regional operator running around 30 mixed-use facilities had, by their own admission, "three different maintenance philosophies depending on who you asked." Chillers were on CBM at some sites, quarterly PM at others. Reliability reporting to the board was largely fiction because nothing was comparable.
They didn't try to fix everything at once. They built the decision matrix, ran a four-site pilot over roughly a quarter, and enforced the criticality sign-off gate consistently. First surprising result: about a third of their assets tagged for condition monitoring didn't actually justify it — sensors were watching low-consequence, predictable-wear equipment. Reclassifying those to PM freed up monitoring budget for switchgear and critical pumps that warranted the attention.
Eighteen months in, unplanned failures on critical assets had declined steadily across the sites they'd reached — not a dramatic overnight number, but a consistent trend they could actually attribute to specific method changes. More importantly, when the reliability lead at one site left, the next person could pick up the program because decisions were documented and gated, not locked in someone's memory. That continuity was arguably worth more than the failure reduction itself.
Bringing the pieces together
An enterprise reliability strategy isn't a method — it's the coordination layer that decides which method, where, when, and by whom. The decision matrix provides consistency. Governance gates provide enforcement and auditability. Staged roll-out produces adoption that's real instead of cosmetic. Competency milestones produce people who can actually run what you built.
A modern CMMS earns its place by holding all of this together — carrying the criticality ratings, enforcing gates as workflow rather than PDF policy, tracking technician competency tiers, and flagging assets due for method re-evaluation before the program quietly fragments. The tooling doesn't make the decisions. It keeps decisions consistent across every site, every shift change, and every year the program has to survive.
Here's a simple workflow view of how the decision matrix, governance gates, staged roll-out, and training milestones connect in practice.
Most reliability programs don't fail because someone picked RCM over CBM. They fail because the logic behind those choices was never written down, never enforced, and never revisited. Build the decision layer, and the method choices largely take care of themselves.
Ready to optimize your maintenance operations?
Join 2,000+ facilities using Openfixit to reduce unplanned outages, extend asset life, and improve operational efficiency.