3 September 2026

Why Does the Same Machine Keep Failing?

A Guide to Identifying and Eliminating Bad Actor Equipment

In every manufacturing facility, some machines are discussed more often than others. “That pump again.” “This motor stopped last month too.” “Didn’t we just replace this bearing?” These comments can slowly become part of the maintenance team’s daily routine.

Over time, the team may become highly familiar with the machine. They may predict which component needs replacement before creating a work order. At first, this may appear to show experience. In reality, it can signal a much deeper problem. The company may not be solving the failure itself. Instead, it may be repairing the same issue repeatedly.

In maintenance and reliability, repeatedly failing equipment is often called a “Bad Actor.” These assets often consume excessive maintenance resources and create disproportionate operational burdens. Identifying Bad Actor equipment means more than finding the most failure-prone machine. The goal is to make chronic problems visible. It also means breaking the cycle that drains maintenance labor and resources. This cycle also consumes spare parts, budget, and production capacity.

What Is Bad Actor Equipment?

A Bad Actor is an asset that performs worse than similar equipment. It may fail more often or cause longer periods of downtime. It may also generate higher maintenance costs or consume excessive maintenance resources.

For example, one pump may fail more often than nine identical pumps. In that case, the pump should be investigated. However, failure frequency alone does not provide the full picture. One machine may fail ten times each year. Each failure, however, may be resolved within only a few minutes. Another machine may fail only twice during the same period. Yet each failure could stop production for several hours.

A true Bad Actor analysis should answer one important question: Which equipment creates the greatest reliability loss for the business? Failure frequency should therefore be evaluated alongside several other factors. These include total downtime, maintenance labor, and spare parts consumption. Total maintenance cost and recurring failure modes should also be considered. MTBF, MTTR, and equipment criticality should also be included.

The Real Risk Is Not the Failure , It Is Normalizing the Failure

A machine failing once may be considered normal. If the same problem occurs for a second time, questions should start being asked. If it happens for the third, fourth, or fifth time, the issue is no longer just a maintenance task. It has become a reliability problem.

One of the biggest traps in maintenance operations is normalization. As technicians become increasingly familiar with a recurring problem, intervention time may become shorter. A maintenance team might eventually say, “That bearing always fails. Replacing it only takes 40 minutes anyway.” This is a dangerous mindset because the question should no longer be “How quickly can we replace the bearing?” but “Why do we keep replacing the same bearing?”

Maintenance performance is not only about restoring equipment quickly. True maintenance maturity means preventing the same failure from occurring again.

Why the “Replace and Restart” Cycle Does Not Solve the Problem

When production is stopped, the maintenance team’s natural priority is to restore the line as quickly as possible. The failed component is identified, replaced, the machine is restarted, and the work order is closed. From an operational perspective, the problem appears to have been resolved, but this approach may only eliminate the symptom of the failure.

Consider a bearing that repeatedly fails. The bearing itself may indeed be damaged, but the underlying cause could be misalignment, poor lubrication, excessive load, incorrect installation, vibration, high operating temperature, incorrect bearing selection, or improper equipment use. If the root cause is not eliminated, the newly installed bearing will be exposed to exactly the same conditions. In other words, a new component is placed back into the same old problem, and the failure cycle begins again.

How Can You Identify Bad Actor Equipment?

If you ask a maintenance manager, “Which machines cause you the most trouble?” they can usually name a few immediately. This operational knowledge is valuable, but investment decisions, maintenance strategy changes, and equipment replacement decisions should not rely on memory alone. Bad Actor analysis should be driven by data.

1. Identify the Equipment with the Most Failure Records

Start by reviewing failure records over a defined period. For example, analyze the last 6 or 12 months and determine which equipment generated the highest number of failures, how many different failure types occurred on the same asset, and how frequently the same failure code appeared. The key point is not only the total number of failures; the recurring failure pattern should be examined closely.

2. Add Downtime to the Analysis

Ten minor failures may create less operational impact than one major failure. For this reason, total downtime should also be included in the assessment.

EquipmentNumber of FailuresTotal DowntimeInitial Assessment
Pump A124 hoursFrequent failures
Conveyor B518 hoursHigh production impact
Motor C811 hoursChronic failure candidate
Fan D32 hoursLower priority

This example demonstrates why looking only at failure frequency can be misleading. Pump A may fail more often, but Conveyor B may create a much greater production loss.

3. Include Maintenance Costs

Some equipment can be repaired quickly but repeatedly consumes expensive spare parts. Therefore, the analysis should also include spare parts used, technician labor hours, external service costs, emergency purchases, overtime, and, where possible, production losses. This allows maintenance teams to move beyond asking “Which machine fails the most?” and instead ask, “Which chronic problem creates the greatest overall cost for the business?”

4. Review the MTBF Trend

MTBF is one of the key indicators used to identify chronic equipment problems. For example, if the time between failures decreases from 180 days to 130 days, then 90 days, and eventually 45 days, this may indicate that equipment reliability is steadily deteriorating.

When individual work orders are reviewed separately, this trend may be difficult to detect. However, once historical maintenance data is analyzed collectively, the behavior of the equipment becomes much more visible.

Pareto Analysis: Focus on the Equipment That Creates the Biggest Problems

Maintenance teams do not have unlimited time or budgets. Therefore, applying the same level of improvement effort to every asset is neither practical nor efficient. A Pareto-based approach can be a strong starting point in Bad Actor management. It helps identify the relatively small group of assets responsible for a significant portion of total failures, downtime, or maintenance costs.

Equipment can be ranked from highest to lowest based on total failure count, total downtime, or total maintenance cost. If the same assets repeatedly appear at the top of these lists, the reliability team has identified a clear area for investigation. However, an important distinction must be made: Pareto analysis shows where the problem is; it does not explain why the problem exists. That requires root cause analysis.

You Found a Bad Actor. What Should You Do Next?

Identifying the most problematic asset is only the beginning. The goal is not to create a list of “problem machines” and close the report, but to permanently remove the source of the problem.

The first step is to determine whether the repeated failures are actually the same failure. Generic descriptions such as “motor failure” are not sufficient. One incident may involve overheating, another a bearing issue, and another an electrical connection problem. Accurate failure classification is therefore essential.

The next step is to identify the physical cause of failure. Which component failed? Was it the bearing, belt, sensor, seal, or electrical connection? The real analysis then begins with another question: Why did that component fail?

For example, a bearing may have failed because of insufficient lubrication. Why was lubrication insufficient? The maintenance interval may have been inappropriate. Why was the maintenance interval inappropriate? The actual operating conditions of the equipment may not have been reflected in the maintenance plan. At this point, the problem is no longer a “bearing issue.” The real problem is the maintenance strategy. In another case, the same investigation may lead to incorrect operation, poor installation, improper equipment selection, insufficient capacity, or a design issue.

More Maintenance Is Not Always the Solution to Chronic Failures

When equipment fails repeatedly, one of the easiest reactions is to increase maintenance frequency. A monthly inspection becomes biweekly, then weekly. But this is not always the right solution.

The underlying cause may have nothing to do with insufficient maintenance. If the equipment was incorrectly sized, more frequent maintenance will not solve a design problem, operators are using the equipment outside its operating limits, additional inspections will not remove the root cause. If the wrong component is being used, extending the maintenance checklist will not change the outcome.

The correct approach is not simply to perform more maintenance. It is to apply the right maintenance intervention to the right problem.

A Simple Bad Actor Prioritization Model

Organizations can create their own scoring model based on their operational priorities. For example, each asset can be scored from 1 to 5 across the following criteria:

Criterion15
Failure frequencyVery lowVery high
Downtime impactMinimalMajor production impact
Maintenance costLowVery high
Recurring failuresNoneContinuous
Equipment criticalityLowCritical

Equipment with the highest total scores can then be prioritized for Bad Actor investigation. This does not need to be a universal or standardized formula. What matters is creating an objective, measurable, and repeatable decision-making model that reflects the organization’s own production structure and risk profile.

Why Is Bad Actor Analysis Difficult Without a CMMS?

Historical data is the foundation of Bad Actor management. But if maintenance records are spread across Excel files, technicians’ personal notes, messaging applications, paper work orders, and separate folders, analyzing chronic failures systematically becomes extremely difficult.

A technician may say, “That machine is always breaking down.” This operational experience is valuable, but management needs more than that. How many times did it fail? What caused the failures? How many hours of downtime did it create? How much did it cost? Which spare parts were used? Is the time between failures decreasing? Did previous maintenance actions actually solve the problem?

Without structured and centralized maintenance data, these questions are difficult to answer reliably.

Next Steps

Have you received sufficient information about “Why Does the Same Machine Keep Failing?”

repairist is here to help you. We answer your questions about the Maintenance Management System and provide information about the main features and benefits of the software. We help you access the repairist demo  and even get a free trial.

Aybit Technology Inc.

Frequently Asked Questions

What is Bad Actor equipment?

Bad Actor equipment refers to assets that fail more frequently than comparable equipment. Assests that create excessive downtime, generate high maintenance costs, or consume a disproportionate amount of maintenance resources.

How can repeatedly failing equipment be identified?

Failure frequency, total downtime, MTBF, MTTR, maintenance costs, spare part consumption, and recurring failure modes should be analyzed together.

Is the machine with the highest number of failures always a Bad Actor?

No. Failure frequency should be considered together with downtime, cost, and production impact. A critical machine that fails less frequently may still create a significantly greater operational loss.