A hoist limit switch failed every quarter, without fail, and each replacement took a shift of production. The switch was fine; the chain that pulled it was not.
This case shows how senior maintenance breaks repeat faults: failure logging, data review, root cause at the system level and redesigns that end the cycle.
Common Mistakes and How to Avoid Them
| Mistake | Why It Happens | Practical Fix |
|---|---|---|
| 1. Replacing the same part again | The cycle simply repeats | Log and review before replacing |
| 2. Failure records kept in memory | Patterns stay invisible | Keep a digital failure log |
| 3. Blaming the component | System causes return | Trace load, motion and environment |
| 4. No MTBF data | Priority becomes guesswork | Calculate MTBF per component type |
| 5. Fix without verification | The next quarter repeats | Verify over two quarters minimum |
| 6. Maintenance without load data | Real stress stays unknown | Measure current, travel, temperature |
| 7. Same spare, same specification | Redesign is never considered | Upgrade the spec after 3 repeats |
| 8. Faults hidden at shift handover | History is lost | Make the log part of handover |
| 9. Treating symptoms as causes | The root cause survives | Ask why five times to the system |
Best Practices
- Log every failure with time, mode and context
- Investigate any fault that repeats three times
- Calculate MTBF per component family
- Trace repeat faults to load, motion and environment
- Verify fixes over two quarters before closing
- Upgrade specification after confirmed repeats
Implementation Roadmap
| Step | Frequency |
|---|---|
| Review the failure log and open actions | Weekly |
| Update MTBF per component | Monthly |
| Investigate 3-repeat faults | Monthly |
| Verify completed fixes | Quarterly |
| Review spares against failure data | Quarterly |
Repeat fault elimination flow
- Log - Time, mode, context. (Unlogged faults repeat invisibly.)
- Count - MTBF per component. (Without counts, priority is opinion.)
- Trace - Load, motion, environment. (The replaced part is rarely the root.)
- Verify - Two quarters of clean run. (Early closing invites a repeat.)
Three repeats of the same fault mean the system is the problem. | MTBF turns maintenance from reaction into planning.
Reliability reference data
Reference values for reliability tracking on plating equipment.
| Parameter | Reference | Why It Matters |
|---|---|---|
| Repeat rule | 3 identical faults = investigation | Prevents chronic failures |
| MTBF target | Per component, e.g. hoist limit switch over 12 months | Below target means redesign |
| Failure log fields | Date, mode, load, downtime | Context beats part name |
| Verification window | 2 quarters clean run | One quarter is not proof |
| Spare coverage | Top 20 failures = 80 percent of spares | Stock follows failure data |
| OEE target | 85 percent or higher | Reliability is the base of OEE |
| Root cause depth | 5 whys to the system | Stops the cycle at the source |
Failure log format
| Date | Component | Mode | Downtime (min) |
|---|---|---|---|
| Jan 12 | Hoist limit switch | Stuck open | 45 |
| Apr 03 | Hoist limit switch | Stuck open | 50 |
| Jul 21 | Hoist limit switch | Stuck open | 48 |
| Oct 09 | Hoist limit switch | Stuck open | 42 |
MTBF board
| Component | Failures per quarter | MTBF (months) | Action |
|---|---|---|---|
| Hoist limit switch | 1.0 | 3 | Redesign |
| Filter pump seal | 0.3 | 10 | Monitor |
| Heater contactor | 0.1 | 30 | OK |
| Rectifier fan | 0.5 | 6 | Upgrade spec |
Case 1 - the limit switch that failed every quarter
Scenario. A hoist limit switch failed every quarter for two years; each replacement cost a shift, and the spare stock was always depleted.
Action. The failure log showed the switch died during high-travel cycles, and a current clamp revealed the hoist motor drawing 30 percent above rating while the chain guide was worn.
Result. The guide was rebuilt and the switch was upgraded to a sealed, higher-rating model; no failure occurred in the next five quarters.
Case 2 - heater contactors welding in a Spanish powder-coating plant
Scenario. A powder-coating oven in Spain had contactor welding every season change; each event shut the oven for a day.
Action. Reliability data linked the welds to cold-start current spikes, so the team added soft-start control and logged oven start current.
Result. Contactors lasted 14 months instead of 3, and the oven's OEE rose from 78 percent to 89 percent.
Frequently Asked Questions
When is a fault a repeat?
When the same component fails the same way three times in a reasonable window, investigate the system.
What is MTBF?
Mean time between failures for a component family; it shows which components deserve redesign budgets.
Why does the same part keep failing?
Usually the load, motion or environment around the part, not the part itself.
What goes in a failure log?
Date, component, failure mode, load at failure and downtime; context is what makes patterns visible.
How do I find the root cause?
Ask why five times, measure the real load with clamps and loggers, and check the surroundings before the component.
How long do I verify a fix?
At least two quarters of clean running; one quarter is often luck.
How does reliability affect OEE?
OEE cannot reach 85 percent while the same fault steals a shift every quarter; reliability is the foundation.
Which spares do I stock?
Stock follows failure data: the components with the lowest MTBF get the most spares and redesign attention.
What is a chronic failure?
A fault that returns despite replacement; the system is producing the fault faster than parts can fix it.
What Would You Like to Solve?
Describe the component that keeps failing, its failure mode and your line's load pattern, and share the last three failure dates.

