The Same Fault Every Quarter: Reliability Tracing, Failure Data and How One Zinc Line Ended the Repeat Cycle
Equipment Maintenance Technician - Senior

The Same Fault Every Quarter: Reliability Tracing, Failure Data and How One Zinc Line Ended the Repeat Cycle

A hoist limit switch failed every quarter, without fail, and each replacement took a shift of production. The switch was fine; the chain that pulled it was not.

This case shows how senior maintenance breaks repeat faults: failure logging, data review, root cause at the system level and redesigns that end the cycle.

MTBF
Tracked per component family
Repeat rule
3 identical faults = investigate
Failure log
Time, mode, context
Root cause
System level, not part level
Verification
2 quarters clean run
The Same Fault Every Quarter: Reliability Tracing, Failure Data and How One Zinc Line Ended the Repeat Cycle
The Same Fault Every Quarter: Reliability Tracing, Failure Data and How One Zinc Line Ended the Repeat Cycle
An industrial setting, likely a factory or manufacturing facility
An industrial setting, likely a factory or manufacturing facility
An industrial machine, specifically a piece of equipment with a round, cylindrical shape that appears to be...
An industrial machine, specifically a piece of equipment with a round, cylindrical shape that appears to be...

Common Mistakes and How to Avoid Them

MistakeWhy It HappensPractical Fix
1. Replacing the same part againThe cycle simply repeatsLog and review before replacing
2. Failure records kept in memoryPatterns stay invisibleKeep a digital failure log
3. Blaming the componentSystem causes returnTrace load, motion and environment
4. No MTBF dataPriority becomes guessworkCalculate MTBF per component type
5. Fix without verificationThe next quarter repeatsVerify over two quarters minimum
6. Maintenance without load dataReal stress stays unknownMeasure current, travel, temperature
7. Same spare, same specificationRedesign is never consideredUpgrade the spec after 3 repeats
8. Faults hidden at shift handoverHistory is lostMake the log part of handover
9. Treating symptoms as causesThe root cause survivesAsk why five times to the system

Best Practices

  • Log every failure with time, mode and context
  • Investigate any fault that repeats three times
  • Calculate MTBF per component family
  • Trace repeat faults to load, motion and environment
  • Verify fixes over two quarters before closing
  • Upgrade specification after confirmed repeats

Implementation Roadmap

StepFrequency
Review the failure log and open actionsWeekly
Update MTBF per componentMonthly
Investigate 3-repeat faultsMonthly
Verify completed fixesQuarterly
Review spares against failure dataQuarterly

Repeat fault elimination flow

  1. Log - Time, mode, context. (Unlogged faults repeat invisibly.)
  2. Count - MTBF per component. (Without counts, priority is opinion.)
  3. Trace - Load, motion, environment. (The replaced part is rarely the root.)
  4. Verify - Two quarters of clean run. (Early closing invites a repeat.)

Three repeats of the same fault mean the system is the problem. | MTBF turns maintenance from reaction into planning.

Reliability reference data

Reference values for reliability tracking on plating equipment.

ParameterReferenceWhy It Matters
Repeat rule3 identical faults = investigationPrevents chronic failures
MTBF targetPer component, e.g. hoist limit switch over 12 monthsBelow target means redesign
Failure log fieldsDate, mode, load, downtimeContext beats part name
Verification window2 quarters clean runOne quarter is not proof
Spare coverageTop 20 failures = 80 percent of sparesStock follows failure data
OEE target85 percent or higherReliability is the base of OEE
Root cause depth5 whys to the systemStops the cycle at the source

Failure log format

DateComponentModeDowntime (min)
Jan 12Hoist limit switchStuck open45
Apr 03Hoist limit switchStuck open50
Jul 21Hoist limit switchStuck open48
Oct 09Hoist limit switchStuck open42

MTBF board

ComponentFailures per quarterMTBF (months)Action
Hoist limit switch1.03Redesign
Filter pump seal0.310Monitor
Heater contactor0.130OK
Rectifier fan0.56Upgrade spec

Case 1 - the limit switch that failed every quarter

Scenario. A hoist limit switch failed every quarter for two years; each replacement cost a shift, and the spare stock was always depleted.

Action. The failure log showed the switch died during high-travel cycles, and a current clamp revealed the hoist motor drawing 30 percent above rating while the chain guide was worn.

Result. The guide was rebuilt and the switch was upgraded to a sealed, higher-rating model; no failure occurred in the next five quarters.

Case 2 - heater contactors welding in a Spanish powder-coating plant

Scenario. A powder-coating oven in Spain had contactor welding every season change; each event shut the oven for a day.

Action. Reliability data linked the welds to cold-start current spikes, so the team added soft-start control and logged oven start current.

Result. Contactors lasted 14 months instead of 3, and the oven's OEE rose from 78 percent to 89 percent.

Frequently Asked Questions

When is a fault a repeat?

When the same component fails the same way three times in a reasonable window, investigate the system.

What is MTBF?

Mean time between failures for a component family; it shows which components deserve redesign budgets.

Why does the same part keep failing?

Usually the load, motion or environment around the part, not the part itself.

What goes in a failure log?

Date, component, failure mode, load at failure and downtime; context is what makes patterns visible.

How do I find the root cause?

Ask why five times, measure the real load with clamps and loggers, and check the surroundings before the component.

How long do I verify a fix?

At least two quarters of clean running; one quarter is often luck.

How does reliability affect OEE?

OEE cannot reach 85 percent while the same fault steals a shift every quarter; reliability is the foundation.

Which spares do I stock?

Stock follows failure data: the components with the lowest MTBF get the most spares and redesign attention.

What is a chronic failure?

A fault that returns despite replacement; the system is producing the fault faster than parts can fix it.

What Would You Like to Solve?

Describe the component that keeps failing, its failure mode and your line's load pattern, and share the last three failure dates.

Downtime Cut by a Third: Maintenance Planning, Spare Parts Discipline and the Work Order System That Ran
Pump Cavitation That Looked Like a Chemistry Problem: Filtration Troubleshooting Before You Blame the Bath