Section Overview
The failure analysis chapter asked why one part broke; this chapter widens the lens to how failure behaves across whole populations and over time — the reliability view of failure (failure-analysis-purpose-and-process). It rests on three ideas. The failure rate is how often a part fails per unit of operating time — the basic measure of how failure-prone something is, and not a fixed number but one that changes across a product's life. The shape of that change is the bathtub curve: the rate runs high at first as weak, defective units die young in infant mortality, falls to a low and roughly constant level through the long useful-life period where failures are random, and climbs again as the population enters wear-out and reaches the end of its physical life (capacitor-and-inductor-failure-modes). The third idea is the most used and most misread number in reliability, MTBF — the mean time between failures, a population statistic for the average operating time between failures across a large population in useful life, and not a promise that any one unit lasts that long. These are tools read back onto the bench: where in the bathtub curve a failure sits — infant mortality of a fresh replacement, a random useful-life event, or the leading edge of wear-out whose same-age siblings are close behind — tells a repairer what diagnosis alone cannot, and a repair should restore a device's reliability rather than quietly degrade it (common-analog-failure-modes).
Why This Matters
This is the section that turns "why did this part fail" into "how reliably was it ever going to last, and what did my repair do to that." This matters because a single failure and the reliability of a population are different questions: failure analysis explains one broken part, but reliability describes how a whole population of parts fails over time, and a technician who can move between the two sees both the failure in front of them and the pattern it belongs to (failure-analysis-purpose-and-process). This matters because the bathtub curve tells a repairer where in life a failure sits: a device that fails within days of a repair may be suffering the infant mortality of a weak replacement part; a device that fails after years may be entering wear-out, in which case the part that failed is likely the first of several same-age siblings to go, and that changes what a thorough repair considers (capacitor-and-inductor-failure-modes). It matters because the single most common reliability error is misreading MTBF: a hundred-thousand-hour mean time between failures is routinely taken to mean a unit will last a hundred thousand hours, when it means something entirely different — one failure per hundred thousand aggregate operating hours across a large population — and acting on the wrong reading leads to wrong expectations about what a device or a part should do. And it matters because a repair can quietly lower a device's reliability: fitting a cheaper part with a worse failure rate than the one it replaces restores function while degrading reliability, so a repairer who understands reliability restores it rather than trading it away (common-analog-failure-modes). Read the failure rate, know the shape of the bathtub curve, and use MTBF for what it actually says — and reliability becomes a working tool rather than a number on a datasheet.
Required Prerequisites
Before starting this section, you should have completed:
- Failure Analysis — Purpose and Process — the investigation of why one part failed, which this section widens into the statistical behavior of failure across populations and time.
- Capacitor and Inductor Failure Modes — concrete component wear and failure, the physical grounding for the wear-out region of the bathtub curve, where aging parts reach the end of their life.
Recommended Consumables
- A datasheet quoting a reliability figure — a FIT number, an hours-between-failures value, or a per-hour failure figure — the specification this section teaches you to read correctly, so the numbers are practiced on a real part rather than in the abstract.
- A set of same-age electrolytic capacitors from one device — the classic wear-out population, so the "when one goes, its siblings are close behind" reasoning is seen on real parts.
- A notebook for a reliability log — recording when and where failures occur, because reliability is a pattern seen across many observations, not a single reading.
Recommended Practice Hardware
- An aging device with a wear-out failure — one old enough to be entering the wear-out region, so a failure can be reasoned about as the leading edge of its population rather than an isolated event.
- A device that failed soon after a repair — an infant-mortality case, so the early-life end of the bathtub curve is seen where a weak replacement part died young.
- Two functionally identical parts with different failure-rate specifications — so the judgment of whether a replacement preserves or degrades reliability is practiced on a real choice.
Real-World Applications
Reliability is the lens that tells a repairer what a single failure belongs to. A technician holding a datasheet MTBF reads it as one failure per that many aggregate operating hours across a population, not as the lifespan of the unit on the bench (failure-analysis-purpose-and-process). A repair that fails within days is recognized as likely infant mortality — a weak replacement part dying young — rather than a fresh mystery (common-analog-failure-modes). A bench facing a wear-out failure in an aging device treats the failed part as probably the first of several same-age siblings to reach wear-out, and considers them rather than only the one that failed (capacitor-and-inductor-failure-modes). And a repairer choosing between two replacement parts weighs their failure rates so the repair restores the device's reliability rather than lowering it. The confusions this prevents: an MTBF mistaken for a lifespan, an infant-mortality failure treated as a random one, a wear-out failure fixed in isolation while its siblings follow, and a repair that silently degrades the reliability it was meant to restore.
Common Challenges
- A mean-time number is read as a lifespan. A large mean-time-between-failures figure is taken as how long one unit will last — when it is a population average, not a single unit's life (failure-analysis-purpose-and-process).
- The rate of failure is assumed constant. A part is treated as equally failure-prone at every age — when the bathtub curve shows the rate high early, low in the middle, and rising at the end (capacitor-and-inductor-failure-modes).
- A wear-out failure is fixed in isolation. The one failed part is replaced and the job called done — when its same-age siblings are near their own wear-out and likely to follow (common-analog-failure-modes).
- A repair degrades reliability unnoticed. A cheaper part restores function — while its worse failure rate quietly lowers the reliability the repair was meant to restore.
Safety Notes
Risk Level: Low. This section is concepts and reading — failure rates, the bathtub curve, and MTBF — and the standing bench law and the earlier chapters' cautions apply to any hardware handled while exploring them, especially aging parts near wear-out.
- Aging parts near wear-out are the hazardous ones — old electrolytics hold charge and vent, stressed parts fail abruptly; discharge, inspect, and handle worn hardware with the diagnostics volumes' care.
- Reliability numbers are not a safety assurance — a low failure rate describes a population, not the safety of the specific part in hand, which is established by physical checks regardless.
- The bench cautions apply to any hardware handled — this section is reading and reasoning, but the parts it points at are governed by the standing bench law.
Professional Tips Before Starting
- Think population, not unit. Reliability describes how a group of parts fails over time — so read every reliability number as a statement about a population, not the unit in your hand (failure-analysis-purpose-and-process).
- Never read a mean-time figure as a lifespan. A mean time between failures is an average across a population — not a promise that one unit lasts that long (common-analog-failure-modes).
- Place the failure along the life curve. Ask whether a failure is infant mortality, a random useful-life event, or wear-out — because where it sits changes what the repair should consider (capacitor-and-inductor-failure-modes).
- When one wear-out part goes, look at its siblings. Same-age parts under the same stress reach wear-out together — so a wear-out failure is rarely the last of its kind on the board.
- Restore reliability, don't trade it away. Weigh a replacement part's failure rate against the original's — so the repair leaves the device as reliable as it found it, or better.
Reading Reliability Over a Product's Life
Failure Rate — How Often, and Why It Changes
The starting idea of reliability is the failure rate: how often a part fails per unit of operating time (failure-analysis-purpose-and-process). It is the basic measure of how failure-prone something is, often quoted for electronics in failures per billion hours — the FIT, or failures-in-time, unit — or folded into a single headline number for a whole device. The essential thing to understand about the failure rate is that it is not a constant. The same part is not equally likely to fail at every age: a population of fresh parts fails at one rate, the survivors settle to a much lower rate through the middle of their lives, and old parts fail at a rising rate as they wear out — so a single failure-rate figure, quoted without saying which part of life it describes, is an average that can hide very different behavior at the two ends. This is why reliability is not captured by one number but by how the number changes over time, and it is that change, plotted across a population's whole life, that gives reliability its most recognizable picture. Understanding the failure rate as a rate that varies with age, rather than a fixed property of a part, is the foundation everything else in this section builds on — the shape of that variation, and the population statistic that summarizes its flat middle, are simply two ways of reading the same underlying truth that a part's proneness to fail is a function of where it is in its life.
The Bathtub Curve — Infant Mortality, Useful Life, Wear-Out
Plot the failure rate of a population against age and it traces a shape reliability engineering calls the bathtub curve, for its high ends and long low middle (capacitor-and-inductor-failure-modes). It has three regions, and each means something different to a repairer. The first is infant mortality: early in life the failure rate is high but falling, as the weak and defective units — a marginal solder joint, a part with a latent manufacturing flaw — die young and are winnowed out of the population. This is why a failure soon after manufacture, or soon after a repair, is often not a random event but a weak unit revealing itself early. The second region is useful life: once the weak units are gone, the survivors settle into a long period where the failure rate is low and roughly constant, and the failures that do occur are random — a surge, a stress, a chance event — rather than systematic. This flat middle is the working life of the product, and it is the region MTBF describes. The third region is wear-out: as the population ages, its members begin to reach the end of their physical life — electrolytics dry out, solder joints fatigue, contacts wear — and the failure rate climbs again (common-analog-failure-modes). The curve's value to a technician is that it places a failure in time: a failure at the left edge is likely a weak part, a failure in the middle is likely random, and a failure at the right edge is likely wear-out — and a wear-out failure carries a warning, that the same-age siblings sharing the part's history are climbing the same rising edge behind it.
MTBF — A Population Statistic, Not a Lifespan
The number most often quoted for reliability, and most often misunderstood, is MTBF — the mean time between failures (failure-analysis-purpose-and-process). Read precisely, it is the average operating time between failures across a large population during its useful-life period — the flat middle of the bathtub curve — and during that period, where the failure rate is roughly constant, it is simply the reciprocal of that rate. A device with a constant failure rate of one failure per hundred thousand operating hours has, by definition, a hundred-thousand-hour MTBF. Here is the misreading that must be undone: a hundred-thousand-hour MTBF does not mean a unit will last a hundred thousand hours. It means that across a large population running in useful life, failures occur on average once per hundred thousand aggregate operating hours — a thousand units running for a hundred hours each would, on average, see one failure, even though not one of them is anywhere near the end of its life. MTBF says nothing about wear-out, which is a different region of the curve entirely; a part can have a very high MTBF and still wear out on a schedule, because the statistic describes the random failures of the flat middle, not the end of physical life. The related figure MTBF has a non-repairable cousin, MTTF — mean time to failure — used for parts that are replaced rather than fixed, but the same caution governs both: they are population averages, not the lifespan of the unit in your hand. Use MTBF for what it says — the random-failure rate of a population in its useful life — and never for what it does not: a promise about how long the part on your bench will last.
Common Mistakes
- Reading a mean-time figure as a guaranteed lifespan. A large mean-time-between-failures number is taken as how long one unit will run before dying — when it is a population's random-failure average in useful life, silent about wear-out (failure-analysis-purpose-and-process).
- Treating the rate of failure as constant over life. A part is assumed equally failure-prone at every age — when the bathtub curve shows a high-falling start, a low-flat middle, and a rising end (capacitor-and-inductor-failure-modes).
- Ignoring the siblings of a wear-out failure. One worn part is replaced and the board declared healthy — when its same-age, same-stress neighbors are on the same rising wear-out edge (common-analog-failure-modes).
- Trading reliability away in a repair. A cheaper part restores function — while its worse failure rate leaves the device less reliable than before the repair.
- Confusing a random failure with a systematic one. A single useful-life failure is treated as a design flaw, or a wear-out pattern as bad luck — when the region of the curve tells which it is.
Troubleshooting Guidance
- A unit failed far sooner than its rated mean time — you are reading MTBF as a lifespan: a mean time between failures is a population's useful-life average, not a single unit's guaranteed run, and an early failure is likely infant mortality or wear-out, neither of which MTBF describes (failure-analysis-purpose-and-process).
- A repaired device failed again within days — suspect infant mortality of the replacement: the left edge of the bathtub curve is where weak parts die young, so a fresh part failing early is often a defective unit rather than a repeat of the original fault (common-analog-failure-modes).
- You fixed one worn part and another failed soon after — the wear-out region was read too narrowly: same-age parts under the same stress reach wear-out together, so a wear-out failure points at its siblings, not just itself (capacitor-and-inductor-failure-modes).
- A device seems less reliable after a repair than before — check the replacement part's failure rate: a part with a worse failure rate than the original degrades the device's reliability even when it restores function, so a repair is judged on reliability, not only on working.
Verification & Testing Methods
Confirm your grasp of component reliability before continuing:
- [ ] I can define the failure rate as how often a part fails per unit time and explain why it is not constant over life.
- [ ] I can describe the bathtub curve — infant mortality, useful life, and wear-out — as the shape of failure rate across a product's life.
- [ ] I can explain what MTBF actually means as a population statistic and why it is not a single unit's lifespan.
- [ ] I can place a failure in the infant-mortality, useful-life, or wear-out region and reason about what it implies.
- [ ] I can judge whether a replacement part preserves or degrades a device's reliability.
Then try the practice exercises below — reading and reasoning only; scenarios differ from the quiz.
Practice Exercises
- Read a mean-time figure correctly (5 minutes, a datasheet). Find a quoted MTBF or FIT figure and state precisely what it says — one failure per so many aggregate operating hours across a population — and, in your own words, what it does not say about the lifespan of a single unit (failure-analysis-purpose-and-process).
- Place failures along the life curve (5 minutes, a set of cases). For several failures — one soon after manufacture or repair, one mid-life, one in an aging device — place each in infant mortality, useful life, or wear-out, and give the reasoning that puts it there (capacitor-and-inductor-failure-modes).
- Reason from a wear-out failure to its siblings (5 minutes, an aging device). For a wear-out failure in an old device, identify the same-age, same-stress parts likely to be on the same rising edge, and state why a thorough repair considers them and not only the one that failed (common-analog-failure-modes).
- Judge a repair's effect on reliability (5 minutes, a part choice). Compare two replacement parts by failure rate and decide which preserves the device's reliability and which would degrade it, so the choice is made on reliability and not only on fit and function.
These core skills — reading an MTBF for what it says, placing a failure on the bathtub curve, reasoning to a wear-out failure's siblings, and judging a repair's effect on reliability — are tested in the Chapter Quiz at the end of this chapter, where a score of 80% is required to continue.
Key Takeaways
- The failure rate is how often a part fails per unit of operating time, and it is not constant — the same part is more failure-prone when young and when old than through the long middle of its life (failure-analysis-purpose-and-process).
- Plotted over life, the failure rate traces the bathtub curve: high and falling in infant mortality as weak units die young, low and roughly constant through useful life where failures are random, and rising through wear-out as parts reach the end of their physical life (capacitor-and-inductor-failure-modes).
- MTBF — the mean time between failures — is a population statistic for the average operating time between failures in useful life, and emphatically not a promise that one unit will last that long: a hundred-thousand-hour MTBF is one failure per hundred thousand aggregate population-hours, not a hundred-thousand-hour lifespan.
- Where a failure sits on the bathtub curve tells a repairer what diagnosis cannot — infant mortality of a fresh part, a random useful-life event, or the leading edge of wear-out whose same-age siblings are close behind (common-analog-failure-modes).
- A repair should restore a device's reliability, not degrade it — which means not fitting a part whose failure rate is worse than the one it replaced, even when the cheaper part restores function.
Skills Learned
After completing this section, you can:
- Explain failure rate, the bathtub curve, and MTBF and how they relate.
- Correct the common misreading of MTBF as a guaranteed lifespan.
- Place a failure in the infant-mortality, useful-life, or wear-out region and reason from it.
- Recognize when a wear-out failure implies its same-age siblings are near their own end.
- Judge whether a replacement part preserves or degrades a device's reliability.
Glossary Additions
New terms introduced in this section:
- failure rate — how often a part or device fails per unit of operating time, the basic measure of how failure-prone it is, often quoted for electronics in failures per billion hours — the FIT, or failures-in-time, unit. Its defining feature for reliability is that it is not constant over a product's life: a population of fresh parts fails at one rate, the survivors settle to a much lower and roughly constant rate through the middle of their lives, and aging parts fail at a rising rate as they wear out. Because of this, a single failure-rate figure quoted without saying which part of life it describes is an average that can conceal very different behavior at the two ends, which is why reliability is captured not by one number but by how the failure rate changes over time — the shape traced by the bathtub curve.
- bathtub curve — the characteristic shape of a population's failure rate plotted against age, named for its high ends and long low middle. It has three regions: infant mortality, early in life, where the failure rate is high but falling as weak and defective units die young and are winnowed out; useful life, the long flat middle, where the failure rate is low and roughly constant and the failures that occur are random rather than systematic; and wear-out, late in life, where the failure rate climbs again as parts reach the end of their physical life through drying electrolytics, solder fatigue, and worn contacts. Its value to a repairer is that it places a failure in time — a failure at the left edge is likely a weak part, one in the middle likely random, and one at the right edge likely wear-out, carrying the warning that the failed part's same-age siblings are on the same rising edge behind it.
- MTBF — mean time between failures, the most quoted and most misread number in reliability: the average operating time between failures across a large population during its useful-life period, and during that period, where the failure rate is roughly constant, simply the reciprocal of that rate. The critical caution is that MTBF is a population statistic, not a single unit's lifespan — a 100,000-hour MTBF means one failure per 100,000 aggregate operating hours across many units, not that any one unit will run 100,000 hours before dying, and it says nothing about wear-out, which is a different region of the bathtub curve entirely. Its non-repairable cousin, MTTF — mean time to failure — is used for parts that are replaced rather than fixed, but both are population averages and neither promises how long the specific part in hand will last.
Suggested Next Sections
Must read next:
- Thermal Cycling and Fatigue — Section 6.2 takes up the dominant wear-out mechanism in real boards: the fatigue that repeated heating and cooling drives into solder joints and components, and why it sets the practical life of much of what a technician sees fail.
Recommended:
- Writing a Failure Analysis Report — the close of the failure analysis chapter, whose single-failure investigation this reliability view widens into the behavior of failure across populations and time.
- Common Analog Failure Modes — the concrete failure modes that populate reliability statistics, the individual ways parts fail that, across a population, sum into a failure rate and a bathtub curve.