Failure Analysis Methodology
Most of this handbook has taught how to find and replace a broken part; this chapter teaches the deeper discipline of understanding why it broke, so that a repair fixes the cause rather than the symptom and does not simply wait for the failure to return. Section 5.1 establishes the purpose and the process — failure analysis as a repeatable, evidence-based method that traces a failure from how it manifests, through the physical mechanism that produced it, to the root cause behind it, and closes with the corrective action that addresses that cause. Section 5.2 takes root cause analysis to its own depth, the disciplined techniques for separating the true originating cause from the symptoms and intermediate failures that mask it. Section 5.3 turns to the analytical toolkit at Professional depth, the distinction between non-destructive methods that preserve the evidence and destructive ones that consume it, and when each is justified. Section 5.4 closes the chapter on communicating the result: writing a failure analysis report that records the evidence, the reasoning, the root cause, and the corrective action in a form another technician or engineer can trust and act on.
4 sections · 92 minutes of reading.
0/4- 5.1Failure Analysis — Purpose and ProcessAlmost everything this handbook has taught up to now has aimed at one thing: find the broken part and replace it. This chapter opens a discipline that goes a layer deeper and asks the question replacement alone never answers — not what failed, but why — because a repair that swaps a failed component without discovering the reason it failed has not fixed the fault, it has only reset the clock on it. Failure analysis is the systematic investigation of why a component or a system failed, and its purpose for a repair technician is intensely practical: a blown part is very often the victim of something else, and a new part dropped into the same conditions meets the same fate, so the difference between a repair that lasts and one that fails again next week is whether the cause behind the symptom was found. The section's central idea is that this investigation is a process, not a guess — a repeatable, evidence-based method that anyone can follow to a defensible conclusion rather than an intuition that works only for the technician who happens to have it. That process moves along a chain the section makes explicit. It begins with the failure mode, the way the failure shows itself — an open, a short, a burned package, a cracked joint, a drift out of tolerance. It works inward to the failure mechanism, the physical process that actually produced that mode — the thermal runaway, the dielectric breakdown, the electromigration, the solder-joint fatigue that is what physically happened to the part. And it traces back to the root cause, the originating reason the mechanism was set in motion — the overstress event, the thin design margin, the manufacturing defect, the wear of age, the environment — which is the thing a repair must actually address. The section frames the disciplined steps that walk this chain — preserving the evidence before it is disturbed, characterizing the failure, determining the mechanism, tracing to the cause, and closing with a corrective action that fixes the cause rather than the symptom — and it sets up the rest of the chapter, which takes root cause analysis, the destructive and non-destructive toolkit, and the written report each to its own depth. The point it leaves the technician with is the one the whole chapter is built on: a failure is a question, and analysis is the discipline of answering why, because only the answer makes a repair permanent.AdvancedLow Risk23 min read
- 5.2Root Cause Analysis TechniquesThe previous section named the root cause as the end of the failure chain and the thing a corrective action must reach; this section is about actually reaching it, because the root cause is the hardest link to find and the easiest to stop short of. The difficulty is that causes chain together, and the first cause a technician finds is almost never the root. A driver transistor failed — why? Because it overheated — why? Because its heatsink was loose — why? Because a screw backed out — why? Because the wrong thread-locking was used at assembly. Each answer is a cause, but only the last is one that, corrected, keeps the failure from returning, and a technician who stops at the first or second answer fixes an immediate cause and leaves the origin in place. Root cause analysis is the set of disciplined techniques for drilling past those immediate and contributing causes to the originating one, and this section teaches the ones a repair technician actually uses. The five whys is the simplest and most powerful: ask why the failure happened, then why that happened, and again, each answer becoming the next question, drilling down the causal chain until the answer is a cause that can be acted on and that going further would take outside the technician's control. Cause categorization keeps the search from tunnel vision, structuring the possible causes across families — the part itself, the design, the process or workmanship, the environment, the way the device was used — so that no whole category of cause is overlooked because the first plausible answer captured all the attention. And fault-tree reasoning works the problem from the failure downward through the logical combinations of conditions that could have produced it, the same top-down fault-isolation logic troubleshooting already uses, turned onto the question of cause. Threaded through all of them is the discipline the previous section insisted on: every step must be supported by evidence rather than assumed, because a chain of whys built on guesses reaches a guessed root, and a corrective action aimed at a guessed root misses. The section also draws the line the whole method turns on — the difference between an immediate cause, a contributing cause, and the root cause — and the stopping rule that tells a technician when they have arrived: when correcting the cause would prevent recurrence and pursuing it further would leave what a repair can change. The point it leaves is that finding the root cause is not luck or intuition but technique, and technique is teachable, repeatable, and defensible in a way a hunch never is.AdvancedLow Risk23 min read
- 5.3Destructive vs. Non-Destructive AnalysisA failure analysis lives on evidence, and this section is the toolkit for gathering it — organized around the single distinction that governs how the whole toolkit is used: whether a method preserves the evidence or consumes it. Every analytical technique falls into one of two families. Non-destructive analysis examines a part while leaving it intact — visual inspection and microscopy, X-ray that sees inside a package without opening it, thermal imaging that finds a hot fault, electrical characterization that measures behavior — and its defining virtue is that the part survives, so the examination can be repeated, checked, and built upon, and the evidence remains available for whatever the analysis needs next. Destructive analysis, by contrast, consumes the part to reveal what non-destructive methods cannot reach — decapsulation that removes a chip's package to expose the die, cross-sectioning that cuts and polishes a part to show its internal structure — and it is powerful precisely because it goes where nothing else can, but it is irreversible: once a part is decapped or cross-sectioned, the original is gone and cannot be un-cut. From that irreversibility follows the one rule that governs the entire toolkit, non-destructive first, always: because a destroyed part cannot be re-examined, every non-destructive method that could answer the question is exhausted before any destructive one is begun, and the analysis climbs a ladder from the least invasive method to the most, going destructive only at the end and only when the answer genuinely requires seeing inside and is worth the part. The section teaches that ladder, and the discipline that comes with it: that a destructive step, when it is finally justified, is aimed by the non-destructive findings rather than taken blind — the X-ray that located the defect telling the cross-section exactly where to cut — because a destructive step gets one attempt and a blind one often destroys the evidence without revealing it. It also treats the real hazards of the destructive methods honestly, the aggressive acids of chemical decapsulation and the cutting and grinding of cross-sectioning, which demand proper facilities and protection and are often best sent to a lab rather than attempted casually. The point it leaves is that most repair-level failure analysis never needs to destroy anything at all — the non-destructive toolkit answers the question — and the professional discipline is to know the whole ladder, to climb it in order, and to reserve the irreversible step for the rare case where the internal answer is both required and worth the part it costs.ProfessionalMedium Risk23 min read
- 5.4Writing a Failure Analysis ReportThe chapter has taught how to find why a part failed — the process, the root cause techniques, the analytical toolkit — and it closes on the step that turns all of that into something of value to anyone but the technician who did it: writing it down as a report others can trust and act on. A failure analysis that lives only in a technician's head and scattered notes is, for every practical purpose, unfinished: it cannot be verified by anyone else, it cannot be reused when the same failure appears again, and its corrective action cannot be reliably carried out by the person who has to do it, so the report is not paperwork after the real work but the deliverable that makes the real work count. The section teaches what a good report contains and, more importantly, the disciplines that make it trustworthy. Its content follows the analysis: the subject and context of the failure; the objective evidence, meaning what was actually observed and measured, stated as fact; the methods used to gather that evidence, so a reader knows how it was obtained and that the non-destructive-first discipline was honored; the analysis proper, the reasoning that traces the failure from its mode through its mechanism to its root cause; the root cause itself, stated plainly and distinguished from the symptoms and contributors around it; the corrective action that addresses that cause; and an honest statement of confidence and limitations. The disciplines are what separate a report that can be trusted from one that merely asserts. The first and most important is the separation of objective evidence from interpretation — recording what was observed apart from what it is taken to mean, so that a reader can check whether the conclusion actually follows from the facts rather than being smuggled in among them. The second is traceable reasoning, every conclusion tied to the evidence that supports it, so the analysis can be followed and verified rather than taken on faith. The third is an honest confidence statement, stating plainly what is proven, what is inferred, and what could not be determined, because a report that overstates its certainty is worse than one that admits its limits. The section closes the chapter, and with it the technician's whole failure-analysis capability: the purpose and process, the root cause techniques, the analytical toolkit, and now the report that communicates the result — because an analysis that is never written down, or written down in a way no one can trust, is an analysis that was, in the end, wasted.AdvancedLow Risk23 min read
- Chapter Quiz28questions · 80% required to continue