Quality & Non-Conformance Management — Edge Cases Reference
Tier 3 reference. Load on demand when handling complex or ambiguous quality situations that don't resolve through standard NCR/CAPA workflows.
These edge cases represent the scenarios that separate experienced quality engineers from everyone else. Each one involves competing priorities, ambiguous data, regulatory pressure, and real business impact. They are structured to guide resolution when standard playbooks break down.
How to Use This File
When a quality situation doesn't fit a clean NCR category — when the data is ambiguous, when multiple stakeholders have legitimate competing claims, or when the regulatory and business implications justify deeper analysis — find the edge case below that most closely matches the situation. Follow the expert approach step by step. Do not skip documentation requirements; these are the situations that end up in audit findings, regulatory actions, or legal proceedings.
Edge Case 1: Customer-Reported Field Failure with No Internal Detection
Situation: Your medical device company ships Class II endoscopic accessories to a hospital network. Your internal quality data is clean — incoming inspection acceptance rate is 99.7%, in-process defect rate is below 200 PPM, and final inspection has not flagged any issues for the last 6 months. Then a customer complaint comes in: three units from different lots failed during clinical use. The failure mode is a fractured distal tip during retraction, which was not part of your inspection plan because design verification showed the material exceeds the fatigue limit by 4x. The hospital has paused use of your product pending investigation.
Why It's Tricky: The instinct is to defend your data. "Our inspection shows everything is within specification. The customer must be using the product incorrectly." This is wrong and dangerous for three reasons: (1) field failures can expose failure modes your test plan doesn't cover, (2) clinical use conditions differ from bench testing, and (3) in FDA-regulated environments, dismissing customer complaints without investigation is itself a regulatory violation per 21 CFR 820.198.
The deeper problem is that your inspection plan was designed around the design verification data, which tested fatigue under controlled, uniaxial loading. Clinical use involves multiaxial loading with torsion, and the fatigue characteristics under combined loading may be significantly different. Your process is "in control" but your control plan has a coverage gap.
Common Mistake: Treating this as a customer-use issue. Sending a "letter of clarification" on proper use without investigating the failure mode. This delays discovery, may worsen patient safety risk, and creates an adverse audit trail if FDA reviews your complaint handling.
The second common mistake: initiating a CAPA that focuses on inspection. "Add fatigue testing to final inspection" solves nothing if the inspection uses the same uniaxial loading condition as the original design verification.
Expert Approach:
- Immediate containment: Place a quality hold on all units of the affected part numbers in your finished goods, distribution, and at the hospital's inventory. Contact the hospital's biomedical engineering department to coordinate the hold — they need serial/lot numbers to identify affected inventory.
- Complaint investigation per 820.198: Open a formal complaint record. Classify for MDR determination — fractured device during clinical use meets the "malfunction that could cause or contribute to death or serious injury" threshold, requiring MDR filing within 30 days.
- Failure analysis: Request the failed units from the hospital for physical failure analysis. Conduct fractographic analysis (SEM if needed) to determine the fracture mode — was it fatigue (progressive crack growth), overload (single-event), stress corrosion, or manufacturing defect (inclusion, porosity)?
- Gap analysis on test coverage: Map the clinical loading conditions against your design verification test protocol. If the failure mode is combined loading fatigue, your uniaxial test would not have detected it. This is a design control gap per 820.30, not a manufacturing control gap.
- Design verification update: Develop a multiaxial fatigue test that simulates clinical conditions. Test retained samples from the affected lots AND from current production. If retained samples fail the updated test, the scope of the problem is potentially every unit shipped.
- Risk assessment per ISO 14971: Update the risk file with the newly identified hazard. Calculate the risk priority based on the severity (clinical failure) and probability (based on complaint rate vs. units in service). Determine whether the risk is acceptable per your risk acceptance criteria.
- Field action determination: Based on the risk assessment, determine whether a voluntary recall, field correction, or enhanced monitoring is appropriate. Document the decision with the risk data supporting it.
- CAPA: Root cause is a gap in design verification testing — the failure mode was not characterized under clinical loading conditions. Corrective action addresses the test protocol, not the manufacturing process.
Key Indicators:
- Complaint rate vs. units in service determines population risk (e.g., 3 failures in 10,000 units = 300 PPM field failure rate)
- Fracture surface morphology distinguishes fatigue, overload, and material defects
- Time-to-failure pattern (all early-life vs. random vs. wear-out) indicates failure mechanism
- If multiple lots are affected, the root cause is likely design or process-related, not material-lot-specific
Documentation Required:
- Formal complaint records per 820.198
- MDR filing documentation
- Failure analysis report with photographs and fractography
- Updated risk file per ISO 14971
- Revised design verification test protocol
- Field action decision documentation (including decision NOT to recall, if applicable)
- CAPA record linking complaint → investigation → root cause → corrective action
Edge Case 2: Supplier Audit Reveals Falsified Certificates of Conformance
Situation: During a routine audit of a casting supplier (Tier 2 supplier to your automotive Tier 1 operation), your auditor discovers that the material certificates for A356 aluminum castings do not match the spectrometer results. The supplier has been submitting CoCs showing material composition within specification, but the auditor's portable XRF readings on randomly selected parts show silicon content at 8.2% against a specification of 6.5-7.5%. The supplier's quality manager initially claims the XRF is inaccurate, but when pressed, admits that their spectrometer has been out of calibration for 4 months, and they've been using historical test results on the CoCs rather than actual lot-by-lot test data.
Why It's Tricky: This is not a simple non-conformance — it's a quality system integrity failure. The supplier did not simply ship nonconforming parts; they submitted fraudulent documentation. The distinction matters because: (1) every shipment received during the 4-month period is now suspect, (2) you cannot trust ANY data from this supplier without independent verification, (3) in automotive, this may constitute a failure to maintain IATF 16949 requirements, and (4) parts from this supplier may already be in customer vehicles.
The containment scope is potentially enormous. A356 aluminum with elevated silicon has different mechanical properties — it may be more brittle. If these castings are structural or safety-critical, the implications extend to end-of-line testing, vehicle recalls, and NHTSA notification.
Common Mistake: Treating this like a normal NCR. Writing a SCAR and asking the supplier to "improve their testing process." This underestimates the severity — the issue is not process improvement but fundamental integrity. A supplier that falsifies data will not be fixed by a corrective action request.
The second common mistake: immediately terminating the supplier without securing containment. If you have weeks of WIP and finished goods containing these castings, cutting off the supplier before you've contained and sorted the affected inventory creates a dual crisis — quality AND supply.
Expert Approach:
- Preserve evidence immediately. Photograph the audit findings, retain the XRF readings, request copies of the CoCs for the last 4 months, and document the supplier quality manager's admission in the audit notes with date, time, and witnesses. This evidence may be needed for legal proceedings or regulatory reporting.
- Scope the containment. Identify every lot received from this supplier in the last 4+ months (add buffer — the calibration may have drifted before formal "out of calibration" date). Trace those lots through your operation: incoming stock, WIP, finished goods, shipped to customer, in customer's inventory or vehicles.
- Independent verification. Send representative samples from each suspect lot to an accredited independent testing laboratory for full material composition analysis. Do not rely on the supplier's belated retesting — their data has zero credibility.
- Risk assessment on affected product. If material composition is out of spec, have design engineering evaluate the functional impact. A356 with 8.2% Si instead of max 7.5% may still be functional depending on the application, or it may be critically weakened for a structural casting. The answer depends on the specific part function and loading conditions.
- Customer notification. In IATF 16949 environments, customer notification is mandatory when suspect product may have been shipped. Contact your customer quality representative within 24 hours. Provide lot/date range, the nature of the issue, and your containment actions.
- Automotive-specific reporting. If the parts are safety-critical and the composition affects structural integrity, evaluate NHTSA reporting obligations per 49 CFR Part 573 (defect notification). Consult legal counsel — the bar for vehicle safety defect reporting is "poses an unreasonable risk to motor vehicle safety."
- Supplier disposition. This is an immediate escalation to Level 4-5 on the supplier ladder. Begin alternate source qualification in parallel. Maintain the supplier on controlled shipping (CS-2, third-party inspection) only for the duration needed to transition. Do not invest in "developing" a supplier that falsified data — the trust foundation is broken.
- Systemic review. Audit all other CoC-reliant incoming inspection processes. If this supplier falsified data, what is the probability that others are as well? Increase verification sampling on other CoC-reliance suppliers, especially those with single-source positions.
Key Indicators:
- Duration of falsification determines containment scope (months × volume = total suspect population)
- The specific spec exceedance determines functional risk (minor chemistry drift vs. major composition deviation)
- Traceability of material lots through your production determines the search space
- Whether the supplier proactively disclosed vs. you discovered impacts the trust assessment
Documentation Required:
- Audit report with all findings, evidence, and admissions
- XRF readings and independent lab results
- Complete lot traceability from supplier through your process to customer
- Risk assessment on functional impact of material deviation
- Customer notification records with acknowledgment
- Legal review documentation (privilege-protected as applicable)
- Supplier escalation and phase-out plan
Edge Case 3: SPC Shows Process In-Control But Customer Complaints Are Rising
Situation: Your CNC turning operation produces shafts for a precision instrument manufacturer. Your SPC charts on the critical OD dimension (12.00 ±0.02mm) have been stable for 18 months — X-bar/R chart shows a process running at 12.002mm mean with Cpk of 1.45. No control chart signals. Your internal quality metrics are green across the board. But the customer's complaint rate on your shafts has tripled in the last quarter. Their failure mode: intermittent binding in the mating bore assembly. Your parts meet print, their parts meet print, but the assembly doesn't work consistently.
Why It's Tricky: The conventional quality response is "our parts meet specification." And technically, that's true. But the customer's assembly process is sensitive to variation WITHIN your specification. Their bore is also within specification, but when your shaft is at the high end of tolerance (+0.02) and their bore is at the low end, the assembly binds. Both parts individually meet print, but the tolerance stack-up creates interference in the worst-case combination.
The SPC chart is not lying — your process is in statistical control and capable by every standard metric. The problem is that capability indices measure your process against YOUR specification, not against the functional requirement of the assembly. A Cpk of 1.45 means you're producing virtually no parts outside ±0.02mm, but if the actual functional window is ±0.01mm centered on the nominal, your process is sending significant variation into a critical zone.
Common Mistake: Dismissing the complaint because the data says you're in spec. Sending a letter citing your Cpk and stating that the parts conform. This is technically correct and operationally wrong — it destroys the customer relationship and ignores the actual problem.
The second mistake: reacting by tightening your internal specification without understanding the functional requirement. If you arbitrarily cut your tolerance to ±0.01mm, you increase your scrap rate (and cost) without certainty that it solves the assembly issue.
Expert Approach:
- Acknowledge the complaint and avoid the "we meet spec" defense. The customer is experiencing real failures. Whether they're caused by your variation, their variation, or the interaction of both is what needs to be determined — not assumed.
- Request the customer's mating component data. Ask for their bore SPC data — mean, variation, Cpk, distribution shape. You need to understand both sides of the assembly equation.
- Conduct a tolerance stack-up analysis. Using both your shaft data and their bore data, calculate the assembly clearance distribution. Identify what percentage of assemblies fall into the interference zone. This analysis converts "your parts meet spec" into "X% of assemblies will have interference problems."
- Evaluate centering vs. variation. If the problem is that your process runs at 12.002mm (slightly above nominal) and their bore is centered low, the fix may be as simple as re-centering your process to 11.998mm — shifting the mean away from the interference zone without changing the variation.
- Consider bilateral specification refinement. Propose a joint engineering review to establish a tighter bilateral tolerance that accounts for both process capabilities. If your Cpk for ±0.01mm around a recentered mean is still > 1.33, the tighter spec is achievable.
- Update your control plan. If the assembly-level functional requirement is tighter than the print tolerance, your control plan should reflect the actual functional target, not just the nominal ± tolerance from the drawing.
- This is NOT a CAPA. This is a specification adequacy issue, not a non-conformance. The correct vehicle is an engineering change process to update the specification, not a CAPA to "fix" a process that is operating correctly per its current requirements.
Key Indicators:
- Your process Cpk relative to the FUNCTIONAL tolerance (not drawing tolerance) is the key metric
- Assembly clearance distribution reveals the actual failure probability
- Shift in customer complaint timing may correlate with a process change on the customer's side (did they tighten their bore process?)
- Temperature effects on both parts at assembly (thermal expansion can change clearances)
Edge Case 4: Non-Conformance Discovered on Already-Shipped Product
Situation: During a routine review of calibration records, your metrology technician discovers that a CMM probe used for final inspection of surgical instrument components had a qualification failure that was not flagged. The probe was used to inspect and release 14 lots over the past 6 weeks. The qualification failure indicates the probe may have been reading 0.015mm off on Z-axis measurements. The affected dimension is a critical depth on an implantable device component with a tolerance of ±0.025mm. Of the 14 lots (approximately 8,400 units), 9 lots have already been shipped to three different customers (medical device OEMs). Five lots are still in your finished goods inventory.
Why It's Tricky: The measurement uncertainty introduced by the probe error doesn't mean the parts are nonconforming — it means you can't be certain they're conforming. A 0.015mm bias on a ±0.025mm tolerance doesn't automatically reject all parts, but it may have caused you to accept parts that were actually near or beyond the lower specification limit.
For FDA-regulated medical device components, measurement system integrity is not optional — it's a core requirement of 21 CFR 820.72 (inspection, measuring, and test equipment). A calibration failure that went undetected means your quality records for 14 lots cannot be relied upon. This is a measurement system failure, not necessarily a product failure, but you must treat it as a potential product failure until proven otherwise.
Common Mistake: Recalling all 14 lots immediately without first analyzing the data. A blind recall of 8,400 implantable device components creates massive supply chain disruption for your customers (who may be OEMs that incorporate your component into a finished device in their own supply chain). If the actual parts are conforming (just the measurement was uncertain), the recall causes more harm than it prevents.
The other common mistake: doing nothing because you believe the parts are "probably fine." Failure to investigate and document constitutes a quality system failure, regardless of whether the parts are actually good.
Expert Approach:
- Immediate hold on the 5 lots still in inventory. Quarantine in MRB area. These can be re-inspected.
- Quantify the measurement uncertainty. Re-qualify the CMM probe and determine the actual bias. Then overlay the bias on the original measurement data for all 14 lots. For each part, recalculate: measured value + bias = potential actual value. Identify how many parts' recalculated values fall outside specification.
- Risk stratification of shipped lots. Group the 9 shipped lots into three categories:
- Parts where recalculated values are well within specification (> 50% of tolerance margin remaining): low risk. Document the analysis but no customer notification needed for these specific lots.
- Parts where recalculated values are marginal (within specification but < 25% margin): medium risk. Engineering assessment needed on functional impact.
- Parts where recalculated values potentially exceed specification: high risk. Customer notification required; recall or sort at customer.
- Customer notification protocol. For medium and high-risk lots, notify the customer quality contacts within 24 hours. Provide: lot numbers, the nature of the measurement uncertainty, your risk assessment, and your recommended action (e.g., replace at-risk units, sort at customer site with your quality engineer present, or engineering disposition if parts are functionally acceptable).
- Re-inspect the 5 held lots. Use a verified, qualified CMM probe. Release lots that pass. Scrap or rework lots that fail.
- Root cause and CAPA. Root cause: probe qualification failure was not flagged by the CMM operator or the calibration review process. Investigate why: was the qualification check skipped, was the acceptance criteria not clear, was the operator not trained on the significance of qualification failure? CAPA must address the system gap — likely a combination of calibration software alerting, operator procedure, and management review of calibration status.
- Evaluate MDR obligation. If any shipped parts are potentially outside specification and the component is in an implantable device, evaluate whether this constitutes a reportable event. Consult with Regulatory Affairs — the threshold is whether the situation "could cause or contribute to death or serious injury." The measurement uncertainty may or may not meet this threshold depending on the functional significance of the affected dimension.
Key Indicators:
- The ratio of measurement bias to tolerance width determines the severity (0.015mm bias on ±0.025mm tolerance = 30% of tolerance, which is significant)
- The distribution of original measurements near the specification limit determines how many parts are truly at risk
- Whether the bias was consistent or variable determines whether the risk analysis is conservative or optimistic
- Customer's use of the component (implantable vs. non-patient-contact) determines the regulatory urgency
Edge Case 5: CAPA That Addresses Symptom, Not Root Cause
Situation: Six months ago, your company closed CAPA-2024-0087 for a recurring dimensional non-conformance on a machined housing. The root cause was documented as "operator measurement technique variation" and the corrective action was "retrain all operators on use of bore micrometer per WI-3302 and implement annual re-certification." Training records show all operators were retrained. The CAPA effectiveness check at 90 days showed zero recurrences. The CAPA was closed.
Now, the same defect is back. Three NCRs in the last 30 days — all the same failure mode (bore diameter out of tolerance on the same feature). The operators are certified. The work instruction has not changed. The micrometer is in calibration.
Why It's Tricky: This CAPA failure is embarrassing and common. It reveals two problems: (1) the original root cause analysis was insufficient — "operator technique variation" is a symptom, not a root cause, and (2) the 90-day effectiveness monitoring happened to coincide with a period when the actual root cause was quiescent.
The deeper issue is organizational: the company's CAPA process accepted a "retrain the operator" corrective action for a recurring dimensional non-conformance. An experienced quality engineer would flag training-only CAPAs for manufacturing non-conformances as inherently weak.
Common Mistake: Opening a new CAPA with a new number and starting fresh. This creates the illusion of a new problem when it's the same unresolved problem. The audit trail now shows a closed CAPA (false closure) and a new CAPA — which is exactly what an FDA auditor looks for when evaluating CAPA system effectiveness.
The second mistake: doubling down on training — "more training, more frequently, with a competency test." If the first round of training didn't fix the problem, a second round won't either.
Expert Approach:
- Reopen CAPA-2024-0087, do not create a new CAPA. The original CAPA was ineffective. Document the recurrence as evidence that the CAPA effectiveness verification was premature or based on insufficient data. The CAPA system must track this as a single unresolved issue, not two separate issues.
- Discard the original root cause. "Operator technique variation" must be explicitly rejected as a root cause. Document why: training was implemented and verified, operators are certified, yet the defect recurred. Therefore, the root cause was not operator technique.
- Restart root cause analysis with fresh eyes. Form a new team that includes people who were NOT on the original team (fresh perspective). Use Ishikawa/6M to systematically investigate all cause categories — the original team likely converged too quickly on the Man category.
- Investigate the actual root cause candidates:
- Machine: Is the CNC spindle developing runout or thermal drift? Check spindle vibration data and thermal compensation logs.
- Material: Has the raw material lot changed? Different material hardness affects cutting dynamics and can shift dimensions.
- Method: Did the tool path or cutting parameters change? Check the CNC program revision history.
- Measurement: Is the bore micrometer the right gauge for this measurement? What's the Gauge R&R? If the gauge is marginal, operators may get variable results even with correct technique.
- Environment: Did ambient temperature change with the season? A 5°C temperature swing in a non-climate-controlled shop can shift dimensions by 5-10μm on aluminum parts.
- Design the corrective action at a higher effectiveness rank. If root cause is machine-related: implement predictive maintenance or in-process gauging (detection control, rank 4). If material-related: adjust process parameters by material lot or source from a more consistent supplier (substitution, rank 2). If measurement-related: install a hard-gauging fixture (engineering control, rank 3). Training is only acceptable as a SUPPLEMENTARY action, never the primary action.
- Extend the effectiveness monitoring period. The original 90-day monitoring was insufficient. For a recurring issue, monitor for 6 months or 2 full cycles of the suspected environmental/seasonal factor, whichever is longer. Define quantitative pass criteria (e.g., zero recurrences of the specific failure mode AND Cpk on the affected dimension ≥ 1.33 for the full monitoring period).
Key Indicators:
- The fact that 90-day monitoring showed zero recurrence but the defect returned suggests the root cause is intermittent or cyclic (seasonal temperature, tool wear cycle, material lot cycle)
- Operator-related root causes are almost never the actual root cause for dimensional non-conformances in CNC machining — the machine is controlling the dimension, not the operator
- Gauge R&R data is critical — if the measurement system contribution is > 30% of the tolerance, the measurement itself may be the root cause of apparent non-conformances
Edge Case 6: Audit Finding That Challenges Existing Practice
Situation: During a customer audit of your aerospace machining facility, the auditor cites a finding against your first article inspection (FAI) process. Your company performs FAI per AS9102 and has a long track record of conforming FAIs. The auditor's finding: you do not perform a full FAI resubmission when you change from one qualified tool supplier to another for the same cutting tool specification. Your position is that the tool meets the same specification (material, geometry, coating) and the cutting parameters are identical, so no FAI is required. The auditor contends that a different tool supplier — even for the same specification — constitutes a "change in manufacturing source for special processes or materials" under AS9102, requiring at minimum a partial FAI.
Why It's Tricky: Both positions have merit. AS9102 requires FAI when there is a change in "manufacturing source" for the part. A cutting tool is not the part — it's a consumable used to make the part. But the auditor's argument is that a different tool supplier may have different cutting performance characteristics (tool life, surface finish, dimensional consistency) that could affect the part even though the tool itself meets the same specification.
The practical reality is that your machinists know different tool brands cut differently. A Sandvik insert and a Kennametal insert with the same ISO designation will produce slightly different surface finishes and may wear at different rates. In aerospace, "slightly different" can matter.
Common Mistake: Arguing with the auditor during the audit. Debating the interpretation of AS9102 in real time is unproductive and creates an adversarial audit relationship. Accept the finding, respond formally, and use the response to present your interpretation with supporting evidence.
The second mistake: over-correcting by requiring a full FAI for every consumable change. This would make your FAI process unworkable — you change tool inserts multiple times per shift. The corrective action must be proportionate to the actual risk.
Expert Approach:
- Accept the audit finding formally. Do not concede that your interpretation is wrong — accept that the auditor has identified an area where your process does not explicitly address the scenario. Write the response as: "We acknowledge the finding and will evaluate our FAI triggering criteria for manufacturing consumable source changes."
- Research industry guidance. AS9102 Rev C, IAQG FAQ documents, and your registrar's interpretation guides may provide clarity. Contact your certification body's technical manager for their interpretation.
- Risk-based approach. Categorize tool supplier changes by risk:
- Same specification, same brand/series, different batch: No FAI required (normal tool replacement)
- Same specification, different brand: Evaluate with a tool qualification run — measure first articles from the new tool brand against the FAI characteristics. If all characteristics are within specification, document the qualification and don't require formal FAI.
- Different specification or geometry: Full or partial FAI per AS9102
- Process change. Update your FAI trigger procedure to explicitly address consumable source changes. Create a "tool qualification" process that is lighter than FAI but provides documented evidence that the new tool source produces conforming parts.
- Corrective action response. Your formal response to the auditor should describe the risk-based approach, the tool qualification procedure, and the updated FAI trigger criteria. Demonstrate that you've addressed the gap with a proportionate control, not with a blanket rule that will be unworkable.
Key Indicators:
- The auditor's interpretation may or may not be upheld at the next certification body review — but arguing the point at the audit is always unproductive
- Your machinist's tribal knowledge about tool brand differences is actually valid evidence — document it
- The risk-based approach is defensible because AS9100 itself is built on risk-based thinking
Edge Case 7: Multiple Root Causes for Single Non-Conformance
Situation: Your injection molding operation is producing connectors with intermittent short shots (incomplete fill) and flash simultaneously on the same tool. SPC on shot weight shows variation has doubled over the last month. The standard 5 Whys analysis by the floor quality technician concluded "injection pressure too low" and recommended increasing pressure by 10%. The problem did not improve — in fact, flash increased while short shots continued.
Why It's Tricky: Short shots and flash are opposing defects. Short shot = insufficient material reaching the cavity. Flash = material escaping the parting line. Having both simultaneously on the same tool is pathological and indicates that the 5 Whys answer ("pressure too low") was oversimplified. Increasing pressure addresses the short shot but worsens the flash. This is a classic case where 5 Whys fails because the failure has multiple interacting causes, not a single linear chain.
Common Mistake: Continuing to adjust a single parameter (pressure) up and down looking for a "sweet spot." This is tampering — chasing the process around the operating window without understanding what's driving the variation.
Expert Approach:
- Stop adjusting. Return the process to the validated parameters. Document that the attempted pressure increase did not resolve the issue and created additional flash defects.
- Use Ishikawa, not 5 Whys. Map the potential causes across all 6M categories. For this type of combined defect, the most likely interacting causes are:
- Machine: Worn platens or tie bars allowing non-uniform clamp pressure across the mold face. This allows flash where clamp force is low while restricting fill where the parting line is tight.
- Material: Material viscosity variation (lot-to-lot MFI variation, or moisture content). High viscosity in one shot → short shot. Low viscosity in next shot → flash.
- Mold (Method): Worn parting line surfaces creating uneven shut-off. Vent clogging restricting gas escape in some cavities (causing short shots) while flash at the parting line.
- Data collection before root cause conclusion. Run a short diagnostic study:
- Measure clamp tonnage distribution across the mold face (platen deflection check with pressure-indicating film between the parting surfaces)
- Check material MFI on the current lot and the last 3 lots
- Inspect the mold parting line for wear, verify vent depths
- Address ALL contributing causes. The corrective actions will likely be multiple:
- Mold maintenance (clean vents, re-stone parting line surfaces) — addresses the flash pathway
- Material incoming inspection for MFI with tighter acceptance criteria — addresses viscosity variation
- Platen deflection correction or mold design modification — addresses the non-uniform clamp force
- The CAPA must capture all three causes. Document that the single defect (short shot + flash) has three interacting root causes. Each cause has its own corrective action. Effectiveness monitoring must track the combined defect rate, not each cause independently.
Key Indicators:
- Combined opposing defects always indicate multiple interacting causes — never a single parameter
- Shot-to-shot weight variation (SPC) distinguishes material variation (random pattern) from machine variation (trending or cyclic pattern)
- Pressure-indicating film between mold halves reveals clamp force distribution problems that are invisible otherwise
- Vent depth measurements should be part of routine mold PM but are commonly skipped
Edge Case 8: Intermittent Defect That Cannot Be Reproduced on Demand
Situation: Your electronics assembly line has a 0.3% field return rate on a PCB assembly due to intermittent solder joint failures on a specific BGA (Ball Grid Array) component. The defect has been reported 47 times across approximately 15,000 units shipped over 6 months. X-ray inspection of returned units shows voiding in BGA solder joints exceeding 25% (your internal standard is <20% voiding). However, your in-process X-ray inspection of production units consistently shows voiding below 15%. The defect is real (47 customer failures is not noise), but your inspection process cannot detect or reproduce it.
Why It's Tricky: The customer failures are real — 47 returns with consistent failure mode across multiple lots rules out customer misuse. But your production inspection shows conforming product. This means either: (1) your inspection is sampling the wrong things, (2) the voiding develops or worsens after initial inspection (during subsequent thermal cycling in reflow for other components, or during customer thermal cycling in use), or (3) the void distribution varies within the BGA footprint and your X-ray angle doesn't capture the worst-case joints.
BGA solder joint voiding is particularly insidious because voids that are acceptable at room temperature can cause failure under thermal cycling — the void acts as a stress concentrator and crack initiation site. The failure mechanism is thermomechanical fatigue accelerated by voiding, which means the defect is present at the time of manufacture but only manifests after enough thermal cycles in the field.
Common Mistake: Increasing the X-ray inspection frequency or adding 100% X-ray inspection. If your current X-ray protocol can't distinguish the failing population from the good population, doing more of the same inspection won't help — you're looking for the defect in the wrong way.
Expert Approach:
- Failure analysis on returned units. Cross-section the BGA solder joints on failed returns. Map the void location, size, and the crack propagation path. Determine if the cracks initiate at voids (they almost always do in BGA thermomechanical fatigue).
- X-ray protocol review. Compare the X-ray imaging parameters (angle, magnification, algorithm) between production inspection and failure analysis inspection. Often, the production X-ray uses a top-down view that averages voiding across the entire joint, while the critical voiding is concentrated at the component-side interface where thermal stress is highest.
- Process investigation using DOE. Solder paste voiding is influenced by: stencil aperture design, paste-to-pad ratio, reflow profile (soak zone temperature and time), pad finish (ENIG vs. OSP vs. HASL), and BGA component pad finish. Run a designed experiment varying the controllable factors against voiding as the response. Use the optimized parameters to reduce the baseline voiding level below the failure threshold.
- Reliability testing. Subject production samples to accelerated thermal cycling (ATC) testing per IPC-9701 (-40°C to +125°C for SnPb, -40°C to +100°C for SAC305). Monitor for failure at intervals. This replicates the field failure mechanism in a controlled environment and allows you to validate that process improvements actually reduce the failure rate.
- SPC on voiding. Implement BGA voiding measurement as an SPC characteristic with limits set based on the reliability test data (not just the IPC-7095 generic guideline). The control limits should be set at the voiding level below which reliability testing shows acceptable life.
Key Indicators:
- 0.3% field return rate in electronics is unusually high for a solder defect — this is a systemic process issue, not random
- Void location within the joint matters more than total void percentage — a 15% void concentrated at the interface is worse than 25% distributed throughout the joint body
- Correlation between void levels and reflow profile parameters (especially time above liquidus and peak temperature) is typically the strongest process lever
Edge Case 9: Supplier Sole-Source with Quality Problems
Situation: Your sole-source supplier for a custom titanium forging (Ti-6Al-4V, closed-die forging with proprietary tooling) has been on SCAR for the third time in 12 months. The recurring issue is grain flow non-conformance — the microstructural grain flow does not follow the specified contour, which affects fatigue life. The forgings are for a landing gear component (aerospace, AS9100). The forgings cost $12,000 each, with 6-month lead time for tooling and 4-month lead time for production. You need 80 forgings per year. There is no other qualified supplier, and qualifying a new forging source would take 18-24 months including tooling, first articles, and customer qualification.
Why It's Tricky: This is the sole-source quality trap. Your supplier quality escalation ladder says you should move to controlled shipping and begin alternate source qualification. But controlled shipping at a sole-source supplier is a paper exercise — you'll inspect the forgings, but if they fail, you have no alternative. And beginning alternate source qualification gives you 18-24 months of continued dependence on a problematic supplier.
The business can't tolerate a supply disruption. Each forging is a $12,000 part with 6+ months of lead time, and you need 80 per year for an active production program. Shutting off the supplier shuts off your production.
Common Mistake: Treating this like a normal supplier quality issue. Following the escalation ladder to the letter (controlled shipping → alternate source qualification → phase-out) without consi
…(truncated)