Skip to content
PQMS
  • Homepage
  • Products
  • About Us
  • Blog
  • Contact us
Login
Book demo
  • FMEA Beyond the Form: Are You There Yet?

    FMEA Beyond the Form: Are You There Yet?

    “What if one could travel briefly into the future, identify a forthcoming error, return to the present before it occurs, and intervene to prevent it?”

    That idea sits at the heart of risk analysis. To identify risks at the start and build on them as new ones emerge.  In any manufacturing setup, small, unnoticed issues can quietly grow into serious consequences.  In an industry like pharmaceutical manufacturing, that would mean patient safety. Risk analysis is therefore a fundamental part of every exercise in pharma, be it buying equipment, developing a new product or setting up its manufacturing process. 

    ICH Q9 lists many risk management tools.  FMEA is one of them.  Yet teams prefer it over others like HAZOP or HACCP, to name some. Its value lies in its simplicity and adaptability across design, manufacturing, maintenance, software, and quality systems, like a Swiss army knife. 

    What is FMEA?

    Failure Mode and Effects Analysis (FMEA) is a systematic method for identifying potential failure modes, their causes, and their effects on a system, product, or process. It focuses on “what could go wrong” at the level of assets/components, process steps, or activities and evaluates how those failures impact safety, quality, and reliability. 

    Regulatory and standards frameworks (for example, ISO 14971 for medical devices or ICH Q9 principles in pharma) expect a structured, documented approach to risk management, and FMEA is a commonly accepted method to support these expectations.

    Core Elements of FMEA

    A typical FMEA analysis involves using a template, typically made using a spreadsheet.  Key column headings in the spreadsheet are:

    1. Component/Step: Break the item you’re analysing into components, process steps, or functions
    2. Potential Failure mode: Define the risk and how the failure will happen.
    3. Effect of failure: The consequence of that failure on the next step, the system, the product, or the patient/user.
    4. Severity (S): How serious the effect of the failure is on the patient, product, process, or business (e.g., from minor inconvenience to critical safety impact)? Teams set up a scale in the spreadsheet, for example, 1–5, with each level defined to avoid confusion.
    5. Cause of failure: Underlying reasons, such as human error, equipment malfunction, design weakness, or inadequate procedure.
    6. Occurrence (O): The estimated frequency or likelihood that a particular cause will lead to the failure mode.  Like the severity scale, teams set up a separate scale in the spreadsheet for occurrence
    7. Current controls: Preventive or detective measures already in place (alarms, SOPs, in‑process tests, automation, training, etc.).
    8. Detection (D): The likelihood that existing controls will detect the failure mode before it causes the effect.  As with the severity and occurrence scales, teams set up one for detection, too.
    9. RPN Score: Teams either multiply the severity, occurrence, and detectability scores to give a single composite score (RPN=S×O×D), or use a traffic light matrix to identify the final risk score.

    a. The first method involves setting up 3-level bands and converting the RPN composite score to a risk level,

    RPN Score RangeRisk levelWhat the team should do
    1-15LowNo immediate action needed
    16-60MediumPlan corrective action
    60-125HighRisk is unacceptable.  Assign owner.  Escalate to design controls.

    b. The second method is where, for a certain severity score and occurrence score, you find the risk class (Low being green, yellow being medium and red being High). You then determine the risk level by comparing the detectability score against the risk class. Various versions of this method exist. 

    Teams can also carry out the same analysis qualitatively, as shown in the figure below.

    Best Practices for Effective FMEA Usage

    FMEA is not without flaws, but knowing its limitations and compensating for them is vital.  In addition, the same simplicity that makes FMEA popular also creates weaknesses when organisations apply it mechanically, score it inconsistently, or treat it as a documentation exercise rather than a decision-making tool. 

    • Risk Template Setup

    It is better to avoid the consolidated RPN Score approach as it masks fundamentally different risk profiles. For example, a failure mode with Severity 5, Occurrence 1, Detectability 5 has the same RPN as one with Severity 5, Occurrence 5, Detectability 1.  Yet the first clearly represents a much higher potential impact.

    In a traffic‑light matrix, the same issue may also happen.  A Severity of 5 with Occurrence 1 can have the same risk class as Severity 1 and Occurrence 5.  Knowing this piece of information allows the user to change the colour of the former to red and the latter to green. 

    • Subjective Scoring

    A quantitative approach is difficult to work with.  Different team members often interpret the same failure differently, leading to scores driven more by opinion than by harmonised criteria or historical evidence. As a result, the same process can receive very different risk ratings across departments, sites, or even separate meetings.  A qualitative scale is easier to work with. 

    • Static Vs Live

    People perform FMEA analysis, and people are fallible; the team will inevitably miss some risks. Teams must therefore link the FMEA to the QMS, feed every newly identified root cause back, and continuously update the analysis.

    In addition, processes change, equipment ages, and new knowledge emerges over time. An outdated FMEA gives the appearance of control without reflecting the current reality. 

    • Scope of Analysis

    Any risk analysis works best when the team defines its scope clearly. Carrying out, for example, risk analysis for an entire facility or a project does not provide focus on each component of the systems/functions that the project covers.

    • Poor linkage to Actions

    Risk analysis loses its purpose if teams don’t link it to design controls or procedural controls when the risk score is high. ISPE C&Q and PDA TR54 provide great examples of risk analysis. 

    conclusion

    Ultimately, FMEA only adds value when it moves beyond numbers and colour codes to actual action taken due to high-risk scores. When teams use FMEA the right way—tightly linked to the QMS, kept current, and revisited as knowledge grows—it becomes far more than a form to complete; it becomes a practical way to ‘visit the future’ long enough to prevent today’s risks from becoming tomorrow’s deviations, recalls, or patient harm.

    Sindu K

    July 1, 2026
    Uncategorized
    FMEA, GMP, ICH Q9, Pharma Manufacturing, Quality Risk Management, RPN score
  • User Requirements Specification — Getting the Sequence Right.

    User Requirements Specification — Getting the Sequence Right.

    Most of us have written a URS with a system already in mind. The vendor had been shortlisted, sometimes already selected, and the document was drafted to formalise what had already been decided. The result is a URS that describes a product rather than defines a need. It cannot be used to challenge whether the selected system is fit for purpose, because that question was never genuinely open.

    This is the most common URS failure in the industry. Not vague language. Not missing signatures. The sequence. A URS written after vendor selection is not a requirements document. It is a justification; those of us who have reviewed enough of them know the difference.

    What the Regulators and ISPE Expect

    The ISPE definition is precise: the URS states what the system must do, not how it should do it. Every major GxP framework shares one expectation underneath the differing language: document what you need before you build it, and prove you did.

    The Annex 11 makes this more explicit than before. URS traceability to testing is no longer implied — it is mandatory. Risk management must be embedded from URS through to system retirement. The one-time IQ/OQ/PQ followed by silence is no longer a defensible position.

    What the URS Must Cover

    Every requirement in it must be verifiable — specific enough that someone can design a test, run it, and return a clear pass or fail. If it cannot be tested, it does not belong.

    The URS must cover, at a minimum:

    • Functional capabilities — calculations, safety functions, access controls, audit trails, and output formats
    • Data handling — data definitions, valid ranges, integrity controls, backup, and archiving
    • Technical performance — capacity, failover behaviour, and disaster recovery
    • Interfaces — with users, other systems, and equipment
    • Environmental and power conditions — the physical operating environment in which the system must function
    • Alarm requirements — including explicit identification of critical alarms
    • System automation — with clearly defined boundaries between systems
    • Data integrity requires its own attention. The URS must define what the system needs to do to satisfy ALCOA+ principles — data that is Attributable, Legible, Contemporaneous, Original, and Accurate, and beyond that, Complete, Consistent, Enduring, and Available. These are not abstract ideals. Each one translates directly into a testable requirement — audit trails, user access controls, timestamps, backup provisions, and data storage conditions.
    • Every requirement should state its category — Quality, Business, or HSE — and its source, whether a CQA, a CPP, a regulatory requirement, a risk assessment output, or an engineering standard. We must explicitly identify quality-critical requirements and write them with precision.

    Finally, not every requirement carries the same weight. Separating mandatory from beneficial and nice-to-haves keeps the qualification effort focused where it actually matters.

    What a GMP-Grade Requirement Looks Like

    The gap between a weak requirement and a defensible one is not technical knowledge. It is discipline, specificity, and process ownership.

    Weak Requirement GMP-Grade Requirement
    The system shall monitor temperature.The system shall continuously monitor storage temperature within 2°C–8°C(CPP), trigger a GMP-critical alarm within 60 seconds of exceedance, and maintain a tamper-proof audit trail per 21 CFR Part 11.
    The system should print labels correctly.The system shall restrict label reprints to QA-authorised personnel with an electronic signature, reason code, and a full audit trail of all print events and template changes.

    A requirement that cannot be tested has no place in a qualification programme.

    Five Failure Modes That Destroy URS Credibility

    These are not theoretical. They are the patterns behind real audit observations, 483s, and warning letters.

    1. Vague, untestable language. Robust, appropriate, user-friendly — these are not requirements. If a test cannot be written for it, it does not belong in the document.
    2. Vendor Brochure Transcription. A URS sourced from supplier documentation belongs to their sales cycle, not your process. Auditors notice immediately.
    3. No CQAs or CPPs identification. You cannot build a credible risk assessment on a document that fails to identify what is critical.
    4. URS Written Post-Selection. A URS written after vendor selection is not a requirements document. It is a justification.
    5. Never Updated After Go-Live. If your team has not updated the document to reflect the current validated state, it is a liability.

    Each of these patterns is recoverable before qualification begins. After a deviation occurs, they become the gap that an investigation cannot close.

    Conclusion

    A URS is only as useful as the process understanding behind it. Written early, classified correctly, and maintained through every system change — it is the document that gives every qualification activity its purpose. Written late, vaguely, or not at all — it is the gap an investigation will eventually fall into.

    Sindu K

    June 15, 2026
    Uncategorized
    ALCOA+, GMP validation, GxP Compliance, Pharmaceutical Validation, QMS, Quality Assurance, Regulatory compliance, User Requirement Specification
  • The Illusion of Root Cause:  5 Whys in QMS

    The Illusion of Root Cause:  5 Whys in QMS

    Most of us have used the 5 Whys during deviation investigations, but have we been using it correctly? It appears on every deviation investigation form, every deviation report, every post-incident review. Five questions. Five answers. Signed and filed as a CAPA starting point. But is this the right approach? 

    The Linear Trap

    Consider the example shown below with a problem statement and associated 5Why analysis :

    Problem Statement: An older version of an SOP was found on the shopfloor that triggered a deviation. 

    LevelQuestionAnswer
    Why 1Why was the older version of SOP-xxxx found on the shopfloor?Operator used a printed copy not updated after the revision.
     Why 2Why was the printed copy not replaced after the new version was approved?No one informed the shopfloor supervisor that SOP-XXXX had been revised.
     Why 3Why was the shopfloor supervisor not notified of the revision?System notifies only the document owner, not active-use areas.
     Why 4Why does the notification process exclude active-use areas?The distribution matrix was not updated when the shopfloor was added as a user area.
     Why 5Why was the distribution matrix not updated?No requirement in change control to review the distribution matrix when adding a user area.

    The above 5 Whys analysis runs as a single straight line — one problem, one Why, one answer, repeated five times. That works when a failure has a single, clear cause.  But real deviations rarely behave in this manner. 

    When you look at the problem statement — older version of SOP-XXXX found on the shopfloor — can we attribute it to just one reason? There may be three distinct causes, each having its own branches. Not all will reach Why 5, and that is fine. The goal is to find the branches that lead to a true systemic cause.

    The Branching Model

    The same analysis done differently reveals three independent causal branches:

    Problem Statement: An older version of an SOP was found on the shopfloor that triggered a deviation. 

    Branch B is important for a different reason: it closes early and clears Document Control of a failure. This is equally valuable — the 5 Whys is not only a tool to find blame, but to eliminate false leads and confirm where controls did hold

    Why This Matters in Pharma QMS

    When the 5 Whys analysis is forced into a single linear chain, investigations routinely stop at the first plausible answer rather than the true systemic cause.  This leads to CAPAs that address symptoms rather than eliminating the underlying system gaps. Examples include retraining the operator, reminding Document Control, etc. The branching model exposes multiple independent failure pathways. In this example, even if Branch A’s CAPA (an HR–Document Control notification procedure) is fully implemented, the deviation could recur via Branch C if the configuration baseline register gap remains unaddressed.

    The branching model depends heavily on the knowledge of the team conducting the analysis.  A fishbone (Ishikawa) diagram can help identify potential causes before drilling down. For high-risk deviations with patient safety implications, FMEA or fault tree analysis provides a more defensible, data-driven structure.

    5Why – The Right Way – Step by Step

    Using 5 Whys correctly in a pharmaceutical QMS context means:

    • Start with a well-defined problem statement — observable, specific, and bounded. Vague statements generate vague chains.
    • Identify all Why 1 possibilities before drilling — brainstorm independently, then build separate branches for each distinct causal pathway.
    • Allow branches to close early when appropriate — a branch that leads to “this control worked correctly” is a valid and useful finding.
    • Do not force five levels — the goal is the systemic root cause, not a count of five. Three Whys that reach a procedural gap are better than five Whys that circle back to a symptom.
    • Validate each root cause with the reverse test — ask: “If this root cause were eliminated, would the deviation have been prevented?” If the answer is no, keep drilling.
    • Link each confirmed root cause to a discrete CAPA action — one root cause, one CAPA. Multiple root causes from multiple branches each require their own corrective action plan.

    The Real Goal

    The 5 Whys was never intended to be a form-filler. Sakichi Toyoda introduced it into Toyota’s production system as a discipline of genuine inquiry, not documentation compliance. In a QMS environment where the pressure is to close deviations quickly and move on, it is easy to treat it as a checkbox. The result is a CAPA system full of retraining records and SOP updates that address the same recurring deviations, investigation after investigation.

    The branching 5 Whys is time-consuming, there is no doubt about that, but it is the version that actually protects product quality and patient safety. The illusion of root cause is easy to create. The root cause itself takes work to find.

    Sindu K

    June 5, 2026
    Uncategorized
    pharmaceutical, quality management system, root cause analysis

Powering Quality from Lab to Label © 2025, Quascenta Pte. Ltd.