Skip to content
PQMS
  • Homepage
  • Products
  • About Us
  • Blog
  • Contact us
Login
Book demo
  • Factory Acceptance and Site Acceptance Testing: Accelerating Equipment Qualification with ValDoc Pro

    Factory Acceptance and Site Acceptance Testing: Accelerating Equipment Qualification with ValDoc Pro

    https://www.pqms.com/wp-content/uploads/2025/12/FAT-SAT-Article-1.mp4

    Synopsis:

    “Suitable for the intended purpose/activities” is a core regulatory expectation for all pharmaceutical manufacturing equipment, and qualification is the documented proof that this is achieved. The qualification lifecycle starts with a User Requirement Specification (URS) document followed by a Functional Specification (FS).  Test Scripts are executed in IQ, OQ and PQ documents to demonstrate that the requirements noted in the URS are met. Because running IQ/OQ/PQ scripts at the manufacturing site is time-consuming and the responsibility falls only on the drug manufacturer, some of the effort can be shared with the vendor via Factory Acceptance Testing (FAT) and Site Acceptance Testing (SAT).  ValDoc Pro can streamline this process and build better compliance.  

    Regulatory Guidelines:

    121 CFR 211.63 states that “Equipment must be of appropriate design, adequate size, and suitably located to facilitate operations for its intended use”[2].  While the FDA recognises the advantages of FAT and SAT (mentioned in presentations by the FDA), there is no guideline that discusses this.   EU GMP Annex 15 & PIC/S explicitly state, in sections 3.4-3.7, that:

    • Equipment, especially if incorporating novel or complex technology, may be evaluated, if applicable, at the vendor prior to delivery.
    • Prior to installation, equipment should be confirmed to comply with the URS/ functional specification at the vendor site, if applicable.
    • Where appropriate and justified, documentation review and some tests could be performed at the FAT or other stages without the need to repeat on-site at IQ/OQ if it can be shown that the functionality is not affected by the transport and installation.
    • FAT may be supplemented by the execution of a SAT following the receipt of equipment at the manufacturing site.

    Leveraging Vendor Testing:

    The Factory Acceptance Test (FAT) is carried out at the vendor’s site and provides documented evidence that the equipment meets the user requirements and the vendor has fulfilled contractual responsibilities. A well‑planned FAT helps save time during qualification by identifying design issues, functional gaps, and integration problems before the equipment leaves the factory.   Sending the equipment back to the vendor to fix issues, especially when they are located across continents, is avoided through FAT.  Fat also provides the ability to check internal components of the equipment that one would not be able to reach at the site unless it is disassembled. 

    It is not necessary that a FAT test be performed for all equipment/systems.  That decision is left to the manufacturing site and its policy.  For a greenfield facility, a commissioning and qualification (C&Q) plan will direct this strategy.  A survey by ECA Academy in 2018 found that FAT is used by 37% of the companies surveyed, while 55% stated “it depends on the complexity of the project”. SAT was found to be used by 49% of those surveyed [5]. 

    Planning and preparation for the Factory Acceptance Test (FAT) are critical because the activity is conducted at the vendor’s site and involves budgeting for extra expenses on both sides. The decision to perform FAT is typically taken very early in the procurement lifecycle, during the development of the User Requirement Specification (URS) for the equipment.  The contract signed between the drug manufacturer and the vendor for an equipment includes scope of the FAT, responsibilities and the commercials.  The details of what is to be covered in FAT are put down in a protocol and frozen before approval of the design stage.  The site and the vendor have to agree on the features/functions to be demonstrated, parameter specification to be met, tests to be carried out, and identify the URS points covered by FAT.  The FAT protocol is sent to the drug manufacturer for review.  Once approved, then only the protocol can be executed.  During execution, the drug manufacturer representative(s) ideally should be present to at least to verify the minimum test requirements and resolve any deviations.  During and after the pandemic, many FATs have been conducted remotely, reducing travel and related expenses and making this an attractive option, particularly when organisations can leverage a wide range of digital tools and platforms to their advantage.

    The FAT protocol, once executed, is compiled along with attachments and sent to the site by the vendor for review.  In most cases, the equipment would have to be disassembled before shipping.  The documentation put together details equipment component labels and the process to follow to reassemble these components. 

    Similar to the FAT, the Site Acceptance Test (SAT) is not a mandatory test to perform.  The advantage of SAT is time savings.   A successful execution of the SAT protocol provides assurance that the equipment works as per specification post-transportation to the site.  While infrastructure, utilities, interfaces, etc, are defined in advance and agreed upon, surprises may crop up.  The SAT protocol also rechecks issues found during FAT.  These should be addressed before IQ begins.  In addition, the presence of the vendor’s team at the site speeds up testing, as that machine is very well known to them. 

    The SAT process follows the same overall approach as the FAT, but is performed at the manufacturing site rather than at the vendor’s facility. Whether a SAT will be performed should be defined in the C&Q plan and, preferably, specified in the URS and subsequently reflected in the contract.  SAT is led by the site, along with the vendor’s project/commissioning/service staff.  The SAT protocol is normally drafted by the vendor or the project engineering/validation team. 

    A decision that the site needs to take is whether they would like to leverage what was tested in FAT and SAT.  Most sites leverage the data from these tests.  The SAT protocol has many tests that overlap with what one would carry out under IQ and OQ protocols. 

    Accelerating FAT and SAT Efficiency with ValDoc Pro:

    ValDoc Pro is an easy to use Qualification management application that assists companies in digitizing their qualification workflow, from URS onwards through the entire lifecycle.  Its role-based access and real-time collaboration features enable vendors and manufacturers to coordinate FAT and SAT seamlessly.  Following are some of the advantages:

    • Vendor Access: Vendor can create, obtain site approval and execute FAT protocols within ValDoc Pro, with controlled access to designated folders/files. 
    • Site Review and Approval: The FAT protocol and FAT report can be submitted for review and approval online.
    • Seamless Flow: Generating the protocol by the vendor, review, execution and compilation all can be carried out seamlessly in ValDoc Pro. 
    • Sequential: ValDoc Pro ensures that the SAT protocol cannot be executed until the executed FAT report has been approved.  This ensures that the SAT protocol captures all relevant tests. 
    • Traceability Matrix: ValDoc Pro maintains a continuous link from URS through FAT/SAT to IQ/OQ/PQ, demonstrating to regulators that all user requirements were tested and all deviations addressed.

    Conclusion:

    Compared with relying solely on traditional IQ, OQ, and PQ, integrating FAT and SAT into the equipment qualification strategy allows issues to be identified earlier at the vendor and manufacturing site, limits risk, and reduces the amount of testing and troubleshooting that must be compressed into already busy commissioning windows. This upstream focus not only improves the quality of delivered assets but also frees site resources to concentrate on value-added verification and routine operation rather than firefighting late-stage gaps. ValDoc Pro-enabled FAT and SAT give pharmaceutical manufacturers a more efficient and compliant path to equipment qualification by moving protocols into a digitized, collaborative environment and enabling options such as Remote FAT, so teams can shorten timelines, reduce on-site disruption, and strengthen data integrity while maintaining full traceability from URS through lifecycle maintenance. In a landscape where each site-based role carries a greater workload, embedding FAT and SAT within a robust application like ValDoc Pro is not just an efficiency gain, but a strategic necessity for sustaining reliable, inspection-ready manufacturing.

    [1] FDA. Process Validation: General Principles and Practices. U.S. Food and Drug Administration.

    [2] “U.S. Food and Drug Administration. 21 CFR Part 211 – Current Good Manufacturing Practice for Finished Pharmaceuticals, §211.63 Equipment design, size, and location.”

    [3] European Commission. EU Guidelines for Good Manufacturing Practice for Medicinal Products for Human and Veterinary Use, EudraLex, Volume 4, Annex 15: Qualification and Validation. In operation since 1 October 2015

    [4] PIC/S. Guide to Good Manufacturing Practice for Medicinal Products, PE 009 (current version), Annex 15: Qualification and Validation

    [5] ECA Academy. FAT & SAT not frequently used as Part of Qualification – ECA Modern Qualification Survey Results. GMP News, gmp-compliance.org (2018)

    ravi

    December 12, 2025
    Uncategorized
    Equipment Qualification, Equipment Testing, GMP Compliance, Pharmaceutical Validation, Process Validation, Quality Assurance
  • Who is actually Responsible for the URS?

    Who is actually Responsible for the URS?

    Synopsis:

    Before purchasing an equipment, software or instrument, regulatory requirement mandates that one documents what is expected of the proposed acquisition. While it may be tempting to rely on vendors for this, as they understand their product, they do not understand the business requirements. That is why the User Requirement Specification (URS) must be authored by the team responsible for this new acquisition. This article explains why a user-driven URS ensures the acquired entity genuinely matches user needs, moving it from basic functionality to real-world fit

    It All Starts with a Need

    Picture this. You have been asked to put together a URS for a new equipment/software/instrument.  If this requirement was identical or similar to one that you have purchased in the past, that makes your task easier.  But what if this is a brand-new requirement.  Where do you start? 

    Appendix D5 from the ISPE GAMP 5 guidance document says it the best when it states that each user requirement should be specific, measurable, achievable, realistic, and testable.   It is also a good practice to prioritize the requirements, typically in two or three levels: (mandatory (high), beneficial (medium), and nice-to-have (low)) Or (mandatory (high) and nice-to-have (low)).

    The Easy Way

    As a vendor, we very often get asked to provide a draft URS that the regulated company then uses as a basis to draft the requirement.  Is this the right approach?  It sounds logical, right? After all, vendors are experts in the products they sell. But here’s the catch — they are experts in their product, not your business process.

    Sometimes we receive a URS that lists a specific brand or model, even including specifications unique to that brand. This is not the intended purpose of a user requirement, which should describe what is the needed functionally, not prescribe a particular vendor or solution.  When this happens, it limits the scope of solutions, undermines competitive assessment, and can introduce bias or compliance risks. Such requirements do not reflect true user needs but rather pre-select a solution, making it harder to compare alternatives and potentially exclude options that better meet functional requirements. GAMP 5 guidelines recommend describing functional needs in clear, objective terms, avoiding references to specific suppliers unless it is absolutely necessary to do so for business reasons

      Understanding the Roles the RACI way

      How do we decide the roles of various stakeholders.  A RACI Matrix, a simple chart that outlines who does what, breaks down the roles:

      Responsible:  Team/end users responsible for the business process.  The end users or process owners are usually responsible for drafting the URS since they know the operational needs best.

      Accountable: Manager(s) responsible for the business process.  Oversees URS creation and ensures it meets desired goals

      Consulted:

      • Other departments to provide their input/expertise as the acquired product may influence their workflows
      • IT team for software/IT hardware requirements to provide input on technical feasibility and rollout implications
      • CSV team for software to include regulatory requirements
      • Vendor – provide input and give technical advice. 
      • Informed-Vendors so as to put together a function requirement specification. 

      What do the Regulators Say?

      The FDA, the EMA and others, have made their stance clear: every system must be “fit for its intended use.”  That phrase — fit for intended use — is powerful. It means the system should perform exactly as needed for your specific products and processes.

      And who defines that intended use? Not the vendor — you do.

      According to the FDA’s General Principles of Software Validation and EMA’s Annex 11 on Computerised Systems, the responsibility for defining and confirming suitability lies with the regulated user organisation for computerised systems.  But that principle doesn’t stop there. Under broader GMP guidance, such as Annex 15 on Qualification and Validation, the manufacturers are responsible for ensuring equipment is suitable for its intended purpose and for defining its requirements in a User Requirements Specification (URS) or functional specification.

      In short, the user determines what’s needed; the vendor shows how their product meets those needs.

      Why Ownership Matters?

      When the customer owns the URS, everything aligns. The equipment fits its real operational needs, and the system passes regulatory scrutiny because it was built for its true purpose. And most importantly, it saves time, money, and frustration down the road.

      When vendors take over the URS, it might seem faster, but it’s like letting someone else write your recipe — you’ll never get the exact flavour you wanted.

      ravi

      November 9, 2025
      Uncategorized
      Equipment Qualification, GMP Compliance, Pharmaceutical Validation, Quality Assurance, URS, User Requirement Specification
    • Demystifying GAMP 5 Software Classification: Understanding Categories 3, 4 and 5

      Demystifying GAMP 5 Software Classification: Understanding Categories 3, 4 and 5
      Demystifying GAMP 5 Software Classifications

      Summary:

      In a GxP environment, the software to be acquired should be fit for use. In order to meet the user requirements, one buys software that can be used as-is, with configuration, or custom-developed. To ensure the appropriate documents are included in the qualification package, it is essential to identify the relevant qualification documents for each scenario. This article explores how the software classifications, defined in GAMP 5, aid in determining the necessary qualification documentation for software in Categories 3, 4, and 5.

      Defining Software Qualification Requirements:

      When a pharmaceutical company, be it a manufacturer or a CRO, purchases software for use in a GxP environment, they must create a clear and detailed User Requirement Specification (URS) that outlines their needs. In response to the URS, the software vendor(s) submit a Functional Specification (FS) that describes how their software will address the user requirements. All URS requirements must be tested and proven through qualification. A Traceability Matrix (TM) is a document that maps each specific requirement within the URS to corresponding testing scripts.

      Every software qualification workflow requires a URS, FS, and TM to ensure that each requirement is properly addressed, tested, and fulfilled. The other documents in the qualification package depend on the category the software falls under per GAMP 5.   

      Category 3 consists of Commercial Off-the-Shelf Software (COTS) that is used as-is; for instance, Microsoft Excel (without macros), or statistical tools like JMP and Minitab. The qualification objective for such software is to ensure proper installation and expected performance. For software installed on a computer or server, Installation Qualification (IQ), Operational Qualification (OQ), and Performance Qualification (PQ) must be completed. For Software as a Solution (SaaS) applications, since the software is pre-installed on the cloud, a simplified IQ may be conducted together with a full OQ to comprise an IOQ, and a PQ test is performed.

      Category 4 encompasses software that is configured without changing the code. Examples include Laboratory Information Management Systems (LIMS) or Manufacturing Execution Systems (MES), where modifications can be made to templates, forms, or workflows without changing the main software. Another example is an Excel sheet that has been configured for a specific use without macros or Visual Basic code. In addition to the qualification documents specified for Category 3 software, configured systems require a Configuration Specification (CS) to document the changes or settings that have been made.

      Category 5 includes software developed from scratch or that has undergone code modifications. One example is an Excel sheet that has macros, scripts, or Visual Basic code added. These systems are unique to a company and may be developed in-house or by an external vendor. Testing for Category 5 is more detailed, involving unit, integration, and full system testing to ensure that all components function as intended. This is critical, as even minor code changes can impact performance. The same documents that are required for Category 4 are also required for Category 5. Additionally, a Design Specification (DS) is needed to explain the software construction, since the logic has been modified or custom-developed.

      Conclusion:

      Identifying a software’s category under GAMP 5 is not always clear.  One does not decide on the basis of the software name; rather, it depends on its usage. The same software can potentially belong to any of the three categories—category 3 if used as-is, category 4 if configuration is required, or category 5 if custom coding is required. Higher the category classification, more extensive is the documentation and testing requirements.

      ravi

      November 5, 2025
      Uncategorized
      Computer System Validation, GAMP 5, GxP Compliance, Pharmaceutical Validation, Software Categories, Software Classification
    • One-Size-Fits-All Paperless Validation – Is It Really Fit for Use?

      One-Size-Fits-All Paperless Validation – Is It Really Fit for Use?

      Synopsis:

      Popular  validation tools have a basic flaw – these applications were designed to only carry out qualification activities.  This article emphasises the differences between qualification and validation, and why for process and cleaning validation/monitoring, such “paperless validation” applications offer little more value than a document management system. An efficient digital validation platform has to be a suite of applications each performing what it has to be designed for, qualification, manufacturing process management, cleaning process management separately,  but still integrated so as to provide a smooth flow of data. 

      Introduction

      No manufacturing sector can survive without a relentless focus on quality—because the very essence of manufacturing is to produce consistent and standardized product.  Any lapse in quality in pharmaceutical industry leads to regulatory scrutiny, fines, damaged reputations, and ultimately business failure. 

      In pharmaceutical manufacturing, the regulatory expectation is that the process be in “a state of control”.  And historically we “validate” processes, e.g., manufacturing, packaging, cleaning or QC method.  The core objective of validation is to demonstrate document consistent, controlled performance over time.  Validation is more than a regulatory requirement—it is the foundation of operational consistency, product quality and most importantly.

      As digital “validation” solutions flood the market, the expectation is that these digital tools do more than merely replicating paper processes in electronic form – that they will enable smarter, more integrated validation approaches that align with regulatory requirements and operational needs. But do they deliver on this promise?

      This article explores the differences between validation and qualification, and why the “one-size-fits-all” approach towards  “validation” is an idea that falls well short of customer expectations. 

      Let’s Clarify the Basics

      Before delving deeper, it is crucial to differentiate between the two terms that are often confused: qualification and validation.

      • The expectation from Qualification is that the equipment, utilities, instruments, software used in a GMP environment have been installed correctly and function as intended.
      • Expectation from Validation is to show that the processes are “in a state of control”. 

      In simple terms, Qualification answers, “Does that entity work as per our requirement?” while Validation dives deeper to explore, “Can we get a reliable and consistent product?”

      Understanding Qualification

      The User Requirements Specification (URS) is usually the first document created when a new entity/asset is being planned, detailing the requirements, including business, compliance, and operational requirements.  This activity is done by stakeholders who are either responsible or accountable for that asset/entity.

      The selected vendor(s), or the internal resource where applicable, respond with a Functional Specification (FS) detailing how the requirements in the User requirement will be met.

      The specific test script document requirements vary depending on what is being qualified. However, the intent remains the same: to demonstrate, through test scripts, that the User Requirement Specification (URS) requirements have been met. These test scripts may be documented in Factory Acceptance Tests (FAT), Installation Qualification (IQ), Operational Qualification (OQ), Site Acceptance Tests (SAT), or Performance Qualification (PQ). The selection of which of these to include depends on the particular asset or system being qualified.

      A Traceability Matrix document is then put together mapping each requirement to the corresponding tests, detailing the user requirement ID and the test script ID. This matrix provides clear visibility of how every user requirement specified in the URS is verified through specific test cases. 

      In this sense, qualification is more likely to be a documentation-driven, structured and checklist-based approach, focusing on verifying that vendor-supplied asset/entity meet predefined specifications. 

      Understanding Validation

      Validation, by contrast, focuses on ensuring that processes consistently produce outcomes that meet predetermined specifications.

      According to the FDA, validation is: “Establishing documented evidence which provides a high degree of assurance that a specific process will consistently produce a product meeting its predetermined specifications and quality attributes.”

      Where qualification proves that an asset/entity is installed and functions correctly, validation demonstrates that the processes using that asset/entity are robust, repeatable, and compliant.  Furthermore, it is even more critical that “effective monitoring and control systems” are put in place for “process performance and product quality, thereby providing assurance of continued suitability and capability of processes”. This is a requirement as per ICHQ10 and continues to be for the last decade one of the most cited observations. The goal in validation and monitoring is not just compliance but smarter, data-driven decisions, continuous improvement being an important pillar. 

      Validation, therefore, is data-driven  and is focussed on data analysis providing the assurance that the process will consistently produce product whose quality meets predetermined specifications.

      Given that qualification and validation have very different approaches and end goals, is it possible to use one software tool to manage both requirements?

      Many software vendors market “all-in-one” applications as a one-stop solution. These apps promise to provide a unified platform for validation.  At first glance, the appeal of a unified system is undeniable. After all, a single tool that manages documentation, qualification, and validation activities across an entire organization sounds like a great solution.

      However, when the design requirement for a qualification activity has absolutely no relationship with what is to be done in validation and ongoing monitoring, how is it possible to combine the use cases?  

      The data type that a qualification management system main data type is textual information. User requirements are primarily captured as detailed text entries within structured tables, and supporting test scripts are also documented largely in text form. While most data centers around descriptive narratives—such as specifications, acceptance criteria, and rationales—numeric data appears mainly when recording process parameters during process qualification (PQ) for equipment. This emphasis on text ensures that every requirement and its verification are clearly documented for traceability, auditability, and regulatory compliance.

      Validation and ongoing monitoring activities are fundamentally centered around the defined product specifications, attribute specifications, equipment process parameter ranges and how these are met. The primary data type handled in this context is numeric, as these processes often involve the collection and analysis of quantifiable results—such as, equipment process parameter range for a noted batch, attribute measured results—to demonstrate compliance with set limits or acceptance criteria. This numerical data provides objective evidence to verify that equipment operate within their approved specifications, and enables continuous, data-driven monitoring to promptly detect deviations or trends. While supporting documentation like SOP may include textual explanations or rationales, it is the structured numeric data that forms the backbone of validation and monitoring efficacy.

      Pharmaceutical regulators expect manufacturers to proactively show processes are in control, detect deviations, implement timely corrections, and continually improve processes to guarantee product safety and efficacy.  While qualification is an integral part of the process validation program and expects all entities/assets used in the product manufacture to be shown to be “fit for use”, the expectation is also timely access to real-time or trending data, making it easy to promptly identify deviations, detect process drifts, or generate automated alerts.

      So an  “all-in-one” application designed to aid in qualification and CSV, can only act as storage repository for files associated with process validation.  Such applications consider validation or ongoing process verification runs as static activities with no ability to let the data speak for itself.   These teams rely on time-consuming, error-prone manual reviews to collate and interpret data, increasing the risk of missed trends and delayed responses to quality issues. Furthermore, static documentation impedes data integration, reduces the potential for cross-functional collaboration, and hinders the effectiveness of continuous improvement initiatives—ultimately compromising regulatory compliance and the robust oversight expected in modern pharmaceutical manufacturing.

      Conclusions

      Ultimately, companies must weigh the true cost of investment, both in terms of efficiency and compliance risk, when selecting validation tools. It may be time for the industry to reconsider whether costly all-in-one solutions are necessary, or if more focused, built for purpose software solutions targeted toward manufacturing process would better serve their needs.

      ravi

      July 27, 2025
      paperless validation
      digital transformation, digitilization, paperless validation, pharmaceutical
    • Can AI help in Validation? Cutting Through the Hype

      Can AI help in Validation? Cutting Through the Hype

      As interest in AI continues to grow, some vendors are promoting AI as tools capable of generating qualification deliverables such as User Requirement Specifications (URS), Functional Specifications (FS), and traceability matrices, with minimal human input. This article provides a grounded, practical look at what AI can and cannot do in the context of URS and FS creation and automating traceability matrix. It explains how LLM models work; what URS and FS documents are, and how they fit in the qualification process; it also looks at tests we ran with custom LLMs to automate reconciliation as in automating generation of a traceability matrix, and the results obtained; and suggests where AI may offer limited support – while highlighting why human ownership and input is still required for critical tasks like URS and FS development.


      Introduction: Opportunity vs. Overpromise

      The rise of AI tools like ChatGPT has sparked broad interest across regulated industries. In pharmaceutical, biotech, medical device, and cosmetics sectors, companies are exploring how these tools might reduce documentation workload, increase consistency, and support compliance efforts.

      However, alongside this enthusiasm, some vendors have begun promoting the idea that large language models (LLMs) can generate qualification documents – such as User Requirements Specifications (URS), Functional Specifications (FS), or even test scripts – with a simple prompt.

      While these claims are designed to attract attention, they risk oversimplifying what these qualification documents really are, how they are developed – and where, if at all, AI actually fits into the process.

      This article explains the role and requirements of key qualification documents, explores how LLMs like ChatGPT work, and examines where AI tools can – and cannot – be appropriately applied to support the qualification process.

      What a URS Actually Does – and Why It Matters

      A URS (User Requirements Specification) is a critical foundation for any qualified entity. It defines what the user needs the system, equipment, utility, or software to do – in terms that are clear, measurable, and auditable.

      A robust URS is not a generic checklist. It must be developed by knowledgeable stakeholders who understand:

      • The intended use of the system,
      • Key process and operational parameters (e.g., throughput, control logic, data handling),
      • The regulatory framework and internal quality expectations,
      • Integration points with other systems (e.g., MES, SCADA, LIMS),
      • Environmental, safety, and compliance constraints.

      A well-defined URS supports:

      • Effective vendor selection and procurement,
      • A clear basis for design and testing,
      • Alignment with qualification protocols and traceability,
      • Compliance with GxP expectations for “fitness for intended use.”

      When a URS is vague or incomplete, the consequences can be significant: unsuitable equipment, misaligned vendor deliverables, inefficient qualification efforts, or gaps that lead to inspection findings and remediation. Poor input at the URS stage can therefore compromise the entire lifecycle of the system – including its qualification.

      How LLMs Like ChatGPT Actually Work

      How LLMs function is often misunderstood. Tools like ChatGPT do not think, reason, or understand. In fact, they generate text by predicting the most statistically likely next word in a sentence, based on patterns learned from massive volumes of publicly available internet text.

      In practice, this means:

      • They do not understand context in the way a human expert does,
      • They do not verify accuracy,
      • And they do not apply regulatory logic or domain-specific judgment.

      Critically, high-quality URSs, FSs, and qualification scripts are typically internal and proprietary – not available in the public domain. As a result, these documents are not well represented in the training data used to build general-purpose language models.

      This creates a risk of ‘hallucination’: that is, AI can sometimes confidently generate text that sounds correct but is factually inaccurate or entirely fabricated. In regulated environments, this presents an obvious risk – especially when such content is used to support compliance activities.

      Can LLMs Write a URS? What’s the Real Limitation?

      When asked to generate a URS, an LLM can produce something that looks well-formatted – but that document is likely to be too generic or vague to be useful in practice.

      For example, asking ChatGPT to generate a URS for a “sifter” will not trigger the LLM to ask clarifying questions on aspects such as:

      • Process requirements: What is the role of the sifter in the process? For example, it could be particle size separation, de-lumping, or scalping. Or which screening motion – centrifugal, vibratory, or gyratory – is best suited for your materials? What are the in-process material characteristics (e.g., moisture content, particle size distribution)? How will the sifter integrate with upstream or downstream equipment, like granulators or during presses – by using gravity-fed or vacuum transfer systems? Is the sifter intended for batch processing or continuous operation?
      • Performance requirements: What throughput is needed? Are there maximum allowable noise or vibration levels to comply with workplace safety standards? What is the acceptable level of material loss or retention during the process?
      • Design and construction requirements: What are the material of construction (MOC) requirements? What specific surface finish (e.g., Ra < 0.8 µm), required? What are the requirements for welding and polish?  What utilities are available, like compressed air pressure, electrical supply, or vacuum systems, to support its operation? Does it need tool-less disassembly for cleaning? What are the cleaning needs (e.g., CIP, SIP, manual)?
      • Controls and automation: Will the sifter be integrated with a Supervisory Control and Data Acquisition (SCADA) system or Programmable Logic Controller (PLC)? Are alarms, interlocks, or data logging required? Should it support electronic batch records or interface with Manufacturing Execution System (MES)? Is Process Analytical Technology (PAT)-driven monitoring required?
      • Environmental or containment needs: Is operator protection (e.g., for potent APIs) required? What classification zone applies?  For example, if OEB 5 containment is required, pneumatic clamping mechanisms have to be put in place.

      These are essential details – and they define far more than the basic equipment type. They shape how the system will be designed, selected, installed, qualified, and operated. Without them, the URS cannot serve its purpose as a regulatory and procurement document.

      Even if LLMs are fine-tuned on customer data, key questions remain:

      • Who validates the quality and compliance of the source material?
      • Is the model version-controlled?
      • Is the data private – or potentially exposed across users or systems?
      • Will the AI output reflect your company’s unique process – or generalize from unrelated inputs?

      And if subject matter experts must still correct or rebuild the draft, it’s worth asking: is the AI truly reducing effort, or simply introducing another layer of review?

      Functional Specifications: A Vendor’s Responsibility

      Once the URS is approved, the FS (Functional Specification) is typically authored by the vendor, service provider, or internal engineering team. The FS outlines how the solution will meet the defined user requirements – through system design, control logic, automation behavior, configuration, and integration.

      Authoring an FS requires:

      • Proprietary system knowledge,
      • Awareness of design constraints,
      • Understanding of operational and regulatory requirements,
      • Collaboration across quality, engineering, IT, and qualification functions.

      The FS is the vendor’s responsibility – it represents their response to the URS. And this raises an important point: why would a vendor delegate that responsibility to an AI tool that doesn’t understand their product? Even with access to a URS, an LLM cannot independently determine how a system meets those requirements without being explicitly told.

      The same applies to qualification test scripts. These are step-by-step instructions used to verify whether the system meets its documented requirements. They must be traceable, reproducible, and inspection-ready – and they require domain knowledge to be meaningful. LLMs cannot reliably generate test descriptions because they don’t understand how the system functions, what needs to be tested, or how to design a test that produces meaningful, validated results.

      If generating test scripts were as simple as prompting an AI, most qualification tools would be fully automated by now. But they are not – because test design remains a technical, judgment-based process.

      Can LLMs Automate the Process of Filling a Traceability Matrix?

      Some vendors have been marketing their ability to automate traceability, namely by using LLMs to identify which test script aligns with a given URS requirement. But is it really possible to perform this task truly reliably using AI?

      One approach would be to use a type of LLM known as a sentence transformer model, a class of AI designed to compare entire sentences rather than individual words. These models can capture semantic relationships between phrases – for example, comparing the intent of a URS requirement with the description in a test script (be it an IQ, OQ, or PQ) – and then rank the closest matches accordingly, to select which script number matches each URS requirement.

      Transformer models are trained on billions of data points and can encode the meaning of phrases or entire sentences, not just individual words. However, their reported accuracy leaves some questions.

      Here are the published accuracy figures for some commonly used transformer models:

      ModelReported Accuracy
      all-distilroberta-v1~85%
      all-MiniLM-L12-v2~80%
      paraphrase-MiniLM-L6-v2~70%

      To assess performance in a qualification context, we tested two of these models – all-distilroberta-v1 and all-MiniLM-L12-v2 – by asking them to match 50 URS statements against 50 corresponding functional requirements.

      Despite both documents being clearly written, the LLMs mismatched 11 of the 50 statements to the wrong requirements – more than 20%. Even when combining different transformer models to improve accuracy, we never obtained an accuracy of 90%. Of note, this occurred even when both input sets were well structured. In real-world scenarios, where URSs and scripts often vary in phrasing, specificity, or formatting, this represents a serious challenge to the idea that LLMs can automate traceability.

      For matching URS requirements to test scripts the accuracy shortfall would become even more pronounced, as scripts are written in less descriptive and more procedural language than FSs – simply describing steps to be executed, and not their intent. Sentence transformer models are not yet capable of reliably inferring test intent from such procedural language. The point then is that if human experts still need to manually verify each AI-suggested match, what, if any, are the efficiency gains?

      It’s also worth noting that these models are more specialized than general-purpose LLMs like ChatGPT or Mistral. Therefore if transformer-based models – trained specifically for sentence matching – struggle with this task, non-specialised models will likely perform worse.

      Where LLMs Might Help

      Despite these limitations, LLMs can still support certain tasks – particularly when used under SME supervision in low-risk contexts. Potential use cases include:

      • Drafting boilerplate content (e.g., introductory or regulatory reference sections),
      • Improving clarity, grammar, or formatting,
      • Suggesting document structures or outlines,

      These kinds of low-risk, repeatable, and time-consuming activities are suitable areas for automation, with AI acting as a support tool. However, even in these cases, outputs need to be reviewed and approved by qualified personnel to ensure regulatory accuracy and contextual appropriateness.

      Final Thoughts

      The importance of accuracy in qualification, and the limitations of LLMs mean human input is still essential. While LLMs can support low-risk, repetitive tasks like drafting boilerplate content or improving formatting, their role in generating qualification documents remains limited.

      By definition, a URS must reflect what human users need from that asset/service to be purchased – something no AI can independently determine. These documents require expert judgment, deep process understanding, and context-specific decisions that LLMs simply cannot replicate, at this level of maturity.

      Similarly, the use of AI for automating traceability – such as matching URS statements to test steps – still falls short. Even advanced sentence transformer models show accuracy rates that are too low for compliance-critical tasks. In regulated environments, a 20% error rate isn’t just inefficient – it’s dangerous.

      Errors in qualification can have serious consequences: from costly equipment failures to regulatory violations or, worse, risks to patient safety. That’s why experienced professionals must remain in control of the process.

      In short, LLMs may enhance productivity in certain areas, but when it comes to URS, FS, and traceability matrix, human expertise, oversight, and responsibility are irreplaceable.

      ravi

      June 18, 2025
      Uncategorized
      AI, pharmaceutical, validation

    Powering Quality from Lab to Label © 2025, Quascenta Pte. Ltd.