Independent and not affiliated with the FDA, MHRA, ISPE, PDA, or any agency. Get the appgoutham@madhadi.com
madhadi.comData Integrity & GxP Quality
Browse all topics → Articles Templates & Procedures Learning paths GlossaryScenariosToolsRegulatory ReferencesLearning PathsTopics About Start here
Checklist Plug-and-play starting point AI & Automation

Checklist: AI Inspection Readiness

A plug-and-play inspection readiness checklist for AI and machine learning in GxP: inventory, intended use, risk assessment, validation evidence, data lineage, model documentation, change and retraining records, human oversight, audit trails, vendor evidence, and monitoring, with pass/fail/NA scoring, a filled specimen, and the regulations it satisfies.

Document type: Checklist

Read and copy the template below into your own quality system. It is a generic starting point for your own internal use, provided as is, with no warranty; see the Terms and License. Adopting it does not by itself create compliance.

This is a ready-to-use inspection readiness checklist for AI and machine learning systems used in a GxP context. Run it per AI system or per deployed model, not once for the whole site, because intended use, risk class, and controls differ model to model. Replace every <<FILL: ...>> placeholder, capture the evidence named in each row, and record a result of Pass, Fail, or N/A. A worked filled specimen follows the template. The AI-regulatory picture is moving, so verify each cited regulation against the current source before you rely on it, and where a reference is draft or proposed, treat it as such.

Document control header

FieldEntry
Document titleAI Inspection Readiness Checklist
Document number<<FILL: CHK-ID, e.g. CHK-QA-021>>
Version<<FILL: version, e.g. 1.0>>
Effective date<<FILL: effective date>>
Document owner<<FILL: role, e.g. Digital Quality Lead / QA>>
Applies to<<FILL: sites / AI system categories in scope>>

How to use this checklist

  1. Pick one AI system or one deployed model. Record its identity in the assessment context below.
  2. Work through each section. For every line, decide Pass, Fail, or N/A, and write down where the evidence is (validation report section, model card, audit trail export, SOP number, record reference). A “Pass” with no evidence is not a pass.
  3. Where a line is a Fail, log it in the gap summary with a risk rating and an owner.
  4. A system is inspection-ready only when every applicable line is Pass and any Fail has a dispositioned remediation plan with interim controls. A single open high-risk fail on a system that drives action means the system is not ready.
  5. Re-run on a defined cycle, before a known inspection, and after any model retrain, architecture change, or change in intended use.

Scoring legend:

  • Pass: control is in place and evidence supports it.
  • Fail: control is missing, partial, or unevidenced. Must be logged and rated.
  • N/A: line does not apply to this system or risk class (state why, for example explainability depth that a low-risk advisory model does not need).

Assessment context

FieldEntry
System / model name and ID<<FILL: SYSTEM NAME / MODEL ID>>
GxP process supported<<FILL: e.g. deviation triage, complaint routing, visual inspection, batch parameter advice>>
AI use patternAdvisory / Automated classification / Process control (<<FILL>>)
Model type<<FILL: e.g. supervised classifier, vision model, LLM via vendor API>>
Build categoryGAMP Category 4 (configured) / Category 5 (custom) (<<FILL>>)
Risk class basis<<FILL: ICH Q9(R1) rationale, High / Medium / Low>>
Model version in production<<FILL: version / hash>>
Assessor (name, date)<<FILL>>

1. AI inventory complete

#CheckPass criterionEvidence to captureResult
IN1System is on the AI inventoryThe model appears on a maintained inventory of AI used in GxP processes, with owner and risk classAI inventory extract<<FILL>>
IN2No hidden AIAny prediction or classification embedded in spreadsheets, vendor features, or scripts is identified and assessed as a model, not overlookedProcess walk-through; inventory completeness review<<FILL>>
IN3Boundary definedThe inventory states where the AI starts and stops in the workflow, and what feeds it and consumes its outputData-flow diagram<<FILL>>
IN4Inventory maintainedThe inventory is reviewed and updated on a defined cycle and on new deploymentReview record; update history<<FILL>>

2. Intended use documented

#CheckPass criterionEvidence to captureResult
IU1Intended use stated preciselyA single approved statement names the model output, the action it triggers, and the accountable roleIntended-use statement, approved<<FILL>>
IU2Use pattern classifiedThe model is classified as advisory, automated classification, or process control, with rationaleRisk class record<<FILL>>
IU3Out-of-scope use statedThe statement says what the model is not to be used forIntended-use statement<<FILL>>
IU4Approved by QAThe intended-use statement is reviewed and approved through document controlApproval signatures<<FILL>>

3. Risk assessment

#CheckPass criterionEvidence to captureResult
RA1Risk assessment performedA documented risk assessment sizes validation effort to patient safety, product quality, and data integrity impactRisk assessment record<<FILL>>
RA2AI-specific risks addressedDrift, explainability limits, training-data bias, automation bias, and model-failure modes are explicitly assessed, not only generic software risksRisk assessment; FMEA where used<<FILL>>
RA3Method documentedThe risk method follows ICH Q9(R1) or an equivalent defined approach, with the formality matched to riskRisk method / SOP reference<<FILL>>
RA4Controls traced to risksEach significant risk maps to a mitigating control (interlock, human review, monitoring, change control)Risk-to-control traceability<<FILL>>

4. Validation and qualification evidence

#CheckPass criterionEvidence to captureResult
VL1Performance spec set before trainingAcceptance metrics and thresholds were defined in the requirements before the model was trained and tested, justified against the consequence of errorURS / performance spec, dated before test report<<FILL>>
VL2Locked test set usedReported performance comes from a version-controlled test set the model never saw in training or tuningTest set version; test report<<FILL>>
VL3Metrics meet specThe reported metrics (for example recall, precision, F1, calibration) meet the approved thresholdsValidation report with confusion matrix or equivalent<<FILL>>
VL4Deterministic controls qualifiedFor process control, the safety interlocks that bound the model are independently validated the conventional wayInterlock qualification record<<FILL>>
VL5Traceability completeTraceability runs from intended use to requirements to test evidence, with rationale recorded where guidance was silentTraceability matrix<<FILL>>

5. Data lineage

#CheckPass criterionEvidence to captureResult
DL1Source and lineage documentedThe origin of the training, tuning, and test data is documented, including whether it was GxP data and how it was controlledData integrity file; source/extract record<<FILL>>
DL2Representativeness assessedThe data is shown to reflect the production population across products, sites, instruments, and edge casesRepresentativeness analysis<<FILL>>
DL3Dataset versionedThe exact training and test datasets are versioned or hashed so the model can be reproducedDataset version / hash record<<FILL>>
DL4Split method recordedThe train, tuning, and test split method is documented, with the test set held outSplit documentation<<FILL>>
DL5ALCOA+ on production data reuseProduction data pulled back for retraining is treated as GxP data under the full ALCOA+ expectationsRetraining data record<<FILL>>

6. Model documentation / model card

#CheckPass criterionEvidence to captureResult
MD1Model card existsA model card or equivalent record describes the model, its purpose, inputs, outputs, and limitsModel card<<FILL>>
MD2Architecture and version recordedModel family, architecture, key hyperparameters, and the production version are documentedModel development record<<FILL>>
MD3Known limitations statedThe model card states known weaknesses, populations it underperforms on, and conditions where it should not be trustedModel card limitations section<<FILL>>
MD4Performance characteristics published to reviewersThe performance profile is available to the operational reviewers who must apply judgment over outputsTraining material; reviewer reference<<FILL>>

7. Change and retraining records

#CheckPass criterionEvidence to captureResult
CH1Change control covers the modelModel retrains, threshold moves, feature changes, and architecture changes are governed under change controlChange control SOP reference<<FILL>>
CH2Predetermined change control planAn approved plan, written before deployment, classifies anticipated changes and the testing each class requiresPredetermined change control plan<<FILL>>
CH3Retrain records completeEach retrain has a record showing the trigger, the data used, the confirmatory test result, and the approvalRetrain change records<<FILL>>
CH4Revalidation threshold definedThe plan states the line above which a change forces full revalidation rather than a confirmatory checkPredetermined change control plan<<FILL>>
CH5Vendor model changes governedFor an API or vendor model, vendor-driven base-model changes trigger a defined re-test and deployment holdVendor change handling record<<FILL>>

8. Human-oversight evidence

#CheckPass criterionEvidence to captureResult
HO1Review step definedThe procedure states what the reviewer does, what information they see, and what decision they ownSOP for the workflow<<FILL>>
HO2Review recordedReviewer conclusions go into a GxP record together with the AI output and the model version reviewedSample reviewed records<<FILL>>
HO3Reviewers trained on weaknessesTraining records show reviewers were trained on the model’s known limitations, not only on the toolTraining records<<FILL>>
HO4Automation bias guardedThe workflow keeps review meaningful (for example reasoning shown, justification on agreement, acceptance-rate monitoring)Workflow design; acceptance-rate metric<<FILL>>
HO5Explainability fits the useThe level of per-decision explanation matches the risk class, and post-hoc explanations are presented as approximations, not literal causesExplanation method note with stated limits<<FILL>>

9. Audit trails

#CheckPass criterionEvidence to captureResult
AT1Inputs, outputs, decisions loggedThe system records the input, the model output, the model version, the reviewer action, and the final decision, time-stamped and attributableAudit trail export sample<<FILL>>
AT2Audit trail cannot be disabled by usersThe audit trail is enabled continuously and ordinary users cannot disable or edit itConfiguration evidence<<FILL>>
AT3Audit trail review definedAudit trail review for the system is defined and performed at a risk-based frequencyAudit trail review records<<FILL>>
AT4Part 11 / Annex 11 controls metAccess control, time synchronization, and electronic signature controls meet 21 CFR Part 11 and Annex 11 where they applyPart 11 assessment<<FILL>>

10. Vendor evidence

#CheckPass criterionEvidence to captureResult
VE1Supplier assessedThe model or platform supplier was assessed for quality and development practices appropriate to the riskSupplier assessment record<<FILL>>
VE2Vendor claims not taken at face valueA vendor “validated AI” claim is not relied on for the site-specific trained instance; site-level evidence existsValidation evidence for the trained instance<<FILL>>
VE3Version pinning and noticeWhere the vendor allows, the model version is pinned, and the vendor’s model-change and notice behavior is documentedVendor configuration; contract / agreement clause<<FILL>>
VE4Vendor responsibilities definedA quality agreement or equivalent defines who is responsible for model maintenance, security, and change noticeQuality agreement<<FILL>>

11. Monitoring

#CheckPass criterionEvidence to captureResult
MO1Monitoring live from day oneA monitoring plan was operating from first production use, not added laterMonitoring plan; start date<<FILL>>
MO2Performance re-checkedPerformance is re-measured against the spec on periodic labeled production samplesPeriodic performance review records<<FILL>>
MO3Drift and override watchedInput drift and the human override / disagreement rate are tracked as leading indicatorsDrift and override metrics<<FILL>>
MO4Triggers and response definedThe plan states what trips a review and the defined response (notify, pause, route to fuller review, investigate)Monitoring plan response section<<FILL>>
MO5Monitoring evidence retainedMonitoring outputs are retained as GxP records and are reviewableRetained monitoring records<<FILL>>

Gap summary

Log every Fail from the lines above here. A system is not inspection-ready while a high-risk fail is open without a dispositioned plan and interim control.

Gap refCheck #DescriptionRisk (H/M/L)OwnerRemediation / due dateInterim controlStatus
<<FILL: G1>><<FILL>><<FILL>><<FILL>><<FILL>><<FILL>><<FILL>><<FILL>>

Overall result

FieldEntry
Lines assessed<<FILL: count>>
Pass<<FILL: count>>
Fail (open)<<FILL: count>>
N/A<<FILL: count>>
Highest open risk<<FILL: H / M / L>>
System verdictReady / Conditional (gaps with plan and interim control) / Not ready
Assessor (name, signature, date)<<FILL>>
QA review (name, signature, date)<<FILL>>

References

21 CFR Part 11 (electronic records and signatures). 21 CFR Part 211 (drug CGMP, including 211.68, 211.180, 211.194) and 21 CFR Part 820 (device quality system; QMSR amendments effective February 2026), as applicable to the process. EU GMP Annex 11 (computerized systems). FDA guidance, “Computer Software Assurance for Production and Quality Management System Software” (current version issued 3 February 2026; the 2022 draft and the 24 September 2025 final were both titled ”…Quality System Software”); it notes the risk-based framework can be applied to AI tools used in production or quality systems. GAMP 5 Second Edition (ISPE, 2022), including its material on machine learning and novel technologies. ICH Q9(R1), Quality Risk Management (Step 4, 2022; FDA final guidance 2023), for sizing effort to risk. FDA AI/ML-Based Software as a Medical Device Action Plan (January 2021), for context. For the predetermined change control plan concept, the source is the FDA final guidance “Marketing Submission Recommendations for a Predetermined Change Control Plan for Artificial Intelligence-Enabled Device Software Functions” (originally issued as final in December 2024, with the current version issued 18 August 2025). Separately, the FDA draft guidance “Artificial Intelligence-Enabled Device Software Functions: Lifecycle Management and Marketing Submission Recommendations” (draft, January 2025) addresses lifecycle management and marketing submissions; treat it as draft. These are device documents; treat them as a source of principles, not as binding on a manufacturing or quality-operations model. EU Artificial Intelligence Act, Regulation (EU) 2024/1689 (in force August 2024, applying in phases). The high-risk system obligations are still being staged in, and a deferral of the high-risk deadline has been proposed but is not yet settled law, so treat the timing as moving. Confirm the current dates and whether your system is in scope before relying on them. IMDRF and the published Good Machine Learning Practice guiding principles, for development practice.

This is general guidance to adapt to your own quality system, not legal or compliance advice. Confirm the current version and clause numbers of each reference, and the current status of any draft or proposed instrument, before issue.

Revision history

VersionDateAuthorSummary of change
<<FILL: 1.0>><<FILL: date>><<FILL: author>>Initial issue.

Approvals

RoleNameSignatureDate
Author<<FILL>>
Reviewer (QA)<<FILL>>
Approver (Quality Head)<<FILL>>

Filled specimen

The following shows part of a completed checklist for an illustrative deviation-triage model that assigns a preliminary criticality tier, which a QA reviewer confirms or overrides. The system, references, and findings are illustrative; replace them with your own.

Assessment context:

FieldEntry
System / model name and IDDeviation Triage Classifier, model DTC-2
GxP process supportedInitial criticality tiering of new deviations
AI use patternAutomated classification with mandatory human confirmation
Model typeSupervised text classifier
Build categoryGAMP Category 5 (custom model on configured platform)
Risk class basisMedium: output sets the investigation timeline; QA confirms within one business day and owns the final tier (ICH Q9(R1) rationale on file)
Model version in productionDTC-2.3, dataset hash 9f1c
Assessor (name, date)A. Patel, 18 June 2026

Sample of completed lines:

#CheckResultEvidence / note
IN2No hidden AIPassProcess walk-through confirmed no other classifier feeds the triage step
IU1Intended use stated preciselyPassApproved statement: model proposes tier, QA confirms or overrides, QA owns final tier
VL1Performance spec set before trainingPassURS dated 03 Feb 2026, threshold recall >= 0.92 on critical class; test report dated 21 Apr 2026
VL2Locked test set usedPassTest set v4, held out, time-split after training window
DL3Dataset versionedPassTraining and test sets hashed; reproduce confirmed
CH2Predetermined change control planFailQuarterly retrain is performed but no approved predetermined change control plan exists; each retrain handled ad hoc
HO4Automation bias guardedFailReviewer acceptance rate is 99.6 percent with no monitoring or justification step, suggesting rubber-stamping
MO3Drift and override watchedPassOverride rate tracked weekly; input drift monitored against training profile

Gap summary:

Gap refCheck #DescriptionRiskOwnerRemediation / due dateInterim controlStatus
G1CH2No predetermined change control plan; retrains handled ad hocMSystem OwnerWrite and approve plan; CR-2026-0188; due 31 Jul 2026Freeze retraining until plan approvedOpen
G2HO499.6 percent acceptance rate suggests rubber-stamp reviewHQA / Process OwnerAdd reasoning display and justification-on-agree; monitor acceptance rate; CAPA-2026-0073100 percent QA second-check on critical-tier deviationsOpen

Overall result: Conditional. One high-risk fail open (G2) with a dispositioned plan and an interim control (full second check on critical-tier deviations) holding the risk while the workflow is redesigned. The verdict, the open high-risk fail, and the interim control are exactly what an inspector expects to see: not a clean sheet, but an honest picture with risk-rated actions and the control gap covered in the meantime.

Common inspection findings this checklist prevents

  • A model drives a GxP decision but never appears on any inventory and was never assessed as a model at all.
  • Performance is reported on data the model trained on, so the headline metric is inflated and unverifiable.
  • The acceptance threshold is dated after the test report, so the criteria were fitted to the result rather than set in advance.
  • Retraining happens with no predetermined change control plan, so the validated state silently lapses each quarter.
  • The system was validated at release with no monitoring, so there is no evidence months later that the model still performs.
  • A defined human review step exists, but a near-100 percent acceptance rate shows the control is exercised on paper only.
  • A vendor “validated AI” claim is relied on for the site-specific trained model with no site-level evidence.
  • Training data has no lineage, no labeling record, and no dataset version, so the model cannot be reproduced or investigated.

How to adapt this checklist

  1. Set your document number, owner, and effective date in the header.
  2. Mark lines N/A by risk class with a stated reason. A low-risk advisory model may N/A the deterministic-interlock line (VL4) and need only light explainability (HO5); a process-control model will not N/A either.
  3. Map your risk ratings to your own ICH Q9(R1) based scoring so the gap summary is consistent across systems.
  4. Point the cross-references (change control, audit trail review, supplier assessment, quality agreement) to your real procedures.
  5. Feed every open fail into your real deviation or CAPA system, with an interim control recorded, not just the table here.
  6. Set the re-run cycle to your risk tiers, and re-run before a known inspection and after any retrain, architecture change, or change in intended use.
  7. Confirm every regulation in the references against the current published version, and check the current status of any draft or proposed instrument, before issue.
Use madhadi.com as an app Full screen, works offline, one tap from your home screen.