This is a ready-to-use checklist for the quality reviewer of an AI model credibility package. It is what QA works through before signing the adequacy judgment, and it is the fastest way to catch a trust claim that will fray under an inspector’s follow-up question. Mark each item Pass, Fail, or N/A, and record a comment for anything that is not a clean pass. Replace every <<FILL: ...>> placeholder with your own specifics. Confirm each cited reference against the current source before you rely on it; AI-specific guidance is still evolving.
Review identification
| Field | Entry |
|---|---|
| Model name and version | <<FILL: name, version, artifact ID>> |
| Credibility assessment reviewed | <<FILL: report doc number>> |
| Context of use (one sentence) | <<FILL>> |
| Reviewer | <<FILL: name, role>> |
| Review date | <<FILL: date>> |
1. Context of use
| # | Item | Pass / Fail / NA | Comment |
|---|---|---|---|
| 1.1 | The context of use is a single sentence naming the output, the decision it feeds, and the accountable role | ||
| 1.2 | The conditions and boundaries (inputs, populations, sites) are stated | ||
| 1.3 | Uses that are out of scope are explicitly listed | ||
| 1.4 | The evidence in the package matches this use and is not borrowed from a different use |
2. Model risk sizing
| # | Item | Pass / Fail / NA | Comment |
|---|---|---|---|
| 2.1 | Model influence is scored with a stated basis | ||
| 2.2 | Decision consequence is scored with a stated basis | ||
| 2.3 | The evidence tier follows from influence and consequence, not from convenience | ||
| 2.4 | The risk conclusion ties to the quality risk management framework (ICH Q9(R1)) |
3. Evidence planned before it was gathered
| # | Item | Pass / Fail / NA | Comment |
|---|---|---|---|
| 3.1 | A credibility plan exists and is dated before the evidence was generated | ||
| 3.2 | The plan names the test population and how it was held out from training and tuning | ||
| 3.3 | Metric thresholds are justified and tied to the consequence of each error type | ||
| 3.4 | The plan specifies uncertainty treatment and a minimum sample size for key metrics |
4. Performance evidence
| # | Item | Pass / Fail / NA | Comment |
|---|---|---|---|
| 4.1 | Performance is measured on a genuinely held-out set, ideally from a later time period | ||
| 4.2 | The evaluation population represents the population the model will face in use | ||
| 4.3 | Metrics are not fooled by class imbalance: recall and precision reported, not accuracy alone | ||
| 4.4 | Performance is reported per relevant subgroup (product, site, rare class), not only in aggregate | ||
| 4.5 | Each point estimate carries an uncertainty statement (interval or sensitivity), not a lone decimal | ||
| 4.6 | Calibration is evidenced where the workflow routes by confidence | ||
| 4.7 | Weak spots are identified honestly rather than averaged away |
5. Human control and generative-specific checks
| # | Item | Pass / Fail / NA | Comment |
|---|---|---|---|
| 5.1 | A human control is defined for anything with real consequence, and the review point is recorded | ||
| 5.2 | For a generative or LLM output, the task is constrained and grounded in a retrievable, checkable source | ||
| 5.3 | The model version is pinned, and a vendor-driven model change is treated as a change requiring re-checking | ||
| 5.4 | Explainability claims are not overstated (post-hoc attribution is not presented as literal causation) |
6. Monitoring and currency
| # | Item | Pass / Fail / NA | Comment |
|---|---|---|---|
| 6.1 | A monitoring plan is live before reliance begins: metrics, thresholds, cadence | ||
| 6.2 | Leading indicators (override rate, input-distribution drift) are watched, not just periodic re-scoring | ||
| 6.3 | A monitoring breach has a defined consequence (re-evaluate, restrict, retrain) | ||
| 6.4 | The evidence is current, not a release-day claim months out of date |
7. Adequacy judgment and records
| # | Item | Pass / Fail / NA | Comment |
|---|---|---|---|
| 7.1 | The adequacy judgment is explicit: adequate, adequate with conditions, or not adequate | ||
| 7.2 | Any conditions or bounds are stated, and watch items are fed into monitoring | ||
| 7.3 | Re-assessment triggers are defined (model change, drift breach, use change) | ||
| 7.4 | The package records model, data, and code versions so a later change is visibly a change | ||
| 7.5 | The assessment, evidence, and review decisions are controlled GxP records (ALCOA+) |
Signoff
| Role | Name | Signature | Date | Decision |
|---|---|---|---|---|
| QA reviewer | <<FILL>> | <<FILL: Approve / Return for rework / Reject>> |
References
ASME V&V 40 (credibility framing); ICH Q9(R1) (risk sizing). FDA draft guidance on AI to support regulatory decision-making for drugs and biological products (January 2025 draft; verify status). FDA/EMA Guiding Principles of Good AI Practice in Drug Development (14 January 2026). 21 CFR Part 11 and EU GMP Annex 11; EU AI Act, Regulation (EU) 2024/1689, where high-risk obligations apply.
Confirm the current version of each reference before use.
Filled specimen
The following shows selected rows completed for the example complaint-screening model, so you can see the level of scrutiny expected.
| # | Item | Pass / Fail / NA | Comment |
|---|---|---|---|
| 1.1 | Single-sentence context of use | Pass | ”Suggests a complaint category, sets initial routing, subject to analyst confirmation.” |
| 2.3 | Evidence tier follows from risk | Pass | Low-Medium influence, Medium consequence, Moderate tier. Consistent. |
| 3.1 | Plan dated before evidence | Pass | Plan approved 03 Mar 2026; test results dated 10 to 14 Mar 2026. |
| 4.4 | Per-subgroup performance | Fail | Two low-volume categories under the 0.85 threshold (0.71, 0.74); not resolved in the package. |
| 4.5 | Uncertainty stated | Pass | 95% CIs given for all per-category metrics. |
| 6.1 | Monitoring live before reliance | Pass | Override rate and input distribution baselined at go-live. |
| 7.1 | Explicit adequacy judgment | Pass | ”Adequate with conditions”; two categories on watch. |
Decision: Approve with conditions. The 4.4 Fail was accepted as a bounded, monitored limitation rather than a blocker, because the two categories are low-volume and non-safety, they are on active watch, and the adequacy judgment states the bound explicitly. Had those been safety-relevant categories, the correct decision would have been Return for rework until the evidence met the threshold.
How to adapt this checklist
- Add or remove rows so the checklist matches the evidence tier the model’s risk demands: a Minimal-tier model does not need every Maximal-tier row.
- Point the references to the guidance versions current at your review date.
- Where an item is N/A, record why, so the N/A is a decision and not an omission.
- Use the comment column as the audit trail of the review; a bare column of “Pass” marks with no comments reads as a rubber stamp.
- Feed every Fail and every accepted condition into the monitoring plan and the re-assessment triggers.