This is a ready-to-use SOP for handling incidents on a GxP computerized system that is already in production, live-system events such as an outage, a data-entry rejection, an audit trail gap, or an unexpected calculation result. Replace every <<FILL: ...>> placeholder with your own specifics, set your document numbers and dates, and route it through your normal document control, review, and approval. A worked filled specimen follows the template. Verify each cited regulation against the current source before you rely on it, and treat this as an educational reference to adapt rather than as legal or regulatory advice.
This procedure covers a different moment than two related documents that are easy to confuse with it. The SOP: Validation Test Incident Management governs failures found while a protocol is being executed, before the system is released to production; this SOP starts only after that release. The general SOP: Deviation Management governs the investigation and disposition once a GxP impact is confirmed; this SOP governs everything before that point, the detection, classification, immediate response, and the impact assessment that decides whether a deviation is opened at all.
Document control header
| Field | Entry |
|---|---|
| Document title | Operational Incident Management for GxP Computerized Systems |
| Document number | <<FILL: SOP-ID, e.g. SOP-IT-024>> |
| Version | <<FILL: version, e.g. 1.0>> |
| Effective date | <<FILL: effective date>> |
| Supersedes | <<FILL: prior version or "New">> |
| Document owner | <<FILL: role, e.g. Head of IT Quality / Head of Validation>> |
| Applies to | <<FILL: sites / systems in scope>> |
| Related procedures | Deviation management <<FILL: SOP-ID>>, change control <<FILL: SOP-ID>>, CAPA <<FILL: SOP-ID>>, business continuity <<FILL: SOP-ID>> |
1. Purpose
This procedure defines how <<FILL: COMPANY NAME>> detects, classifies, responds to, investigates the GxP impact of, and closes incidents affecting GxP computerized systems that are in operational, production use. The objective is that every incident, regardless of how it was first noticed, is captured promptly, assessed for whether it touched a GxP record, escalated to a deviation where warranted, and closed with evidence that the system and its data can still be relied on.
2. Scope
This procedure applies to unplanned events affecting any computerized system recorded as GxP in the system inventory maintained under <<FILL: SOP-ID or register reference>>, at the sites named in the header, from the point the system enters production use until it is formally retired. It covers system unavailability, functional malfunctions, data-quality anomalies, audit trail failures, and security events that present operationally (unexpected lockouts, suspicious activity) rather than as a planned security investigation.
It does not cover:
- Test discrepancies found during protocol execution before production release, governed by
<<FILL: SOP-ID for validation test incident management>>. - Deliberate, planned modifications to a system, governed by
<<FILL: SOP-ID for change control>>. Where an incident is caused by a change gone wrong, the change record and this incident record are cross-referenced, not merged into one. - The investigation and disposition activity that follows once a deviation is opened; that work is governed by
<<FILL: SOP-ID for deviation management>>. This procedure hands off to it at the point in section 5.5 where a deviation is confirmed.
3. Responsibilities
| Role | Responsibility |
|---|---|
| Discoverer (any user or support staff) | Reports the event immediately on noticing it, through the defined channel, without attempting an undocumented fix first. |
| System administrator / IT on-call | Performs first response, restores service, and supplies the technical facts (logs, timestamps, what was affected). |
| System owner | Owns the incident record end to end: classification, GxP impact assessment, decision on workaround invocation, and closure. Decides whether a deviation is warranted, subject to QA confirmation. |
| Quality Assurance | Independently confirms or overrides the GxP impact assessment and the deviation decision, and approves closure of every critical and major incident. |
| Process owner / affected department head | Confirms whether GxP activities can continue, authorizes invoking a documented workaround, and owns the reconciliation of any manual data captured during the outage. |
| Vendor / supplier (where applicable) | Investigates vendor-side causes, supplies technical root cause for product defects, and confirms when a fix is available. |
4. Definitions
- Incident: an unplanned event affecting a GxP computerized system’s availability, functionality, or data quality. It is distinct from a deviation: an incident is what happened; a deviation is the quality record opened once GxP impact is confirmed.
- Severity: the operational classification (critical, major, minor) assigned at detection and reassessed as facts emerge, described in section 5.2.
- GxP impact: whether the incident affected the creation, integrity, availability, or reliability of any GxP record, as distinct from mere inconvenience or downtime with no data consequence.
- Workaround: a pre-validated fallback procedure, typically paper-based, that keeps a GxP activity running while the system is unavailable, invoked and reconciled per section 5.7.
- Restoration: the point the system returns to normal operation, which is not the same point as the incident being closed; restoration ends the outage, closure ends the record.
- Reconciliation: the documented, independently verified transfer of any manual or workaround data into the system once it is restored.
5. Procedure
5.1 Detect and report
- Any user, administrator, or monitoring alert may be the first detection of an incident. The discoverer reports it immediately through
<<FILL: channel, e.g. the IT service desk, an incident hotline, the monitoring platform's alert queue>>. - The discoverer does not attempt an undocumented workaround or repeated retries before reporting. A brief factual note of what was observed (what was tried, what happened, the time) is captured at the point of detection, not reconstructed afterward.
- An incident record is opened with a unique identifier within
<<FILL: e.g. 1 hour for critical, same business day for major and minor>>of detection.
5.2 Classify severity
Assign one of three severities at intake, and re-assess as facts emerge. A severity may be raised or lowered as the picture becomes clearer, with the reason recorded.
| Severity | Definition | Example | Response target |
|---|---|---|---|
| Critical | The system is unavailable, or is generating or holding data that cannot be trusted | A LIMS rejecting all result entry during an active release window; an MES that lost data mid-batch; an audit trail that silently stopped recording | Immediate response; restoration and GxP impact assessment both start without delay |
| Major | The system is degraded, or a specific function is down, but GxP activities can continue with a documented workaround | A report producing incorrect output for one configuration; an interface failing but with a manual fallback available | Response and resolution within <<FILL: e.g. 8 business hours>> |
| Minor | Cosmetic issues or performance degradation with no GxP data-quality consequence | A slow-loading screen with no functional impact | Normal support queue |
5.3 Respond and restore
- The system administrator or IT on-call performs first response: contain the issue, prevent further data loss or corruption, and work toward restoration.
- For a critical or major incident, the process owner decides, with the system owner, whether a documented workaround is invoked per section 5.7 while restoration is in progress.
- Restoration is confirmed and timestamped. Restoration ends the outage; it does not close the incident record, which still requires the impact assessment, any deviation, and reconciliation to be complete.
5.4 Escalate and notify
- Follow the notification chain defined for the system’s criticality tier: at minimum, the system owner and QA are notified of every critical incident on detection, and of every major incident within the shift.
- For a critical incident during a release window, batch-critical activity, or another time-sensitive GxP process, notification also reaches whoever is authorized to approve a workaround, so that decision does not wait on the standard escalation chain.
- Provide status updates at a defined cadence while the incident remains open,
<<FILL: e.g. every 30 to 60 minutes for critical>>, so the record shows the response was active rather than reconstructed after the fact.
5.5 Assess GxP impact
- Every incident on a GxP system gets a documented impact assessment, regardless of severity, answering: did this affect any GxP record, existing or in progress, and if so, which records and how.
- Where the answer is no, meaning no data already in the system was altered, lost, or made unreadable, and no record created during the event is in question, document that conclusion with its basis, restore service, and proceed to closure without opening a deviation.
- Where the answer is yes, or where an audit trail, timestamp, or access control gap means the normal evidence needed to answer the question is itself missing, open a deviation under
<<FILL: SOP-ID for deviation management>>immediately. Do not wait for the incident record to be finished. Cross-reference the deviation number on the incident record and the incident number on the deviation record. - The impact assessment is written at the time, not reconstructed weeks later. A conclusion drafted from memory after the fact is treated as unreliable evidence at periodic review.
5.6 Assess scope
For any incident that opens a deviation, assess whether the same cause could affect other systems, other functions of the same system, or data already produced before the incident was noticed. An audit trail gap discovered in one module often applies to a shared logging service used elsewhere; check before assuming the finding is contained.
5.7 Invoke and reconcile a workaround
- A workaround is invoked only where a validated fallback procedure exists for the activity; an improvised paper substitute invented during the outage is not a workaround, it is a new, undocumented process and is itself a finding.
- The workaround record (paper form, manual log) meets the same ALCOA+ expectations as the system record it temporarily replaces: attributable, legible, contemporaneous, original, and accurate, with the person and time captured at the point of entry.
- Once the system is restored, the workaround data is transcribed or entered into the system and independently verified by someone other than the person who captured it, with the original workaround record retained as source data, not discarded once transcribed.
- Reconciliation is itself documented: what was transferred, by whom, verified by whom, and any discrepancy found and how it was resolved.
5.8 Investigate root cause
- For a critical or major incident, or any incident that opened a deviation, determine root cause proportionate to risk. A vague cause such as “system was slow” is not acceptable; the cause must be specific enough that a corrective action can act on it.
- Where the incident is a vendor product defect, the vendor’s own root cause analysis is reviewed and assessed by the system owner or SME rather than accepted without review.
- Raise a CAPA under
<<FILL: SOP-ID for CAPA management>>where the cause could recur or affects more than the single instance.
5.9 Close the incident
An incident closes only when all of the following are complete and recorded:
- Restoration is confirmed and timestamped.
- The GxP impact assessment is documented with its conclusion and basis.
- Where a deviation was opened, it is cross-referenced and its status noted (open incidents do not wait for the deviation to fully close, but the link must exist).
- Any workaround used is reconciled and independently verified.
- Root cause and CAPA references are recorded where applicable.
- The system owner signs the closure; QA approves closure for every critical and major incident.
5.10 Feed the operational record
- The incident is logged in the system’s operational incident register so it is visible at the next periodic review, not just in a ticketing system that periodic review does not look at.
- Trend the incident log at the cadence used for operational monitoring: a repeated incident on the same function is a pattern, not a series of unrelated coincidences, and is escalated as such even where each individual occurrence was minor.
6. Acceptance criteria
- Every incident is logged with a timestamp at or near detection, not reconstructed afterward.
- Every incident carries a severity classification with a stated basis, re-assessed if the facts changed.
- Every incident has a documented GxP impact assessment with an explicit conclusion, written at the time.
- Every incident where the answer to the impact assessment is yes, or where the evidence needed to answer it is itself missing, has a cross-referenced deviation.
- Any workaround used is a pre-validated procedure, not an improvisation, and its data is reconciled and independently verified before the incident closes.
- Root cause and CAPA are recorded for every critical or major incident and for any recurring minor incident.
- QA has independently reviewed and approved closure of every critical and major incident.
- The incident register is a visible input to the next periodic review.
7. Records generated
| Record | Owner | Retention |
|---|---|---|
| Incident record (detection through closure) | System Owner | <<FILL: retention period>> |
| GxP impact assessment | System Owner / QA | <<FILL>> |
| Workaround / fallback records and reconciliation evidence | Process Owner | <<FILL>> |
| Deviation and CAPA records raised from an incident | Quality Assurance | Per deviation and CAPA procedures |
| Operational incident register (trending input to periodic review) | System Owner | Life of the system plus <<FILL: period>> |
8. References
21 CFR 211.68(b) (equipment and records shall be maintained and protected), for the underlying expectation that a system failure affecting records is addressed and documented. 21 CFR 211.192 (production record review; investigation of unexplained discrepancies), applied by analogy to a computerized system incident that touches a GxP record. EU GMP Annex 11, Computerised Systems, clause 13 (incident management: incidents, not only system failures, should be reported and assessed, with root cause identified for critical incidents and corrective and preventive action taken). A substantially expanded draft revision was published for consultation in July 2025 and had not been finalized as of this writing; confirm the current in-force version before citing clause numbers. ICH Q9(R1), Quality Risk Management, for the risk basis of severity classification and investigation depth. ISPE GAMP 5, Second Edition (2022), for the operational-phase treatment of incident management within the system lifecycle. Reference only; do not reproduce its text.
Confirm the current version and clause numbers of each reference before issue.
9. Revision history
| Version | Date | Author | Summary of change |
|---|---|---|---|
<<FILL: 1.0>> | <<FILL: date>> | <<FILL: author>> | Initial issue. |
10. Approvals
| Role | Name | Signature | Date |
|---|---|---|---|
| Author | <<FILL>> | ||
| Reviewer (IT / Infrastructure) | <<FILL>> | ||
| Reviewer (System Owner representative) | <<FILL>> | ||
| Approver (Quality Assurance) | <<FILL>> |
Filled specimen
The following shows an incident on a clinical Electronic Data Capture platform. The company, system, names, and numbers are illustrative; replace them with your own.
| Field | Entry |
|---|---|
| Incident ID | OPS-INC-2026-0208 |
| System | EDC production instance, inventory ID EDC-PROD-01 |
| Detected | 2026-08-18 14:12, by site coordinator reporting a form submission error |
| Description | Randomization form at Site 014 submitted twice within 40 seconds after a browser timeout; second submission created a duplicate subject visit record |
| Classification | Major, one site affected, data-quality concern, no system-wide outage |
| Immediate response | IT on-call confirmed the duplicate record at 14:40; the form was locked from further edits at 14:52 pending review |
| Notification | System owner and QA notified 14:45, within the shift target |
| GxP impact assessment | Yes. Two records exist for one subject visit, one of which is not a genuine event. Deviation opened same day |
| Deviation | DEV-2026-0157, opened 2026-08-18 |
| Workaround | None required; no downtime, EDC remained otherwise available |
| Root cause | Client-side retry on browser timeout resubmitted the form without server-side idempotency check; a known interaction with the sponsor’s proxy configuration at that site |
| Scope check | Reviewed submission logs across all sites for the prior 90 days; two further duplicate pairs found at a different site, same proxy configuration, both flagged for the same review |
| CAPA | CAPA-2026-071, add server-side idempotency check to the form submission handler, vendor confirmed fix in next release; interim compensating control, a daily duplicate-record report reviewed by the data management team |
| Closure | 2026-08-27, duplicate record disposition confirmed by unblinded data manager per DEV-2026-0157, system owner and QA signed |
What made this defensible: the coordinator reported instead of quietly re-entering data, the system was not treated as merely “glitchy,” the impact assessment reached a clear yes and a deviation opened the same day, the scope check went looking for the same pattern elsewhere instead of assuming one site was unique, and the interim compensating control covered the gap until the permanent fix shipped.
Common inspection findings this SOP prevents
- An outage or malfunction is logged only as an IT help-desk ticket, with no GxP impact assessment anywhere in the record.
- An incident’s severity is downgraded after the fact to avoid the escalation and QA review a higher severity would require.
- A workaround was used but was never a validated procedure, so the fallback data itself cannot be relied on.
- Workaround data was transcribed into the system but never independently verified, so a transcription error propagated silently.
- The GxP impact assessment is written weeks after the event, from memory, once an audit or inspection prompts someone to produce one.
- An incident is closed on restoration alone, with the impact assessment, deviation cross-reference, or reconciliation still outstanding.
- The same incident recurs three or four times, each logged and closed independently, with no one connecting the pattern until an inspector does it for them.
- The operational incident register and the deviation log carry different counts for the same period, with no reconciliation between the two.
How to adapt this SOP
- Set your document number, owner, effective date, and the related-procedure references in the header and in section 2.
- Point every cross-reference to your real change control, deviation, CAPA, and business continuity procedures.
- Set your severity definitions and response targets in section 5.2 to match your actual operating capability; a target you cannot meet is worse than a longer one you can.
- Name your actual notification channels and escalation roles in sections 5.1 and 5.4.
- Confirm your workaround procedures are validated and referenced by SOP number before relying on section 5.7; a workaround with no prior validation is not usable under this procedure as written.
- Confirm every regulation in section 8 against the current published version before issue, including the in-force status of the Annex 11 revision.