GCP · Blog
Back to journal

Clinical Trial Data Integrity: What the Data Life Cycle Actually Requires

The most useful definition comes from MHRA, and its second half is the part that does the work.

GCP 9 min read
A

Aileen

Aileen writes practical guidance for clinical trial teams at GCP Blog.

On this page · 15 sections
  1. 01 At a glance
  2. 02 What data integrity actually means
  3. 03 The eight stages
  4. · 1. Capture
  5. · 2. Metadata and audit trails
  6. · 3. Review of data and metadata
  7. · 4. Corrections
  8. · 5. Transfer, exchange and migration
  9. · 6. Finalisation before analysis
  10. · 7. Retention and access
  11. · 8. Destruction
  12. 04 Data you did not capture yourself
  13. 05 Governance is what holds it together
  14. 06 Where this goes wrong in practice
  15. 07 Sources

At a glance

  • Data integrity is not a property of a record. It is a property of a life cycle, and the definition says so: the characteristics have to be maintained throughout it.
  • ICH E6(R3) names eight life-cycle stages, from capture to destruction. Each has its own obligation and its own way of failing.
  • Verification effort is meant to scale with how critical the data is, not to be applied uniformly.
  • Review of data and audit trails is a planned, risk-based activity. Treating it as something done when someone has time is the most common gap.
  • Migration between systems is a data-integrity event, not an IT task. Integrity has to survive the move.

What data integrity actually means

The most useful definition comes from MHRA, and its second half is the part that does the work.

Data integrity is the degree to which data are complete, consistent, accurate, trustworthy, reliable, and that these characteristics of the data are maintained throughout the data life cycle (MHRA GxP Data Integrity Guidance, definitions).

Read the two halves separately. The first is a list of attributes, and it is where most writing on this subject stops. The second is a durability requirement, and it is where trials actually fail. A value that was accurate when a coordinator typed it, and is unverifiable three systems and two years later, has not partially satisfied the definition. It has failed it.

That reframing matters because it moves the question from “is this record good?” to “can this record still be trusted, and shown to be trustworthy, at every point between entry and archive?”

The attributes themselves, commonly abbreviated ALCOA, are covered separately and in detail. This page is about the life cycle that has to preserve them.

The eight stages

ICH E6(R3) is unusually concrete here. Procedures should be in place to cover the full data life cycle (ICH E6(R3) §4.2), and the guideline then enumerates the stages rather than leaving “life cycle” as a gesture. Taking them in order gives you a map of where integrity is won or lost.

1. Capture

When data captured on paper or in an electronic health record are manually transcribed into a computerised system, the need for and the extent of data verification should take the criticality of the data into account (ICH E6(R3) §4.2.1).

The instruction is proportionality, not uniformity. Transcription of a primary endpoint warrants verification that a non-critical field does not. Teams that verify everything to the same standard usually end up verifying nothing well, and teams that verify nothing have no basis for trusting the fields that matter.

The failure mode is undocumented transcription: a number that moved from paper to system with no record of who moved it or whether anyone checked.

2. Metadata and audit trails

The approach used by the responsible party for implementing, evaluating, accessing, managing and reviewing relevant metadata associated with data of higher criticality should entail evaluating the system for the types and content of metadata available (ICH E6(R3) §4.2.2).

Note “the responsible party” and “higher criticality” again. This is not a requirement to hoard every audit-trail entry equally. It is a requirement to know what metadata your systems produce, decide which of it matters, and be able to get at it.

The failure mode is a system whose audit trail nobody has ever looked at, discovered during an inspection to log changes without capturing why they were made.

3. Review of data and metadata

Procedures for review of trial-specific data, audit trails and other relevant metadata should be in place. It should be a planned activity, and the extent and nature should be risk-based, adapted to the individual trial and adjusted based on experience during the trial (ICH E6(R3) §4.2.3).

Three requirements are packed into that. Review is planned, not opportunistic. It is risk-based, not uniform. And it is adjusted based on experience, meaning a plan written at startup and never revisited does not satisfy it.

This is the stage most often left undefined. Sites and sponsors frequently have the capability to review audit trails and no documented statement of when they will, on what, or why.

4. Corrections

There should be processes to correct data errors that could impact the reliability of the trial results. Corrections should be attributed to the person or computerised system making the correction, justified, and supported by source records around the time of original entry (ICH E6(R3) §4.2.4).

Three conditions: attributed, justified, and supported by contemporaneous source records. The last one is the one people miss. A correction made months later, justified only by recollection, cannot be supported in the way this asks for.

Corrections are not failures of integrity. Undocumented corrections are.

5. Transfer, exchange and migration

Validated processes, or other appropriate processes such as reconciliation, should be in place to ensure that electronic data including relevant metadata transferred between computerised systems retains its integrity and preserves its confidentiality (ICH E6(R3) §4.2.5).

Every trial moves data: lab to EDC, EDC to statistical environment, one vendor’s platform to another’s. Each move is a point where integrity can be lost silently, most often by the metadata failing to travel with the data. Content arrives intact and the audit trail explaining it does not.

The guideline offers validation or reconciliation, which is a practical concession: you either prove the process is sound in advance, or you check the result afterwards. What is not offered is doing neither.

6. Finalisation before analysis

Data of sufficient quality for interim and final analysis should be defined and achieved by implementing timely and reliable processes for data capture, verification, validation, review and rectification of errors (ICH E6(R3) §4.2.6).

“Defined” is the operative word. What counts as sufficient quality for analysis is something you decide in advance, not something you conclude after looking at the data.

7. Retention and access

The trial data and relevant metadata should be archived in a way that allows for their retrieval and readability and should be protected from unauthorised access and alterations throughout the retention period (ICH E6(R3) §4.2.7).

Retrievable, readable, and protected from alteration, for the entire retention period. How long that period is, and how it differs by jurisdiction, is a subject of its own and is covered separately. What matters here is that integrity obligations do not end when the trial does.

8. Destruction

The trial data and metadata may be permanently destroyed when no longer required as determined by applicable regulatory requirements (ICH E6(R3) §4.2.8).

Destruction is a stage of the life cycle, not the absence of one. “When no longer required as determined by applicable regulatory requirements” means the destruction date is a regulatory calculation, and calculating it wrongly in the early direction is the one error in this entire life cycle that cannot be corrected.

Data you did not capture yourself

A large share of trial data never passes through the site’s hands: central laboratory results, centrally read imaging, device output, eCOA responses, records from another institution. The integrity obligation follows the data anyway, and two provisions make that explicit.

FDA’s guidance on electronic source data enumerates who can originate data in a trial, and the list is deliberately broad: clinical investigators and delegated staff, participants or their legally authorised representatives, consulting services such as a radiologist reporting on a CT scan, medical devices such as an electrocardiograph or blood pressure machine, electronic health records, and automated laboratory reporting (FDA Electronic Source Data in Clinical Investigations, 2013). Each of those is a point of origin whose output has to carry its provenance forward.

That guidance also sets the durability test for the resulting record: data element identifiers should allow sponsors, FDA and other authorised parties to examine the audit trail of the eCRF data, and that audit trail should be readily available in a human readable form (FDA Electronic Source Data in Clinical Investigations, 2013). Human readable is doing real work there. An audit trail that exists only as a vendor-proprietary export nobody can interpret does not satisfy it.

On the site side, the investigator is responsible for the timely review of data including relevant data from external sources that can have an impact on participant eligibility, treatment or safety, such as central laboratory data, centrally read imaging data and other institutions’ records (ICH E6(R3) §2.12.3).

Read that against the life cycle above and the implication is uncomfortable but clear. The investigator owes review of data they did not generate, cannot directly correct, and often receive on someone else’s schedule. That is exactly why §3.16.1(k) obliges the sponsor to provide timely access: without it the site is accountable for reviewing data it cannot see in time to act on.

The practical consequence for laboratory and eCOA data specifically is that integrity questions have to be settled in the vendor arrangement, before first participant in. Who holds the audit trail, in what format, for how long, and how does the site get at it? Left to the end of the trial, those questions have already been answered badly.

Governance is what holds it together

Eight stages with eight separate owners is not a system. MHRA puts the connective tissue plainly: data governance must be applied across the whole data life cycle to provide assurance of data integrity (MHRA GxP Data Integrity Guidance, data governance).

Governance also carries an externally facing obligation. Data governance systems should ensure that data are readily available and directly accessible on request from national competent authorities (MHRA GxP Data Integrity Guidance, data governance). Readily available and directly accessible are tests your arrangements either pass or fail, and arrangements that depend on a departed employee or a lapsed vendor contract fail them.

The sponsor-side duty sits alongside this. The sponsor should ensure the integrity and confidentiality of data generated and managed, and should apply quality control to the relevant stages of data handling to ensure the data are of sufficient quality to generate reliable results (ICH E6(R3) §3.16.1). At site level the same obligation lands on the investigator: in generating, recording and reporting trial data, the investigator should ensure the integrity of data under their responsibility, irrespective of the media used (ICH E6(R3) §2.12.1).

That last phrase is worth keeping. Irrespective of the media used. Paper and electronic records carry identical obligations, and a well-run paper trial has better data integrity than a badly run electronic one.

Where this goes wrong in practice

The failures cluster in predictable places, and none of them are technology problems.

No defined review. The capability exists; the plan does not. §4.2.3 asks for a planned, risk-based activity, and “we can pull the audit trail if we need to” is not one.

Metadata left behind at a boundary. Data survives a migration and its audit trail does not, so everything before the move becomes unverifiable.

Corrections without contemporaneous support. The correction is attributed and justified, but the source records that would substantiate it were never captured at the time.

Uniform effort. Verification applied evenly across critical and trivial fields, which reads as diligence and functions as dilution.

Integrity treated as ending at database lock. Stages 7 and 8 are part of the life cycle, and an archive nobody can read is a failure of the same obligation that governs data entry.

The unifying test across all eight stages is simple to state and uncomfortable to apply: could someone who was not there reconstruct what this data was, who touched it, and why it changed? If yes at every stage, you have integrity. If the answer breaks down at any single stage, you do not have it anywhere downstream of that point.

Sources

  • ICH E6(R3) Good Clinical Practice, version R3
  • MHRA ‘GxP’ Data Integrity Guidance and Definitions (2018)
  • FDA Guidance for Industry: Data Integrity and Compliance With Drug CGMP (2018)
A

Written by

Aileen

Aileen writes practical guidance for clinical trial teams at GCP Blog.