Méthodologie de contrôle qualité / Audit de base de données
Mise à jour le
03 août 2015 (CBA-GM)
Methodology based on GCDMP, Chapter "Measuring Data Quality", last version
1 - Principles
The database audit or quality control is aimed at assessing the quality of the database and detecting the following type of errors:
- Data entry error
- Electronic data acquisition error (e.g. power glitch, back up that didn’t run, cord not attached securely)
- Data linked to wrong subject
- Database updated incorrectly from data clarification form or query
- Programming error in user interface or database or data manipulations
Data quality is quantified using error rate to guard against
misinterpretation of error counts, facilitate comparison of data quality
across database tables and trials.
2 - Methodology
- Data to be checked:
- either: all data of selected CRF
- or: only key data (e.g. inclusion/exclusion criteria, safety data,
primary criteria), depending on the contract
- Critical variables must be agreed with the client prior to the QC / audit
- Sample size:
- either: sample size algorithm "square root plus one" (√n +1) of the total study population
- or: ten percent (10%) of the total study population
- Random CRF / data selection:
- CRF / data are randomly selected by Sunnikan / a Biostatistician /
other
- using a validated tool
- if done by Sunnikan: use relevant randomization tables
[to be
identified by GM - some are provided in appendices of statistical books]
which cannot be refuted by the audited team
- One Auditor reading the CRF data (source) and one Auditor checking the database
- Check documentation:
- Scenario 1: patient data listings available on a paper format: tick
checked data on the listings (green: no discrepancy / red: discrepancy)
and record discrepancies in a table
- Scenario 2: patient data listings not available on a paper format,
and discrepancies are exhaustively reported in a table
3. Counting the checks - Calculating the error rate
Definitions
- A "field" or a "variable" is a particular area (as of a record in a database) in which the same type of information is regularly recorded.
- A "check" is defined as a comparison of a value (field value) between the database and the CRF (or equivalent document).
- A "default field" is a field that does not require data
entry (e.g. Protocol Number, Site Number, Sponsor Number, Patient Number).
- A "derived field" is a field that is not systematically
populated in a CRF because it depends on the answer to a previous question
(e.g. female / male ? – if female menopausal status : yes/no).
- An "error" is defined as a discrepancy between the dataset and the CRF that is not explained by
Data Handling Conventions, site signed Data Clarification Forms, or programming conventions defined in the Data Validation Plan.
- An "error rate" is defined as the number of errors detected divided
by the total number of checks (number of fields inspected). The error rate is expressed as the number of errors per 10,000 fields.
- An “acceptable quality level” means a rate of 50 errors per 10,000
fields overall, with less than 10 errors per 10,000 fields for critical
variables, and less to 100 errors per 10,000 fields for noncritical
variables.
Counting the checks and the errors
- The derived fields are checked (one check per derived field).
- Allocation to the default fields that are electronically populated
throughout the CRF pages (e.g. header) are counted as one check per visit
(or per CRF, depending on the data entry methodology defined for these
particular fields).
- Derived data automatically calculated by programing are neither checked,
nor counted
(e.g. automatic calculation of the BMI, based on height and weight).
- Any Investigator's comments written outside CRF fields do not need
to be checked unless the client has decided to include them in
the statistical analysis. However, they will be available
for information
3 - Points to consider regarding fields counting
The methodology defined in section 2 should be agreed by the Sponsor prior to
QC conduct, in particular regarding the counting of default fields that may
impact the final error rate. An example is given below.
"There are many ways to quantify data quality and calculate an error rate.
While the differences among the methods can be subtle, the differences among the
results can be by a factor of two or more.
For example, consider the hypothetical situation of two lab data vendors
calculating error rates on the same database with three panels. The Protocol
Number, Site Number, and Sponsor Number are default fields that do not require
data entry, in all of three database panels.
Vendor 1 includes each of these default fields in the field count as
fields inspected, which results in a denominator of 100,000 fields inspected in
the error rate calculation. Vendor 2 does not include them in the field count
since they are default fields, for a denominator of 50,000 fields inspected.
Both vendors do a data quality inspection and both vendors find 10 errors.
When they calculate the error rates, Vendor 1 has an error rate half that of
Vendor 2 only because they did not follow the same algorithm for field counts.
This example illustrates how important it is for a common algorithm to be
followed by all parties calculating error rates.
It is imperative that the units in the numerator and denominator be the same.
Some other examples of algorithm details that could skew results are:
- Should data errors involving derived fields be counted?
- Is an error in the month and year fields of a derived date one error or two?
- How should errors be counted in a header that are entered one time then electronically populated throughout the study pages?"
(Good Clinical Data Management Practices, SCDM "Measuring Data Quality - revised Sep. 2008)
4 - Preparation
Main documents & materials required
The database owner or client must provide:
- Read-only access to database or patient datalisting with a layout allowing a rapid and easy check versus CRF
- Original copy of the selected paper CRF including all queries
- Queries should not be filed separately from the CRF
- Data Entry Manual
- Data Management Plan / Annotated CRF / Data Validation Plan
- Self-Evident Corrections (SEC) Manual
- CRF history
- Study Protocol
Agenda
Once the methodology is agreed with the client, the agenda
(v. française) is sent to inform the audited staff of the audit schedule and methodology.
Tool
Using the documents provided by the client /database owner, the audit tools
are prepared to collect discrepancies.
5 - QC/audit report
- Audit /QC Report should be written in English except when a client requests a French version.
- Audit Certificate is issued for audited staff and client.
_____