Share

On January 14, 2026, FDA and the European Medicines Agency published ten joint guiding principles for good AI practice in drug development. The principles added another piece to a regulatory framework already taking shape: FDA's draft guidance on AI to support regulatory decision-making, EMA's reflection paper on AI across the medicinal product lifecycle, and ICH E6(R3), whose second annex arrived in June 2026.

Most of the discussion has focused on submissions. How will regulators assess AI-generated evidence? How should model credibility be established for a defined context of use?

Preparation that starts at the submission stage starts too late. The first hard questions about AI may not wait for submission. They can arrive through a different channel: the GCP inspection.

The first hard questions about AI may not wait for submission. They can arrive through a different channel: the GCP inspection.

Inspections run on the expectations that already exist

Guidance documents get the press releases. Inspections run on a slower, older clock. An inspector walking into a sponsor or CRO environment can already apply many of the same expectations that govern computerized systems today: data integrity, validated systems, audit trails, documented oversight of service providers. ICH E6(R3), now FDA final guidance, sharpened those expectations with proportionate, risk-based quality management and explicit sponsor accountability for computerized systems and the vendors behind them. Annex 2 extends the same logic into decentralized elements and real-world data.

AI already sits inside trial conduct, whether or not anyone has labeled it that way. Electronic outcome assessment platforms embed automated scoring and anomaly flags. Central monitoring systems generate risk signals. Medical coding tools propose MedDRA terms. Safety databases run signal-detection algorithms. Authoring tools draft protocol text that humans edit. Some of these uses fall outside the scope of FDA's draft AI guidance, which excludes certain operational-efficiency applications that do not affect patient safety, drug quality, or the reliability of study results. That does not make them automatically irrelevant to GCP. Where an AI-enabled function affects trial conduct, participant protection, data integrity, traceability, or the reliability of trial results, existing computerized-system and GCP expectations can already become relevant.

Six questions worth preparing for

I spend much of my professional time helping organizations prepare for inspectional scrutiny of computerized systems, data flows, and vendor oversight. If I were asked to write the inspector's playbook for AI tomorrow, it would contain six questions.

1. "Where is AI used in this trial?"

Most organizations cannot answer this today. AI capability arrives embedded in vendor platforms, sometimes switched on in routine updates, described in release notes that nobody in clinical operations reads. A defensible answer starts with a living register: every model or AI-enabled function that touches trial conduct, its intended purpose, its accountable owner, its risk classification. One page per system, updated when the vendor updates. Without the register, every answer that follows is improvisation.

2. "How do you know it works for the way you use it?"

FDA's draft guidance organizes credibility around context of use: what the model is for, in what setting, with what consequence if it errs. Even where the draft guidance is not directly in scope, its context-of-use logic is consistent with the risk-based, fit-for-purpose validation principles already embedded in modern GCP. A vendor's validation package may establish that the platform performs as designed, but it does not automatically establish fitness for your configured use, your data environment, or the decisions you make from its outputs. A good answer is proportionate and documented: confirmation in your own environment, acceptance criteria tied to the decision the output informs, and change control that notices when the model, its inputs, or its use drift from what was assessed.

3. "Show me how this output was produced."

This is the lineage question, and it is where inspections get uncomfortable. The inspector picks a data point, a coded adverse event term or a site flagged for atypical data, and walks backward to source. Every automated transformation along the way needs the same discipline as any manual step: attributable, legible, contemporaneous, original, accurate. If a model recoded, translated, summarized, or scored the data, the audit trail should show what went in and what came out, along with the model version that did the work. Many organizations have lineage diagrams in validation binders. Far fewer can produce one for the study sitting in front of the inspector.

4. "Who reviewed this before it was used?"

The joint principles put human-centric design first, and E6(R3) keeps responsibility with the sponsor regardless of how the work gets done. A model can hold the task. The accountability stays with you. A good answer names roles, shows training records, and provides evidence that review actually happens: sampled quality checks, documented sign-offs, recorded overrides where a human disagreed with the system. "The vendor monitors performance" answers a different question.

5. "How do you oversee the companies behind it?"

E6(R3) makes the subcontracting issue explicit: sponsor oversight extends to important trial-related activities that a service provider further subcontracts. With AI, the technology stack may reach beyond the vendor named in your contract, extending to model providers, cloud providers, data processors, and other technology partners. Your outcome-assessment vendor may build on a foundation model from a provider you have never heard of. Your monitoring platform may change models between your study's interim and final analyses. Contracts need audit access, notice of material model changes, and clarity on which sub-processors touch subject data. The inspector will ask what your agreement does about your vendor's vendor. Have an answer before the question arrives.

6. "What happened the last time it failed?"

Every system fails. Inspectional interest sits in what your organization did about it: incidents logged, deviations assessed for impact on data and subjects, corrective and preventive actions where warranted, and a fallback procedure that staff can execute under pressure. Here is a contrarian point from experience: for a heavily used, evolving AI-enabled system, a spotless incident log after two years can look less like perfection than under-reporting. Mature organizations make failures visible and show how they were handled.

A spotless incident log after two years can look less like perfection than under-reporting. Mature organizations make failures visible and show how they were handled.

Write for the reader you will actually have

Most AI validation documentation is written by technical teams for technical reviewers. GCP inspectors are a different audience. They read Part 11-style evidence: intended use, specifications, test results, change history, training records. If the quality unit cannot explain a system without the data science team in the room, the documentation has a gap, and the inspection will find it.

The fix is translation. Each system in the register deserves a short, plain summary an inspector can read in five minutes: what the system does in one sentence, what decision it informs, its risk classification, what was tested, what changed since, where the records live. Write it for a smart reader who has never heard of a transformer architecture. That same summary will serve boards, auditors, and partners, who ask the same questions in different vocabulary.

Where to start

Four actions fit inside a quarter:

  1. Build the register for one ongoing study. Every AI touchpoint, end to end, including the ones embedded in vendor tools. Then scale what you learn.
  2. Trace one data point from source to submission through every automated step. Do it live, in the systems, with the people who run them.
  3. Add one AI scenario to your next mock inspection. Question 3 or question 6 will exercise the most muscle.
  4. Reconcile SOPs with reality. Procedures should name the tools actually in use, and tool version changes should trigger documented review.

Companion Resource

The 4-Layer AI-Readiness Scorecard

A free, five-page self-assessment across the four readiness layers: quality-by-design, structured protocols, governed data, and AI-enabled execution, including the intended-use, lineage, human-review, audit-trail, and validation questions an inspection will touch.

  • Thirteen questions, mapped to each layer's anatomy
  • Aligned with ICH E6(R3), ICH M11, and CDISC USDM
  • Print-ready · No form, no gate
Open the Scorecard

Self-Assessment Scorecards

Inspection Readiness Is the Practical Test

The guidance will keep moving. FDA's draft will finalize, E6(R3) implementation will mature, and the joint principles will acquire annexes and companions. The six questions above will stay stable, because they are the questions GCP has always asked of computerized systems, now pointed at systems that learn. Organizations that answer them calmly will hold something more valuable than a compliance position. They will know where their AI is, what it does, and when to doubt it.

A readiness question, then, in the tradition of this blog: if an inspector asked tomorrow where AI touches your ongoing trials, who in your organization would answer, and how long would it take them?

Further reading on kushdhody.com

Selected sources

About the author

Kush Dhody, M.D., M.S. is a physician-scientist and clinical development executive with more than 20 years of experience leading global clinical programs, protocol design, regulatory strategy, and clinical operations across multiple therapeutic areas. He currently serves as President of Amarex Clinical Research, LLC, An NSF Company, and is involved in AI-enabled regulatory and quality workflow innovation, including the NSF/Microsoft Azure initiative featured as a Microsoft customer story.

DISCLAIMER: The views expressed in this blog are those of the author and do not necessarily represent the official position of Amarex Clinical Research, LLC, An NSF Company, NSF, any sponsor or partner, or any regulatory authority. This post reflects the author's interpretation of publicly available information and emerging developments in artificial intelligence in drug development, GCP inspection practices, computerized system expectations, and regulatory modernization. Adoption of any approach discussed here should be evaluated in the context of the specific product, study design, therapeutic area, regulatory jurisdiction, organizational capabilities, and applicable health authority expectations. It is intended for informational and educational purposes only and should not be construed as regulatory, legal, compliance, or medical advice.

Share this post