Agentic AI for patient data processing and programming in clinical trials

By Sofía Sánchez González
Clinical trials generate large volumes of patient-level data across multiple systems, formats and stages of the study lifecycle. Before that data can support statistical analysis or regulatory submissions, it must be cleaned, standardized, transformed, validated and programmed according to predefined specifications.
Much of this work still depends on highly structured but manual processes.
Agentic AI introduces a different approach. Instead of using AI only to generate an isolated output, AI agents can execute and coordinate multiple steps across a clinical data workflow, while operating within predefined rules and maintaining human oversight.
For clinical data teams, the opportunity is not simply to process data faster. It is to make complex data workflows more automated, traceable and reproducible.
What does patient data processing involve in clinical trials?
Clinical data rarely moves directly from collection to analysis.
Data may originate from electronic data capture (EDC) systems, laboratory systems, ePRO platforms, medical devices and other sources. It then needs to be reviewed, transformed and structured before it can be used for analysis or regulatory reporting.
Standards such as CDISC SDTM provide a common structure for organizing clinical study data, while ADaM supports analysis-ready datasets and traceability between analysis results and source study data.
Typical activities include:
- Mapping source data to standardized structures
- Applying transformation and derivation rules
- Identifying missing or inconsistent values
- Generating SDTM and ADaM datasets
- Producing or assisting with statistical programming code
- Validating datasets and outputs
- Maintaining metadata and traceability between transformations
These activities are highly structured, which makes them particularly relevant for agentic AI.
How can agentic AI process clinical data?
Traditional automation usually follows a predefined sequence: when a specific condition occurs, the system executes a predetermined action.
An AI agent can work differently.
Given a defined objective, access to authorized data and an established set of rules, an agent can interpret the task, determine which steps are required, use appropriate tools and evaluate intermediate outputs before continuing.
For example, an agent supporting SDTM preparation could:
- Inspect incoming source datasets and metadata
- Identify relevant variables and domains
- Compare them against predefined mapping specifications
- Propose or execute transformations
- Identify inconsistencies or missing information
- Run validation checks
- Flag exceptions requiring human review
- Document the actions performed
The important distinction is that the agent is not simply generating a dataset. It is participating in the workflow used to create and verify that dataset.
Agentic AI and statistical programming
Programming is another area where agentic workflows could change how clinical data teams operate.
Statistical programmers frequently work with repeatable tasks: dataset creation, derivations, validation programs, tables, listings and figures (TLFs), quality checks and updates following specification changes.
AI agents can support these activities by generating or modifying code according to controlled specifications and then performing additional steps around that code.
An agent could, for example, interpret an analysis specification, generate the required program, execute validation checks, compare the resulting dataset with expected structures and identify discrepancies for review.
This moves AI beyond simple code generation.
The objective becomes a controlled workflow in which programming, execution, checking and documentation can be connected.
Traceability becomes more important, not less
Greater automation does not reduce the need for traceability.
It increases it.
FDA‘s current Study Data Technical Conformance Guide includes specific expectations around study data validation and traceability. CDISC similarly emphasizes traceability between analysis results, ADaM datasets and their SDTM inputs.
If an AI agent transforms patient data or generates programming code, teams need to understand:
- What data the agent accessed
- Which instructions and specifications it followed
- Which transformations it performed
- What code or output it generated
- What validation checks were executed
- What changed between versions
- Which actions required human approval
The goal should therefore not be autonomous processing without visibility. It should be controlled automation with an auditable path from input to output.
Patient data requires additional controls
Clinical data processing also involves sensitive information.
ICH E6(R3) places explicit emphasis on data governance across the clinical trial data lifecycle, including data integrity, traceability, security, confidentiality and appropriate management of computerized systems. It also states that data transfers and migrations should preserve integrity and confidentiality and be documented to ensure traceability.
The EMA similarly recommends a human-centric approach to AI in the medicinal product lifecycle and emphasizes data protection, governance and appropriate oversight.
For agentic AI, this means capabilities need to be accompanied by controls such as role-based access, audit trails, version management, validation and human review.
An agent should only access the data and tools required for its defined task, and critical decisions should remain subject to appropriate oversight.
From individual tasks to data workflows
The most significant change introduced by agentic AI may not be faster code generation or faster data transformation.
It is the ability to connect these activities.
A clinical data workflow can involve dozens of interdependent steps across data management, biostatistics, statistical programming, medical writing and regulatory operations. Changes made at one stage can affect multiple downstream outputs.
Agentic systems can potentially coordinate these dependencies, execute defined tasks and identify when human intervention is required.
That is a different model from asking an AI tool to generate code or summarize a dataset.
It is AI operating as part of the clinical data workflow itself.

Anonymization as part of the patient data workflow
Processing patient-level clinical data also requires organizations to consider how sensitive information is protected throughout the workflow.
Anonymization can become particularly relevant when clinical data or documents need to be shared, reused or processed across different systems and teams. Instead of treating anonymization as a separate step at the end of the process, agentic AI can help integrate it directly into the broader clinical data workflow.
AI agents can support the identification of potentially sensitive information, apply predefined anonymization rules and flag cases that require human review. This can reduce the manual effort involved while maintaining a controlled and traceable process.
Within Narrativa Navigator, anonymization capabilities can be incorporated into these workflows. Redaction Scout supports the identification and redaction of sensitive information, helping teams protect patient data while maintaining traceability and human oversight.
This means that data processing, programming and anonymization do not necessarily need to operate as disconnected activities. They can form part of the same controlled workflow, with appropriate permissions, audit trails and review points.
For regulated environments, the objective is not simply to remove identifiers. Organizations need to understand what information was anonymized, according to which rules, when the action occurred and whether human review was required.
Agentic AI can help make that process more systematic while keeping the final control with clinical and regulatory teams.
What should pharma teams consider before using agentic AI for patient data?
Organizations evaluating agentic AI for clinical data processing should start with the workflow rather than the technology.
They should define what the agent is allowed to do, which data it can access, where human approval is required, how outputs are validated and how every relevant action is recorded.
For regulated clinical data, automation and control cannot be separated.
The real opportunity for agentic AI is therefore not simply to remove manual work. It is to create clinical data workflows that are more connected, reproducible and traceable while keeping experts in control.
Frequently asked questions
Can agentic AI process patient data in clinical trials?
AI agents can support activities involving clinical trial data when deployed within an appropriately controlled environment. Access controls, data protection, validation, traceability and human oversight should reflect the intended use and associated risk.
Can AI agents generate SDTM and ADaM datasets?
AI agents can support mapping, transformation, programming and validation activities involved in producing standardized datasets. However, outputs still need to comply with applicable specifications, standards and organizational validation procedures.
Can agentic AI replace statistical programmers?
Agentic AI is better understood as a way to automate and coordinate parts of the programming workflow. Statistical expertise remains necessary for defining specifications, reviewing outputs, handling exceptions and ensuring that analyses are scientifically and statistically appropriate.
Why is traceability important when using AI for clinical data?
Clinical data must remain traceable through transformations and analyses. When AI participates in these processes, organizations need sufficient records to understand how inputs were transformed into outputs and which actions were performed by systems or humans.
About Narrativa
Narrativa® Agentic AI solutions unlock a faster, smarter future for life sciences organizations, helping them to efficiently produce complex, high-volume documentation for regulatory and commercialization workflows. By automating content creation, Narrativa® delivers greater speed, accuracy, and consistency—while ensuring full compliance in highly regulated environments.
The Narrativa® Navigator platform provides secure and specialized Agentic AI-powered automation features. It includes complementary user-friendly tools such as Clinical Atlas for CSR and Protocol generation, Narrative Pathway, TLF Voyager, and Redaction Scout, which operate cohesively to transform clinical data into submission-ready documents for regulatory and commercialization. From database to delivery, pharmaceutical sponsors, biotech firms, and contract research organizations (CROs) rely on Narrativa® to streamline workflows, decrease costs, and reduce time-to-market across the clinical lifecycle and, more broadly, throughout their entire businesses.
Explore www.narrativa.com and follow on LinkedIn, Facebook, Instagram, and X.





