
As AI-powered adverse event processing becomes established within pharmacovigilance, organisations are seeing the same pattern repeatedly. Even when models have been shown to achieve increasingly strong levels of precision and recall, reviewers are continuing to reopen and review cases, despite the required quality checks having been met.
Even when the evidence shows that an AI system is performing reliably, too often reviewers are continuing to check its work as though they don’t quite trust it. This is holding them back from extracting the technology’s promised benefits. Until reviewers feel able to stop routinely rechecking cases that validated AI systems have already cleared, highly trained PV specialists will remain occupied with work that adds little value.
When accuracy isn’t The Limiting Factor
As AI becomes part of routine pharmacovigilance, attention is shifting from proving the technology to making the best use of it. The first priority for AI in pharmacovigilance was to demonstrate that it could perform routine case-processing tasks accurately enough for operational use. For many routine activities, that evidence now exists. The next step is not better algorithms, but greater confidence in using validated AI systems as they were intended to be used.
Some of the difficulty comes down to expectations. AI systems are often expected to produce exactly the same output every time they receive the same input. In practice, AI doesn’t work in quite the same way as conventional rules-based software. Against an expectation of perfect consistency, even a model performing at 95 per cent accuracy can appear to fall short. While the temptation might be to add more rules and constraints in pursuit of greater consistency, every additional constraint narrows the range of cases the technology can handle reliably, reducing rather than enhancing its operational value.
The real challenge for pharmacovigilance, then, isn’t validating every individual AI decision, but developing enough confidence in a system’s overall performance for teams to rely on it appropriately in routine case processing.
This distinction comes to life in AI-generated case narratives. Many of the edits that reviewers make to AI-generated case narratives reflect local writing conventions rather than factual or clinical errors. Where every edit has been recorded as a correction, it could seem that the AI has underperformed, when in reality it has produced a clinically adequate narrative that simply differs from an organisation’s preferred style.
Making Trust Measurable
Replicating blanket quality control makes little sense if organisations are to realise AI’s promised gains in PV. The opportunity the technology presents is not merely to process more cases more efficiently, but also to release experienced specialists from repeatedly verifying work that evidence already shows to be reliable.
That shifts attention to trust. This can’t be gauged by simply asking reviewers whether they feel confident in an AI system. It has to be inferred from behaviour — whether scientists are continuing to recheck outputs that no longer require routine verification, or whether they are willing to rely on the system’s assessments and conclusions within agreed governance boundaries.
One way of addressing this would be through a formal trust coefficient. Rather than measuring AI performance alone, a coefficient would bring together multiple operational dimensions, including AI performance, human reviewer behaviour, validation evidence and governance, task risk and organisational readiness to rely on AI outputs. The purpose is not to prescribe a fixed calculation, but to establish a structured way of thinking about appropriate reliance on validated AI systems.
The aim would not be to measure confidence as a subjective feeling, but to determine how much independent human verification a validated AI system requires within a defined context of use. As operational evidence accumulates that the technology is doing a reliable job, the coefficient should support progressively lighter oversight while maintaining appropriate governance.
Although, elsewhere in PV operations, there are service-level measures that govern uptime or turnaround time, there is currently no equivalent industry metric for managing the transition from universal AI-powered workflow checking to evidence-based technology oversight. Unlike existing operational KPIs, which measure performance, efficiency or compliance, a trust coefficient would represent a new class of operational governance metric. Its purpose would be to measure the appropriate level of reliance that can be placed on a validated AI system under defined conditions of use. While trust has become an important design principle for AI platforms, it has yet to be translated into an operational measure that regulators, technology providers and pharmaceutical organisations can apply consistently.
This is likely to be the next phase of AI operationalisation – not simply validating AI systems, but reconciling trust with governance so organisations can reduce routine review where the evidence supports it. Technology alone won’t solve this problem, not least because the goal for AI in a PV context has never been full automation but rather proportionate, risk-based review — flagging cases with genuinely ambiguous or high-consequence content for human attention while allowing routine, well-supported cases to move through with lighter oversight.
The ambition, then, should be highly streamlined, AI-powered AE processing that remains rigorous and explainable, staffed by PV specialists who are able to rely appropriately on the systems they oversee (and who don’t fear their roles’ obsolescence). Ultimately, this is about applying human oversight more intelligently, so that expert capacity can actually be released for higher-value use.
AI Changes The Work, Not The Accountability
Agentic AI signals the next stage in the technology-enabled advancement of case processing. Unlike earlier AI tools that support individual tasks, agentic systems can complete sequences of activities with limited human intervention. This doesn’t remove the need for human judgement in pharmacovigilance; it changes where that judgement is applied.
Currently, reviewers spend much of their time preparing information — compiling data, checking fields and assembling evidence — before they can reach a conclusion. Agentic AI moves more of that preparatory work to the system, allowing specialists to focus on interpreting evidence, making decisions and resolving genuinely uncertain cases.
The differentiation is subtle but important: between keeping humans in the loop (reviewing every step) and keeping them on the loop (intervening where risk, uncertainty or clinical significance requires it). Some production systems are already moving in this direction, using one AI model to review another’s output before it reaches a human reviewer. Accountability, though, remains unchanged. Agentic AI changes how work is allocated, not who is responsible for the outcome. AI can undertake more of the groundwork, but responsibility for patient safety still lies with qualified PV professionals.
Redesign Before Automation
Experience from early AI implementations suggests that successful deployment begins with redesigning the process, rather than simply automating it. Many pharmacovigilance workflows have accumulated additional hand-offs, duplicate checks and manual interventions over years of incremental change. Simply adding AI to those processes risks preserving inefficiencies rather than removing them.
Experience so far suggests that organisations gain most when they treat process redesign as a continuous activity, rather than a one-off implementation project. As AI capabilities evolve, the challenge becomes less about selecting the right technology and more about building operating models that can evolve alongside it. This approach also aligns closely with emerging regulatory thinking.
A More Proportionate Regulatory Approach
Regulatory guidance is increasingly reflecting this proportionate, risk-based approach to AI governance. This year has seen regulators and international expert groups publish new guidance that emphasises lifecycle governance, strong data stewardship, continuous performance monitoring and transparency about system limitations.
In January 2026, the European Medicines Agency (EMA) and the US Food and Drug Administration (FDA) jointly published 10 guiding principles for the use of AI across the medicines lifecycle, spanning evidence generation from early research through to post-marketing safety monitoring. The principles emphasise proportional validation based on intended use, strong data governance, continuous performance monitoring and transparency about system limitations.
The Council for International Organizations of Medical Sciences (CIOMS)’s Working Group XIV report on AI in pharmacovigilance, published in December 2025, is the product of similar thinking. Rather than prescribing rules for specific technologies, it advocates governance frameworks that can adapt as AI capabilities evolve.
Together, these publications reinforce the idea that governance should be designed into AI systems from the outset, rather than added later as a compliance exercise.
As medicines become more complex and volumes of adverse event reports continue to grow, organisations will find routine rechecking increasingly difficult to sustain. VigiBase, the World Health Organization’s global database of individual case safety reports, contained more than 40 million reports by the end of 2024, with around 70% submitted in the past decade.
The limiting factor is no longer whether AI can process this volume, but whether organisations can determine when validated AI systems can appropriately be relied upon.
Trust cannot be assumed from benchmark scores alone. The next frontier in AI governance is not improving the accuracy of validated systems, but developing evidence-based ways to determine when those systems can appropriately require less routine human verification.
Exactly how such a coefficient should be defined remains open to debate. Whatever form it eventually takes, it should give organisations a practical way of deciding when validated AI systems can be relied upon with less routine human verification, while maintaining appropriate oversight.
Foot note: This article builds on themes discussed in a recent life sciences industry podcast, which can be accessed in full at https://www.arisglobal.com/podcast/
References:
1. European Medicines Agency and U.S. Food and Drug Administration, ‘EMA and FDA set common principles for AI in medicine development’, 14 January 2026. Available at: https://www.ema.europa.eu/en/news/ema-fda-set-common-principles-ai-medicine-development-0
2. Council for International Organizations of Medical Sciences (CIOMS), ‘Artificial Intelligence in Pharmacovigilance’, CIOMS Working Group XIV report, Geneva, December 2025. Available at: https://cioms.ch/working_groups/working-group-xiv-artificial-intelligence-in-pharmacovigilance/Brand JS, Gauffin O, Sartori D, Fusaroli M, Sköld H, Bergvall T, Sandberg L, Wallberg M, Hjelmström P, Norén GN. ‘VigiBase: Resource Profile Update with a Summary of Global Patterns and Trends in Adverse Event Reports for Medicines and Vaccines’. Drug Safety. 2026;49(6):613–629. DOI: 10.1007/s40264-025-01642-6
Read the full article — it's free
Register with Pharma Focus Asia to unlock expert insights, research articles and in-depth industry analysis.