It Is Time To Require Human-Relevant Efficacy Evidence Before First-In-Human Trials
By Michael Phelan, Founder, InnovApproach Consulting

Ask a severely ill patient in a clinical trial why they’ve signed up. Most will answer, "I’m hoping to find a cure." They might be devastated to learn that not only do most clinical drugs fail due to lack of efficacy, but regulators don’t even require a meaningful demonstration of efficacy before a clinical trial is approved. This is despite the fact that advanced human-relevant disease models can now generate meaningful evidence of a new drug’s ability to treat some of humanity’s most complex and challenging diseases.
This paradigm isn’t an accident. The code that defines U.S. investigational new drug (IND) applications (21 CFR 312)1 is quite clear:
“FDA's primary objectives in reviewing an IND are, in all phases of the investigation, to assure the safety and rights of subjects, and, in Phase 2 and 3, to help assure that the quality of the scientific evaluation of drugs is adequate to permit an evaluation of the drug's effectiveness and safety. Therefore, […] FDA's review of Phase 1 submissions will focus on assessing the safety of Phase 1 investigations…”
The global benchmark for nonclinical study requirements, ICH M3(R2), echoes this thinking:2
“The goals of the nonclinical safety evaluation generally include a characterisation of toxic effects […] … Human clinical trials are conducted to investigate the efficacy and safety of a pharmaceutical […].”
The consequence of this paradigm is staggering. A 2026 systematic synthesis spanning six decades of clinical drug development found that the failure rate is approximately the same as it was in the 1960s. The failure rate sits at approximately 91% for the 2000s and 2010s. Around 77% of those failures trace to biological factors, chiefly efficacy and safety.3 Lack of efficacy alone accounts for a larger share of late-stage failures than any other single cause.4 These failures concentrate in Phase 2 and Phase 3, the most expensive and time-consuming stages of development with the highest ethical burden on patients. We are systematically sending drugs into human trials without adequate evidence they will work and then acting surprised when they do not.
Why, when human-relevant evidence of likely efficacy can be generated before a trial begins, is it not evaluated before patients are exposed?
Efficacy Failure Is Not A Victimless Problem
The regulatory focus on safety makes sense. Regulators exist, in large part, because of historic drug-related catastrophes. The motivation to prevent direct harm to patients is absolutely justified. But safety risks are not the only source of harm clinical trials expose patients to. Efficacy failure can cause real harms in four specific ways.
First, efficacy failures do cause direct harm to clinical trial patients. Many clinical trial participants are seriously ill. Many enroll with the expectation that the experimental medicine might help them. Research consistently shows that trial participants hold high therapeutic expectations despite informed consent processes designed to manage them.5 When a drug fails for lack of efficacy, those expectations are shattered. That is a harm. And if a validated human-relevant model could have predicted that failure before a single patient was dosed, do we not have an ethical imperative to check?
Second, lack of efficacy evaluation exposes trial participants to an inappropriate risk-benefit ratio. Even with robust safety testing, clinical trial participation is never risk free. If a validated efficacy model indicated a near-zero probability of benefit, the risk-benefit ratio for that trial would be unacceptably high. No review committee presented with that data would authorize enrollment. But without the efficacy data, the question is never appropriately asked.
Third, waiting to review efficacy until the end squanders critical resources. Drug development is constrained by capital, time, and the availability of trial participants. Efficacy failures typically appear after Phase 2 or Phase 3, the stages where the largest investments have already been committed. Every donation, grant, and research dollar and every patient-year spent on a drug that was never likely to work is a dollar and a patient-year not spent on one that might. This directly delays access to effective therapies for patients who need them.
Fourth, the frequency of efficacy failure has created an inefficient economic norm that favors volume over human health. When the barrier to entering clinical trials is safety alone, high failure rates naturally follow. Companies face an incentive to push as many candidates as possible into the clinic on the chance that one succeeds. Economics is suddenly in opposition to patient welfare. This model also drives sponsors toward jurisdictions with lower trial costs and faster start-up times, not because those jurisdictions produce better science, but because the economics of a high-failure-rate model favor speed and volume over rigor. This is particularly true when many start-ups plan on leveraging successful IND and Phase 1 outcomes to secure badly needed capital or to out-license their asset. A requirement for human-relevant efficacy data could shift the economic calculus. It would reward programs with stronger preclinical efficacy evidence, regardless of where the trial is initiated or who finishes it.
Animal Disease Models Are At The Core Of This Problem
Naturally, drug developers do check efficacy before clinical trials and often include that evidence in their Investigator’s Brochure. That is not the issue. The issue is how they test it, how it is submitted to regulators, and the level of scrutiny it receives. Discovery and preclinical efficacy assessment still relies heavily on animal disease models. The FDA itself has acknowledged the limitations of those models. Its April 2025 Roadmap to Reducing Animal Testing states plainly that animal-based data have been particularly poor predictors of drug success for several common diseases, naming cancer and Alzheimer’s among them.6 This also may explain why efficacy data are not treated as a stand-alone requirement. Historic efficacy models have not been predictive enough to be relevant to regulatory review.
In some therapeutic areas, the failure of animal models to predict human efficacy is catastrophic. Alzheimer’s disease has had a clinical trial failure rate above 99%.7 Neurology drugs, broadly, have among the lowest success rates of any therapeutic category.8 Oncology drugs have an overall likelihood of approval from Phase 1 of just over 5%,8 and lack of efficacy is the leading cause of late-stage oncology failure.9
Neurodegenerative disease and cancer represent some of the largest unmet medical needs in the world and the largest shares of the clinical trial pipeline. If animal disease models were solving the efficacy prediction problem, the numbers would look very different.
The specifics are damning. Standard transgenic Alzheimer’s mouse models are built on familial early-onset mutations that account for only a few percent of human cases. The sporadic late-onset disease that accounts for up to 99% of Alzheimer’s patients is not what these animals reproduce.10 Similar biological gaps exist in many of the models used to investigate our most complex diseases.
Human-Relevant Disease Models Now Exist
There has always been a practical argument against requiring efficacy evidence: there were no reliable, validated human-relevant tools to generate it. That is changing rapidly.
New approach methodologies (NAMs), including microphysiological systems, patient-derived organoids, iPSC-based disease models, and computational approaches, are producing human-relevant efficacy data that animal models cannot. The same 2026 synthesis describes human-on-a-chip and organoid systems that recapitulate disease biology more faithfully than animal models, including a case where such a model supplied efficacy data supporting an IND.3 Many of these technologies are well characterized, commercially available, and have claimed or published success rates well above those of animal models.
Regulatory infrastructure is catching up. The FDA’s ISTAND program provides a formal qualification pathway for NAMs as drug development tools, and the March 2026 FDA draft guidance on NAMs establishes a framework for integrating their data into regulatory submissions with a focus on validation and weight of evidence.11 The EU, U.K., and U.S. have all published road maps toward reducing reliance on animal testing. But there is a gap between what is happening and what is required. Human-relevant efficacy data is already being used to support individual IND applications, accepted case by case at the discretion of a given reviewer. However, ISTAND submissions to date overwhelmingly target safety endpoints. None have yet sought qualification for a preclinical efficacy context of use. Consequently, acceptance depends on individual sponsors persuading specific reviewers and proof of efficacy remains the exception, not the rule.
A Call To Action
The tools exist. It is time for regulators and industry to catch up.
Review committees must ask the efficacy question. When reviewing a first-in-human application, the question should not be limited to "is this drug safe enough to test in humans?" It should include "is there a reasonable, human-relevant basis to believe this drug might work?" Where validated or qualified models exist and are available, failure to consider them is a gap in the scientific and ethical assessment. For indications with the highest failure rates (>95%), regulators should consider mandating that validated human-relevant tools be included in the submission.
NAMs developers must pursue efficacy qualification. The ISTAND program and equivalent pathways in other jurisdictions are open. Disease-specific efficacy models, particularly for indications with the highest clinical failure rates, should be among the first candidates for formal regulatory qualification. The safety-first approach in current ISTAND submissions is understandable but must not become permanent.
Sponsors must integrate efficacy NAMs into preclinical strategy. Even without formal qualification, sponsors can and should use human-relevant, efficacy-focused NAMs to de-risk their own pipelines. Human-relevant disease models can identify programs with low probability of clinical benefit before Phase 2 spending begins. Equally, investors should demand proof of efficacy in human-relevant models before funding human clinical trials.
The preclinical landscape is changing faster than the regulatory expectations that govern it. We have spent decades refining the science of predicting whether a drug will harm a patient. It is time to apply the same rigor to predicting whether it will help one.
References
- 21 CFR 312.22(a). Code of Federal Regulations, Title 21, Part 312. U.S. Food and Drug Administration.
- ICH M3(R2). Guidance on Nonclinical Safety Studies for the Conduct of Human Clinical Trials and Marketing Authorization for Pharmaceuticals. International Council for Harmonisation. 2009.
- Filipova D, et al. Need for NAMs: a systematic evidence synthesis revealing over half a century of drug development failure. NAM Journal. 2026. doi:10.1016/j.namjnl.2026.100119.
- Harrison RK. Phase II and Phase III failures: 2013-2015. Nature Reviews Drug Discovery. 2016;15:817-818.
- Pentz RD, et al. Therapeutic misconception, misestimation, and optimism in participants enrolled in phase 1 trials. Cancer. 2012;118(18):4571-4578.
- U.S. Food and Drug Administration. Roadmap to Reducing Animal Testing in Preclinical Safety Studies. April 2025.
- Asher S, Priefer R. Alzheimer’s disease failed clinical trials. Life Sciences. 2022;306:120861.
- Hay M, Thomas DW, Craighead JL, Economides C, Rosenthal J. Clinical development success rates for investigational drugs. Nature Biotechnology. 2014;32(1):40-51.
- Jardim DL, Groves ES, Breitfeld PP, Kurzrock R. Factors associated with failure of oncology drugs in late-stage clinical development. Cancer Treatment Reviews. 2017;52:12-21.
- Mullane K, Williams M. Preclinical models of Alzheimer’s disease: relevance and translational validity. Current Protocols in Pharmacology. 2019;84(1):e57.
- FDA Draft Guidance: General Considerations for the Use of New Approach Methodologies in Drug Development. March 2026.
About The Author
Michael Phelan, Ph.D., is the founder and principal consultant of InnovApproach Consulting. He advises on the qualification and regulatory acceptance of new approach methodologies in drug development across global regulatory frameworks.
Phelan led the first technology accepted into the FDA ISTAND program to qualify NAMs and serves as an external program consultant to the NIH Complement-ARIE program, which includes the Validation and Qualification Network (VQN). He contributed to the Australian Bellberry/ASCEPT NAMs HREC Review Checklist and was invited as an expert speaker to a European Parliament roundtable on the use of NAMs in pharmaceutical development. Phelan holds a Ph.D. in bioengineering, earned as an intramural fellow at the NIH’s National Eye Institute. His graduate research focused on retinal organoid development. He now resides in Melbourne, Australia, and works with international clients.