New Imaging Approach Could Improve IPF Trial Design
A conversation between Professor Peter George, Ph.D., consultant pulmonologist, and Clinical Leader Executive Editor Abby Proch

Timing is everything, as they say. And for diagnostic imaging, the precision and speed with which it’s executed can make all the difference — not only to patients but to researchers.
So, what if idiopathic pulmonary fibrosis (IPF) trials could use advancing imaging to spot a treatment effect before traditional lung-function measures do?
Pulmonologist Peter George is attempting just that, and in this Q&A, he explains how AI-powered quantitative CT imaging could strengthen IPF clinical trials by revealing structural changes that forced vital capacity (FVC) may miss. He also discusses the challenges of studying a rare, heterogeneous, and progressive disease; the evidence needed to validate imaging-derived endpoints; and why pairing them with functional and structural measures could improve early development decisions.
For sponsors, the more sensitive, reliable endpoints could provide a clearer read on whether a therapy is working — and a better chance of avoiding costly late-stage failures.
Clinical Leader: For readers who are less familiar with IPF drug development, how are clinical trials in the disease typically designed, and which endpoints most often determine whether a therapy appears promising?
Peter George: Most IPF trials still follow a fairly conventional path: randomized, double-blind, placebo-controlled studies, usually layered on top of standard-of-care antifibrotics, such as nintedanib or pirfenidone. Early-phase studies (Phase 2) typically run for between 12 and 26 weeks, primarily focusing on safety and tolerability while aiming to generate a signal of biological activity and determine whether the drug is affecting the disease before committing to a larger, longer Phase 3 program.
The endpoint that has anchored the field for decades is change in FVC, measured by spirometry. Regulators have accepted FVC decline as a proxy for disease progression and mortality, and it remains the primary endpoint in all pivotal trials. Alongside it, trials typically track secondary measures such as hospitalization and mortality, PROs such as breathlessness or quality-of-life scores, and increasingly, high-resolution CT (HRCT) imaging to assess structural changes in the lung. A therapy is generally considered "promising" at the early stage if it shows a favorable trend in FVC decline, ideally with corroborating HRCT structural signals, without being definitively powered to prove efficacy.
What makes IPF difficult to study compared with other chronic respiratory diseases?
A few factors make IPF particularly challenging to study. First, it is a relatively rare and heterogeneous disease, meaning that patients can experience disease progression at different rates. Consequently, trial cohorts can differ from one study to the next even with similar enrollment criteria. This means that it is difficult to accurately model the behavior of a placebo arm, and trials must be large enough to capture the complexity of the disease.
Second, the disease is progressive and irreversible. By the time a patient is symptomatic and diagnosed, meaningful structural damage has often already occurred, which narrows the window in which a therapy can demonstrate benefit.
Third, while the disease is relentlessly progressive and is associated with reduced survival, the mortality rate within the period of a 52-week clinical trial is (thankfully) too low for it to be used as an endpoint. We therefore use decline in FVC as a surrogate for mortality, but this is a challenge. While it is the endpoint currently accepted by the regulators, it is imperfect, and meaningful FVC decline may take many months to become statistically detectable, especially now that patients are frequently already on background antifibrotic therapies that slow progression. The combination of a heterogeneous irreversible disease and a relatively coarse functional endpoint is what makes IPF trials so challenging to power and interpret, particularly in early phase.
Forced vital capacity is a central endpoint in many IPF trials. What does FVC measure well, and what might it miss?
FVC is a well-validated, fairly reproducible measure of overall lung function, with a long track record that regulators trust, and its change correlates with meaningful outcomes such as mortality over the long term. That is why it has remained the backbone of IPF trials.
But FVC is also a global, effort-dependent measure, with meaningful day-to-day and lab-to-lab variability. It tells us how much lung function has changed, but not where or how the underlying disease is changing. It cannot distinguish fibrosis from other causes of volume loss, localize disease within the lung, or provide much sensitivity to relatively short-term structural change.
Clinically, we sometimes see patients whose FVC appears stable while they are becoming more breathless. Quantitative CT can give us another window into what is happening structurally, measuring changes in fibrosis extent, interstitial lung disease burden, and lung volume that may not yet be reflected in spirometry. That is the gap quantitative imaging is beginning to address.
What are the challenges in executing IPF trials?
Practically, IPF trials face several intersecting challenges. Recruitment is a challenge because the patient population is limited and often elderly, with comorbidities that complicate enrollment. Attrition is a real issue, too. IPF carries significant mortality and morbidity, so patients can be lost to acute exacerbations, transplant, or death during a trial, which erodes statistical power over time.
On the measurement side, spirometry itself has meaningful test-retest variability, and background antifibrotic use, now standard of care for most patients, reduces the amount of decline available to detect a drug effect on top of it. That means trials increasingly need larger sample sizes or longer durations to reach statistical significance using FVC alone, both of which raise cost, timeline, and patient burden. And because the disease has historically had very few approved therapies, there's real pressure to find endpoints that de-risk the go/no-go decision earlier, before committing to an expensive multiyear Phase 3 program.
Why have encouraging signals in some early-phase IPF studies failed to translate into successful later-stage trials?
Part of the challenge is also understanding what an observed change actually represents. If an early trial shows a favorable trend in FVC without corroborating structural or mechanistic evidence, it can be difficult to know whether the signal reflects a genuine effect on the underlying disease or variability in a relatively coarse, effort-dependent measure.
This is where quantitative imaging can add another layer of information. In analysis of taladegib, for example, the CT measures didn’t simply reproduce the FVC finding, they showed changes in different anatomical components of the lung, including lung volume, interstitial lung disease extent, and fibrosis extent. That gives us greater insight into the structural changes that may be driving a functional outcome.
Pairing functional and structural endpoints therefore has the potential to give investigators greater confidence about whether an early signal represents a genuine biological effect worth taking forward.
How does AI-powered imaging differ from conventional radiologist review of CT scans? And how can it add value?
Conventional radiologist review of HRCT remains essential but is largely qualitative. An experienced reader can characterize disease pattern and severity, but subtle interval change can be difficult to identify consistently, particularly when disease is already extensive or is on the other end of the spectrum, i.e., very limited. This creates potential for inter- and intra-reader variability when trying to track progression over time.
AI-powered quantitative analysis applies standardized algorithms to the same CT data to measure structural features such as lung volume, fibrosis extent, and overall ILD burden. Importantly, this doesn't mean replacing the radiologist's judgement. The quantitative outputs can highlight where and how the imaging has changed, giving the radiologist an additional objective layer of information that they can interrogate against the underlying scan.
For clinical trials, the combination of objective measurement and expert interpretation could be particularly valuable when trying to detect relatively subtle changes consistently across timepoints and sites.
What evidence is needed to show that an AI-derived imaging measure is reproducible, clinically meaningful, generalizable, and suitable for use as an endpoint?
There is a well-established bar to reach here. Imaging protocols need to be standardized across a trial so that scanner manufacturer or acquisition parameters do not change within a site midway through a trial. Prior to embedding a QCT metric into a clinical trial, AI-derived imaging measures need to have been extensively validated in real-world cohorts to prove that changes identified in a clinical trial are clinically meaningful and correlate with outcomes patients and physicians care about, such as functional decline, symptoms, or mortality.
Generalizability means validating the algorithm across diverse patient populations, disease severities, and scanner types, not just the data set it was trained on. And to be accepted as a trial endpoint, particularly a primary or co-primary one, a measure typically needs regulatory engagement, ideally through qualification pathways, along with a track record of use across multiple independent studies. That's part of why it's meaningful that e-Lung has already been used across several Phase 2 and Phase 3 programs and was selected as a co-primary endpoint in a Phase 3b pulmonary fibrosis trial. That kind of repeated cross-program use is exactly the evidence base regulators and sponsors want to see before relying on a novel endpoint.
Another important consideration with AI-derived measures is whether the algorithm remains stable when it is applied to new data sets. For a clinically deployed locked algorithm, the model is not continually retrained on the data sets to which it is subsequently applied. That provides an important safeguard against simply fitting to the data being evaluated. Ultimately, however, confidence comes from the broader evidence base: demonstrating reproducibility and clinical relevance across independent data sets, populations, scanners, and clinical programs.
Looking ahead, what would be an ideal IPF trial endpoint strategy?
FVC remains a validated, regulator-accepted measure with a long history, and will likely remain part of the picture. I believe the future is a multimodal strategy that combines functional measures such as FVC with objective quantitative imaging biomarkers, particularly in early-phase and proof-of-concept studies.
The value is twofold. First, imaging may be able to detect structural change earlier than a meaningful change becomes apparent in FVC. In the INBUILD analysis, for example, changes in quantitative CT measures at 24 weeks were able to predict FVC decline at 52 weeks, suggesting the possibility of bringing forward some clinical development decisions.
Second, imaging can tell us more about what is happening biologically. Rather than simply telling us that lung function has changed, quantitative CT can help characterize changes in fibrosis, interstitial lung disease burden, and other structural compartments.
If these measures can consistently provide earlier and more mechanistic evidence of treatment effect, they could help sponsors make better-informed go/no-go decisions before committing to large multiyear Phase 3 programs.
About The Expert:
Professor Peter George, Ph.D., is a consultant respiratory physician and clinical lead for interstitial lung disease at Royal Brompton Hospital. Dr. George qualified at Imperial College School of Medicine in London in 2005 and has trained at the most prestigious centers in London. In 2009, he was awarded an academic clinical fellowship in respiratory medicine. He completed his Ph.D. at Imperial College, London, focusing on the role of the innate immune system in pulmonary vascular inflammation. Dr. George has trained in interstitial lung disease at both Hammersmith and Royal Brompton Hospitals and spent time in the interstitial lung disease unit at the globally renowned National Jewish Health Hospital in Denver. He is heavily involved in basic science and clinical research, supervising research degrees and regularly publishing in high impact journals.