
Interpreting real-world evidence in ARPI treatment decisions, with Martin Schoen, MD, PhD
Key Takeaways
- ARPIs are being deployed from localized high-risk disease through androgen pathway modulation–resistant states, reflecting durable activity of androgen signaling suppression across clinical contexts.
- Clinicians frequently choose among ARPIs based on indication-specific trial data, insurance coverage, and interaction profiles, with many viewing enzalutamide, apalutamide, and darolutamide as largely mechanistically interchangeable.
Martin Schoen, MD, PhD, explains how real-world evidence can help compare ARPIs and why study design and patient selection are critical to interpretation.
As real-world evidence (RWE) continues to expand the evidence base for
In this Q&A, Martin Schoen, MD, PhD, a medical oncologist at the VA Medical Center in St. Louis, Missouri, and a professor at Saint Louis University, discusses the role of androgen receptor pathway inhibitors (ARPIs) across the prostate cancer treatment spectrum and how RWE can help inform treatment selection. Schoen also examines the importance of study design, patient selection, statistical adjustment, and target trial emulation when evaluating comparative RWE, highlighting factors that can help clinicians distinguish robust findings from hypothesis-generating analyses.
Urology Times: Please provide a brief overview of where ARPIs currently fit into the treatment landscape.
Schoen: ARPIs are one of the best therapies that have been created in the last 20 years for prostate cancer. They are available across both high-risk localized disease as well as metastatic and castration-resistant disease, now called androgen pathway modulation resistant [disease]. With the newest trials at ASCO this year, they're even being used in the neoadjuvant and adjuvant setting prior to a prostatectomy. So now, ARPIs are being used across the entire spectrum from localized to resistant disease to try and augment responses to therapy, as we know that prostate cancer is driven primarily by androgens and modification of androgen signaling has been shown to be an effective method of controlling disease in any disease state.
Urology Times: With 4 ARPIs now widely used in practice, what factors most commonly drive a clinician's initial selection among them in the absence of head-to-head trial data?
Schoen: That's an excellent question because the 4 ARPIs have been tested in different scenarios, so there is data for certain ARPIs in certain situations. In castration-resistant disease, there are actually only 2 ARPIs that are approved by the FDA, which includes enzalutamide [(Xtandi)] and abiraterone [(Zytiga)]. Although, I would say that most medical oncologists and urologists would use any of them because we really do not feel that that there is a large difference between them. I think commonly insurance purposes drive the decision or drug-drug interactions. In the localized disease setting, abiraterone has been the drug that has shown the most efficacy in combination with radiation. And then now, apalutamide has been shown to have efficacy in the neoadjuvant and adjuvant setting after a prostatectomy. So, a lot of times, we are choosing ARPI based off of the clinical trial and the setting in which they are tested. But I do think that most clinicians would feel that they are interchangeable because the mechanisms are very similar, especially among the -amides, which is enzalutamide, apalutamide [(Erleada)], and darolutamide [(Nubeqa)]. So commonly the choice is made off the scenario, but also a patients' drug-drug interactions and their insurance status.
Urology Times: Where does real-world evidence fit within the broader body of evidence clinicians use when selecting an ARPI?
Schoen: I think that is a key question, because there have not been head-to-head trials, so it has limited our ability to distinguish between the ARPIs. Real-world evidence can fill that gap. In a place where we don't have head-to-head [data], we can use data from observational studies and large practice networks to answer that question, especially if they are used in a very similar treatment setting or if the drugs have the same indications. By using patients in areas such as the Veterans Health Administration or other similar situations where there's less heterogeneity [between patients], real-world evidence can give hints to one drug's effectiveness in that population compared to another, because we have so many patients and we are able to adjust or control for some baseline differences.
Urology Times: How can differences in study methodology influence the conclusions of a comparative real-world evidence analysis?
Schoen: That is key to when you do a real-world evidence analysis: the patient selection and the methods used to analyze the patients are 100% important. Real-world evidence is not much different than prospective trials, and any clinician that's been involved in prospective trials knows that patient selection is key. You have to know who are the patients that would benefit from this therapy and then do an appropriate analysis of their outcomes, whether it's overall survival or progression-free survival, that you have to use the best methods. In real-world analysis, it's the same thing; we have to use patients that have a similar clinical scenario and then think about it as though we were doing a randomized trial, with appropriate inclusion and exclusion [criteria]. The methods that we use should be almost like a target trial emulation, where we're taking the example of a prospective trial, the trial that we would want to do to compare one ARPI to another, and we'd apply that to our evidence so that the patients have very similar clinical scenarios at baseline and the outcomes that we analyze are very similar to what a trial would analyze, such as time to progression or overall survival. That, I think, is the key to the study methodology. Studies that don't have that careful selection of patients, similar to the careful selection that's done in randomized clinical trials, sometimes have less influence or should be taken with a grain of salt. It is both a combination of patient selection and outcome measurements that makes real-world evidence more robust.
Urology Times: As you mentioned, statistical adjustment methods are often used to account for confounding. What are the limits of what these adjustments can actually correct for?
Schoen: That's incredibly important. We can only account for things that can be measured. A lot of times, things like frailty or performance status are not always available in the medical record and also can be assessed differently by different people. Things that are very common in a prospective trial like performance status are not available, so we can't adjust for that type of thing. Even if you do adjust for known factors, there can be unknown factors that could be imbalanced. Sometimes patients can develop castration resistance, and there's not a code for castration resistance, or the labs or another analysis was done outside of our data and we're not aware of it. That is a limit of real-world evidence; we can only account for the things that we know, and there are definitely limits on what we know in observational trials because the data is not collected prospectively.
Urology Times: Please walk through what distinguishes study designs such as an indirect treatment comparison or a matching-adjusted indirect comparison.
Schoen: There is a lot of confusion about what some of the real-world methods are and how we do them. An indirect treatment comparison is what we commonly call a cross-trial comparison, where we look at the outcomes of one trial and we compare them to another trial. There are lots of methods to do that, [one being] a synthetic analysis of the Kaplan-Meier curve that's been published, [where we] overlay one Kaplan-Meier curve onto another, assuming that the that the characteristics of each population are similar at baseline. That is now more commonly done.
The difference with a cross-trial comparison and a matching-adjusted indirect comparison is that you take into account some of the baseline characteristics of the patients in the trial. Because you typically will have the characteristics of one trial that is run by a sponsor or another organization such as ECOG or Alliance, you can then use the characteristics of that trial to match to existing trial data and adjust for some of the baseline characteristics. Now, you don't always have all the characteristics of the second trial, but you try your best to match and then do the same indirect comparison or cross-trial comparison.
Urology Times: How do you decide when a piece of real-world evidence is robust enough to influence a treatment discussion vs when it should be treated mainly as hypothesis-generating?
Schoen: That's really important. I would say that the more data that is available and the more that the data is similar to a prospective trial with all of the appropriate controls, the better the analysis can be. One thing that we didn't mention that is sometimes even better is if you have a completely individual patient data comparison, where you have the individual patient data from both trials, and you're able to recreate the trials by a complete match. That's probably one of the best methods. In observational data, those are also typically more reliable, using optimal selection criteria so that the patients are properly selected for inclusion and exclusion. These analyses can be very helpful when you have an individual patient data level analysis.
The other thing that can happen is that we can do a network meta-analysis. Those do not typically have all of the individual patient data, but if a trial is done very similarly and we have multiple trials that are done in similar manners, a network meta-analysis can be very helpful because you're able to increase the number of patients and the power of your data analysis. Those would be the 2 methods that I think are optimal and very robust, when you're able to have individual patient data with reliable inclusion/exclusion data in a similar setting, or you're able to have multiple studies that show a similar finding and are able to do a meta-analysis.
Urology Times: What is the most important takeaway that you would want clinicians to keep in mind when they're interpreting comparative real-world evidence on ARPIs?
Schoen: These data are super helpful in a setting where we do not have head-to-head data to compare agents. Being able to interpret real-world data is important, and I think that this is going to continue to happen in other drug classes. The principles of understanding who are the patients and how are they selected is important, as well as if [the study design] fits into a target trial method, that we are emulating what a clinical trial would be. Were patients selected at a similar state and did they have similar rules? An easy example of this is that patients typically start on androgen deprivation therapy and then start on an ARPI within 3 months in a clinical trial. If some patients were started 14 months later, that's not exactly what a clinical trial would do. So, was the trial run very similarly to a randomized trial? Were appropriate inclusion/exclusion criteria used? Is the analysis similar to what you would want for phase 3 studies or for a regulatory body? Those things can help clinicians as we have more real-world evidence to help make these decisions between different treatments where we don't have head-to-head [data].
Urology Times: Is there anything else that you wanted to add?
Schoen: One last thing that I would add is that for a very long time in prospective randomized trials, we required 2 separate studies that had a similar finding in order to approve or give a category 1 indication. I would say a similar idea should be thought of for real-world evidence. You want to have a finding in that is robust—it occurs in multiple different populations and done in different ways. For example, you could have a study of veterans that receive a therapy, and then a study done in Medicare or a non-veteran population. If you have similar studies that are done similar ways in different populations, it increases the reliability of the data. That's very similar to prospective trials. If you have multiple trials that are showing a similar finding, then it increases the reliability of the data and the outcomes that that we use to make decisions.











