Cathy Yuan: Thank you.
Ashish Kamat: Cathy, you've been now a integral member of the International Bladder Cancer Group for over a year. You are well versed in working with clinical trials and with the ERTC and EUA, and of course, in everything that you do.
And one of the things that I asked you to tackle, which I want to thank you in advance for tackling, is this very important question. Because back in 2016, it was proposed to the FDA by several groups, including the International Bladder Cancer Group and the AUA and GU ASCO to consider single-arm trials for BCG-unresponsive disease. And then what has happened is there's been an explosion of clinical trials. And a lot of these drugs have been approved. People look at CR rates, they look at hazard ratios, and many people try to understand it, but it's confusing. It's confusing for those that may not be doing this statistical analysis and clinical trials for a living.
So again, I threw this challenge at you, you accepted this, and I'd love for you to share with all of us your thoughts on single-arm studies and hazard ratios. So stage is yours.
Cathy Yuan: Okay. Thank you, Professor Kamat, and thank you, UroToday. So today's topic is very unique. I give the title as Single-Arm Trials and Hazard Ratio in Non-Muscle-Invasive Bladder Cancer: Why Singles, and What Are the Hazards? So there's two different topic here. One is single-arm trials. I think it's very unique for oncologists, and probably we don't see much in internal medicines, and also how to interpret the hazard ratio. This is also very commonly seen in cancer-related trials. So of course, it's quite challenge, but I want to make it easier to let the urology understand why and what happened.
So before we talking about clinical trial, we, of course, need to understand that for different condition, we got different treatment goals. So for example, for bladder cancer, we have to separate for non-muscle-invasive and muscle-invasive and metastatic bladder cancer, so how we should treat them. So this is the first goal we have to consider what we are talking about.
So today we are focusing on the NMIBC. So the topic I'm talking, which is this focus on what the outcome we want to achieve for NMIBC, according to the publication from Professor Kamat. So we want to look at the how we can prevent the recurrence, how we can prevent the progression, and how we can preserve the function. And of course, at the end, when the patient are not always staying in the same status, so how we should consider them when they go to different stage, and what happened if they have recurrence, for example, after they got improval, and what happened if they progress. So what's the next step?
So when we talk about evidence, of course, clinical trial evidence. So this is a very most commonly seen, the evidence-based pyramid in our lots of evidence-based handbook. So we saw at least this triangle, and we can see the lower level should be the primary study, for example, study from case report, and then we go up to case control, cohort study, RCT.
And the higher level, we have seen a lot of publication relate to systematic review and meta-analysis. But you can see from here that we don't see any trial called single-arm trials. So you only see, okay, single-arm study, people usually refer to case series, but control cohort study. They usually refer to comparative cohort study. So you can see this is something missing a pieces here, and this is what we are today talking about where we should put these single-arm trials and how we should consider their role would be, and what are the limitation, and why, and why we should have them.
The right up corner, we can see there is a small enlarged pieces of this triangle. You can see there's a no peril cutoff here. It means that sometimes you think that, oh, RCT is better than core study, but no, actually it's a well-performed low risk of biased cohort study, probably they provide a higher level certainty of evidence compared to RCT.
So remember, this is not a truly cutoff level. We shouldn't say, "Okay, one arm-single RCT is always worse than observational study." This is not the case.
So next go to the single-arm design. So as I mentioned that this is probably more commonly seen in cancer study, but not that common seen in other area. For example, in our internal medicine, we don't see.
So why single-arm design? First is we need to consider about the ethical issues. For patients who have been unresponsive, BCG-unresponsive in NMIBC is unethical to re-treat them with the BCG because they already failed it. On the other hand, probably it's not feasible to have them to undergo the surgical comparison because surgical treatment, probably the patient already refused it or not eligible to be considered as a surgical treatment. So in this case, we come up with a single-arm clinical trial design.
So as Professor Kamat mentioned that FDA already approve it, and some of the new drugs we have been witnessed, they got approval and they got the data based on a single-arm design.
So what I mean is that single-arm study, they have lots of limitation, but we should think about why we want it. It doesn't mean that we want to have a shortcut around the rigor, but there is some time we should consider it because of the ethic and feasibility.
So how do we think about a single-arms trial should be a good or not good? Because this is not a case theory study, it's not simple observational study and let the doctor decide who should be treated, who should not be treated. So we should consider the methodology of the single-arm trial.
At least a few key points here to be considered. We call it a trials, means that everything should be designed prospectively and should be rigorously. So first, you should have a very rigorous definition about the eligible population. So which patient should be considered eligible? For example, the BCG-unresponsive criteria, we should follow the FDA criteria, and we cannot just pick the patient and say, "Okay, this is easy to check," and then we put it in the single-arm trial.
So this is very similar to all the other clinical trial. We should have a rigorous inclusion criteria, and we should have a very rigorous criteria how we should assess the endpoint. For instance, we should consider the multimodal methods, and we should have a central probably blind assessment. For the outcome, we should not only consider one time point as clinical responses, but we also consider the duration of the response. So that's why we also have the DOR in all the clinical trial. We not only consider the clinical responses.
On the other hand, we should also consider a time-to-event endpoint. So from the time to survival, instead of just a cutoff point, say that, "Okay, how many patients call responses?" Probably the time to responses and time to survival also should be considered. The different endpoints, not only one.
And the other thing we should consider is how we should analyze the patient. What's the sample size of the patient we should analysis? We call it the denominators of the analysis. Should we consider everybody who... including a child? Or some people might only want to report the patient who been evaluate, which is not rigorous if in most of the patient who failed, and you probably have very loose criteria to how to treat them as a treatment failure or how to treat them as a protocol violation. And this is the case, you might have a small number of patient you got assessed, which it didn't reflect the truly population who received the treatment.
Again, I mentioned about the time point. For example, your clinical response time point, should you look at any time point you call it as the outcome, successful outcome failure? Or whether you should have a very predefined landmark, for example, three months, six months, 12 months. So this is all we should need to consider. I didn't mean that which one should we taken, but I mean that it should be considered when you design trial.
And the other issue is how should we define what the patient will failure or what the patient will success? Did you just say anybody who failed at certain time has failure or you allow the clinical trial is to re-treat them when it's born, and how many time you allow them to re-treat and then you call them failure? This is our other issues we should consider when we design a trial.
So as a urologist, how would you interpret a single-arm trial? Because I mentioned that there's no any comparative arm, so therefore the control always we talk about compared to the historical control. So when we do this interpretation, we need to consider, okay, are you consider the patient comparable? Are you taking the so-called historical-control arm very high risk? I would say probably a different type of population have a very different selection criteria, a very different assessment than what I mentioned before.
So if this is the case, then you compare with the control and probably more likely you compare orange and apple, which is not comparable. And secondly, when you review a single-arm trial, need to consider, do they have any selection bias or enrollment bias? For example, you even have a very rigorous inclusion criteria. What happened if many of the patient, they refuse to enter this trial?
So this patient probably could be very different than patient who entered the trial. So that's why we always recommend the trial should have report the patient's deposition or we call it the full chart, how many patient got screened, and how many patient did include it in the trial, and what are the reason that they didn't include it? Is because the doctor's decision, because they're too sick or any patient's willingness? So this should be considered as well.
So for the assessment, I mentioned already we should have essential blinding or try our best to reduce the risk of bias. So when you read the paper, probably you need to check it whether this trial's data will collect rigorously or the physician who assess it, whether they were blinded or not blinded, how about the outcome assessed, whether they try to blind the possible participant or the personnel.
I do have one key point I want to mention here, it said that we shouldn't compare directly the number of response or number of the outcome between trials or between drugs because these are not comparable because single-arm trial, they have no control arm, they only compare with the historic control. But on the other hand, if you collect lots of single-arm trial and then you compare them, say, "Okay, drug A is better than drug B because the complete response rate is higher in drug A compared to drug B," this is not correct because they're not comparable. But we do have method that can help you to compare them, which is I will mention later. So just not a direct compare of the headline number, say, 80%, 70%.
And again, when you consider the complete response rate, we should consider also what's the duration of the maintain effect? Is it a long durable drug-free effect? Or whether when you report about the duration, are the patients were also on drugs?
The last point is that we should consider you shouldn't use patient data from a subgroup to generalize it to the general population. Not general, but I mean the patient population because we know that different subgroup have different response rate, different outcome responses data. So when we interpret a single-arm study, we try not to use the small group of data, a subgroup data, and we expand it to the general population.
Again, I just mentioned about there's a lower level part here. I just mentioned about how we can compare. If the data come from a single-arm study, we should not try to direct compare them, but we do have statistical method. You can use it to compare, which is this is not the major topic for today, but we do keep in mind that you should not compare directly between drugs, but we do have method to compare.
And secondly, we should also consider when the trial says no comparator, what's the percentage you should achieve your outcome you called it good. How good is good enough to guide you design the clinical trial, the single-arm clinical trial? So there are some others that this go model, they do simulation. Of course, this might not be accurate, but they do give some guidance, say, that maybe what we expect in the clinical practice might be different than what the model that hope we can achieve. So there still also should be considered when you design a trial, how high the number you achieve, and you call it is a good enough the complete response rate.
So that's the first section I talk about the single-arm clinical trials. Again, I want to emphasize even the single arm, we can do it as a good clinical trial. Even the single arm, we should still design it according to all the preset criteria and do it rigorous.
The second topic I want to mention is about the hazard ratio, because again, in the oncologist in trial related to cancer, we do often see a hazard ratio report. And this looks like a very good number, especially when you report in headlines.
We know that in two recently trials in urology, actually is in bladder cancer, they report a hazard ratio both coincident report is 0.68, which is across lots of discussions. How do we interpret this? What are these hazard ratio of 0.68? Is it good or not good?
So I do have a mention here, there's a paper called Hazards of the Hazard Ratio. So it means that hazard ratio is not... we shouldn't interpret directly is the number or the face value, we should interpret... several consideration when we interpret the hazard ratio. This is what we should keep in mind. This is my major topic.
The first I want to mention is that the 0.68 hazard ratio, they have both the same number, but they actually is not report the same results. The both clinical trial, they have two different outcome to measure. One is called event-free survival. The other one is called disease-free survival. By definition, of course, they're different.
And second, they have different follow-up. So we cannot just say that, oh, because their hazard ratio is 0.68, so the two drugs are the same effect, which is not.
Second, hazard ratio is a relative effect. Sometimes very hard for the clinician to understand what you mean that the hazard of the patient who survived is 0.68% compared to control. So in this case, sometime we need to convert it to the absolute effect. So later I will show you the cell formulation we can convert it from the hazard ratio become related risk, and then we can calculate a number to treat.
But look at for these two trial, even they both have the 0.68 hazard ratio, but the difference, the clinical response rate between two arm actually is very different. One is 7%, one's the other, 4%. So the number need to treat when we calculate it, one is 14, the other one's 20.
A recent meta-analysis report, a number need to treat for multiple study when they combine together, it's only 25. It means that not only you should say you have to treat at least 25 patients to see one patient get clinical response. So how do we interpret this number? You got a 0.68, but you got a number need to treat 25. So I will explain this to you later for this interpretation.
The other point is that we always remember hazard ratio actually is a way based on a proportional hazard assumption. It means that the related risk of event actually is occurring between two groups that remain constant over time. So this was the assumption. So if the assumption doesn't hold, then you report the one hazard ratio that is an average number across the different time point could be very misleading. So in this case as a urologist, probably you need to think about, "Oh, is it this assumption whole? If not, then how should we interpret the data?"
Again, I summarize why we need hazard ratio because it's a time-to-event outcome. We handle the data censoring. And also we can report a confidence interval based on the adjustable summary Cox model, and we can report the data which is we can summarize in a meta-analysis. We can pull a data in from a different trial.
But hazard ratio, they do have limitation, I mentioned. First, it's a related measure, not easy to interpret it by the physician. Hazard ratio is not exactly the same as a related risk. We cannot just directly interpret as a risk or risk ratio. We need the calculation to do the conversion.
Also, it's based on proportional hazard assumption. And sometimes you might see the KM curve show that they violate the assumption. That's why we recommend the offer, not only report the hazard ratio, you should also report the KM curve, because the KM curve can tell us whether the assumption might be violated, whether the curve will separate very lay or whether they cross. And this is also another topic of statistic.
Also, we should consider whether the trial follow-up time is very different, and we call it a depression of susceptible. So I want to give a very quick simple explanation. So for example, we have two group. The endpoint is survival. Actually, the survival is based on your mortality. You convert to mortality.
So interesting, you will find out that the longer you follow up, the each arm very close to one. So the people say, "Okay, everybody will die anyway. So if you follow the patients long enough the time, everybody will die." So you have that hazard ratios one compared to go.
So how we can translate it to the hazard ratio follow-up time? Because the longer the follow-up, the patient who stay in the trial is different than the patient who already die, who already drop out, because the most serious patient or most patient who's sensitive or who are responsive to the treatment, they already reached the endpoint, they already left the trial. So the longer you follow up the patient, the patient who probably become more homogenous between the two arms. So it means that the drugs have impact on these two group probably will be different.
So this is another type of bias. It means that the hazard ratio itself is subject to some kind of... We call it depression of susceptibility bias. So in this case, it doesn't mean that the longer the follow-up, the better you can tell the truth of the impact of the drug, which is not the case. So you will see in certain time we should prefer to report a certain endpoint. For example, your endpoint with one year, five years, seven years, how the hazard ratio changed. In this case, report a hazard ratio by different time points more important. You report a one mean average hazard ratio, which could be misleading.
Another bias in hazard ratio, we call it a non-collapsibility bias. It means that sometimes when you adjust it for so-called confounding, even your adjustable dose variable, even though those variable are not confounding, but you will see adjusted hazard ratio could be very different than your unadjusted hazard ratio. So we should avoid not over-interpret those data.
Again, we should report the hazard ratio with the KM curve and then check the assumptions, and sometimes will pair with the absolute outcome data reporting. For example, the number need to treat or the restrict means survival time.
Again, there's another summary. So what question you should ask when you see a hazard ratio? What's the control arm? Because this help us to calculate the absolute effect. And does the KM curve stay proportional? This is help us to check the hazard ratio assumption. So what are the outcomes? Because lots of time we can see composite outcome. For example, event-rate-free survival or progression-free survival. So this can lead to different endpoint and different interpretation, and you cannot just combine them or put them together.
So how wide a confidence interval would be? I show you the 0.68 hazard ratio from the two trial, but actually, you will find out that both trial have a very high wide confidence interval. Both of the upper-level confidence interval very close to one. So it means that the data might not be as precise as you would believe. It might be some consideration when you rate the level of certainty of the evidence.
So another viewpoint is that do we have any confounding? So this is a very unique term. They called it backbone confounding in the BSG and non-responses NMIBC trials. And I'm not the clinical expert, but I know that there's some controversy happening. So people discuss about whether because of the backbone treatment that patient received might have impact on the effect of the treatment between two arms. So this is also the information we should consider.
Again, we should report both the harm and benefit. I earlier mentioned about a number need to treat is 25, but actually in the same meta-analysis also show you that number need to harm is 5% for severe adverse event. It means that you treat 25 patients to get one clinical response, but every five patient, actually they got severe adverse event. So this is a four to five time difference. So is it worth to do this trial of say, "Okay, I treat 25 patients, and five of them got serious event, but only one of them get improvement"? This is also a consideration we should consider the balance between benefit and harm.
And also again, we should identify the bias when we decide a trial. There's a very good information online. And when I attend last year, the year you, Dr. Kamat, host a debate between the two expert about the 0.68, how we should interpret it. So one expert say, "Oh, this is a new standard," and the other one probably consider this is the overreach.
So again, I emphasize one HR should not be only hazard ratio, which could be misleading. And again, statistical significance here doesn't mean that can be translated to clinical significance because we should consider both benefits and harm. And the true interpretation will be depend on lots of information, the toxic and benefit, and relate the risk of bias in the design, and also the other, the certainty of evidence.
So when we talk about certainty of evidence, this is the last slide I want to mention. So before that, we already mentioned about the clinical certainty. For example, we should check the study design, the risk of bias, and whether the data were consistent, and whether we have imprecision. So this is only the first part about our evidence.
So how we can translate to the recommendation. I do mention that we do have a method to convert the hazard ratio become a related risk. Related risk is absolute risk difference. So we use that to easier for the clinician to interpret, especially when we make the recommendation.
For example, hazard ratio 87, how do we interpret it? So we can interpret, say, okay, for the control arm, you got 55 patient got responses, for the tumor arm, you got 59 patients who got response, so the doctor can easier to interpret the data.
The last slide would be what happened after we have all those hazard ratio and related risks? And when we treat the patient, this is not only the numbers we consider. Again, it's not only the statistical significance we consider, there's many other factor we should consider when we treat the patient.
For example, the benefit harm, we already consider the certainty of the evidence we already show in the last slide, but we do have to consider the resource use and, of course, the equality and feasibility. And even the newly recently introduced in the GRADE evidence profile also have a primary health. And we also discuss about whether in the intervention provided to the patient have impact on the global health or whether they're green.
So again, if we have a recommendation to the patient, it's very hard to just say 0.68 hazard ratio, should we recommend, not recommend the drug? Should you have strong or conditional recommend? Or at this time point, probably you will say, "Okay, I need a recommend." I will say probably it's a conditional recommendation for either one, either the intervention or either the control. So this, of course, I will leave it to the consensus group by Dr. Kamat, but thank you so much for giving me the opportunity to share some of my thoughts.
Ashish Kamat: So Dr. Yuan, as I suspected, you did a phenomenal job. I mean, there were so much that you covered in a short time.
I do want to highlight, I know we are running out of time, but I still want to highlight a couple things. Number one, it's very important what you said, that the hazard ratio by itself is a meaningless number to the patient because it depends. If your event rate is 80%, 90% compared to an event rate of 30, 40%, the hazard ratio really has a much different connotation for the patient as far as the benefit. So that's a very important point that you made.
And then the whole list that you provided, and we'll obviously have your slides, and we'll even maybe even put a PDF to your slides if you're okay with it. I think anybody, whether they're a junior investigator, whether they're a student or even a seasoned clinician, needs to recognize that single-arm studies are fraught with problems. And if you're going to look at single-arm studies, you have to make sure that all those points that you mentioned have been evaluated because otherwise there's a lot of bias that gets built in. So I, again, can't thank you enough for taking the time, and looking forward to seeing you again soon.
Cathy Yuan: Thank you.