Alternative statistical method could improve clinical trials(phys.org)
phys.org
Alternative statistical method could improve clinical trials
https://phys.org/news/2021-12-alternative-statistical-method-clinical-trials.html
5 comments
The main point IMO is that 99.9% of active researchers in clinical medicine should never have been allowed to conduct research. I should know, I'm one of them. The state of the field is plainly abysmal and most clinical researchers don't know a mean from a median, yet happily design full-scale clinical trials.
There is no point.
The fragility index is essentially a repackaged p-value, and serves only to introduce even more confusion on the top of this already often misunderstood metric. See very good discussion here: https://academic.oup.com/eurheartj/article/38/5/346/2422087
Also, trials are generally designed to recruit the minimal number of patients we can get away with, and this is done for good reason (cost, feasibility, and the ethical concern of "using" the lowest possible number of human beings, while allowing any beneficial treatment to benefit the most as soon as possible). If someone's then pointing out that this trial is "fragile" -- well, yes, it was designed so in advance! Would you propose to expose 100 more patients to an inferior treatment just to improve your fragility index?
The fragility index is essentially a repackaged p-value, and serves only to introduce even more confusion on the top of this already often misunderstood metric. See very good discussion here: https://academic.oup.com/eurheartj/article/38/5/346/2422087
Also, trials are generally designed to recruit the minimal number of patients we can get away with, and this is done for good reason (cost, feasibility, and the ethical concern of "using" the lowest possible number of human beings, while allowing any beneficial treatment to benefit the most as soon as possible). If someone's then pointing out that this trial is "fragile" -- well, yes, it was designed so in advance! Would you propose to expose 100 more patients to an inferior treatment just to improve your fragility index?
Honestly? Yes. The issue I'd have with said minimums is that, at least from my experience in biomedical research is that they're often done with exceedingly poor statistical grounding and pathetic sample sizes. Admittedly, perhaps things are better in the clinical space, though my gut argues otherwise.
So, to re-answer your question
>Would you propose to expose 100 more patients to an inferior treatment just to improve your fragility index
Absolutely, especially if the alternative is exposing a million more patients to a different, more expensive inferior treatment because the statistical analysis was garbage.
So, to re-answer your question
>Would you propose to expose 100 more patients to an inferior treatment just to improve your fragility index
Absolutely, especially if the alternative is exposing a million more patients to a different, more expensive inferior treatment because the statistical analysis was garbage.
Garbage statistics will not get better if you insist on some arbitrary fragility index threshold (and no accepted threshold of an "acceptable" fragility index does exist to the best of my knowledge). Also, using a lower alpha will automatically give you a larger sample size, without invoking a completely superfluous fragility index.
Nobody should interpret clinical studies in isolation; they only have meaning in a "qualitative" Bayesian framework which integrates physiological plausibility, other available trials, and risk/benefit ratio. The fragility index only muddies waters, as clinicians misinterpret it even more frequently than the much maligned p-value, all while not delivering any more information than the p-value itself.
Nobody should interpret clinical studies in isolation; they only have meaning in a "qualitative" Bayesian framework which integrates physiological plausibility, other available trials, and risk/benefit ratio. The fragility index only muddies waters, as clinicians misinterpret it even more frequently than the much maligned p-value, all while not delivering any more information than the p-value itself.
Yeah, Full agreement on the arbitrary fragility index. One of the main issues with the more squishy sciences is the cargo-culting around specific breakpoints, which happens due to a lack of statistical chops to begin with.
At least at the labs I've been at, the vast majority have no clue what those p-values they get mean, nor do they realize that the 0.05 cutoff that they so often target is entirely arbitrary.
At least at the labs I've been at, the vast majority have no clue what those p-values they get mean, nor do they realize that the 0.05 cutoff that they so often target is entirely arbitrary.
Side note about an advertisement on the linked page. It's a google vpn ad and the behavior is new to me. On a mobile device while swipe scrolling, I scrolled on the ad itaelf. This resulted in the ad actually opening a new tab, even when I tried it multiple more times. I didn't tap on accident.
The ad is actually opening a new tab when you're just scrolling the page while touching the ad area and never actually tap it or intend to open the ad
The ad is actually opening a new tab when you're just scrolling the page while touching the ad area and never actually tap it or intend to open the ad
So they conduct a trial, ideally they pre-registered their trial, and nobody asked them about test power, beta, and type II error? Seriously?
The authors introduce a measure called the "fragility index" (rather, a modification of an existing fragility index) that's supposed to diagnose a lack of robustness in clinical trials satisfying the usual standards of p < 0.05.
First, the problem with clinical trial results failing to replicate or discover true effects goes far beyond the statistical methods used. For example, "virtually all major RCTs funded by an NIH institute (NHLBI) before 2000 were false positives. Once hypothesis preregistration is required in 2000, everything becomes a null." [0] Arguing about statistical methods is just rearranging deck chairs on the Titanic until all that institutional stuff is taken care of.
Continuing anyway: The introduced measure is extremely ad hoc and not supported by any theoretical framework (e.g. proofs of good asymptotic properties). I don't know how to interpret it, especially because (as the authors note) it can easily flag a study as "fragile" when it isn't. Say what you want about p-values, but they at least have mathematical guarantees when used properly.
Further, there's no discussion of how this compares to existing approaches. Why not just do careful Bayesian modeling? Or, why not try to learn something from the amazing success of machine learning researchers in predicting out of sample generalization (e.g. see [1])?
Maybe I'm missing something, but I really don't see what the contribution is here.
[0] https://twitter.com/paulnovosad/status/1427332860902584329 [1] http://jakewestfall.org/publications/Yarkoni_Westfall_choosi...