Does the sentence imply that they aren't costly? Would you write "strongly averse" instead of "somewhat averse" or "it is the end of the world" instead of "it isn't the end of the world"? If so, what language would you use to convey that strongly negative changes are that much worse than neutral ones?
I've been called worse things, and at this stage in life it is hard to be offended by people who don't know me well enough to give an insightful insult. But I hadn't responded to the earlier comment because it was more generally antagonistic and seemed to reflect a (possibly intentional?) misinterpreted and negative reading of the blog. "Don't feed the trolls", as the saying goes.
I started writing that particular blog with different experimental methods in mind, but wrote so much as a prerequisite that I wanted to stop and make that first part a standalone post. My last paragraph was supposed to make it clear that this was a launching off point, rather than a summation of everything I know about decision theory or experimentation.
Thanks for writing more substance into this comment. These opinions are sensible, I agree with much that you wrote here. On one: I do use MABs and see them as a feasible method for many organizations, and at some point I'd like to write about some challenges with those too.
Wishing you the best in your techbro or ftxbro or <whatever>bro life too!
Totally agree about visualization, and that those authors are great advocates for it. Confidence intervals are definitely much more informative and intuitive than p-values.
Would the policy be "look at our confidence intervals later and then decide what to do"? One remaining issue is how to have consistent decision criteria, and to convey it ahead of time. Imagine a context with 10-50 teams at a company that run experiments, where the teams are implicitly incentivized to find ways to report their experiments as successful. Quantified criteria can be helpful in minimizing that bad incentive.
It all depends on what you're experimenting on. I do think there are many teams out there who are (in effect) making decisions with less certainty than this, but they wouldn't want to actually quantify it.
Thanks for that. I'll give [1] a read. I'm familiar with [2], and cited one of those papers in the blog.
About the stupid or harmful nature of null hypothesis testing in general, what do you recommend instead for decision making and for summarization of uncertainty? In the scenario of large (yet fast moving) organizations where most people will have little stats background.
The "stasis" and "arbitrarily adjustments" regimes that I wrote about are certainly ones that I've seen, which don't rely solely on p < 0.05 but are still pretty suboptimal. Furthermore, it's not only about whether 0.05 is the sole criteria, but also about whether it's a useful criteria for us to highlight at all, depending on whether the anchoring effect of it is damaging relative to alternatives.
But let me turn that around and ask: what product decision regime do you see most often or think would be the most relevant to use as an example? I'd be happy to hear your perspective and make sure I keep it in mind for future blogs.
The right choice between the "get something done and working" and "your systems should be going down" paths might depend on the situation. Sometimes the right call for the company is to actually have "relative success" despite the downsides if it's better than the downsides of the early failure.
Implicitly, this article implies some communication failures in an organization, where visible failures are used as as attention-getting substitute for what should be healthy communication and prioritization channels.