How do you answer? Notice that this is a problem regardless of whether you are a big company or a small company.
b) 3 months later, your client comes back and asks: “we are having trouble with customer support. How do we know that it’s not related to this change we made?” With your superior experience working with hundreds of startups, you are able to tell them if it is or isn’t after some investigation. Your client asks you: “how can we do that for ourselves without calling on you every time we see something weird?”
How do you answer?
(My answers are in the WBR essay and the essay that comes immediately before that, natch)
It is a common excuse to wave away these ideas with “oh, these are big company solutions, not applicable to small businesses.” But a) I have applied these ideas to my own small business and doubled revenue; also b) in 1992 Donald Wheeler applied these methods to a small Japanese night club and then wrote a whole book about the results: https://www.amazon.sg/Spc-Esquire-Club-Donald-Wheeler/dp/094...
Wheeler wanted to prove, (and I wanted to verify), that ‘tools to understand how your business ACTUALLY works’ are uniformly applicable regardless of company size.
If anyone reading this is interested in being able to answer confidently to both questions, I recommend reading my essays to start with (there’s enough in front of the paywall to be useful) and then jump straight to Wheeler. I recommend Understanding Variation, which was originally developed as a 1993 presentation to managers at DuPont (which means it is light on statistics).
If you are interested in these ideas, you should know that this essay kicks off a series of essays that culminates, a year later, with an examination of the Amazon-style Weekly Business Review:
(It took that long because of a) an NDA, and b) it takes time to put the ideas to practice and understand them, and then teach them to other business operators!)
The ideas presented in this particular essay are really attributed to W. Edwards Deming, Donald Wheeler, and Brian Joiner (who created Minitab; ‘Joiner’s Rule’, the variant of Goodhart’s Law that is cited in the link above is attributed to him)
Most of these ideas were developed in manufacturing, in the post WW2 period. The Amazon-style WBR merely adapts them for the tech industry.
I hope you will enjoy these essays — and better yet, put them to practice. Multiple executives have told me the series of posts have completely changed the way they see and run their businesses.
And how are you going to tell that when you have a) variation (that is, every metric wiggles wildly)? And also b) how are you able to tell if it has or hasn’t impacted other parts of your business if you do not have a method for uncovering the causal model of your business (like that aquarium quote you cited earlier?)
Reality has a lot of detail. It’s nice to quote books about goals. It’s a different thing entirely to achieve them in practice with a real business.
I should note that this essay kicks off an entire series that eventually culminates in a detailed examination of the Amazon Weekly Business Review (which takes some time to get to because of a) an NDA, and b) it took some time to test it in practice). The Goodhart’s Law essay uses publicly available information about the WBR to explain how to defeat Goodhart’s Law (since the ideas it draws from are five decades old); the WBR itself is a two decades-old mechanism on how to actually accomplish these high-falutin’ goals.
Over the past year, Roger and I have been talking about the difficulty of spreading these ideas. The WBR works, but as the essay shows, it is an interlocking set of processes that solves for a bunch of socio-technical problems. It is not easy to get companies to adopt such large changes.
As a companion to the essay, here is a sequence of cases about companies putting these ideas to practice:
The common thing in all these essays is that it doesn’t stop at high-falutin’ (or conceptual) recommendation, but actually dives into real world application and practice. Yes, it’s nice to say “let’s have a re-evaluation date.” But what does it actually look like to get folks to do that at scale?
Well, the WBR is one way that works in practice, at scale, and with some success in multiple companies. And we keep finding nuances in our own practice: https://x.com/ejames_c/status/1849648179337371816
What would you say if I told you Bryar has lots of stories of this style of thinking applied in early Amazon? This is pre-AWS Amazon, mind you — where they were trying to figure out how to build e-commerce web software at scale, from scratch. Granted, the bulk of their process control was directed at customer-facing controllable input metrics, but the software engineers were as much a part of it as the operational folks.
(To be fair to you, you are adamant that SPC does not apply to software development — which I take to mean measuring the productivity or act of building software. And I think we are all in agreement there! (That said, like kqr and jacques_chester, I want to believe that this has not been sufficiently explored) But it's not true that SPC has no place in software development — one way I've used this is that because XmR charts detect changes in variation, you can use it in a customer-facing software context to see if a feature change has resulted in user behaviour change without running an A/B test. Naturally, it makes sense to have the software engineer be responsible for observing this behaviour change themselves, since XmR charts are easy enough for the layman to use, and it gives them a sense of ownership for the feature or change. Some detail (on usage vs A/B tests) here: https://commoncog.com/two-types-of-data-analysis/)
To be fair to OP, Wheeler never claims that for stable/in-control/predictable processes roughly half of the measurements will lie above the average. The only claim he makes is that 97% of all data points for a stable process (assuming the process draws from a J-curve or single-mound distribution) will fall between the limit lines.
He can't make this claim (about ~half falling above/below the average line), because one of the core arguments he makes is that XmR charts are usable even when you're not dealing with normal distributions. He argues that the intuition behind how they work is that they detect the presence of more than one probability distribution in the variation of a time series.
So I've got a dumb question here: what happens when you use vanilla XmR charts with J-curve shaped or sub-exponential distributions?
My current simplistic (and very dumb!) solution that I've used for power-law type distributions — like HN virality, for instance — is to count the number of days between viral events, and then subject that to process control.[1] I basically take Wheeler's approach to chunky data and use that for J-curve type data, which tells me if the behaviour of my 'HN virality process' has changed.
I'd be very interested to learn of other approaches.
[1] HN traffic for commoncog.com displays routine variation most weeks with an Upper Process Limit of 192 and a Lower Process Limit of 0, unless one of my articles hit the front page, at which point I get 11-16k additional uniques).
A couple of quick notes, from someone who has actually put this to practice — and in a non-manufacturing context, to boot!
(From a brief reading of this thread, it seems like kqr, jacques_chester, and I are the only ones who have put this to practice in non-manufacturing contexts — though correct me if I'm wrong.)
The bulk of the debate in this HN thread seems to be centred around what is or isn't a 'stable process'. I think this is partially a terminology issue, which Donald Wheeler called out in the appendix of Understanding Variation. He recommends not using words like 'stable' or 'in-control', or even 'special cause variation', as the words are confusing ... and in his experience lead people to unfruitful discussions.
Instead, he suggests:
- Instead of calling this 'Statistical Process Control', call this 'Methods of Continual Improvement'
- Use the term 'routine variation' and 'exceptional variation' whenever possible. In practice, I tend to use 'special variation' in discussion, not 'exceptional variation', simply because it's easier to say.
- Use the term 'process behaviour chart' instead of 'process control chart' — we use these charts to characterising the behaviour of a process, not merely to 'control' it.
- Use 'predictable process' and 'unpredictable process' (instead of 'stable'/'in-control' vs 'unstable'/'out-of-control' processes) because these are more reflective of the process behaviours. (e.g. a predictable process should reliably show us data between two limit lines).
Using this terminology, the right question to ask is: are there processes in software development that display routine variation? And the answer is yes, absolutely. kqr has given a list in this comment: https://news.ycombinator.com/item?id=39638491
In my experience, people who haven't actually tried to apply SPC techniques outside of manufacturing do not typically have a good sense for what kinds of processes display routine variation. I would urge you to see for yourself: collect data, and then plot it on an XmR chart. It usually takes you only a couple of seconds to see if it does or does not apply — at which point you may discard the chart if you do not find it useful. But you should discover that a surprisingly large chunk of processes do display some form of routine variation. (Source: I've taught this to a handful of folk by now — in various marketing/sales and software engineering roles —and they typically find some way to use XmR charts relatively quickly within their work domains).
[Note: this 'XmR charts are surprisingly useful' is actually one of the major themes in Wheeler's Making Sense of Data — which was written specifically for usage in non-manufacturing contexts; the subtitle of the book is 'SPC for the Service Sector'. You should buy that book if you are serious about application!]
I realise that a bigger challenge with getting SPC adopted is as follows: why should I even use these techniques? What benefits might there be for me? If you don't think SPC is a powerful toolkit, you won't be bothered to look past the janky terminology or the weird statistics.
So here's my pitch: every Wednesday morning, Amazon's leaders get together to go through 400-500 metrics within one hour. This is the Amazon-style Weekly Business Review, or WBR. The WBR draws directly from SPC (early Amazon exec Colin Bryar told me that the WBR is but a 'process control tool' ... and the truth is that it stems from the same style of thinking that gives you the process behaviour chart). What is it good for? Well, the WBR helps Amazon's leaders build a shared causal model of their business, at which point they may loop on that model to turn the screws on their competition and to drive them out of business.
But in order to understand and implement the WBR, you must first understand some of the ideas of SPC.
If that whets your interest, here is a 9000 word essay I wrote to do exactly that, which stems from 1.5 years of personal research, and then practice, and then bad attempts at teaching it to other startup operator friends: https://commoncog.com/becoming-data-driven-first-principles/
I don't get into it too much, but the essay calls out various other applications of these ideas, amongst them the Toyota Production System (which was bootstrapped off a combination of ideas taught by W Edwards Deming — including the SPC theory of variation), Koch Industries's rise to powerful conglomerate, Iams pet foods, etc etc.
Yes, rest assured the previous article talks about exactly this. (I mean, the title of this piece should give you a hint: “Pay attention to deviations from mainstream incentives” because the mainstream incentives in F&B are just so, so bad.)
Discourse is one of the best pieces of software I've ever used — I self-host, and it's never given me any problems. There's built-in backups to S3, I can upgrade the software 80% of the time from a web interface (even on my phone), and the community and moderation features are incredibly well thought out. So much so that if you take the time to slowly go through all the options in the admin panels, you'd realise that the team has thought through 90% of all community interaction models.
I'm looking forward to trying Discourse chat — if the level of engineering and thoughtfulness is the same as the rest of the feature-set, it'll be a great addition to my community.
This is actually predicted by the COR developmental model presented in the article. The terminology might be a bit wonky, but the basic idea is that ‘loss’ produces a stress response. In other words, if you passionately pursue a project, only to have management kill the project a year or so later, you’re going to feel ‘loss’ (wonkily defined) and be at risk for burnout past a certain loss threshold.
> In response, some folk back into the weak form of the argument. The weak form is what you get when you hear someone say “history doesn’t repeat itself, but it rhymes” — that is, perhaps the lessons of history are too specific to a particular context to generalise, but that some … similarity is consistent over time
Well, it seems you're drawing on the same argument that has already been addressed (or at least mentioned in passing) in the article.
So ... pop quiz: how is reading for concept instantiations different from reading for patterns?
The problem is that Dr Ung walks back certain elements of his report, as is summarised in the final judgment:
> the High Court found that the appellant had borderline intellectual functioning; not that he was suffering from mild intellectual disability. This was conceded by the appellant’s own psychiatrist, Dr Ung Eng Khean (“Dr Ung”). Further, Dr Ung also accepted (see Nagaenthran (CM) at [76]) that borderline intellectual functioning is not a mental “disorder” as set out in the American Psychiatric Association, Diagnostic and Statistical Manual of Mental Disorders (American Psychiatric Association Publishing, 5th Ed, 2013). Further, in Nagaenthran (CA) (at [34]–[41]), we held that even assuming the appellant suffered from an abnormality of mind, any such abnormality did not substantially impair his mental responsibility, because he did not lose his ability to tell right from wrong.
I then think Nagaenthran's lawyer made a grievous concession:
> [The appellant’s counsel, Mr Thuraisingam] eventually conceded that this was a case of a poor assessment of the risks on the appellant’s part. But, as the Minister stated in Singapore Parliamentary Debates, Official Reports (14 November 2012) vol 89 … ‘[g]enuine cases of mental disability are recognised [under s 33B(3)(b) of the MDA], while, errors of judgment will not afford a defence’. To put it quite bluntly, this was the working of a criminal mind, weighing the risks and countervailing benefits associated with the criminal conduct in question. The appellant in the end took a calculated risk which, contrary to his expectations, materialised. Even if we accepted that his ability to assess risk was impaired, on no basis could this amount to an impairment of his mental responsibility for his acts. He fully knew and intended to act as he did. His alleged deficiency in assessing risks might have made him more prone to engage in risky behaviour; that, however, does not in any way diminish his culpability.
I'm not sure if your reading of 'intellectually disabled man' aligns with the court's reading, and I don't think it's as clear cut as "oh, Singapore executed a man with no mental responsibility over what he is doing, according to the rule of law there, and so LHL is evil"
> Malaysia currently has a moratorium on the death penalty, due to serious national debate (in which the government has been an active participant) on whether to abolish it, or at least significantly narrow its scope
You're right on the moratorium. 'Due to healthy debate' is a rather charitable reading, though. There's a depressing account of the entire moratorium in Chapter 45, 'Law Reform', from former AG Tommy Thomas's autobiography My Story: Justice in the Wilderness, which lays out the background machinations of the former administration's attempt to repeal the death penalty, which was ultimately a failure. The moratorium was implemented by the Prisons Board under the instructions of acting AG Engku Nor Faizah Engku Atek, and supported by Thomas. More surprisingly, (if Thomas is to be believed) the Prisons Board themselves had no objection to abolishing the death penalty! I won't try to summarise the complex political machinations here, but saying that 'strong National debate' is a factor for the moratorium would not be Tommy Thomas's reading of the situation. I don't expect the moratorium to last — though I pray that it will. But it's difficult to say which side might use it as a political football given the current state of Malaysian politics.
My overall point: using the moratorium on the death penalty as an example of Malaysia's 'enlightened approach' vs Singapore is grossly mistaken; it's more accurate to describe it as a political football used to score points against one's opponents (made complicated by the fact that there is quiet support for the penalty from both the conservative Malay power base as well as the conservative Chinese political power base, on both sides of the parliament, as Thomas found out the hard way). I say that the chapter is depressing because 'national debate' seems to have very little to do with it.
I'd actually urge you to look into the details of the case — Nagaenthran had two psychiatrists and a psychologist evaluate him on separate occasions, each of them certifying that he was of borderline intelligence, but not mentally ill (or 'mentally retarded', per DSM-5). The third examination established that Nagaenthran knew he was carrying drugs, that he was the member of a gang, that he was carrying out such an act because of 'a sense of loyalty and gratitude to his boss', that he had not been coerced into delivering the drugs, and that he was aware of the sentence for drug trafficking in Singapore.
Nagaenthran's lawyer disagreed with these assessments, but, unable to find a psych willing to certify Nagaenthran's intellectual disability, got so desperate that he signed an affidavit declaring himself convinced that Nagaenthran was intellectually disabled. The court rejected this affidavit after establishing that Nagaenthran's lawyer had no medical expertise and after he admitted that he was essentially 'speculating on what the appellant's mental age was'. Then the AG's office got a fourth medical assessment in 2021, which was blocked from submission to the court by Nagaenthran's lawyer, arguing that the 'appellant was interested in 'medical confidentiality'. At that point the court sort of threw up its hands and went 'you claim he's intellectually disabled but now you want to block medical reports on that exact issue, due to privacy reasons?'
Keep in mind that at this point, the case had been in Singapore's courts since 2010, and the various assessments and appeals had lasted more than a decade.
Nagaenthran's family, with the help of activists, requested appeals from both Malaysian's prime minister and a direct appeal to Singapore's president, who had the power to stay the execution by issuing a pardon. Singapore's president said she looked into it and was satisfied that the case was sound.
Finally, Nagaethran's family — again with the help of activists — filed a final appeal, arguing that the sentence was unconstitutional because Chief Justice Sundaresh Menon, who presided over Nagaenthran's previous failed appeals, was also the Attorney-General during his conviction.
At this point the panel of three judges threw the application out on the grounds that it was a 'calculated attempt' to diminish the finality of the court process. "No court in the world would allow an applicant to prolong matters ad infinitum" by filing such applications, said the judges. They'd been pushing such applications over the course of a full decade by that point, switching arguments three times.
I'm sympathetic to the goals of the activists — they want to push a test case to repeal the death penalty. The problem is that it doesn't seem to be a particularly good test case. Nagaenthran had successfully stayed his execution multiple times over the past decade, and he had pretty decent legal representation throughout. (Notwithstanding the affidavit — which probably harmed his case more than it helped).
You may argue that Singapore is a 'rather extreme outlier among developed countries' — but its policies on drug enforcement is 100% consistent with just about every other country in its vicinity. In fact, had Nagaenthran been caught in his home country of Malaysia, he would be sentenced to death just the same, since Malaysia has pretty much the same laws on drug trafficking. The activists have been raising international awareness of the case, but Malaysia's government is pretty blase about the whole thing, since they have similar laws and a similar justice system, and Malaysian anti-narcotics officers have a long history of working with Singaporean anti-narcotics officers to stop the drug trade from trickling into their countries from the Golden Triangle to their north.
Edit to add: I was actually prepared to get angry over this case, but then I read the full judgment and now I think ... "hmm, this is a more complicated case than I thought."
Author here. Quite the contrary, actually — I’ve read all the above sources. On top of that, with the exception of Chambliss (which is a landmark work, to be clear), and Dreyfus (which I prefer to look at adaptations of; Accelerated Expertise has several updated or derivative scales that expertise researchers in the military use), I’ve talked to collaborators of these researchers, and in the case of Klein, am actively in contact with and am following the rest of the work of his community (https://naturalisticdecisionmaking.org/).
I’ve updated the bottom of the post to link to the many other articles I’ve written on expertise. These are either summaries of papers, or summaries of branches of the research, or pointing out obvious holes in the popsci representations of the research, or synthesise various (mostly military-funded) approaches to accelerating expertise.
This particular essay is a note from personal application.
I disagree with this assessment. (Or, more accurately, I’d like to see a more comprehensive counter argument).
The article you link to is essentially a response to 2 pages in the book, where Hoffman et al mention, almost in passing, that CLT is a silly theory when you want to train for real world scenarios (the intuition is that if you’re training marine fire squad commanders to plan on the battlefield, perhaps it helps to simulate shooting at them during training?) Hoffman et al use this as an example of a learning theory that doesn’t seem to map to real world requirements.
This reads like a disagreement over one particular dismissal in the book, perhaps because CLT is a pet theory of the article’s authors. The problem: this argument is not core to the book!
The article does not, for instance,
a) Deal with the many examples of successful real world accelerated training programs with no curriculum design (as is commonly understood; ordering of simulations isn’t really designing a syllabus) in Chapter 9 (some of which were designed by some of the authors)
b) Have a rejoinder to the two learning theories presented in Chapter 11 that the authors claim underpins their training approach (if there were something to attack, this would be it!)
c) Nor have a rejoinder to a more central claim in the book, (and to my mind a more controversial claim) that atomisation of concepts impedes rapidised training.
And, perhaps most surprisingly to me, your claim that
> Reading it I got the distinct impression that the authors did not understand a great deal of the research they cited, either when supporting or dismissing it.
is remarkable, given that one of the authors of Accelerated Expertise is Paul J Feltovich, one of the founders of the field of expertise research, and a contemporary of Ericsson’s.
1. Deliberate practice only works for skills with a history of good pedagogical development. If no such pedagogical development exists, you can’t do DP. Source: read Peak, or any of Ericsson’s original papers. Don’t read third party or popsci accounts of DP.
2. Once you realise this, then the next question you should ask is how can you learn effectively in a skill domain where no good pedagogical development exists? Well, it turns out a) the US military wanted answers to exactly this question, and b) a good subsection of the expertise research community wondered exactly the same thing.
3. The trick is this: use cognitive task analysis to extract tacit knowledge from the heads of existing experts. These experts built their expertise through trial and error and luck, not DP. But you can extract their knowledge as a shortcut. After this, you use the extracted tacit knowledge to create a case library of simulations. Sort the simulations according to difficulty to use as training programs. Don’t bother with DP — the pedagogical development necessary for DP to be successful simply takes too long.
Broadly speaking, DP and tacit knowledge extraction represent two different takes on expertise acquisition. For an overview of this, read the Oxford Handbook of Expertise and compare against the Cambridge Handbook of Expertise. The former represents the tacit knowledge extraction approach; the latter represents the DP approach. Both are legitimate approaches, but one is more tractable when you find yourself in a domain with underdeveloped training methods (like most of the skill domains necessary for success in one’s career).
My two questions (a) and (b) were not rhetorical. Let’s get concrete.
a) You are advising a company to “check back after a certain period”. After the certain period, they come back to you with the following graph:
https://commoncog.com/content/images/2024/01/prospect_calls_...
“How did we do? Did we improve?”
How do you answer? Notice that this is a problem regardless of whether you are a big company or a small company.
b) 3 months later, your client comes back and asks: “we are having trouble with customer support. How do we know that it’s not related to this change we made?” With your superior experience working with hundreds of startups, you are able to tell them if it is or isn’t after some investigation. Your client asks you: “how can we do that for ourselves without calling on you every time we see something weird?”
How do you answer?
(My answers are in the WBR essay and the essay that comes immediately before that, natch)
It is a common excuse to wave away these ideas with “oh, these are big company solutions, not applicable to small businesses.” But a) I have applied these ideas to my own small business and doubled revenue; also b) in 1992 Donald Wheeler applied these methods to a small Japanese night club and then wrote a whole book about the results: https://www.amazon.sg/Spc-Esquire-Club-Donald-Wheeler/dp/094...
Wheeler wanted to prove, (and I wanted to verify), that ‘tools to understand how your business ACTUALLY works’ are uniformly applicable regardless of company size.
If anyone reading this is interested in being able to answer confidently to both questions, I recommend reading my essays to start with (there’s enough in front of the paywall to be useful) and then jump straight to Wheeler. I recommend Understanding Variation, which was originally developed as a 1993 presentation to managers at DuPont (which means it is light on statistics).