Re: 'root causes' -- I find the words somewhat important. Like, if you say you're looking for a root cause, then people tend to be in the sort of moral mindset, and have a harder time seeing it as a collection of contingent events.
Also, in the (truly amazing) "How Complex Systems Fail", he's pretty down on "root cause":
Like the idea of focus on a single core progression for a startup. Very much like the list of Gotchas and Edge Cases. My favorite:
>5. Getting Test Users. I often hear people rationalize a PR push as the only way they can get enough users to test product market fit.
...
The solution isn’t PR, it’s go to some events and make some friends in that market.
Yo, the author here. Thanks for the feedback. I totally meant idempotency, drat. (In fact, on Hadoop, thanks to speculative execution of reduce tasks, you also have to worry a bit about reentrancy, but what I was talking about was, in fact, idempotency).
Shutting down the pipeline: I hear you on prod/non-prod. For our setup, the pipeline ends up writing to a datastore, so if we kill the pipeline, the datastore is still up, it just stops updating. Which is working so far. May end up flagging suspect data as you suggest, instead of the full stop (or only full stop if more than a very small percentage of the data is suspect).
Because I have that same problem -- if someone has nice visual slides, the Slideshare is often kind of useless.