What makes software engineering for climate models different? (2010)(easterbrook.ca)
easterbrook.ca
What makes software engineering for climate models different? (2010)
http://www.easterbrook.ca/steve/2010/03/what-makes-software-engineering-for-climate-models-different/
3 comments
The impact of errors in climate model software engineering that causes the software to fail safely (crash, return no output, fail to compile, etc.) is zero, which is different than other industries—effectively, a climate model is only needed insofar as it is predicting something with higher precision than any other model we already have, and, when it fails (in essence, when it provides output with zero confidence), we simply fall back on what we were doing before we had it—sourcing our predictions from some other model, or perhaps relying on the "model" of our own intuition.
Of course, there are also software engineering errors that cause incorrect outputs from the model to be provided, and incorrectly high confidence to be attached to those predictions. But, given the kinds of errors commonly seen in similar fields (e.g. physics engines), I would expect the probability of those kinds of error to be far outweighed by the other kind, the 'clear errors' where the model encounters out-of-tolerance data, and can then either bail out or "taint" all computations going forward.
Of course, there are also software engineering errors that cause incorrect outputs from the model to be provided, and incorrectly high confidence to be attached to those predictions. But, given the kinds of errors commonly seen in similar fields (e.g. physics engines), I would expect the probability of those kinds of error to be far outweighed by the other kind, the 'clear errors' where the model encounters out-of-tolerance data, and can then either bail out or "taint" all computations going forward.
The author of the article seems to think that those involved are all unique butterflies simply because they write software in an academic environment. Sure, writing software in academia is a lot different than in a business setting. But guess what?! There's a TON of software written in academia!
The vast majority of all models that ever get written are done in academia like nuclear simulation models or geological models or whatever. In college I worked in high performance computing which led me to write software in academia for blast simulations and for genetic sequence alignment. Both share lots of the traits that the author outlined, but neither had to do with climate.
Just because your situation isn't a business doesn't mean you're unique.
The vast majority of all models that ever get written are done in academia like nuclear simulation models or geological models or whatever. In college I worked in high performance computing which led me to write software in academia for blast simulations and for genetic sequence alignment. Both share lots of the traits that the author outlined, but neither had to do with climate.
Just because your situation isn't a business doesn't mean you're unique.
The nice thing about working within science specifically, though, is that one paper standing on its own is pretty meaningless. If you get your model wrong, that'll show up when a meta-analysis of your paper (with its model) against several others (with their own models) drops yours as an outlier.
The author explains exactly what he means by that in a bullet point.
"Because there are many other modeling groups, and scientific results are filtered through processes of replication, and systematic assessment of the overall scientific evidence, the impact of software errors on, say, climate policy is effectively nil."
The commenter seemed to completely ignore the explanation, and took the quote out of context. The rebuttal was answered in the original statement.
"Because there are many other modeling groups, and scientific results are filtered through processes of replication, and systematic assessment of the overall scientific evidence, the impact of software errors on, say, climate policy is effectively nil."
The commenter seemed to completely ignore the explanation, and took the quote out of context. The rebuttal was answered in the original statement.
I don't find the author's explanation very reassuring for a couple of reasons:
1. There's probably a considerable amount of overlap between the software used by different research groups. The article says:
"Single Site Development – virtually all coupled climate models are managed and coordinated at a single site, once they become sufficiently complex, usually a government lab as universities don’t have the resources"
Which implies that several university research groups use the software being maintained by a handful of large labs. They might be running models with different parameters, but the underlying modeling software is the same. So an error in the software might be widely propagated across hundreds or thousands of research papers. Also new scientific results frequently build upon previously published research, so anyone who cited a buggy result as evidence may have a weaker case.
2. If other scientists have similar attitudes toward writing software (i.e., that climate software is "different" from other software), they're also not likely to be using the best practices of software development, and much of the modeling software in the world is likely to be unreliable. In that situation, how would you be able to figure out which modeling software is giving the correct result?
As that commenter wrote in another comment:
"Unfortunately, I have encountered the following argument a distressing number of times:
1. As a good scientist, I am automatically a good engineer.
2. As a good engineer, I am automatically a good programmer.
3. As a good programmer, I am automatically a good Software Quality Assurance Analyst (whatever that is, nothing significant I would guess).
What hubris. Programming is a domain in its own right. It’s the most complicated domain there is because if people could make their programs any more complicated they would. Programming is also an art. Even a computer science degree does not make you a good programmer."
1. There's probably a considerable amount of overlap between the software used by different research groups. The article says:
"Single Site Development – virtually all coupled climate models are managed and coordinated at a single site, once they become sufficiently complex, usually a government lab as universities don’t have the resources"
Which implies that several university research groups use the software being maintained by a handful of large labs. They might be running models with different parameters, but the underlying modeling software is the same. So an error in the software might be widely propagated across hundreds or thousands of research papers. Also new scientific results frequently build upon previously published research, so anyone who cited a buggy result as evidence may have a weaker case.
2. If other scientists have similar attitudes toward writing software (i.e., that climate software is "different" from other software), they're also not likely to be using the best practices of software development, and much of the modeling software in the world is likely to be unreliable. In that situation, how would you be able to figure out which modeling software is giving the correct result?
As that commenter wrote in another comment:
"Unfortunately, I have encountered the following argument a distressing number of times:
1. As a good scientist, I am automatically a good engineer.
2. As a good engineer, I am automatically a good programmer.
3. As a good programmer, I am automatically a good Software Quality Assurance Analyst (whatever that is, nothing significant I would guess).
What hubris. Programming is a domain in its own right. It’s the most complicated domain there is because if people could make their programs any more complicated they would. Programming is also an art. Even a computer science degree does not make you a good programmer."
The context is also highly political. The results can be misconstrued and used to fit a storyline. The author states it has a huge societal impact, which is true - but I feel like climate change models are more important to politians than the majority of citizens.
“developers are domain experts – they do not delegate programming tasks to programmers, which means they avoid the misunderstandings of the requirements common in many software projects”
This also likely means that these systems are poorly designed, poorly documented, poorly implemented, and poorly understood.
This is only my assumption based upon my reading of this article, but this assumption does come with some experience. I have worked in biotech my entire career, often alongside domain experts who write code. This code is typically, to be kind, not very good. It is rife with bugs, makes too many assumptions, is difficult to understand (and, as a consequence, its limitations and assumptions are not well understood by its users), full of copied code, etc. etc. etc. I have rewritten a few such systems.
Also, comments like the following are scary and I can only hope that they are not indicative of the general attitude in the field (though, based upon my own experience, may very well be):
"The software has huge societal importance, but the impact of software errors is very limited."
Yikes. The entire statement seems like a non-sequitur to me, but that attitude leads down a dangerous road. So we have potentially buggy systems which output data used in studies which have "huge societal impact"? How can the author make the claim that software errors don't have an appreciable impact on the result if the system was not developed using standard, accepted engineering practices? How do the users know how to correctly interpret the data?
As an analogy, I recently rewrote an imaging and image processing system used by my company. This system was designed and implemented by academics, and exhibits all of the problems we in the software industry typically associate with such code.
While rewriting it from scratch, I had no documentation to rely upon. I found many implicit and explicit assumptions that the users where not aware of. Most importantly, the system was originally designed for enumeration of certain types of cells, but not for any sort of quantitative interpretation.
However, down the road, the users in the lab realized that they raw output of the image analysis process could contain useful information. So, they started mining it. They began comparing samples using various measurements taken during analysis. They began making even more assumptions about what that data meant, but they were often wrong.
On the surface, it seemed as though their work made sense, but only if one did not understand how those numbers were gathered and under what circumstances their interpretation was valid. Some of the statements made by the author show a striking resemblance to the opinions of the original authors of the system I had to rewrite.
These people were smart, very smart, but not engineers. They didn't have the discipline, training, or experience required to write a system that would stand up to scrutiny. It was a research vehicle, and it did what it was originally intended to do, but as time passed, warts appeared.
I find it very hard to believe that climate model programming has even one single characteristic which would cause an engineer to think that a different engineering model was required or even warranted. To me this sounds like people in the research/academic camp making statements about an aspect of engineering that they do not understand.
This also likely means that these systems are poorly designed, poorly documented, poorly implemented, and poorly understood.
This is only my assumption based upon my reading of this article, but this assumption does come with some experience. I have worked in biotech my entire career, often alongside domain experts who write code. This code is typically, to be kind, not very good. It is rife with bugs, makes too many assumptions, is difficult to understand (and, as a consequence, its limitations and assumptions are not well understood by its users), full of copied code, etc. etc. etc. I have rewritten a few such systems.
Also, comments like the following are scary and I can only hope that they are not indicative of the general attitude in the field (though, based upon my own experience, may very well be):
"The software has huge societal importance, but the impact of software errors is very limited."
Yikes. The entire statement seems like a non-sequitur to me, but that attitude leads down a dangerous road. So we have potentially buggy systems which output data used in studies which have "huge societal impact"? How can the author make the claim that software errors don't have an appreciable impact on the result if the system was not developed using standard, accepted engineering practices? How do the users know how to correctly interpret the data?
As an analogy, I recently rewrote an imaging and image processing system used by my company. This system was designed and implemented by academics, and exhibits all of the problems we in the software industry typically associate with such code.
While rewriting it from scratch, I had no documentation to rely upon. I found many implicit and explicit assumptions that the users where not aware of. Most importantly, the system was originally designed for enumeration of certain types of cells, but not for any sort of quantitative interpretation.
However, down the road, the users in the lab realized that they raw output of the image analysis process could contain useful information. So, they started mining it. They began comparing samples using various measurements taken during analysis. They began making even more assumptions about what that data meant, but they were often wrong.
On the surface, it seemed as though their work made sense, but only if one did not understand how those numbers were gathered and under what circumstances their interpretation was valid. Some of the statements made by the author show a striking resemblance to the opinions of the original authors of the system I had to rewrite.
These people were smart, very smart, but not engineers. They didn't have the discipline, training, or experience required to write a system that would stand up to scrutiny. It was a research vehicle, and it did what it was originally intended to do, but as time passed, warts appeared.
I find it very hard to believe that climate model programming has even one single characteristic which would cause an engineer to think that a different engineering model was required or even warranted. To me this sounds like people in the research/academic camp making statements about an aspect of engineering that they do not understand.
The author seems to have a rather cavalier attitude about the correctness of code, which one of the commenters (George Crews) picked up on:
"Then there is the statement: 'The software has huge societal importance, but the impact of software errors is very limited.' I don’t see how it can be both ways. How can something be of great importance whether or not it is correct? IMHO, the most serious consequence of a climate software being defective would be to then use it to make a defective political decision costing trillions of dollars to society."