There are a few things that really bother me about this post and many of the comments regarding it.
First, let me address a common theme that I see underlying this discussion and many others: theory vs reality. The common claim seems to be that scientists or other members of the academic community are often disconnected with "the real world" and blindly hold to models that don't quite work. This is both a damaging idea and a false one.
Science and math are two different things and science has always been about the real world. The key difference between science and math (or theory as some like to call it) is empirical data. In the physical sciences the fact that collected data contains errors or biases is a fundamental part of the process. Characterizing and understanding the error or other problems with collected data is critical to doing good science.
It seems that people who study computer science often forget this. When does computer science change from mathematics to a proper science? When you start applying it to real data or to real systems that contain imperfections. Situations precisely like the one being described. In this case there is nothing wrong with Bayes' theory and as many have pointed out the theory is not being tweaking at all. Rather, change in the mathematics is to address a shortcoming of the data. This is simply how science should work. Science is the real world and if you think otherwise you are not understanding or practicing good science.
My second criticism is with the suggestion that we should not care why this change improves results. Now, John Cook is not actually advocating this, but a casual reading of his blog post could certainly be read that way in error. He is saying that it's ok to use an algorithm in practice even though we don't fully understand it. In so far as it's been properly tested, I doubt most people nor most academics would argue with this.
But it's still important to try to understand what is going wrong in the application of the mathematics such that a fudge factor is required. Very often, understanding the root cause can lead to better methodology or a better modelling of the data.
First, let me address a common theme that I see underlying this discussion and many others: theory vs reality. The common claim seems to be that scientists or other members of the academic community are often disconnected with "the real world" and blindly hold to models that don't quite work. This is both a damaging idea and a false one.
Science and math are two different things and science has always been about the real world. The key difference between science and math (or theory as some like to call it) is empirical data. In the physical sciences the fact that collected data contains errors or biases is a fundamental part of the process. Characterizing and understanding the error or other problems with collected data is critical to doing good science.
It seems that people who study computer science often forget this. When does computer science change from mathematics to a proper science? When you start applying it to real data or to real systems that contain imperfections. Situations precisely like the one being described. In this case there is nothing wrong with Bayes' theory and as many have pointed out the theory is not being tweaking at all. Rather, change in the mathematics is to address a shortcoming of the data. This is simply how science should work. Science is the real world and if you think otherwise you are not understanding or practicing good science.
My second criticism is with the suggestion that we should not care why this change improves results. Now, John Cook is not actually advocating this, but a casual reading of his blog post could certainly be read that way in error. He is saying that it's ok to use an algorithm in practice even though we don't fully understand it. In so far as it's been properly tested, I doubt most people nor most academics would argue with this.
But it's still important to try to understand what is going wrong in the application of the mathematics such that a fudge factor is required. Very often, understanding the root cause can lead to better methodology or a better modelling of the data.