Showing posts with label cognitive bias. Show all posts
Showing posts with label cognitive bias. Show all posts

Wednesday, 12 January 2011

Reviewing medical literature, part 3: Threats to validity

You have heard this a thousand times: no study is perfect. But what does this mean? In order to be explicit about why a certain study is not perfect, we need to be able to name the flaws. And let's face it: some studies are so flawed that there is no reason to bother with them, either as a reviewer or as an end-user of the information. But again, we need to identify these nails before we can hammer them into a study's coffin. It is the authors' responsibility to include a Limitations paragraph somewhere in the Discussion section, in which they lay out all of the threats to validity and offer educated guesses as to the importance of these threats and how they may be impacting the findings. I personally will not accept a paper that does not present a coherent Limitations paragraph. However, reviewers are not always, as, shall we say, hard assed about this as I am, and that is when the reader is on her own. Let us be clear: even if the Limitations paragraph is included, the authors do not always do a complete job (and this probably includes me, as I do not always think of all the possible limitations of my work). So, as in everything, caveat emptor! Let us start to become educated consumers.

There are four major threats to validity that fit into two broad categories. They are:
A. Internal validity
  1. Bias
  2. Confounding/interaction
  3. Mismeasurement or misclassification
B. External validity
  4. Generalizability
Internal validity refers to whether the study is examining what it purports to be examining, while external validity, synonymous with generalizability, gives us an idea about how broadly the results are applicable. Let us define and delve into each threat more deeply.

Bias is defined as "any systematic error in the design, conduct or analysis of a study that results in a mistaken estimate of an exposure's effect on the risk of disease" (the reference for this is Schlesselman JJ, as cited in Gordis L, Epidemiology, 3rd edition, page 238). I think of bias as something that artificially makes the exposure and the outcome either occur together or apart more frequently than they should. For example, the INTERPHONE study has been criticized for its biased design, in that it defined exposure as at least one cellular phone call every week. Now enrolling such light users can really result in such a small exposure as not to be able to detect any increase in adverse events. This is an example of a selection bias, by far the most common form that bias takes. Another example of a frequent bias is encountered in retrospective case-control studies where people are asked to recall distant exposures. Take for example middle-aged women with breast cancer who are asked to recall their diets when they were in college. Now, ask the same of similar women without breast cancer. What you are likely to get is the effect, absent in women without cancer, of seeking an explanation for the cancer that expresses itself in a bias in what women with cancer recall eating in their youth. So, a bias in the design can make the association seem either stronger or weaker than it is in reality.

I want to skip over confounding and interaction at the moment, as these threats deserve a post of their own, which is forthcoming. Suffice it to say here that a confounder is a factor related to both, the exposure and the outcome. An interaction is also referred to as effect modification or effect heterogeneity. This means that there may be population characteristics that alter the response to the exposure of interest. Confounders and effect modifiers are probably the trickiest concepts to grasp. So, stay tuned for a discussion of those.

For now, let us move on to measurement error and misclassification. Measurement error, resulting in misclassification, can happen at any step of the way: it can be in the primary exposure, a confounder, or the outcome of interest. I run into this problem all the time in my research. Since I rely on administrative coding for a lot of the data that I use, I am virtually certain that the codes routinely misclassify some of the exposures and confounders that I deal with. Take Clostridium difficile as an example. There is an ICD-9 code to identify it in administrative databases. However, we know from multiple studies that it is not all that sensitive or all that specific; it is merely good enough, particularly for making observations over time. But even for laboratory values there is a certain potential for measurement error, though we seem to think that lab results are sacred and immune to mistakes. And need I say more about other types of medical testing? Anyhow, the possibility of error and misclassification is ubiquitous. What needs to be determined by the investigator and the reader alike is the probability of that error. If the probability is high, one needs to understand whether it is a systematic error (for example, a coder always more likely than not to include C. diff as a diagnosis) or a random one (a coder is just as likely to include as not to include a C diff diagnosis). And while a systematic error may result in either a stronger or a weaker association between the exposure and the outcome, a random, or non-differential, misclassification will virtually always reduce the strength of this association.

And finally, generalizability is a concept that helps the reader understand what population the results may be applicable to. In other words, will the data be applied strictly to the population represented in the study? If so, is it because there are biological reasons to think that the results would be different in a different population? And if so, is it simply the magnitude of the association that can be expected to be different or is it possible that even the direction could change? In other words, could something found to be beneficial in one population be either less beneficial or even more harmful in another? The last question is the reason that we perseverate on this idea of generalizability. Typically, a regulatory RCT is much less likely to give us adequate generalizability than a well designed cohort study, for example.

Well, these are the threats to validity in a nutshell. In the next post we will explore much more fully the concepts of confounding and interaction and how to deal with them either at the study design or study analysis stage.            

Monday, 3 January 2011

Gaol fever and intercessory prayer: Redefining the role of p-value?

Happy 2011, everyone! I hope that it is everything you want it to be. Sorry for a brief hiatus in blogging -- needed to recharge my batteries and read others' writing for a change. Well, back now. And thanks to you all for coming back too.

I want to resume our recent discussions of statistical testing in the context of biologic plausibility. We discussed the latter at length a few months ago here, and came to the conclusion that our mere impression of biologic plausibility is not a good litmus test for an association. The oft-cited discovery of H. pylori as the cause of peptic ulcer disease is a tried and true example of the knowledge we would be missing today if we used biologic plausibility as the only yardstick for measuring the prospects of research.

At the same time, we spent a fair bit of time and energy talking about p values and how they need to be used in a Bayesian manner. To review, Bayes theorem relies on pre-test probability of an association to help us understand how much stock we need to put into a finding of an association. That is, the lower the pre-test probability, the more suspicious we should be of an observed association. To put it in concrete terms, for example the finding that intercessory prayer is associated with improved health outcomes requires a much greater amount of scrutiny than one that treating a bacterial infection with an antibiotic improves survival. There is a certain mechanistic elegance to the latter that is missing in the former, unless higher powers are invoked. Here is a quote from the Cochrane meta-analysis of intercessory prayer -- I especially love the last sentence [emphasis mine]:
REVIEWER'S CONCLUSIONS: Data in this review are too inconclusive to guide those wishing to uphold or refute the effect of intercessory prayer on health care outcomes. In the light of the best available data, there are no grounds to change current practices. There are few completed trials of the value of intercessory prayer, and the evidence presented so far is interesting enough to justify further study. If prayer is seen as a human endeavour it may or may not be beneficial, and further trials could uncover this. It could be the case that any effects are due to elements beyond present scientific understanding that will, in time, be understood. If any benefit derives from God's response to prayer it may be beyond any such trials to prove or disprove.
At the same time, just because we do not have a mechanistic explanation at the ready does not mean that we should discount an association. In a rather lengthy post in October I wrote about my own conflicted feelings about applying Bayesian versus frequentist (this refers to all associations standing on similar probabilistic ground prior to testing) thinking in research. Although more Bayesian in my own thinking, I recognize metacognitively that it may at times be a trap:
Yet, there is something to be said about the frequentist approach, even though it is not my way generally. The frequentist approach, which is what underlies the bulk of our traditional clinical research, does not rely on differential prior probabilities for different possible associations, but treats them all equally. Despite many disadvantages, one obvious advantage is that we do not discount potential associations that do not have biologic plausibility, given our current understanding of biology, and sometimes help us stumble on brand new hypotheses. So, clearly, there is a tension here, and I am still working on what is the better way, if any.
The last sentence here implies that there is a right and a wrong way, but having spent the last several months exploring these issues, I am beginning to think that this is incorrect. In fact, all of the p value discussions are leading me to believe that both approaches are useful, and it is the nuances of when either should predominate that need to be worked out.

Consider my examples above -- those of intercessory prayer and antibiotic treatment of a bacterial infection. Let us transport ourselves to, say 18th century England, where typhus, known as "gaol fever", killed more prisoners than the executioners did. How improbable would it have seemed to the medical profession of those days that a). the disease was caused by a microorganism, and b). it could be eradicated with an antibiotic? Why, I would guess that these assertions either would appear heretical or else confirm for the religious the divine presence. Either way, the biology was lacking and the plausibility was simply not there. Yet, this does not change the reality as we understand it today. What explanations will we have 200 years from now for the occasionally observed success of intercessory prayer? And more importantly, what do we do in the meantime to tread most sensibly that purgatory between accepting absurd associations and missing the unlikely ones that are nevertheless real?

The answer may be in the p value after all. Let us model qualitatively what things might look like for intercessory prayer. Let us pretend that we have just conducted the very first randomized controlled trial of the impact of intercessory prayer on the development of post-operative infection following coronary bypass surgery among 1,200 patients. We have found that there is indeed a lowered risk of infection in the intervention group, and the difference has the p value of 0.04. Great, right? We can walk away congratulating ourselves on a positive study. Well, of course this is absurd. Even though we can come up  with some remotely plausible mechanism for this potentially causal association, our pre-test probability is still minuscule. The answer at this point should obviously be what has been suggested for genome-wide interaction studies: a much lower alpha level as the significance threshold. How low? This I cannot answer yet; while the rationale is, similar to genome-wide studies, a fishing expedition without much understanding of why we should find what we should find, here we are not merely engaging in multiple hypotheses testing, the number of which could help determine the appropriate significance level. No, here we are testing a single hypothesis whose mechanism is either absent or highly biologically implausible. So, how to determine the adequate threshold for significance under these circumstances remains unclear to me at this time. I can only say that the traditional 0.05 is highly inappropriate under the circumstances early in the research efforts.

As more studies are performed, their quality and directionality of results should impact how much stock we put in the results. That is, if well done studies consistently continue to demonstrate a positive association of intercessory prayer with clinical outcomes, despite inadequate mechanistic understanding, our level of skepticism should diminish, and commensurately the acceptable alpha can creep higher. In short, the more evidence and the stronger it is, despite poor understanding of why, the more liberal we can afford to be with what we consider a significant result.

So, my point? How we interpret the significance of results needs to be fluid. A p value is not a p value is not a p value. This much embattled and misunderstood statistic may yet be the bridge between Bayesian and frequentist approaches. If we get smarter about setting its thresholds, perhaps we can keep the baby while getting rid of the rancid bath water at the same time. Of course, I am not even attempting to address all of the cognitive biases that derail us in our pursuit of scientific truths. Incorporating them into our inference testing is definitely a discussion for another day.  

Monday, 20 December 2010

Why we need collaborations across healthcare sectors

I want to digress from our recent focus on methods and talk a bit about conflict of interest (COI for short). There has been a lot in the press lately about doctors taking money from the biopharmaceutical manufacturers, and doctors inserting unnecessary hardware into patients' hearts and spines. All of this has been happening against the background of a low hum of an ongoing discussion of what constitutes a COI, how much is too much and for what (for example, can a doc who takes research and education dollars from a manufacturer with an interest in anticoagulation sit on a committee that develops the guidelines for prevention of thromboembolic disease?), and how to mitigate these ubiquitous and pesky COIs.

In some ways watching this discussion has been amusing, while in others it has been downright sad. Medical journals, while insisting that advertising money is OK to take (presumably because the editorial and marketing offices are separated by some sort of a fire wall), though professional societies should not be able to take this tainted education money. Professional societies, on the other hand, are running away from the accusations by tightening their continuing medical education (CME) criteria and scrambling to replace the lavish budgets derived from pharma to develop their coveted evidence-based practice guidelines. And while all the pots are calling all the kettles black, academic researchers are being barred from collaborating with the industry on research projects, and industry researchers are being precluded from presenting their data at professional society meetings. While all the time the public is being whipped into lather about these alleged systematic transgressions, and forced to cheer for the ensuing retribution.

But, like many things in life, and especially stuff that we discuss on this blog, this issue is neither black nor white. Don't take me wrong: I am not condoning the egregious excesses of greed demonstrated by some members of my hallowed profession. If you have been reading my blog for some time, you know that I do not dispute the shameful reality of many breeches of public trust. I am an ardent supporter of exposing these breeches and of harsh punishments that they deserve. This is not what I am talking about here.

I am much more concerned about the one-sided story that we have been hearing about pharma-academic collaborations. Because of the persecutory nature of public opinion, some institutions are now shying away from such collaborations. This attitude is akin to navigating a treacherous road while looking in the rearview mirror. Yes, there have been transgressions, yes there has been greed and even scientific fraud in the name of money. Does this mean that we need to stop everything and come up with an entirely new way of managing these risks? Absolutely! Does this mean that we have to get rid of all pharma-academic collaborations? Absolutely not! In my humble opinion, erecting non-scaleable walls between these two groups is a big mistake. Here is why.

First, let me make a disclaimer: I do have active ongoing collaborations with multiple manufacturers. I do not take speaking or other promotional money, but limit myself to consulting and research grant funding. I also do a good deal of unfunded research, and I have never taken a penny for any of my blogging or blogging-related activities. And here is the crux of the matter: In this world of über-subspecialization, with the expertise being demographically and geographically diffuse, how can we afford not to collaborate across different types of organizations with different types of capabilities? Can we really afford to leave all of therapeutic development in the hands of organizations whose overarching purpose is to make money? And equally importantly, can we afford to continue this fragmented model of medical development without any thought to integration of the needs of all of the stake holders? I think not. Just as we are reaping the fruit of electronic medical record development in isolation from the end-user, so this isolation of research effort will lead to even less coherence in medicine. And unless we are ready to socialize our entire healthcare system, it seems naïve to expect that this one sector will acquiesce and start working outside of our coveted free market for the good of humankind alone.

My readers know that I am not an industry apologist. On the contrary, I have said many times that there has been bad behavior across all the sectors of healthcare, starting with biopharma. But if we want to advance rather than stagnate and regress, we need robust collaborations. We also need higher ethical standards and greater professionalism to keep public's health as our top priority.

There is COI everywhere, and, while financial COI is most visible, it is the more hidden COI that is most insidious. An hidden COI can be intellectual, reputational, ego-driven, career-mediated, etc. It is incumbent on us all in this complex world to ask questions and mitigate any ill effects of any cognitive biases, including those created by COI. Ultimately, as I have begun to realize of late, nothing will replace an educated and empowered patient: This is the only model that can provide appropriate checks and balances for our oftentimes misaligned and perverse incentives, both academic and economic.