Tuesday, 11 January 2011

Do private ICU rooms really reduce HAIs?

We have known for quite some time now that the patient's environment in a hospital matters to his/her outcomes. The concept of biophilia was applied by Roger Ulrich back in the 1980s to surgical patients in a series of experiments. Famously, this work showed that looking out your hospital room's window on a bunch trees is associated with better and less eventful post-operative recovery than staring at a brick wall, for example. We have also known for some time that some of the hospital-associated delirium can be mitigated by having the patient dwell in a room with a window and be exposed to the diurnal light changes.

Another, perhaps even more tangible outcome that can be modified by hospital design is the spread of hospital-acquired infections. This week a paper in the Archives of Internal Medicine from the group in Quebec, who brought us detailed reports of the devastating multihospital hypervirulent Clostridium difficile outbreak in the last decade, generally confirms the effectiveness of private ICU rooms in containing the spread of HAIs. There are some interesting details to point out.

For example, the intervention hospital appears to have had a higher proportion of medical patients than the control institution. Why is this important? Well, medical patients generally experience more chronic and therefore longer stays in the ICU. This gives them a greater opportunity for exposure to HAIs than their surgical counterparts. On the other hand, we know that VAP, for example, an infection very likely to be caused by one of the resistant organisms listed in Table 2 of the paper, happens much more frequently in trauma ICUs than medical ICUs.

Second, the unadjusted ICU length of stay shows some interesting results, depicted in the graph below:
So, while at the intervention hospital the raw ICU LOS has remained stable, at the comparator institution it has been slowly creeping up. Of course, the investigators adjusted for all kinds of factors that may influence this outcome, and showed that there may be a (marginal) reduction in the ICU LOS in association with the switch to private rooms. The authors note that the adjusted average ICU LOS fell by 10%, though under similar circumstances in other similar investigations there is a 95% chance that this would fall somewhere between 0% and 19% reduction. So, under the best of circumstances, if we get a 20% reduction in the 5-day ICU LOS, this translates to about 1 day. And given that transfer timing is more likely to be driven by the availability of ward beds than by the patient's clinical readiness, I question whether this is truly a staggering reduction. Additionally, if you read on, you will realize that there is very little reason to believe that this maximal reduction in ICU LOS is unlikely to be achieved by an average institution. In fact, even the 10% seen on average in this investigation may be a bar that is too high in other less well organized ICUs.  

It is important to remember a couple of things: 1). In some circumstances there is unlikely to be any reduction in the ICU LOS; 2). Since LOS is not a normally distributed function, the mean value underestimates the true measure of central tendency in this outcome (this is due to the typically long right tail present in this distribution); and 3). This investigation, though not strictly speaking experimental, was done at 2 academic institutions with highly organized infrastructure and what looks like closed model ICUs (a dedicated specialized team of critical care professionals caring for all ICU patients). For this reason, a similar intervention at a less stringently streamlined institution is unlikely to produce the same magnitude of results.

But the mere fact that the rates of exogenous transmission of pathogenic organisms were reduced is itself encouraging. At the same time, by focusing on carriage rates and not just clinical infections, the authors may be overstating the clinical significance of the observed reduction. Additionally, one of the issues that does not appear to have been addressed explicitly has to do with the availability of sinks: In the intervention unit there was a plethora of sinks, missing in the pre- period and also not available in the comparator hospital. Is it possible then that simply putting in more sinks would accomplish the same for a lot less money?

And this brings me to my next issue with the paper -- cost effectiveness. Now, according to the AHA annual survey of US hospitals, the average age of the physical plant is on the order of 10 years. Given the rapid pace of change in medicine, this may well signal a time for capital investments in plant improvements. And surely from the patient's and family's perspective, private rooms are preferable. However, one must ask the pesky question of the return on such an investment in this era of much needed fiscal restraint in medicine. If the same outcomes of reducing the spread of infectious organisms can be achieved with merely adding sinks, this may be a less drastic and more immediately feasible intervention well worth considering.                

Monday, 10 January 2011

Reviewing medical literature, part 2b: Study design continued

To synthesize what we have addressed so far with regard to reading medical literature critically:
1. Always identify the question addressed by the study first. The question will inform the study design.
2. Two broad categories of studies are observational and interventional.
3. Some observational designs, such as cross-sectional and ecological, are adequate only for hypothesis generation and NOT for hypothesis testing.
4. Hypothesis testing does not require an interventional study, but can be done in an appropriately designed observational study.

In the last post, where we addressed at length both cross-sectional and ecologic studies, we introduced the following scheme to help us navigate study designs:
Let's now round out our discussion of the observational studies and move on to the interventional ones.

Case-control studies are done when the outcome of interest is rare. These are typically retrospective studies, taking advantage of already existing data. By virtue of this they are quite cost-effective. Cases are defined by the presence of a particular outcome (e.g., bronchiectasis), and controls have to come from a similar underlying population. The exposures (e.g., chronic lung infection) are identified backwards, if you will. In all honesty, case-control studies are very tricky to design well, analyze well and interpret well. Furthermore, it has been my experience that many authors frequently confuse case-control with cohort designs. I cannot tell you how many times as a peer-reviewer I have had to point out to the authors that they have erroneously pegged their study as a case-control when in reality it was a cohort study. And in the interest of full disclosure, once, many years ago, an editor pointed out a similar error to me in one of my papers. The hallmark of case-control is that the selection criteria are the end of the line, or the presence of a particular outcome, and all other data are collected backwards from this point.

Cohort studies, on the other hand, are characterized by defining exposure(s) and examining outcomes occurring after these exposures. Similar to case-control design, retrospective studies are opportunistic in that they look at already collected data (e.g., administrative records, electronic medical records, microbiology data). So, although retrospective here means that we are using data collected in the past, the direction of the events of interest is forward. This is why they are named cohort studies, to evoke a vision of Caesar's army advancing on their enemy.

Some of the well known examples of prospective cohort studies are The Framingham Study, The Nurses Study, and many others. These are bulky and enormously expensive undertakings, going on over decades, addressing myriad hypotheses. But the returns can be pretty impressive -- just look at how much we have learned about coronary disease, its risk factors and modifiers from the Framingham cohort!

Although these observational designs have been used to study therapeutic interventions and their consequences, the HRT story is a vivid illustration of the potential pitfalls of these designs to answer such questions. Case-control and cohort studies are better left for answering questions about such risks as occupational, behavioral and environmental exposures. Caution is to be exercised when testing hypotheses about the outcomes of treatment -- these hypotheses are best generated in observational studies, but tested in interventional trials.

Which brings us to interventional designs, the most commonly encountered of which is a randomized controlled trial (RCT). I do not want to belabor this, as RCT has garnered its (un)fair share of attention. Suffice it to say that matters of efficacy (does a particular intervention work statistically better than the placebo) are best addressed with an RCT. One of the distinct shortcomings of this design is its narrow focus on very controlled events, frequently accompanied by examining surrogate (e.g., blood pressure control) rather than meaningful clinical (e.g., death from stroke) outcomes. This feature makes the results quite dubious when translated to the real world. In fact, it is well appreciated that we are prone to see much less spectacular results in everyday practice. What happens in the real world is termed "effectiveness", and, though ideally also addressed via an RCT, is, pragmatically speaking, less amenable to this design. You may see mention of pragmatic clinical trials of effectiveness, but again they are pragmatic in name only, being impossibly labor- and resource-intensive.

Just a few words about before-and after studies, as this is the design pervasive in quality literature. You may recall the Keystone project in Michigan, which put checklists and Peter Pronovost on the map. The most publicized portion of the project was aimed at eradication of central line-associated blood stream infections (CLABSI) (you will find a detailed description in this reference, Pronovost et al. N Engl J Med 2006;355:2725-32). The exposure was a comprehensive evidence-based intervention bundle geared ultimately at building a "culture of safety" in the ICU. The authors call this a cohort design, but the deliberate nature of the intervention arguably puts it into an interventional trial category. Regardless of what we call it, the "before" refers to measurement of CLABSI rates prior to the intervention, while the "after", of course, is following it. There are many issues with this type of a design, ranging from confounding to Hawthorne effect, and I hope to address these in later posts. For now, just be aware that this is a design that you will encounter a lot if you read quality and safety literature.

I will not say much about the cross-over design, as it is fairly self-explanatory and is relatively infrequently used. Suffice it to say that subjects can serve as their own controls in that they get to experience both the experimental treatment and the comparator in tandem. This is also fraught with many methodologic issues, which we will be touching upon in future posts.

The broad category of "Other" in the above schema is basically a wastebasket for me to put designs that are not amenable to being categorized as observational or interventional. Cost effectiveness studies frequently fall into this category, as do decision and Markov models.

Let's stop here for now. In the next post we will start to address threats to study validity. I welcome your questions and comments -- they will help me to optimize this series' usefulness.                

Friday, 7 January 2011

Reviewing medical literature, part 2a: Study design

It is true that the study question should inform the study design. I am sure you are aware of the broadest categorization of study design -- observational vs. interventional. When I read a study, after identifying the research question I go through a simple 4-step exercise:
1. I look for what the authors say their study design is. This should be pretty easily accessible early in the Methods section of the paper, though that is not always the case. If it is available,
2. I mentally judge whether or not it is feasible to derive an answer to the posed question using the current study design. For example, I spend a lot of time thinking about issues of therapeutic effectiveness and cost-effectiveness, and a randomized controlled trial exploring efficacy of a therapy cannot adequately answer the effectiveness questions.
If the design of the study appears appropriate,
3. I structure my reading of the paper in such a way as to verify that the stated design is, in fact, the actual design. If it is, then I move on to evaluate other components of the paper. If it is not what the authors say,
4. I assign my own understanding to the actual design at hand an go through the same mental list as above with the current understanding in mind.

Here is a scheme that I often use to categorize study designs:
As already mentioned, the first broad division is between observational studies and interventional trials. An anecdote from my course this past semester illustrates that this is not always a straight-forward distinction to make. In my class we were looking at this sub-study of the Women's Health Initiative (WHI), that pesky undertaking that sank the post-menopausal hormone replacement enterprise. The data for the study were derived from the 3 randomized controlled trials (RCT) of HRT, diet and calcium and vitamin D, as well as from the observational component of the WHI. So, is it observational or interventional? The answer to this is confusing to the point of pulling the wool over even experienced clinicians' eyes, as became obvious in my class. To answer the question, we need to go back to definitions of "interventional" and "observational". To qualify as an interventional, a study needs to have the intervention be a deliberate part of the study design. A common example of this type of a study is the randomized controlled trial, the sine qua non of drug evaluation and approval process. Here the drug is administered as a part of the study, not as a background of regular treatment. In contradistinction, an observational study is just that: an opportunistic observation of what is happening to a group of people under ordinary circumstances. Here no specific treatment is predetermined by the study design. Given that the above study looked at multivitamin supplementation as the main exposure, despite its utilization of the data from RCTs, the study was observational. So, the moral of this tale is to be vigilant and examine the design carefully and thoroughly.

We often hear that observational designs are well suited to hypothesis generation only. Well, this is both true and false. Some studies actually can test hypotheses, while others are relegated to generation only. For example, cross-sectional and ecological studies are well suited to generating hypotheses to be tested by another design. To take a recent controversy as an example, the debunked link between vaccinations and autism initially gained steam from the observation that as the vaccination rates were rising, so was the incidence of autism. The type of a study that shows two events changing at the group/population level either in the same or in the opposite direction is called "ecologic". Similar types of studies gave rise to the vitamin D and cancer association hypothesis, showing geographic variation in cancer rates based on the availability of sun exposure. But, as demonstrated well by the vaccine-autism debacle, running with the links from ecological studies is dangerous, as they are prone to a so-called "ecological fallacy". It occurs when, despite the finding in groups of a linked change of the two factors under investigation, there is absolutely no connection between them at the individual level. So, don't let anyone tell you that they tested an hypothesis in an ecological study!

Similarly in cross-sectional studies an hypothesis cannot be tested, and, therefore, causation cannot be "proven". This is due to the fundamental property of "a snapshot in time" that defines a cross sectional study. Since all events (with few minor exceptions) happen at the same time, it is not possible to assign causation to the exposure-outcome couplet. These studies can merely help us think of further questions to test.

So, to connect the design back to the question, if a study purports to "explore a link between exposure X and outcome Y", either an ecologic or a cross-sectional design is OK. On the other hand, if you see one of these designs used to "test the hypothesis that exposure X causes outcome Y", run the other way screaming.

We will stop here for now, and in the next post will continue our discussion of study designs. Not sure yet if we can finish it in one more post, or if it will require multiple postings. Start praying to the goddess of conciseness now!

    

Reviewing medical literature, part 1: The study question

Let's start at the beginning. Why do we do research and write papers? No, not just to get famous, tenured or funded. The fundamental task of science is to answer questions. The big questions of all time get broken down into infinitesimally small chunks that can be answered with experimental or observational scientific methods. These answers integrated together provide the model for life as we understand it.

Clearly, the question is the most important part of the equation, and this is why in my semester-long graduate epidemiology course on the evaluative sciences we spend fully the first four to five weeks talking about how to develop a valid and answerable question. The cornerstone of this validity is its importance. Hence, the first question that we pose is: Is the study question important?

This is a bit of a loaded question, though. Important to whom? How is "important" defined? This is somewhat subjective, yet needs to be scrutinized nevertheless. In the context of an individual patient, the question may become: Is the study question important to me? So, importance is dependent on perspective. Nevertheless, there are questions upon whose importance we can all agree. For example, the importance of the question of whether our current fast-food life style promotes obesity and diabetes is hard to dispute.

Regardless of how we feel about the importance of the question, we must first identify the said research question. At least some of the time you will be able to find it in the primary paper, buried in the last paragraph of the Introduction section. Most of the questions we ask relate to etiologic relationships ("etiology" is medicalese for "causation"). Now, you have heard many times that an observational study cannot answer a causal question. Yet, why do we bother with the time, energy and money needed to run observational studies? Without getting too much into the weeds, philosophers of science tell us that no single study design can give us unequivocal evidence of causality. We can merely come close to it. What does this mean in practical terms? It means that, although most observational studies are still interested in causality rather than a mere association, we have to be more circumspect in how we interpret the results from such studies than from interventional ones. But I am jumping ahead.

Once we have identified and established the importance of the question, we need to evaluate its quality. A question of high quality is 1). clear, 2). specific, and 3). answerable. The question that I posed above regarding fast food and obesity possesses none of these characteristics. It is too broad and open to interpretation. If I were really posing a question in this vein, I would choose a single well defined exposure (consuming 3 cans of soda per day) influencing a single outcome (10% body weight gain) over a specific period of time (over 30 weeks). While this is a much narrower question that the one I proposed above, it is only by answering bundles of such narrow questions and putting the information together that we can arrive at the big picture.

A general principle that I like to teach to my student is the PICO or PECOT model (I did not come up with it, but am its avid user). In PICO, P=population, I=intervention or exposure, C=comparator, and O=outcome. The PECOT model is an adaptation of the PICO for observations over time, resulting in P=population, E=exposure, C=comparator, O=outcome, T=time. These models can help not only pose the question, but to unravel the often mysterious and far from transparent intent of the investigators.

Once you have identified the question and dealt with its importance, you are ready to move on to the next step: evaluating the study design as it relates to the question at hand. We will discuss this in the next post.

Series launch: Critical review of medical literature

Today I am launching a series of posts on how to read medical literature critically. The series should provide a solid foundation for this task and dove-tail nicely with some of the more dense methods themes that occur on this blog. Who should read the series? Everyone. Although the current model of dissemination of medical information relies on a layer of translators (journalists and clinicians), it is my belief that every educated patient must at the very least understand how these interpreters of medical knowledge (should) examine it to arrive at the information imparted to the public. At the same time, both journalists and clinicians may benefit from this refresher. Finally, my own pet project is to get to a better place with peer reviews -- you know how variable the quality of those can be from my previous posts. So, I particularly encourage new peer reviewers for clinical journals to read this series.  

First, a conflict of interest statement. What comes first -- the chicken or the egg? What comes first -- expertise in something or a company hiring you to develop a product? Well, in my case I would like to think that it was the expertise that came first and that Pfizer asked me to develop this content based on what I know, not on the fact that they funded the effort. At any rate, this is my disclaimer: I developed this presentation about three years ago with (modest) funding from Pfizer, and they had it on a web site intended for physician access. Does this mere fact invalidate what I have to say? I don't think so, but you be the judge.

Roughly, the series will examine how to evaluate the following components of any study:
1. Study question
2. Study design
3. Study analyses
4. Study results
5. Study reporting
6. Study conclusions
I am not trying to give you a comprehensive course on how all of this is done, but merely make the reader aware of what entails a critical review of a paper.

Look for the first installment of the series shortly.

Thursday, 6 January 2011

National Healthcare Expenditures, 2009 (In pictures)

Well, it's that time of the year again: CMS has given us the accounting of our National Healthcare Expenditures (NHE) in a paper published in Health Affairs. I am sure you have already heard that the spending only went up by 4% this year over last, an all-time low.

At the same time, we have achieved the highest ever NHE as a proportion of the GDP (17.6%) and as expenditures per capita ($8,086). But the GDP proportion is a somewhat deceptive number on the one hand, as the GDP has suffered a substantial drop from its 2008 value of $14.4 trillion to $14.1 trillion in 2009. On the other hand, this implies that healthcare is eating into the rest of our expenditures on life. At the same time the per capita expenditures have continued their relentless rise.

Let us look at the components of the NHE individually and see what they can tell us.



As usual, the bulk of the expenditures went to personal health care (85%). Public health got a measly 3% of the total NHE, and this continues to be one of our gravest misappropriations. You may recall that about a year ago I did a post where I cited some startling statistics about some broad categories of causes of premature death in the US. Access to medical care accounted for a measly 10% of those, and the rest were attributable to behavior, genetics, environment and social factors. So, while, by inference, fixing medicine may impact 10% of these premature deaths, in reality 97% of the entire NHE goes to medicine rather than to potentially more impactful public health interventions. And the real travesty is that, despite these astronomical expenditures, we are still losing 1,000 lives per day to our broken healthcare system.

Looking a bit more closely at the "personal health" category, we see that, just as in years past, hospital costs and professional services comprise the bulk of this spending.

The "professional services" category, 81% of which is physician and other clinical services, is a bit murky. Yet, without too many leaps of faith we can say that if this expenditure buys us better preventive care, it may be a cost-effective area. At the same time we know that we can make this area a lot more efficient by streamlining and realigning incentives to promote better health rather than more care. Hospital expenditures, on the other hand, are a juggernaut that without a doubt requires containing. It is very likely that exchanging our inflated personal healthcare budgets for well placed public health funding along with reimbursement reform and improved end-of-life decisions, could substantially alter this category of spending.

One final data point that interested me was the breakdown of what are considered investments in the healthcare system. This broadly includes government-funded research and allocations for structures and equipment. Now, I am not sure what "structures and equipment" means, so, if any of my readers know, please, enlighten me. I do know what "research" means, however, and am rather disappointed about this breakdown. What I do not understand is, given that structures and equipment should have some kind of a half-life and not be replaced annually, how it is that this budget also grows consistently year-over-year at a steady rate? Would love to get more details on this.

To be sure, the total research expenditure of $45 billion is nothing to sneeze at. The big question is, however, are we spending it on the right research. I am not at all sure that the answer is yes, given that we still struggle with the same issues at the bedside that we have been struggling with for over a decade. But more on this later.  

Wednesday, 5 January 2011

Radium, dopamine and innovation: Name your poison

Reading Deborah Blum's "The Poisoner's Handbook" is an intellectual treat. Although non-fiction, it paints in understated sepia tones the crevices of New York City at the dawn of the Industrial Revolution, where bootlegged booze and poisons were fare of the day, homicides went unpunished and the corrupt coroner system basked in the glow of its own willful ignorance and political approval. That is until Charles Norris and Alexander Gettler, two single-minded and tireless men, brought science into the lagging American medical jurisprudence and created the now burgeoning field of forensic medicine.

The chapter on radium in particular sparked my interest. Blum describes in vivid detail the well-known misadventure of the "radium girls", a label given to young women in a watch factory in Orange, NJ, in the early 1920s. She sets up the story with the fascinating background of radium discovery by the Curies and Marie Curie's penchant for carrying a "pet" bottle of radium in her skirt pocket, exhibiting its breathtaking beauty in a circus-like fashion (she died a horrible death from aplastic anemia induced by radiation exposure). Once its tumor shrinking properties became known, it did not take long for entrepreneurs, backed by the medical establishment, to create and sell all kinds of tonics and pills containing radium to the clueless public searching for the fountain of youth. The tragic tale of the radium girls, who, because of occupational ingestion of radium used for painting numbers on the faces of watches (according to Blum's account, the girls were encouraged to lick the paint brushes to make them pointy), and playful applications of this glow-in-the-dark paint on their lips and faces, developed debilitating jaw necrosis and other bony complications and early deaths, delivered a dose of sobriety to the public and policy makers about this new health panacea. Even the gifts of radium to Marie Curie were now delivered in a thick lead shield to contain its homicidal particles.

The story of radium raised all sorts of questions for me. When the element was first discovered, even the scientists could not conceive of its deadly health effects on human tissues. And for this reason there was no caution exercised in its use. What I puzzle over, as you may have guessed from many previous posts, is how we can balance our adoption of new glittering technologies, about which we do not have complete information, and keeping a modicum of caution about their currently unknown potentially adverse effects. I particularly wonder about this in the context of how our brains are wired and of our prevailing concerns for the economy even at the expense of our health.

Humans are seekers. I recently read Jonah Lehrer's "How We Decide", and it made me appreciate just how susceptible we are to the pleasurable effects of dopamine, and how craving its effects drives us to perform irrational acts that will soothe our neurons in a bath of dopamine bubbles. Addiction, the ultimate seeking-and-never-finding behavior, is, at least in part, mediated by dopamine. Does this addiction fuel our drive for innovation as well? And does it also make us throw caution to the wind when a desirable new object, like, say, glow-in-the-dark radium or a smart phone, is within grasp?

On the same side of this equation is the corporate voice, thundering in the background about the importance of innovation, injecting doubt about the potential for untoward effects and invoking the reigning rhetoric of Queen Economy as the ultimate justification. You don't believe me? Just look at the tobacco history, rife with denials, manipulation and lies. And this is exactly what our consumer brain wants to hear. So we paint caution as unscientific alarm and walk away from it, shaking our heads, filled with self-righteousness.

This balance that I am describing is once again the baby and bath water problem. We encounter it in every aspect of our modern lifestyle: the environment and the threat of climate change; the healthcare system with its record technology spending without commensurate results in health; our food system and obesity and superbug epidemics; the galloping pace of technological development, far outpacing our cognitive abilities to incorporate these technologies sensibly into our lives. Simply put, the question becomes, how do we harness innovation without demanding corpses (literally and figuratively) as proof of its potential untoward effects?

The first step is clearly understanding our history, and for this read Blum's book -- you won't be sorry! Next we need awareness of how our brains operate and how these biological principles set well known traps in our reasoning. Using metacognition to understand these pitfalls in thinking may at least put us on a smarter course walking this fine line. Finally, as I have advocated before, we need to stop shouting at each other and start listening. Perhaps we are not so drastically different in our views as the press and politicians will have us believe. After all, we are all susceptible to the same poisons. And dopamine.                

Tuesday, 4 January 2011

Guest post: How our brains are wired to advance science

We have a treat today. Today I am featuring a guest post from my brilliant 17-year-old niece Katherine Dana. She is currently applying to colleges, and this is one of her brief essays. Kathy is interested in animal communication specifically, but, as you can see, also spends a lot of time thinking about science in general. And oddly, she seems to be contemplating similar themes to the ones we address here. 
While it is hard for me to stop waxing poetic about how proud I am of her, I will now cut myself short, so that you can enjoy her lucid commentary.

By Katherine E. Dana

Marcel Proust once wrote, "The real voyage of discovery consists not in seeking new landscapes, but in having new eyes." Thus goes the song of science, humanity's great unifier. Science is not merely the means for collecting random information—it is the means through which we make sense of our world. It is messier than mathematics, less exact. And yet in some ways, it is this very inexactitude that gives science its potency, and allows it to cut to the very heart of nature's chaotic randomness. It works by taking the givens of nature and churning out elegant guesses, which predict as effectively as they describe.

One quality that distinguishes mind from machine is that leap of thought that psychologists term "heuristics"—mental shortcuts, expressly designed to help us connect the dots without having to consciously traverse the spaces between. This is our organic advantage.

While today's machines, no matter how complex, are restricted to lengthy algorithms, we may leap from branch to branch. Nowhere in human endeavors is this cognitive edge more apparent than in the combined efforts of humans seeking to find new truth. For before we can know, we must question; and this is where insight is most crucial. It is not enough to investigate the familiar. We must find the courage to ask uncomfortable questions, and be willing to uproot even our most cherished beliefs, all in the name of a deeper understanding.

Monday, 3 January 2011

Gaol fever and intercessory prayer: Redefining the role of p-value?

Happy 2011, everyone! I hope that it is everything you want it to be. Sorry for a brief hiatus in blogging -- needed to recharge my batteries and read others' writing for a change. Well, back now. And thanks to you all for coming back too.

I want to resume our recent discussions of statistical testing in the context of biologic plausibility. We discussed the latter at length a few months ago here, and came to the conclusion that our mere impression of biologic plausibility is not a good litmus test for an association. The oft-cited discovery of H. pylori as the cause of peptic ulcer disease is a tried and true example of the knowledge we would be missing today if we used biologic plausibility as the only yardstick for measuring the prospects of research.

At the same time, we spent a fair bit of time and energy talking about p values and how they need to be used in a Bayesian manner. To review, Bayes theorem relies on pre-test probability of an association to help us understand how much stock we need to put into a finding of an association. That is, the lower the pre-test probability, the more suspicious we should be of an observed association. To put it in concrete terms, for example the finding that intercessory prayer is associated with improved health outcomes requires a much greater amount of scrutiny than one that treating a bacterial infection with an antibiotic improves survival. There is a certain mechanistic elegance to the latter that is missing in the former, unless higher powers are invoked. Here is a quote from the Cochrane meta-analysis of intercessory prayer -- I especially love the last sentence [emphasis mine]:
REVIEWER'S CONCLUSIONS: Data in this review are too inconclusive to guide those wishing to uphold or refute the effect of intercessory prayer on health care outcomes. In the light of the best available data, there are no grounds to change current practices. There are few completed trials of the value of intercessory prayer, and the evidence presented so far is interesting enough to justify further study. If prayer is seen as a human endeavour it may or may not be beneficial, and further trials could uncover this. It could be the case that any effects are due to elements beyond present scientific understanding that will, in time, be understood. If any benefit derives from God's response to prayer it may be beyond any such trials to prove or disprove.
At the same time, just because we do not have a mechanistic explanation at the ready does not mean that we should discount an association. In a rather lengthy post in October I wrote about my own conflicted feelings about applying Bayesian versus frequentist (this refers to all associations standing on similar probabilistic ground prior to testing) thinking in research. Although more Bayesian in my own thinking, I recognize metacognitively that it may at times be a trap:
Yet, there is something to be said about the frequentist approach, even though it is not my way generally. The frequentist approach, which is what underlies the bulk of our traditional clinical research, does not rely on differential prior probabilities for different possible associations, but treats them all equally. Despite many disadvantages, one obvious advantage is that we do not discount potential associations that do not have biologic plausibility, given our current understanding of biology, and sometimes help us stumble on brand new hypotheses. So, clearly, there is a tension here, and I am still working on what is the better way, if any.
The last sentence here implies that there is a right and a wrong way, but having spent the last several months exploring these issues, I am beginning to think that this is incorrect. In fact, all of the p value discussions are leading me to believe that both approaches are useful, and it is the nuances of when either should predominate that need to be worked out.

Consider my examples above -- those of intercessory prayer and antibiotic treatment of a bacterial infection. Let us transport ourselves to, say 18th century England, where typhus, known as "gaol fever", killed more prisoners than the executioners did. How improbable would it have seemed to the medical profession of those days that a). the disease was caused by a microorganism, and b). it could be eradicated with an antibiotic? Why, I would guess that these assertions either would appear heretical or else confirm for the religious the divine presence. Either way, the biology was lacking and the plausibility was simply not there. Yet, this does not change the reality as we understand it today. What explanations will we have 200 years from now for the occasionally observed success of intercessory prayer? And more importantly, what do we do in the meantime to tread most sensibly that purgatory between accepting absurd associations and missing the unlikely ones that are nevertheless real?

The answer may be in the p value after all. Let us model qualitatively what things might look like for intercessory prayer. Let us pretend that we have just conducted the very first randomized controlled trial of the impact of intercessory prayer on the development of post-operative infection following coronary bypass surgery among 1,200 patients. We have found that there is indeed a lowered risk of infection in the intervention group, and the difference has the p value of 0.04. Great, right? We can walk away congratulating ourselves on a positive study. Well, of course this is absurd. Even though we can come up  with some remotely plausible mechanism for this potentially causal association, our pre-test probability is still minuscule. The answer at this point should obviously be what has been suggested for genome-wide interaction studies: a much lower alpha level as the significance threshold. How low? This I cannot answer yet; while the rationale is, similar to genome-wide studies, a fishing expedition without much understanding of why we should find what we should find, here we are not merely engaging in multiple hypotheses testing, the number of which could help determine the appropriate significance level. No, here we are testing a single hypothesis whose mechanism is either absent or highly biologically implausible. So, how to determine the adequate threshold for significance under these circumstances remains unclear to me at this time. I can only say that the traditional 0.05 is highly inappropriate under the circumstances early in the research efforts.

As more studies are performed, their quality and directionality of results should impact how much stock we put in the results. That is, if well done studies consistently continue to demonstrate a positive association of intercessory prayer with clinical outcomes, despite inadequate mechanistic understanding, our level of skepticism should diminish, and commensurately the acceptable alpha can creep higher. In short, the more evidence and the stronger it is, despite poor understanding of why, the more liberal we can afford to be with what we consider a significant result.

So, my point? How we interpret the significance of results needs to be fluid. A p value is not a p value is not a p value. This much embattled and misunderstood statistic may yet be the bridge between Bayesian and frequentist approaches. If we get smarter about setting its thresholds, perhaps we can keep the baby while getting rid of the rancid bath water at the same time. Of course, I am not even attempting to address all of the cognitive biases that derail us in our pursuit of scientific truths. Incorporating them into our inference testing is definitely a discussion for another day.  

Tuesday, 21 December 2010

The changing language of medicine

A very close friend of mine has breast cancer. It is a very small tumor, diagnosed on an annual mammogram, requiring confirmation with a breast MRI. She had a lumpectomy today, and I was with her at the hospital. This proved to be an enlightening experience.

To put things in perspective, when I was in training and in practice (yes, in the dark ages when we were expected to stay awake AND care for patients for over 48 hours at a time every 3 days), we had not heard of patient-centered medicine. I learned that my role was to diagnose, come up with a plan of action and convince the patient at any cost that my plan was the correct one. To be sure, I always tried to do this in a nice way, but would get a bit impatient when my judgment was questioned. This is the behavior modeled for me by my elders and others whom I respected.

Well, that was then. Having had quite a few years to reflect on the practice of medicine in the context of our healthcare system, I have learned just how misguided this attitude is. And, being a Sagittarius, I cannot fathom how this universal truth is escaping others. Yet escaping it is. This became obvious to me today.

My friend had to have a nuclear medicine test prior to her lumpectomy to define the extent of axillary nodal involvement. She had been told that this is an arduous and painful experience that cannot be mitigated with pre-medication. She was also informed that asking the radiologist to deliver the radionuclide slowly rather than as a rapid push might reduce the sensation. So, my friend, who is herself a physician, was prepared for a civilized and simple conversation with the practitioner. Yet, this is not what transpired. You would think that being asked to deliver the chemical slowly is not such a big and unreasonable request. Well, if you thought this, you were wrong: evidently this was such a big ego blow to the radiologist that she felt compelled to respond snidely, "Well, OK, I am not going to fight with you about it". Now, this is off-putting under the best of circumstances. Imagine being about to go to the OR to have a cancer removed from your breast, and having this snide come-back thrown at you. And why? What is the harm in going along with the patient's request if it makes no difference in the end-result of the test? Is it really necessary to diminish her in such a blatant way?

Well, this physician was of a similar vintage to me, and I can only imagine that she came into practice before patient-centered care became the standard. In her mind, as in mine in those distant days, my involvement with the patient's care was not really about the patient necessarily, unless they fell in line with my recommendation. The shameful fact is that my ego was much too fragile to allow a discussion or questions about my considered course of action. How could they go against my years of training, deep knowledge and their best interests? I cannot say for sure, but it is likely that my friend's radiologist was cut from similar cloth. And what is so obvious to me today has not yet been assimilated by so many of my colleagues, including this person.

As I have said before, the new direction for medicine cannot be what I am used to in real estate: "I do not have what you need, but I will show what I do have". The new direction in medicine must undoubtedly be one where the patient is the center of the encounter, and it is the patient's interest rather than the doctor's ego that must be protected assiduously.

Lest you think that the entire hospital experience was negative, let me be clear: of all the people taking care of my friend, the radiologist was the sole disappointing exception. Her surgeons, anesthesiologists, nurses and ancillary personnel went above and beyond my expectations. I was amazed by the level of civility, good humor, politeness and real involvement everyone exhibited -- it was truly different from my days on the wards and pleasantly eye-opening. It even gave me some hope for the future of medicine in the midst of my normally nihilistic ruminations.            

The great poet Rumi said that changing language can change our life. Well, when the recovery room nurse said to my friend "Let me know when you feel that you would rather rest at home than here", I was overcome with warmth and good will. The language of medicine does seem to be changing. And if it continues in this vein, perhaps it will change our lives.

Monday, 20 December 2010

Why we need collaborations across healthcare sectors

I want to digress from our recent focus on methods and talk a bit about conflict of interest (COI for short). There has been a lot in the press lately about doctors taking money from the biopharmaceutical manufacturers, and doctors inserting unnecessary hardware into patients' hearts and spines. All of this has been happening against the background of a low hum of an ongoing discussion of what constitutes a COI, how much is too much and for what (for example, can a doc who takes research and education dollars from a manufacturer with an interest in anticoagulation sit on a committee that develops the guidelines for prevention of thromboembolic disease?), and how to mitigate these ubiquitous and pesky COIs.

In some ways watching this discussion has been amusing, while in others it has been downright sad. Medical journals, while insisting that advertising money is OK to take (presumably because the editorial and marketing offices are separated by some sort of a fire wall), though professional societies should not be able to take this tainted education money. Professional societies, on the other hand, are running away from the accusations by tightening their continuing medical education (CME) criteria and scrambling to replace the lavish budgets derived from pharma to develop their coveted evidence-based practice guidelines. And while all the pots are calling all the kettles black, academic researchers are being barred from collaborating with the industry on research projects, and industry researchers are being precluded from presenting their data at professional society meetings. While all the time the public is being whipped into lather about these alleged systematic transgressions, and forced to cheer for the ensuing retribution.

But, like many things in life, and especially stuff that we discuss on this blog, this issue is neither black nor white. Don't take me wrong: I am not condoning the egregious excesses of greed demonstrated by some members of my hallowed profession. If you have been reading my blog for some time, you know that I do not dispute the shameful reality of many breeches of public trust. I am an ardent supporter of exposing these breeches and of harsh punishments that they deserve. This is not what I am talking about here.

I am much more concerned about the one-sided story that we have been hearing about pharma-academic collaborations. Because of the persecutory nature of public opinion, some institutions are now shying away from such collaborations. This attitude is akin to navigating a treacherous road while looking in the rearview mirror. Yes, there have been transgressions, yes there has been greed and even scientific fraud in the name of money. Does this mean that we need to stop everything and come up with an entirely new way of managing these risks? Absolutely! Does this mean that we have to get rid of all pharma-academic collaborations? Absolutely not! In my humble opinion, erecting non-scaleable walls between these two groups is a big mistake. Here is why.

First, let me make a disclaimer: I do have active ongoing collaborations with multiple manufacturers. I do not take speaking or other promotional money, but limit myself to consulting and research grant funding. I also do a good deal of unfunded research, and I have never taken a penny for any of my blogging or blogging-related activities. And here is the crux of the matter: In this world of über-subspecialization, with the expertise being demographically and geographically diffuse, how can we afford not to collaborate across different types of organizations with different types of capabilities? Can we really afford to leave all of therapeutic development in the hands of organizations whose overarching purpose is to make money? And equally importantly, can we afford to continue this fragmented model of medical development without any thought to integration of the needs of all of the stake holders? I think not. Just as we are reaping the fruit of electronic medical record development in isolation from the end-user, so this isolation of research effort will lead to even less coherence in medicine. And unless we are ready to socialize our entire healthcare system, it seems naïve to expect that this one sector will acquiesce and start working outside of our coveted free market for the good of humankind alone.

My readers know that I am not an industry apologist. On the contrary, I have said many times that there has been bad behavior across all the sectors of healthcare, starting with biopharma. But if we want to advance rather than stagnate and regress, we need robust collaborations. We also need higher ethical standards and greater professionalism to keep public's health as our top priority.

There is COI everywhere, and, while financial COI is most visible, it is the more hidden COI that is most insidious. An hidden COI can be intellectual, reputational, ego-driven, career-mediated, etc. It is incumbent on us all in this complex world to ask questions and mitigate any ill effects of any cognitive biases, including those created by COI. Ultimately, as I have begun to realize of late, nothing will replace an educated and empowered patient: This is the only model that can provide appropriate checks and balances for our oftentimes misaligned and perverse incentives, both academic and economic.

Sunday, 19 December 2010

How e-patients can fix our healthcare system

We got a little into the weeds last week about significance testing and test characteristics. Because information is power, I realized that it may be prudent to back up a bit and do a very explicit primer on medical testing. I am hoping that this will provide some vocabulary for improved patient-clinician communication. But, alas, please do not be surprised if your practitioner looks at you as if you were an alien -- it is safe to say that most clinicians do not think in these terms in their everyday practices. So, educate them!

Let's dig a little deeper into some of the ideas we batted around last week, specifically those pertaining to testing. Let's start by explicitly establishing the purpose of a medical test. The purpose of a medical test is to detect disease when such disease is present. This fact alone should underscore the importance of your physician's ability to arrive at the most likely reasons for your symptoms. This exercise that every doc should go through as he/she is evaluating you is called "differential diagnosis". When I was a practicing MD, my strategy was to come up with 3-5 most likely and 3-5 most deadly if missed potential diagnoses and explore them further with appropriate testing. Arranging these possible diagnoses as a hierarchy can help the clinician to assign informal probabilities to each, a task that is central to Bayesian thinking. From this hierarchy then follows the tactical sequential work-up, avoiding the frantic shotgun approach.

So, having established a hierarchy of diagnoses, we now engage in adjunctive testing. And here is where we really need to be aware not only of our degree of suspicion for each diagnosis, but also the test characteristics as they are reported in the literature and the test characteristics as they exist in the local center where the testing takes place. Why do I differentiate between the literature and practice? We know very well that the mere fact of observation, not to mention experimental cleanliness of trials, often tends to exaggerate the benefits of an intervention. In other words, real world is much messier than the laboratory of clinical research (which of course itself is messy enough). So, it is this compounded messiness that each clinician has to contend with when making testing decisions.

OK, so let us now deconstruct test characteristics even further. We have used the terms sensitivity, specificity, positive and negative predictive values. We've even explored their meanings to an extent. But let's break them down a bit further. Epidemiologists find it helpful to construct 2-by-2 (or 2 x 2) tables to think through some of these constructs, so, let's engage in that briefly. Below you see a typical 2 x 2 table.
In it we traditionally situate disease information in columns and test information in rows. A good test picks up signal when the signal is there while adding minimal noise. The signal is the disease, while the noise is the imprecise nature of all tests. Even simple blood tests, whose "objective accuracy" we take for granted, are subject to these limitations.

It is easiest to think of sensitivity as how well the test picks up the corresponding disease. In the case of mammography from last week, this number is 80%. This means that among 100 women who actually harbor breast cancer a mammogram will recognize 80. This is the "true positive" value, or disease actually present when the test is positive. What about the remaining 20? Well, those real cancers will be missed by this test, and we call them a "false negative". If you look at the 2 x 2 table, it should become obvious that the sum of the true positives and the false negatives adds up to the total number of people with the disease. Are you shocked that our wonderful tests may miss so much disease? Well, stay tuned.

The flip side of sensitivity is "specificity". Specificity refers to whether or not the test is identifying what we think it is identifying. The noise in this value comes from the test in effect hallucinating disease when the person does not have the disease. A test with high specificity will be negative in the overwhelming proportion of people without the disease, so the "true negative" cell of the table will contain almost the entire group of people without disease. Alas, for any test we develop we walk the tight-rope between sensitivity and specificity. That is, depending on our priorities for testing, we have to give up some accuracy in either the sensitivity or the specificity. The more sensitive the test, the higher our confidence that we will not miss the disease when it is there. Unfortunately, what we gain in sensitivity we usually lose in specificity, thus creating higher odds for a host of false positive results. So, there really is no free lunch when it comes to testing. In fact, it is this very tension between sensitivity and specificity that is the crux of the mammography debate. Not as straight-forward as we had thought, right? And this is not even getting into pre-test probabilities or positive and negative predictive values!

Well, let's get into these ideas now. I believe that the positive and negative predictive values of the test are fairly well understood at this point, no? Just to reiterate, a positive predictive value, which is the ratio of true positives to all positive test results (the latter is the sum of the true and false positives, or the sum of the values across the top row of the 2 x 2 table), tells us how confident we can be that a positive test result corresponds to disease being present. Similarly, the negative predictive value, the ratio of true negative test results to all negative test results (again, the latter being the sum across the second row of the 2 x 2 table, or true and false negatives), tells us how confident we can be that a negative test result really represents the absence of disease. The higher the positive and negative predictive values, the more useful the test becomes. However, when one is likely to be quite high but the other quite low, it is a pitfall of our irrationality to rush head first into the test in hopes of obtaining the answer with a high value (as in the case of the negative predictive value for mammography in women aged 40-50 years), since the opposite test result creates a potentially difficult conundrum. This is where pre-test probability of disease comes in.

Now, what is this pre-test probability and how do we calculate it? Ah, this is the pivotal question. The pre-test probability is estimated based on population epidemiology data. In other words, given the type of a person you are (no, I do not mean nice or nasty or funny or droll) in terms of your demographics, heredity, chronic disease burden and current symptoms, what category of risk you fit into based on these population studies of disease. This approach relies on filing you into a particular cubby hole with other subjects whose characteristics are most similar to yours. Are you beginning to appreciate the complexity of this task? And the imprecision of it? Add to this barely functioning crystal ball the clinician's personal cognitive biases, and is it any wonder that we do not do better? And need I even overlay this with another bugaboo, that of the overwhelming amount of information in the face of the incredible shrinking appointment, to demonstrate to you just how NOT straightforward any of this medicine stuff is?

OK, get your fingers out of that Prozac bottle -- it is not all bad! Yes, these are significant barriers to good healthcare. But guess what? The mere fact that you now know these challenges and can call them by their appropriate names gives you more power to be your own steward of your healthcare. Next time a doc appears certain and recommends some sexy new test, you will know that you cannot just say OK and await further results. Your healthcare is a chess match: you and your healthcare provider need to plan 10 moves ahead and play out many different contingencies.

On our end, researchers, policy makers and software developers all need to do better developing more useful and individualizable information, integrating this information into user-friendly systems, and encouraging thoughtful healthcare encounters. I am convinced that patient empowerment with this information followed by teaming up with providers in advocacy in this vein is the only thing that can assure course correction for our mammoth, unruly, dangerous and irrational healthcare system.
           

Top 5 this week

Thursday, 16 December 2010

P-values, Bayes and Ioannidis, oh my!

When I graduated from college in the early 1980s, much to my parents' chagrin, I was not sure what to do with my career. So, instead of following some of my more wizened classmates to Wall Street, I got a job in a very well regarded molecular endocrinology laboratory in Boston. When I look back on that time, almost 30 years ago, so much seems surreal. For example, in those days it took us well over a year to sequence a gene! How about that? Something that takes hours with today's technology took an equivalent of eternity. And what a pain it was, running those cumbersome sequencing gels, with the glass cracking in the night, negating days of work. Oh, well, times have changed and for the better, I believe. But these changes are bringing with them some odd incongruities in our prior thinking.

A few days ago, I picked up a link from Michael Pollan's twitter feed to a story that startled me: It was insinuating that all these decades of sequencing human genome to hunt for targets of disease susceptibility have come up nearly empty-handed (with a few well known exceptions). The story recounted a tale of decades of vigorous funding, meteoric career growth and very few results to date. Yet, the scientists involved are not ready to give up on human genes as the primary locus for disease susceptibility. In fact, the strong rescue bias and the fear of losing all that beautiful funding are colluding to generate some creative and complex hypotheses. I came upon one such hypothesis earlier today, following a link from Ed Yong, that British science journalist extraordinaire. But the hypothesis itself, though fascinating, is not what interested me most, no. Care to guess what did? You are correct if you said the p-value.

What grabbed me is the calculation that to identify interactions of significance, the adjusted alpha level has to be set at 10^ -12; that's 0.000000000001! Why is this, and what does it mean in the context of our recent ruminations on the topic of p-value? Well, oddly enough it ties in very nicely with Bayesian thinking and John Ioannides' incendiary assertion of lies in research. How? Follow me.

The hallmark of identifying "significant" relationships (remember that the word "significant" merely means "worthy of noting") in genetic research is searching for statistical associations between the exposure (in this case a mutation in the genetic locus) and the outcome (the specific disease in question). When we start this analysis, we have absolutely no idea which gene(s) mutation(s) is(are) likely to be at play. This means that we have no way of establishing... can you guess what? Yes, that's right, the pre-test probability. We know nothing about the pretest probability. This shotgunning screening of thousands of genes is in essence a random search for matching pairs of mutation and disease. So, why should this matter?

The reason that it matters resides in the simple, albeit overused, coin toss analogy. When you toss a coin, what are your chances of getting heads on any given toss? The odds are 1:1 every time whether you will get heads or tails, meaning that the chance of getting heads is 50% (unless the coin is rigged, in which case we are talking about a Bayesian coin, not the focus of what we are getting at here). So, if you call heads prior to any given attempt, you will be correct half the time. Now, what does this have to do with the question at hand? Well, let's go back to our definition of the p-value: a p-value represents the probability of obtaining the result (in our gene case it is the association with disease) of the magnitude observed or greater under the conditions that no real association exists. So, our customary p-value threshold of 0.05 can be verbalized as follows: "There is a 5% probability that the association of the magnitude observed or greater could happen by chance if there is in reality no association". Now, what happens if we test 20 associations? In other words, what are the chances that one of them will come up "significant" at the p=0.05 level? Yes, there is a 1 in 20 (or 5 in 100) chance that this association will hit our level of significance preset at 0.05. And it follows that the chances of getting a "significant" result grow with more testing.

This is the very reason that gene scientists have asked (and now answered) the question of what level of statistical significance (also called alpha, signified by the p-value) is acceptable when exploring hundreds of thousands of potential associations without any prior clue of what to expect. And that level is a shocking 0.000000000001!

This reinforces a couple of ideas that we have been discussing of late. First, to say simply that the p-value of 0.05 was reached is to say nothing. The p-value needs to be put into the context of a). prior probability of association and b). the number of tests of association performed. Second, as a corollary, the p-value threshold needs to be set according to these two parameters: the lower the prior probability and/or the higher the number of tests, the lower the p-value needs to be in order to be noteworthy. Third, if we pay attention only to the p-value, and particularly if we fail to understand how to set the threshold for significance, what we get is junk, and, as Ioannidis aptly points out, lies. Fourth, and final, genetic data are tougher to interpret that most of us appreciate. It is possibly even less precise than clinical data we usually discuss on this blog. And we know how uncertain that can be!

So, as sexy as the emerging field of genetics was in the 1980s, I am pretty happy that, after four years in the lab, I decided to go first to medical and then to epidemiology schools. Dealing with conditions where we at least have some notion of the pre-test probabilities makes this quantitative window through which I see healthcare just a little less opaque. And these days, I am even happier about having bucked my classmates' Wall Street trend. But that's a story for another day.