Sunday, 13 February 2011
Friday, 11 February 2011
CMS never events: Evidence of smoke in mirrors?
Let me tell you a fascinating story. In 1999, I was still fresh out of my Pulmonary and Critical Care Fellowship, struggling for breath in the vortex of private practice, when a cute little paper appeared in the Lancet from a great group of researchers in Spain, describing a study performed in one large academic urban medical center's two ICUs: one respiratory and one medical. Its modest aim was to see if semi-recumbent (partly sitting up) compared to supine (lying flat on the back) positioning could reduce the incidence of that bane of the ICU, ventilator-associated pneumonia (VAP). The study was a well done randomized controlled trial, and the investigators even went so far as to calculate the power (the number needed to enroll in order to detect a pre-determined magnitude of effect [in this case an ambitious 50% reduction in clinically suspected VAP]), and this number was 182 based on the assumption of a 40% VAP prevalence in the control (supine) group. The primary endpoint was the prevalence (percentage of all mechanically ventilated [MV] patients developing) and the secondary the incidence density (number of cases among all MV patients spread over all the cumulative days of MV [patient-days of MV]) of clinically suspected VAP, based on the CDC criteria, while microbiologically confirmed VAP (also rigorously defined) served as the secondary endpoint.
Here is what they found. The study was stopped early due to efficacy (this means that the intervention was so superior to the control in reaching the endpoint that it was deemed unethical after the interim look to continue the study), enrolling only 86 patients, 39 in the intervention and 47 in the control groups. And here are the results for the primary and secondary outcomes:
So, this is great! No matter how you slice it, VAP is reduced substantially; there is a microbiologically confirmed prevalence reduction of nearly 6-fold (this is unadjusted for potential differences between groups; and there were differences!). Well, you know what's coming next. That's right, the "not so fast" warning. Let's examine the numbers in context.
First of all, if we look at the evidence-based guideline on HCAP, HAP and VAP from the ATS and IDSA, the prevalence of VAP is generally between 5 and 15%; in the current study the control group exceeds 20%. Now, for the incidence density, for years now the CDC has been keeping and reporting these numbers in the US, and the rate in patients comparable to the ones in the study should be around 2-4 cases per 1,000 MV days. In this study, no matter how you slice it, clinically or microbiologically, the incidence density is exceedingly high, more in line with some of the ex-US numbers reported in other studies. So, they started high and ended high, albeit with a substantial reduction.
Second of all, there is a wonderful flow chart in the paper that shows the enrollment algorithm. One small detail has always been somewhat obscure to me: the 4 patients in the semi-recumbent group that were excluded from analysis due to reintubation (this means that they were taken off MV, but had to go back on it within a day or two), which was deemed a protocol violation. Now, you might think that 4 patients is a pretty small number to worry about. But look at the total number of patients in the group: 39. If the excluded 4 all had microbiologically confirmed VAP, that would bring our prevalence from 5% to 14% (6 out of 43). This would certainly be a less than 6-fold reduction in VAP.
Thirdly, and this I think is critical, the study was not blinded. In other words, the people who took care of the patients knew the group assignment. So what, you ask. Well remember that VAP is a pretty difficult, elusive and unclear diagnosis. So, let us pretend that I am a doc who is also an investigator on the study, and I am really invested in showing how marvelous semi-recumbent positioning is for VAP prevention. I am likely to have a much lower threshold for suspecting and then diagnosing VAP in the comparator group than in my pet intervention group. And this is not an indictment of anyone's judgment or integrity; it is just how our brains are wired.
Next, there were indeed important differences between groups in their baseline risk factors for VAP. For example, more patients in the control (38%) than in the intervention (26%) group were on MV for a week or longer, the single most important risk factor for developing VAP. Likewise, the baseline severity of illness was higher in the control than the intervention group. To be sure, the authors did statistical analyses to adjust these differences away, and still found an adjusted odds ratio of VAP among the supine group to be 6.8, with the 95% confidence interval between 1.7 and 26.7. This is generally taken to mean that, on average, the risk of VAP increases nearly 7-fold for supine position as opposed to semi-recumbent, and if the trial was repeated 100 times, 95 of those times this estimate would fall between a 1.7 and a 26.7-fold increase. OK, so we can accept this as a possible viable strategy, right?
But wait, there is more. Remember what we said about the odds ratio? When the event happens in more than 10% of the sample, the odds ratio vastly overestimates the risk of this event. 28.4% anyone?
Now, let's put it all together. A single center study from a Spanish academic hospital, among respiratory and medical ICU patients, with a minuscule sample size, yet halted early for efficacy, an exceedingly high baseline rate of VAP, a substantial number of patients excluded for a nebulous reason, unblinded and therefore prone to biased diagnosis, reporting an inflated reduction in VAP development in the intervention group. It would be very easy to write this off as a flawed study (like all studies tend to be in one way or another) in need of confirmatory evidence, if it were not so critical in the current punitive environment of quality improvement. (By the way, to the best of my knowledge, there is no study that replicates these results). The ATS/IDSA guideline includes semi-recumbent positioning as a level I (highest possible level of evidence) recommendation for VAP prevention, and it is one of the elements of the MV bundle, as promoted by the Institute for Healthcare Improvement, which demands 95% compliance with all 5 elements of the bundle in order to get the "compliant" designation. And even this is not the crux of the matter. The diabolical detail here is that CMS is creeping up on making VAP into one of their magical "never" events, and the efforts by hospitals will most assuredly be including this intervention. So, ICU nurses are already expected to fall in step with this deceptively simple yet not-so-easily executable practice.
And this is what is under the hood of just one simple level I recommendation by two reputable professional organizations in their evidence-based guidelines. One shudders to think...
Here is what they found. The study was stopped early due to efficacy (this means that the intervention was so superior to the control in reaching the endpoint that it was deemed unethical after the interim look to continue the study), enrolling only 86 patients, 39 in the intervention and 47 in the control groups. And here are the results for the primary and secondary outcomes:
So, this is great! No matter how you slice it, VAP is reduced substantially; there is a microbiologically confirmed prevalence reduction of nearly 6-fold (this is unadjusted for potential differences between groups; and there were differences!). Well, you know what's coming next. That's right, the "not so fast" warning. Let's examine the numbers in context.
First of all, if we look at the evidence-based guideline on HCAP, HAP and VAP from the ATS and IDSA, the prevalence of VAP is generally between 5 and 15%; in the current study the control group exceeds 20%. Now, for the incidence density, for years now the CDC has been keeping and reporting these numbers in the US, and the rate in patients comparable to the ones in the study should be around 2-4 cases per 1,000 MV days. In this study, no matter how you slice it, clinically or microbiologically, the incidence density is exceedingly high, more in line with some of the ex-US numbers reported in other studies. So, they started high and ended high, albeit with a substantial reduction.
Second of all, there is a wonderful flow chart in the paper that shows the enrollment algorithm. One small detail has always been somewhat obscure to me: the 4 patients in the semi-recumbent group that were excluded from analysis due to reintubation (this means that they were taken off MV, but had to go back on it within a day or two), which was deemed a protocol violation. Now, you might think that 4 patients is a pretty small number to worry about. But look at the total number of patients in the group: 39. If the excluded 4 all had microbiologically confirmed VAP, that would bring our prevalence from 5% to 14% (6 out of 43). This would certainly be a less than 6-fold reduction in VAP.
Thirdly, and this I think is critical, the study was not blinded. In other words, the people who took care of the patients knew the group assignment. So what, you ask. Well remember that VAP is a pretty difficult, elusive and unclear diagnosis. So, let us pretend that I am a doc who is also an investigator on the study, and I am really invested in showing how marvelous semi-recumbent positioning is for VAP prevention. I am likely to have a much lower threshold for suspecting and then diagnosing VAP in the comparator group than in my pet intervention group. And this is not an indictment of anyone's judgment or integrity; it is just how our brains are wired.
Next, there were indeed important differences between groups in their baseline risk factors for VAP. For example, more patients in the control (38%) than in the intervention (26%) group were on MV for a week or longer, the single most important risk factor for developing VAP. Likewise, the baseline severity of illness was higher in the control than the intervention group. To be sure, the authors did statistical analyses to adjust these differences away, and still found an adjusted odds ratio of VAP among the supine group to be 6.8, with the 95% confidence interval between 1.7 and 26.7. This is generally taken to mean that, on average, the risk of VAP increases nearly 7-fold for supine position as opposed to semi-recumbent, and if the trial was repeated 100 times, 95 of those times this estimate would fall between a 1.7 and a 26.7-fold increase. OK, so we can accept this as a possible viable strategy, right?
But wait, there is more. Remember what we said about the odds ratio? When the event happens in more than 10% of the sample, the odds ratio vastly overestimates the risk of this event. 28.4% anyone?
Now, let's put it all together. A single center study from a Spanish academic hospital, among respiratory and medical ICU patients, with a minuscule sample size, yet halted early for efficacy, an exceedingly high baseline rate of VAP, a substantial number of patients excluded for a nebulous reason, unblinded and therefore prone to biased diagnosis, reporting an inflated reduction in VAP development in the intervention group. It would be very easy to write this off as a flawed study (like all studies tend to be in one way or another) in need of confirmatory evidence, if it were not so critical in the current punitive environment of quality improvement. (By the way, to the best of my knowledge, there is no study that replicates these results). The ATS/IDSA guideline includes semi-recumbent positioning as a level I (highest possible level of evidence) recommendation for VAP prevention, and it is one of the elements of the MV bundle, as promoted by the Institute for Healthcare Improvement, which demands 95% compliance with all 5 elements of the bundle in order to get the "compliant" designation. And even this is not the crux of the matter. The diabolical detail here is that CMS is creeping up on making VAP into one of their magical "never" events, and the efforts by hospitals will most assuredly be including this intervention. So, ICU nurses are already expected to fall in step with this deceptively simple yet not-so-easily executable practice.
And this is what is under the hood of just one simple level I recommendation by two reputable professional organizations in their evidence-based guidelines. One shudders to think...
Wednesday, 9 February 2011
Evidence and profit: An unhealthy alliance
My JAMA Commentary came out this week, and I am getting e-mail about it. It seems to have resonated with many docs who feel that the research enterprise is broken and its output fails them at the office. But what I want to do is tie a few ideas together, ideas that I have been exploring on this blog and elsewhere, ideas that may hold the key to our devastating healthcare safety problem.
The last four decades can be viewed as a nexus between the growth of evidence-based medicine (EBM) on the one hand, and the unbridled proliferation of the biopharmaceutical industry and its technologies. The result has been rapid development, maximization of profit, and a juggernaut of poorly thought-out and completely uncoordinated research geared initially at regulatory approval and subsequently to market growth. It is not that the clinical research has been of poor quality, no. It is that our research tools are primitive and allow us to see only slivers of reality. And these slivers are prone to many of our cognitive biases to boot. So, the drive to produce evidence and the drive to grow business colluded to bring us to where we are today: inundated with evidence of unclear validity, unbalanced with regard to where the biggest difference to public health can be made. Yet we are constantly poked and prodded by the eager bureaucracy to do better at implementing this evidence, while the system continues to perform in a devastatingly suboptimal fashion, causing more deaths every year than strokes.
A byproduct of this technological and financial race has been the rapid escalation of healthcare spending, with the consequent drive to contain it. The containment measures have, of course, had the "unintended consequence" of increased patient volume for providers and of the incredible shrinking appointment, all just to make a living. The end-result for clinicians and patients is the relentless pressure of time and the straight jacket of "evidence-based" interventions in the name of quality improvement. And in this mad race against the clock and demoralization, very few have had the opportunity to think rationally and holistically about the root causes of our status quo. The reality is that we are now madly spinning our wheels at the margins, getting bogged down in infinitesimal details and losing the forest for the trees (pardon all of the metaphor mixing). Our evidence-based quality improvement efforts, while commendable, are like trying to plug holes in a ship's hull with bandainds: costly and overall making little if any difference.
But if we step back and stop squinting, we can see the big picture: stagnated and outdated research enterprise still rewarding spending over substance, embattled clinicians trying to stay afloat, and a $2.5 trillion healthcare gorilla feeding the economy at the expense of human lives. Will technology fix this mess? Not by itself, no. Will more "evidence" be the answer? No, not if we continue to generate it as usual. Is throwing more money at the HHS the solution? I doubt it. A radical change of course is in order. Take profit out of evidence generation, or at least blunt its influence (this will reduce the clutter of marginal, hair-splitting technologies occupying clinicians' collective consciousness), develop new tools for better patient care rather than for maximizing the bottom line, give clinicians more time to think about their patients' needs rather than about how to maintain enough income to pay for the overhead, these are some of the obvious yet challenging solutions to the current crisis. Challenging because there needs to be political will to implement them. And because we are currently so invested in the path we are on that it is difficult and perhaps impossible to stray without losing face. But what is the alternative?
The last four decades can be viewed as a nexus between the growth of evidence-based medicine (EBM) on the one hand, and the unbridled proliferation of the biopharmaceutical industry and its technologies. The result has been rapid development, maximization of profit, and a juggernaut of poorly thought-out and completely uncoordinated research geared initially at regulatory approval and subsequently to market growth. It is not that the clinical research has been of poor quality, no. It is that our research tools are primitive and allow us to see only slivers of reality. And these slivers are prone to many of our cognitive biases to boot. So, the drive to produce evidence and the drive to grow business colluded to bring us to where we are today: inundated with evidence of unclear validity, unbalanced with regard to where the biggest difference to public health can be made. Yet we are constantly poked and prodded by the eager bureaucracy to do better at implementing this evidence, while the system continues to perform in a devastatingly suboptimal fashion, causing more deaths every year than strokes.
A byproduct of this technological and financial race has been the rapid escalation of healthcare spending, with the consequent drive to contain it. The containment measures have, of course, had the "unintended consequence" of increased patient volume for providers and of the incredible shrinking appointment, all just to make a living. The end-result for clinicians and patients is the relentless pressure of time and the straight jacket of "evidence-based" interventions in the name of quality improvement. And in this mad race against the clock and demoralization, very few have had the opportunity to think rationally and holistically about the root causes of our status quo. The reality is that we are now madly spinning our wheels at the margins, getting bogged down in infinitesimal details and losing the forest for the trees (pardon all of the metaphor mixing). Our evidence-based quality improvement efforts, while commendable, are like trying to plug holes in a ship's hull with bandainds: costly and overall making little if any difference.
But if we step back and stop squinting, we can see the big picture: stagnated and outdated research enterprise still rewarding spending over substance, embattled clinicians trying to stay afloat, and a $2.5 trillion healthcare gorilla feeding the economy at the expense of human lives. Will technology fix this mess? Not by itself, no. Will more "evidence" be the answer? No, not if we continue to generate it as usual. Is throwing more money at the HHS the solution? I doubt it. A radical change of course is in order. Take profit out of evidence generation, or at least blunt its influence (this will reduce the clutter of marginal, hair-splitting technologies occupying clinicians' collective consciousness), develop new tools for better patient care rather than for maximizing the bottom line, give clinicians more time to think about their patients' needs rather than about how to maintain enough income to pay for the overhead, these are some of the obvious yet challenging solutions to the current crisis. Challenging because there needs to be political will to implement them. And because we are currently so invested in the path we are on that it is difficult and perhaps impossible to stray without losing face. But what is the alternative?
Tuesday, 8 February 2011
Medical decision making: More signal less noise, please!
February 08, 2011
decision support, EMR, false positive, HIT, medical decision making, methods
No comments
It's official, I'm a country bumpkin! Driving in Boston last week I was distracted, annoyed, made anxious and confused by the constant traffic, billboards and signs. Even highway markings confused me, particularly one indicating a detour to Storrow Drive East, which never materialized. Despite the fact that I know the geography of Boston like the back of my hand, I nearly went down the wrong streets multiple times, including driving the wrong way on some one-way roads. Yes, I am now the menace I used to save my prize driving language for in my younger days.
But it seems that over the years of my living away, there has been a sharp increase in the information thrown at me from all directions, accompanied by a decline in places to rest my gaze without suffering the perseveration of conscious processing. And while the value of this information is at best questionable, the sum total of this overstimulation is clearly confusion, wrong road choices and possibly a reduction in the safety of my driving. This whole experience reminded me of Thomas Goetz's distaste for how medical results are reported. If you have not seen him preach about it, you really should. Here is his excellent TED talk on the subject.
It is ironic that during this overwhelming city visit I also had the chance to speak to a doctor about "routine" preoperative testing and its value. Before surgery, it is recommended that a patient get a screening evaluation. Yet the components of this evaluation vary widely, and may include blood work, urinalysis, electrocardiogram, a chest X-ray and the like. Although evidence suggests that most of the points of this evaluation are useless at best, many institutions continue to order a shotgun panel of preoperative testing for everyone. This one-size-fit-all medicine results in reams of useless and distracting information, a high frequency of abnormal findings of questionable significance, a potential for harm, worry and needless healthcare spending. In my particular conversation I asked the anesthesiologist what the pre-test probability for someone with my characteristics was for a useful chest X-ray result, for example, and whether the fancy electronic medical record used by the hospital could help her determine this. While the answer to the former question was "probably exceedingly low", the answer to the latter was a definitive "no." So, given some elementary thinking, it became clear that a patient like me should not in fact be subjected to a chest X-ray, since any pathology found on one would likely represent a false positive finding, which would nevertheless require potentially invasive follow-up. And guess what? By focusing on the particular individual in the office, rather than all comers, we could have gone through the entire menu of the possible preoperative tests "routinely" ordered and eliminated most if not all of them. But my bet is that not all patients, not even all e-patients, either know or are able to initiate this type of a critical discussion. And yet what tests to obtain, if any, should always be a thoughtful and individualized decision. To approach testing in any other way is to risk generating noise, distraction and harm.
And this brings me back to Thomas Goetz's idea of redesigning how test results are reported. I love his idea. But to me what needs to happen before making the data patient-friendly, is making the decision-making provider-friendly. So, great idea, Mr. Goetz, but let us move it upstream, to the office, where the decision to get chest X-rays, cholesterols and urinalyses is made, and help the doctor visualize her patient's risk for a disease being present, the characteristics of the test about to be ordered, the probability of a positive test result, and all the downstream probabilities that stem from this testing, so as to put a positive test result in the context of the individual's risk for having the disease. Because getting the results of tests that perhaps should never have been obtained in the first place is following the GIGO principle. It is generating noise, distraction and detours going wrong way down one-way roads. And when applied to medicine, these are definitely unwelcome metaphors.
But it seems that over the years of my living away, there has been a sharp increase in the information thrown at me from all directions, accompanied by a decline in places to rest my gaze without suffering the perseveration of conscious processing. And while the value of this information is at best questionable, the sum total of this overstimulation is clearly confusion, wrong road choices and possibly a reduction in the safety of my driving. This whole experience reminded me of Thomas Goetz's distaste for how medical results are reported. If you have not seen him preach about it, you really should. Here is his excellent TED talk on the subject.
It is ironic that during this overwhelming city visit I also had the chance to speak to a doctor about "routine" preoperative testing and its value. Before surgery, it is recommended that a patient get a screening evaluation. Yet the components of this evaluation vary widely, and may include blood work, urinalysis, electrocardiogram, a chest X-ray and the like. Although evidence suggests that most of the points of this evaluation are useless at best, many institutions continue to order a shotgun panel of preoperative testing for everyone. This one-size-fit-all medicine results in reams of useless and distracting information, a high frequency of abnormal findings of questionable significance, a potential for harm, worry and needless healthcare spending. In my particular conversation I asked the anesthesiologist what the pre-test probability for someone with my characteristics was for a useful chest X-ray result, for example, and whether the fancy electronic medical record used by the hospital could help her determine this. While the answer to the former question was "probably exceedingly low", the answer to the latter was a definitive "no." So, given some elementary thinking, it became clear that a patient like me should not in fact be subjected to a chest X-ray, since any pathology found on one would likely represent a false positive finding, which would nevertheless require potentially invasive follow-up. And guess what? By focusing on the particular individual in the office, rather than all comers, we could have gone through the entire menu of the possible preoperative tests "routinely" ordered and eliminated most if not all of them. But my bet is that not all patients, not even all e-patients, either know or are able to initiate this type of a critical discussion. And yet what tests to obtain, if any, should always be a thoughtful and individualized decision. To approach testing in any other way is to risk generating noise, distraction and harm.
And this brings me back to Thomas Goetz's idea of redesigning how test results are reported. I love his idea. But to me what needs to happen before making the data patient-friendly, is making the decision-making provider-friendly. So, great idea, Mr. Goetz, but let us move it upstream, to the office, where the decision to get chest X-rays, cholesterols and urinalyses is made, and help the doctor visualize her patient's risk for a disease being present, the characteristics of the test about to be ordered, the probability of a positive test result, and all the downstream probabilities that stem from this testing, so as to put a positive test result in the context of the individual's risk for having the disease. Because getting the results of tests that perhaps should never have been obtained in the first place is following the GIGO principle. It is generating noise, distraction and detours going wrong way down one-way roads. And when applied to medicine, these are definitely unwelcome metaphors.
Wednesday, 2 February 2011
Intervention in ICU reduces hospital mortality, but by how much?
Addendum #2, 12:09 PM EST, 2/2/11:
So, here is the whole story. Stephanie Desmon, the author of the JH press release, e-mailed me back and pointed me to Peter Pronovost as the source for the 10% reduction information. I e-mailed Peter, and he got back to me, confirming that
And speaking of details, I must admit to an error of my own. If you look at the figure reproduced below, I called out the wrong points. For adjusted data, you need to look at the open circles (for the intervention group) and squares (for the control group). In fact, the adjusted mortality went from about 20% at baseline to 16% in the 13-22 months interval for the Keystone cohort, while for the control group it went from a little over 20% to a little under 18%. This makes the absolute reduction a tad more impressive, though there is still less than a 2% absolute difference between the reduction seen in the intervention vs. the control group, leaving all of my other points still in need of addressing.
Addendum #1, 11:00 AM EST, 2/2/11:
I just found what I think is the origin of the 10% mortality reduction rumor in this press release from Johns Hopkins. I just e-mailed Stephanie Desmon, the author of the release, to see where the 10% came from. Will update again should I hear from either Maggie Fox or Stephanie Desmon.
Remember the Keystone project? A number of years ago when we started to pay close attention to healthcare-associated infections (HAI), and hospitals started to take introspective looks at their records, it turned out the the ICUs in the state of Michigan for one reason or another had very high rates of HAIs. As this information percolated through our collective consciousness, the stars aligned in such a away as to release funding from the AHRQ in Washington, DC, for a group of ICU investigators at the Johns Hopkins University School of Medicine in Baltimore, MD, headed by Peter Pronovost, to design and implement a study employing IHI-style (Boston, MA) bundled interventions to prevent catheter-associated blood stream infections (CABSI) and ventilator-associated pneumonia (VAP) across the consortium of ICUs in MI. Whew! This poly-geographic collaboration resulted in a landmark paper in 2006 in the New England Journal of Medicine, wherein the authors showed that the bundled interventions directed by a checklist aimed at CABSI were indeed associated with a satisfying reduction of CABSI. Since 2006 the ICU community has been eagerly awaiting the results of the VAP intervention from Keystone, but none has come out. When there is a void of information, rumors fill this void, and plenty of rumors have circulated about the alleged failure of the VAP trial.
I do not want to belabor here what I have written before with regard to VAP and its prevention, and what makes the latter so difficult, and how little evidence there really is that the IHI bundle actually does anything. You can find at least some of my thoughts on that here. But why am I bringing up the Keystone project again anyway? Well, it is because Pronovost's group has just published a new paper in BMJ, and this time their aim was even more ambitious: to show the impact of this state-wide QI intervention on hospital mortality and length of stay. This is a really reasonable question, mind you, since, we could argue that, if the intervention reduces HAI, it should also do something to those important downstream events that are driven by the particular HAI, namely mortality and LOS. But here are a couple of issues that I found of great interest.
First, as we have discussed before, whether or not VAP itself causes death in the ICU population (that is patients die from VAP), or whether VAP tends to attack those who are sicker and therefore more likely to die anyway (patients die with VAP) remains unclear in our literature. There is some evidence that late VAP may be associated with an attributable increase in mortality, but not early, and these data need to be confirmed. Why is this important? Because if VAP does not impart an increase in mortality, then trying to decrease mortality by reducing VAP is just swinging at windmills.
So, let's talk about the study and what it showed as reported in the BMJ paper. You will be pleased that I will not here go through the traditional list of potential threats to validity, but take the data at face value (well, almost). The authors took an interesting approach of comparing the performance of all eligible ICUs regardless of whether they actually chose to take part in the project. Of all the admissions examined in the intervention group, 88% came from Keystone participants. This is a really sound way to define the intervention cohort, and it actually biases the data away from showing an effect. So, kudos to the investigators. The comparator cohort came from ICUs in the hospitals surrounding Michigan, those that were not eligible for Keystone participation. One point about these institutions also requires clarification: I did not see in the paper whether the authors actually looked at the control hospitals' QI initiatives. Why is this important? Well, if many of the comparator hospitals had successful QI initiatives, then one could expect to see even less difference between the Keystone intervention and the control group. So, again, good on them that they biased the data against themselves.
This is the line of thinking that brings me to my second point. Reuters' Maggie Fox covered this paper in an article a couple of days ago, an article whosebyline lede (thanks for the correction, @ivanoransky) floored me:
There are multiple places to look for the mortality data. One is found in this figure:
So, here is the whole story. Stephanie Desmon, the author of the JH press release, e-mailed me back and pointed me to Peter Pronovost as the source for the 10% reduction information. I e-mailed Peter, and he got back to me, confirming that
"The 10 percent is the rounded differences in differences in odds ratios"Moral of the story: The devil is in the details.
And speaking of details, I must admit to an error of my own. If you look at the figure reproduced below, I called out the wrong points. For adjusted data, you need to look at the open circles (for the intervention group) and squares (for the control group). In fact, the adjusted mortality went from about 20% at baseline to 16% in the 13-22 months interval for the Keystone cohort, while for the control group it went from a little over 20% to a little under 18%. This makes the absolute reduction a tad more impressive, though there is still less than a 2% absolute difference between the reduction seen in the intervention vs. the control group, leaving all of my other points still in need of addressing.
Addendum #1, 11:00 AM EST, 2/2/11:
I just found what I think is the origin of the 10% mortality reduction rumor in this press release from Johns Hopkins. I just e-mailed Stephanie Desmon, the author of the release, to see where the 10% came from. Will update again should I hear from either Maggie Fox or Stephanie Desmon.
Remember the Keystone project? A number of years ago when we started to pay close attention to healthcare-associated infections (HAI), and hospitals started to take introspective looks at their records, it turned out the the ICUs in the state of Michigan for one reason or another had very high rates of HAIs. As this information percolated through our collective consciousness, the stars aligned in such a away as to release funding from the AHRQ in Washington, DC, for a group of ICU investigators at the Johns Hopkins University School of Medicine in Baltimore, MD, headed by Peter Pronovost, to design and implement a study employing IHI-style (Boston, MA) bundled interventions to prevent catheter-associated blood stream infections (CABSI) and ventilator-associated pneumonia (VAP) across the consortium of ICUs in MI. Whew! This poly-geographic collaboration resulted in a landmark paper in 2006 in the New England Journal of Medicine, wherein the authors showed that the bundled interventions directed by a checklist aimed at CABSI were indeed associated with a satisfying reduction of CABSI. Since 2006 the ICU community has been eagerly awaiting the results of the VAP intervention from Keystone, but none has come out. When there is a void of information, rumors fill this void, and plenty of rumors have circulated about the alleged failure of the VAP trial.
I do not want to belabor here what I have written before with regard to VAP and its prevention, and what makes the latter so difficult, and how little evidence there really is that the IHI bundle actually does anything. You can find at least some of my thoughts on that here. But why am I bringing up the Keystone project again anyway? Well, it is because Pronovost's group has just published a new paper in BMJ, and this time their aim was even more ambitious: to show the impact of this state-wide QI intervention on hospital mortality and length of stay. This is a really reasonable question, mind you, since, we could argue that, if the intervention reduces HAI, it should also do something to those important downstream events that are driven by the particular HAI, namely mortality and LOS. But here are a couple of issues that I found of great interest.
First, as we have discussed before, whether or not VAP itself causes death in the ICU population (that is patients die from VAP), or whether VAP tends to attack those who are sicker and therefore more likely to die anyway (patients die with VAP) remains unclear in our literature. There is some evidence that late VAP may be associated with an attributable increase in mortality, but not early, and these data need to be confirmed. Why is this important? Because if VAP does not impart an increase in mortality, then trying to decrease mortality by reducing VAP is just swinging at windmills.
So, let's talk about the study and what it showed as reported in the BMJ paper. You will be pleased that I will not here go through the traditional list of potential threats to validity, but take the data at face value (well, almost). The authors took an interesting approach of comparing the performance of all eligible ICUs regardless of whether they actually chose to take part in the project. Of all the admissions examined in the intervention group, 88% came from Keystone participants. This is a really sound way to define the intervention cohort, and it actually biases the data away from showing an effect. So, kudos to the investigators. The comparator cohort came from ICUs in the hospitals surrounding Michigan, those that were not eligible for Keystone participation. One point about these institutions also requires clarification: I did not see in the paper whether the authors actually looked at the control hospitals' QI initiatives. Why is this important? Well, if many of the comparator hospitals had successful QI initiatives, then one could expect to see even less difference between the Keystone intervention and the control group. So, again, good on them that they biased the data against themselves.
This is the line of thinking that brings me to my second point. Reuters' Maggie Fox covered this paper in an article a couple of days ago, an article whose
(Reuters) - A U.S. program to help make sure hospital staff maintain strict hygiene standards lowered death rates in intensive care units by 10 percent, U.S. researchers reported on Monday.Mind you, I read the article before delving into the peer-reviewed paper, so my surprise came out of just knowing how supremely difficult it is to reduce ICU mortality by 10% with any intervention. In the ICU we celebrate when we see even a 2% absolute mortality reduction. So, it became obvious to me that something got lost in translation here. And indeed, it did. Here is how I read the data.
There are multiple places to look for the mortality data. One is found in this figure:
Now, look at the top panel and focus on the solid circles -- these depict the adjusted mortality in the Keystone intervention group. What do you see? I see mortality going from about 14% at the baseline to about 13.5% at implementation phase to about 13% at 13-22 months post implementation. I do not see a 10% reduction, but at best about a 1% mortality advantage. What is also of interest is that the adjusted mortality in the control group (solid squares) also went down, albeit not by as much. But almost at every point of measurement it was lower already than in the intervention group.
Then there is this table, where the adjusted odds ratios of death are given for the two groups at various time points:
And this is where things get interesting. If you look at the last line of the table, the adjusted odds ratios indeed look impressive, and, furthermore, the AOR for the intervention group is lower than that for the control group. And this is pleasing to any investigator. But what does it mean? Well it means that the odds of death in the intervention group went down roughly by 24% (give-or-take the 95% confidence interval) and by 16% in the control group,each compared to itself at baseline. This is impressive, no?
Well, yes, it is. But not as impressive as it sounds. A relative reduction of 24% with the baseline mortality of 14% means an absolute reduction in mortality of 14% x 24% = 3.4%. But, you notice that we did not actually observe even this magnitude of mortality reduction in the graph. What gives? There is an excellent explanation for this. It is a little known fact to the the reader (and only slightly more so to the average researcher and peer reviewer) that the odds ratio, while a fairly solid way to express risk when the absolute risk is small (say, under 10%), tends to overestimate the effect when the risk is higher than 10%. I know we have not yet covered the ins and the outs of odds ratios, relative risks and the like in the "reviewing literature" series, but let me explain briefly. The difference between odds and risk is in the denominator. While the denominator for the latter is the entire cohort at risk for the event (here all patients at risk for dying in the hospital), that for the former is that part of the cohort that did not experience the event. See the difference? By definition, the denominator for the odds ratio is smaller than for the relative risk calculation, thus yielding a more impressive, yet inaccurate, reduction in mortality.
Bottom line? Interesting results. Not clear if the actual intervention is what produced the 1% mortality reduction -- could have been secular trends, regression to the mean or Hawthorne effect, to name just a few alternatives. But regardless, preventing death is good. The question is were these improvements in mortality sustained after hospital discharge, or were these patients merely kept alive so that they could die elsewhere? Also, what is the value balance here in terms of resources expended on the intervention versus the results that may not even be due to the particular intervention in question?
All of this is to say that I am really not sure what the data are showing. What I am sure of is that I did not find any evidence of a 10% reduction in mortality reported by Reuters (I did e-mail Maggie Fox and at this time still awaiting a reply; will update if and when I get it). In this time of aggressive efforts to bend the healthcare expenditures curve we need to pay attention to what we invest in and the return on this investment, even if the intervention is all "motherhood and apple pie."
Tuesday, 1 February 2011
The beautiful uncertainty of science
I am so tired of this all-or-nothing discussion about science! On the one hand there is a chorus singing praises to science and calling people who are skeptical of certain ideas unscientific idiots. On the other, with equal penchant for eminence-based thinking, are the masses convinced of conspiracies and nefarious motives of science and its perpetrators. And neither will stop and listen to the other side's objections, and neither will stop the name-calling. So, is it any wonder we are not getting any closer to the common ground? And if you are not a believer in the common ground, let me say that we are only getting farther away from the truth, if such a thing exists, by retreating further into our cognitive corners. These corners are comfortable places, with our comrades-in-arms sharing our, shall we say, passionate opinions. Yet this is not the way to get to a better understanding.
Because I spend so much time contemplating our larger understanding of science, the title "Are We Hard-Wired to Doubt Science" proved to be a really inflammatory way to suck me into thinking about everything I am interested in integrating: scientific method, science literacy and communication and brain science. The author, on the heels of doing a story on the opposition to smart meters in California, was led to try to understand why we are so quick to reject science:
I happen to think that the author missed an opportunity to educate her readers about why we need to come to a better understanding and how to get there. The public (and even some of my fellow scientists) needs to understand what science is and, even more importantly, what it is not.
First, science is not dogma. Karl Popper had a very simple litmus test for scientific thinking: He asked how you would go about disproving a particular idea. If you think that the idea is above being disproved, then you are engaging in dogma and not science. The essence of scientific method is developing an hypothesis from either a systematically observed pattern or from a theoretical model. The hypothesis is necessarily formulated as the null, making the assumption of no association the departure point for proving the contrary. So, to "prove" that the association is present you need to rule out any other potential explanation for what may appear to be an association. For example, if thunder were always followed by rain, it might be easy to engage in the "post hoc ergo propter hoc" fallacy and conclude that thunder caused rain. But before this could become a scientific theory, you would have to show that there was no other explanation that would disprove this association.
So, the second point is that science is driven by postulating and then disproving the null hypotheses. By definition, an hypothesis can only be disproved if we 1). the association exists, and 2). the constellation of phenomena is not explained by something else. And here is the third and critical point, the point that produces equal parts frustration and inspiration to learn more: That "something else" as the explanation of a certain association is by definition informed only by what we know today. It is this very quality of knowledge production, the constancy of the pursuit, that lends the only certain property to science, the property of uncertainty. And our brains have a hard time holding and living with this uncertainty.
The tension between uncertainty and the need to make public policy has taken on a political life of its own. What started out as a modest storm of subversion of science by politics in the tobacco debate, has now escalated into a cyclone of everyday leveraging of the scientific uncertainties for political and economic gains. After all, how can we balance the accounting between the theoretical models predicting climate doom in the future and the robust current-day economic gains produced by the very pollution that feeds these models? How can we even conceive that our food production system, yielding more abundant and cheaper food than ever before, is driving the epidemic of obesity and the catastrophe of antimicrobial resistance? And because we are talking about science, and because, as that populist philosopher Yogi Berra famously quipped, "Predictions are hard, especially about the future," the uncertainty of our estimates overshadows the probability of their correctness. Yet by the time the future becomes present, we will be faced with potentially insurmountable challenges of a new world.
I have heard some scientists express reluctance about "coming clean" to the public about just how uncertain our knowledge is. Nonsense! What we need under the circumstances is greater transparency, public literacy and engagement. Science is not something that happens in the bastions of higher education or behind the thick walls of corporations. Science is all around and within us. And if you believe in God, you have to believe that God is a scientist, a tinkerer, always looking for a more elegant solution. The language of science that may seem daunting and obfuscatory. Yet do not be afraid -- patterns of a language are easy to decipher with some willingness and a dictionary. Our brains are attuned to the most beautiful explanations of the universe. Science is what provides them.
Self-determination is predicated upon knowledge and understanding. Abdicating our ability to understand the scientific method leaves us subject to political demagoguery. Don't be a puppet. We are all born scientists. Embrace your curiosity, tune out the noise of those at the margins who are not willing to engage in a sensible dialogue, leave them to their schoolyard brawling. And likewise, leave the politicians, corporate interests, and, alas, many a journalist, and start learning the basics of scientific philosophy and thought. Allow the uncertainty of knowledge excite and delight you. You will not be disappointed.
Because I spend so much time contemplating our larger understanding of science, the title "Are We Hard-Wired to Doubt Science" proved to be a really inflammatory way to suck me into thinking about everything I am interested in integrating: scientific method, science literacy and communication and brain science. The author, on the heels of doing a story on the opposition to smart meters in California, was led to try to understand why we are so quick to reject science:
She goes on to think about the different ways of perceiving risk, and how our brains play tricks on us by perpetuating our many cognitive biases. In essence, new data are unable to sway our opinion because of rescue bias, or our drive to preserve what we think we know to be true and to reject what our intuition tells us is false. If we follow this argument to its logical conclusion, it means that we just need to throw our hands up in the air and accept the status quo, whatever it is.But some very intelligent people I interviewed had little use for the existing (if sparse) science. How, in a rational society, does one understand those who reject science, a common touchstone of what is real and verifiable?The absence of scientific evidence doesn’t dissuade those who believe childhood vaccines are linked to autism, or those who believe their headaches, dizziness and other symptoms are caused by cellphones and smart meters. And the presence of large amounts of scientific evidence doesn’t convince those who reject the idea that human activities are disrupting the climate.
I happen to think that the author missed an opportunity to educate her readers about why we need to come to a better understanding and how to get there. The public (and even some of my fellow scientists) needs to understand what science is and, even more importantly, what it is not.
First, science is not dogma. Karl Popper had a very simple litmus test for scientific thinking: He asked how you would go about disproving a particular idea. If you think that the idea is above being disproved, then you are engaging in dogma and not science. The essence of scientific method is developing an hypothesis from either a systematically observed pattern or from a theoretical model. The hypothesis is necessarily formulated as the null, making the assumption of no association the departure point for proving the contrary. So, to "prove" that the association is present you need to rule out any other potential explanation for what may appear to be an association. For example, if thunder were always followed by rain, it might be easy to engage in the "post hoc ergo propter hoc" fallacy and conclude that thunder caused rain. But before this could become a scientific theory, you would have to show that there was no other explanation that would disprove this association.
So, the second point is that science is driven by postulating and then disproving the null hypotheses. By definition, an hypothesis can only be disproved if we 1). the association exists, and 2). the constellation of phenomena is not explained by something else. And here is the third and critical point, the point that produces equal parts frustration and inspiration to learn more: That "something else" as the explanation of a certain association is by definition informed only by what we know today. It is this very quality of knowledge production, the constancy of the pursuit, that lends the only certain property to science, the property of uncertainty. And our brains have a hard time holding and living with this uncertainty.
The tension between uncertainty and the need to make public policy has taken on a political life of its own. What started out as a modest storm of subversion of science by politics in the tobacco debate, has now escalated into a cyclone of everyday leveraging of the scientific uncertainties for political and economic gains. After all, how can we balance the accounting between the theoretical models predicting climate doom in the future and the robust current-day economic gains produced by the very pollution that feeds these models? How can we even conceive that our food production system, yielding more abundant and cheaper food than ever before, is driving the epidemic of obesity and the catastrophe of antimicrobial resistance? And because we are talking about science, and because, as that populist philosopher Yogi Berra famously quipped, "Predictions are hard, especially about the future," the uncertainty of our estimates overshadows the probability of their correctness. Yet by the time the future becomes present, we will be faced with potentially insurmountable challenges of a new world.
I have heard some scientists express reluctance about "coming clean" to the public about just how uncertain our knowledge is. Nonsense! What we need under the circumstances is greater transparency, public literacy and engagement. Science is not something that happens in the bastions of higher education or behind the thick walls of corporations. Science is all around and within us. And if you believe in God, you have to believe that God is a scientist, a tinkerer, always looking for a more elegant solution. The language of science that may seem daunting and obfuscatory. Yet do not be afraid -- patterns of a language are easy to decipher with some willingness and a dictionary. Our brains are attuned to the most beautiful explanations of the universe. Science is what provides them.
Self-determination is predicated upon knowledge and understanding. Abdicating our ability to understand the scientific method leaves us subject to political demagoguery. Don't be a puppet. We are all born scientists. Embrace your curiosity, tune out the noise of those at the margins who are not willing to engage in a sensible dialogue, leave them to their schoolyard brawling. And likewise, leave the politicians, corporate interests, and, alas, many a journalist, and start learning the basics of scientific philosophy and thought. Allow the uncertainty of knowledge excite and delight you. You will not be disappointed.
Monday, 31 January 2011
Reviewing medical literature, part 5: Inter-group differences and hypothesis testing
Happy almost February to everyone. It is time to resume our series. First I am grateful for the vigorous response to my survey about interest in a webinar to cover some of this stuff. Over the next few months one of my projects will be to develop and execute one. I will keep you posted. In the meantime, if anyone has thoughts or suggestions on the logistics, etc., please, reach out to me.
OK, let's talk about group comparisons and hypothesis testing. Scientific method that we generally practice demands that we articulate an hypothesis prior to conducting a study which will test this hypothesis. The hypothesis is generally advanced as the so-called "null hypothesis" (or H0), wherein we express our skepticism that there is a difference between groups or an association between the exposure and outcome. By starting out with this negative formulation, we set the stage for "disproving" the null hypothesis, or demonstrating that the data support the "alternative hypothesis" (HA, or the presence of the said association or difference). This is where all the measures of association that we have discussed previously come in, and most particularly the p value. The definition of the p value once again is "the probability that the found inter-group difference, or one that is greater than what was found, would have been found under the condition of no true difference." Following through on this reasoning, we can appreciate that the H0 can never be "proven." That is, the only thing that can be said statistically when no difference is found between groups is that we did not disprove the null hypothesis. This may be because there truly is no difference between the groups being compared (that is the null hypothesis approximates reality) or because we did not find the difference that in fact exists. The latter is referred to as the Type II error, and can be present for various reasons, the most common of which is a sample size that is too small to detect statistically significant difference.
This is a good place to digress and talk a little about the distinction between "absence of evidence" and "evidence of absence." The distinction, though ostensibly semantic, is quite important. While "evidence of absence" implies that studies to look for associations have been done, done well, published, and have consistently shown the lack of association between a given exposure and outcome or a difference between two groups, "absence of evidence" means that we have just not done a good job looking for this association or difference. Absence of evidence does not absolve the exposure from causing the outcome, yet so often it is confused with the definitive evidence of absence of an effect. Nowhere is this more apparent than in the history of the tobacco debate, which is the poster child for this obfuscation. And we continue to rely on this confusion in other environmental debates, such as chemical exposures and cell phone radiation. One of the most common reasons for finding no association when one exists, or the type II error, is, as I have already mentioned, a sample size that is too small to detect the difference. For this reason, in a published study that fails to show a difference between groups it is critical to assure that the investigators performed the power calculation. This maneuver, usually found in the Methods section of the paper, lets us know that the sample size is adequate to detect a difference if one exists, thus minimizing the probability of type II error. The trouble is that, as we know, there is a phenomenon called "publication bias." This refers to the scientific journals' reluctance to publish negative results. And while it may be appropriate to reject studies prone to type II error due to poor design (although even these studies may be useful in the setting of a meta-analysis, where pooling of data overcomes small sample sizes), true negative results must be made public. But this is a little off topic.
I will ask you to indulge me in one other digression. I am sure that in addition to "statistical significance" (this is simplistically represented by the p value), you have heard of "clinical significance." This is an important distinction, since even a finding that is statistically significant may have no clinical significance whatsoever. Take for example a therapy that cuts the risk of a non-fatal heart attack by 0.05% in a certain population. This means that in a population at a 10% risk for a heart attack in one year, the intervention will bring this risk on average to 9.95%. And though we can argue whether or not this is an important difference, at the population level, this does not seem all that clinically important. So, if I have the vested interest and the resources to run the massive trial that will give me this minute statistical significance, I can do that and then say without blushing that my treatment works. Yet, statistical significance always needs to be examined in the clinical context. This is why it is not enough to read the headlines that tout new treatments. The corollary to this is that the lack of statistical significance does not equate to the lack of clinical significance. Given what I just said above about type II error, if the difference appears significant clinically (e.g., reducing the incidence of fatal heart attacks from 10% to 5%), but does not reach statistical significance, the result should not be discarded as negative, but examined as to the probability of the type II error. This is also where Bayesian thinking must come into play, but I do not want to get into this now, as we have covered these issues in previous posts on this blog.
OK, back to hypothesis testing. There are several rules to be aware of when reading how the investigators tested their hypotheses, as different types of variables require different methods. A categorical variable (one characterized by categories, like gender, race, death, etc.) can be compared using the chi square method if there is an abundance of events or the Fisher's exact test when values are scant. A normally distributed continuous variable (e.g., age is a continuum that is frequently distributed normally) can be tested using the Student's t-test, while one that has a skewed distribution (e.g., hospital length of stay, costs), requires testing with the Mann-Whitney U-test or the Wilcoxon rank-sum test or the Kruskall-Wallis test. Each of these "non-parametric" tests is appropriate in the setting of a skewed distribution. You do not need to know any more than this: the test for the hypothesis depends on the variable's distribution. And recognizing some of the situations and test names may be helpful to you in evaluating the validity of a study.
One final frequent computation you may encounter is survival analysis. This is often depicted as a Kaplan-Meier curve, and does not have to be limited to examining survival. This is a time-to-event analysis, regardless of what the event is. In studies of cancer therapies we frequently talk about median disease-free survival between groups, and this can be depicted by the K-M analysis. To test the difference between times to event, we employ the log-rank test.
Well, this is a fairly complete primer for most common hypothesis testing situations. In the next post we will talk a little more about measures of association and their precision, types I and II errors, as well as measures of risk alteration.
OK, let's talk about group comparisons and hypothesis testing. Scientific method that we generally practice demands that we articulate an hypothesis prior to conducting a study which will test this hypothesis. The hypothesis is generally advanced as the so-called "null hypothesis" (or H0), wherein we express our skepticism that there is a difference between groups or an association between the exposure and outcome. By starting out with this negative formulation, we set the stage for "disproving" the null hypothesis, or demonstrating that the data support the "alternative hypothesis" (HA, or the presence of the said association or difference). This is where all the measures of association that we have discussed previously come in, and most particularly the p value. The definition of the p value once again is "the probability that the found inter-group difference, or one that is greater than what was found, would have been found under the condition of no true difference." Following through on this reasoning, we can appreciate that the H0 can never be "proven." That is, the only thing that can be said statistically when no difference is found between groups is that we did not disprove the null hypothesis. This may be because there truly is no difference between the groups being compared (that is the null hypothesis approximates reality) or because we did not find the difference that in fact exists. The latter is referred to as the Type II error, and can be present for various reasons, the most common of which is a sample size that is too small to detect statistically significant difference.
This is a good place to digress and talk a little about the distinction between "absence of evidence" and "evidence of absence." The distinction, though ostensibly semantic, is quite important. While "evidence of absence" implies that studies to look for associations have been done, done well, published, and have consistently shown the lack of association between a given exposure and outcome or a difference between two groups, "absence of evidence" means that we have just not done a good job looking for this association or difference. Absence of evidence does not absolve the exposure from causing the outcome, yet so often it is confused with the definitive evidence of absence of an effect. Nowhere is this more apparent than in the history of the tobacco debate, which is the poster child for this obfuscation. And we continue to rely on this confusion in other environmental debates, such as chemical exposures and cell phone radiation. One of the most common reasons for finding no association when one exists, or the type II error, is, as I have already mentioned, a sample size that is too small to detect the difference. For this reason, in a published study that fails to show a difference between groups it is critical to assure that the investigators performed the power calculation. This maneuver, usually found in the Methods section of the paper, lets us know that the sample size is adequate to detect a difference if one exists, thus minimizing the probability of type II error. The trouble is that, as we know, there is a phenomenon called "publication bias." This refers to the scientific journals' reluctance to publish negative results. And while it may be appropriate to reject studies prone to type II error due to poor design (although even these studies may be useful in the setting of a meta-analysis, where pooling of data overcomes small sample sizes), true negative results must be made public. But this is a little off topic.
I will ask you to indulge me in one other digression. I am sure that in addition to "statistical significance" (this is simplistically represented by the p value), you have heard of "clinical significance." This is an important distinction, since even a finding that is statistically significant may have no clinical significance whatsoever. Take for example a therapy that cuts the risk of a non-fatal heart attack by 0.05% in a certain population. This means that in a population at a 10% risk for a heart attack in one year, the intervention will bring this risk on average to 9.95%. And though we can argue whether or not this is an important difference, at the population level, this does not seem all that clinically important. So, if I have the vested interest and the resources to run the massive trial that will give me this minute statistical significance, I can do that and then say without blushing that my treatment works. Yet, statistical significance always needs to be examined in the clinical context. This is why it is not enough to read the headlines that tout new treatments. The corollary to this is that the lack of statistical significance does not equate to the lack of clinical significance. Given what I just said above about type II error, if the difference appears significant clinically (e.g., reducing the incidence of fatal heart attacks from 10% to 5%), but does not reach statistical significance, the result should not be discarded as negative, but examined as to the probability of the type II error. This is also where Bayesian thinking must come into play, but I do not want to get into this now, as we have covered these issues in previous posts on this blog.
OK, back to hypothesis testing. There are several rules to be aware of when reading how the investigators tested their hypotheses, as different types of variables require different methods. A categorical variable (one characterized by categories, like gender, race, death, etc.) can be compared using the chi square method if there is an abundance of events or the Fisher's exact test when values are scant. A normally distributed continuous variable (e.g., age is a continuum that is frequently distributed normally) can be tested using the Student's t-test, while one that has a skewed distribution (e.g., hospital length of stay, costs), requires testing with the Mann-Whitney U-test or the Wilcoxon rank-sum test or the Kruskall-Wallis test. Each of these "non-parametric" tests is appropriate in the setting of a skewed distribution. You do not need to know any more than this: the test for the hypothesis depends on the variable's distribution. And recognizing some of the situations and test names may be helpful to you in evaluating the validity of a study.
One final frequent computation you may encounter is survival analysis. This is often depicted as a Kaplan-Meier curve, and does not have to be limited to examining survival. This is a time-to-event analysis, regardless of what the event is. In studies of cancer therapies we frequently talk about median disease-free survival between groups, and this can be depicted by the K-M analysis. To test the difference between times to event, we employ the log-rank test.
Well, this is a fairly complete primer for most common hypothesis testing situations. In the next post we will talk a little more about measures of association and their precision, types I and II errors, as well as measures of risk alteration.
Friday, 28 January 2011
Soliciting contributions: "Healthcare professional as e-patient" series
I am contemplating a series of posts arising from my own recent experience as an e-patient to help the broader e-patient community navigate the stormy medical waters with a bit more comfort. I am looking for other healthcare professionals who have had their own experiences as an e-patient that may be instructive for non-healthcare professionals as patients to contribute to the series. Namely, I am most interested in helping people establish better communication lines and channels with their healthcare providers. I am not looking for a comprehensive description of every aspect of your encounter, but rather one specific point that may be particularly instructive. If you have more than one to share, that is great too, we can do that as well. No bitching or moaning, just lessons that we can all learn from.
Issues I would like to touch upon range from how to bring out risk-benefit balance to how to feel OK about confronting your physician with dissenting information to how best to communicate (not everyone is good on e-mail, for example), given our individual styles and time constraints.
I think contributions from healthcare professionals may be very valuable, as we can see both sides of the coin, so to speak.
Would love feedback from both, healthcare professionals and e-patients on what would be valuable. If you are interested in contributing, please, either let me know in the comments section or e-mail me at healthcareetcblog@gmail.com. If you have an idea for a post, please, be very specific about what your theme is, as I will make decisions based on how relevant it is to what I am envisioning.
This is jut a thought at this point, but seems like there may be something to it. Looking forward to your ideas.
Issues I would like to touch upon range from how to bring out risk-benefit balance to how to feel OK about confronting your physician with dissenting information to how best to communicate (not everyone is good on e-mail, for example), given our individual styles and time constraints.
I think contributions from healthcare professionals may be very valuable, as we can see both sides of the coin, so to speak.
Would love feedback from both, healthcare professionals and e-patients on what would be valuable. If you are interested in contributing, please, either let me know in the comments section or e-mail me at healthcareetcblog@gmail.com. If you have an idea for a post, please, be very specific about what your theme is, as I will make decisions based on how relevant it is to what I am envisioning.
This is jut a thought at this point, but seems like there may be something to it. Looking forward to your ideas.
Thursday, 27 January 2011
The price of marginal thinking in healthcare policy
January 27, 2011
health economics, healthcare, lung cancer, obesity, Pareto principle, screening
No comments
I find it fascinating how our brains have this propensity to latch on to what is at the margins at the expense of seeing the bulk of what sits in the center. This peripheral only vision is in part responsible for our obscene healthcare expenditures and underwhelming results.
I have blogged ad nauseam about the drivers of early mortality in the US. In one post I reproduced a pie chart from the Rand Corporation, wherein they show explicitly that a mere 10% of all premature deaths in the US can be attributed to being unable to access medical care. The other 90% is split nearly evenly between behavioral, social-environmental and genetic factors, of which 60%, the non-genetic drivers, can be modified. Yet instead of investing the bulk of our resources in this big bucket of behavioral-environmental-social modification, we put 97% of all healthcare dollars towards medical interventions. This investment can at best produce marginal improvements in premature deaths, since the biggest causes of the effect in question are being all but ignored.
A couple of other striking examples of this marginal magical thinking have surfaced in a few recent stories covered with gusto in the press. One of the bigger ones is the obesity epidemic (oh, yes, you bet it was intended), and its causes. This New York Times piece with its magnetic headline "Central Heating May Be Making Us Fat" entertains the possibility that because of the more liberal use of heat in our homes we are no longer engaging our brown fat, which is a furnace for burning calories. And this is all well and good and fascinating, in a rounding out sort of a way. And it is just as interesting to hear that lack of sleep may be contributing to our expanding waistlines. But it is also baffling that we are still expending these enormous amounts of energy (OK, this one was not intended) on finding the silver bullet, when the target is not a supernatural being, but a super-sized expectation. Is it really that mysterious that we are fatter now than we were 20 years ago, when our current portion sizes are 70% bigger and we spend our days worshipping at the temple of the screen, in all its manifestations? While I am all for learning as much as we can, what we need right now is immediate action to abrogate this escalating epidemic, and I think we can all agree that the way to do it is not through lowering house temperatures. Plenty of behavioral research is available to inform our strategies to get people to eat less and move more. Let's start translating it into practice rather than latch on to one marginal magical idea after another.
And finally, I have to touch upon lung cancer, of course. The current fodder for this was provided by the Washington Post with this story about the growing advocacy among lung cancer patients for early detection. You may recall that recently I did several posts on the heels of the large NCI-sponsored study National Lung Screening Trial (NLST) whose purpose was to understand whether early detection of lung cancer in heavy smokers may improve lung cancer survival. I do not wish to go into all of the specifics of this study and my interpretation of the results -- you can find my thoughts on this study in particular and on screening in general here. What I do want to reiterate is that 85% of all lung cancer is caused by a single exposure: smoking. And guess what? The same behavioral strategies that can help people stop overeating can be deployed towards smoking cessation. Yet, instead of spending 85% of all expenditures on smoking cessation efforts, we prefer to allocate it to early detection. My point is that we need both, but the balance has to be informed by pragmatism, not the marginal magical thinking.
And so it goes that the Pareto principle is bleeding into our healthcare policy decisions -- this is the steep price of the marginal magical thinking. What will it take to get the blinders off and face up to the idea that some intervention points are just more impactful than others? Marginal panaceas will improve our lives, but only at the margins. And without being addressed, the big elephants in the room are likely to stampede us.
I have blogged ad nauseam about the drivers of early mortality in the US. In one post I reproduced a pie chart from the Rand Corporation, wherein they show explicitly that a mere 10% of all premature deaths in the US can be attributed to being unable to access medical care. The other 90% is split nearly evenly between behavioral, social-environmental and genetic factors, of which 60%, the non-genetic drivers, can be modified. Yet instead of investing the bulk of our resources in this big bucket of behavioral-environmental-social modification, we put 97% of all healthcare dollars towards medical interventions. This investment can at best produce marginal improvements in premature deaths, since the biggest causes of the effect in question are being all but ignored.
A couple of other striking examples of this marginal magical thinking have surfaced in a few recent stories covered with gusto in the press. One of the bigger ones is the obesity epidemic (oh, yes, you bet it was intended), and its causes. This New York Times piece with its magnetic headline "Central Heating May Be Making Us Fat" entertains the possibility that because of the more liberal use of heat in our homes we are no longer engaging our brown fat, which is a furnace for burning calories. And this is all well and good and fascinating, in a rounding out sort of a way. And it is just as interesting to hear that lack of sleep may be contributing to our expanding waistlines. But it is also baffling that we are still expending these enormous amounts of energy (OK, this one was not intended) on finding the silver bullet, when the target is not a supernatural being, but a super-sized expectation. Is it really that mysterious that we are fatter now than we were 20 years ago, when our current portion sizes are 70% bigger and we spend our days worshipping at the temple of the screen, in all its manifestations? While I am all for learning as much as we can, what we need right now is immediate action to abrogate this escalating epidemic, and I think we can all agree that the way to do it is not through lowering house temperatures. Plenty of behavioral research is available to inform our strategies to get people to eat less and move more. Let's start translating it into practice rather than latch on to one marginal magical idea after another.
And finally, I have to touch upon lung cancer, of course. The current fodder for this was provided by the Washington Post with this story about the growing advocacy among lung cancer patients for early detection. You may recall that recently I did several posts on the heels of the large NCI-sponsored study National Lung Screening Trial (NLST) whose purpose was to understand whether early detection of lung cancer in heavy smokers may improve lung cancer survival. I do not wish to go into all of the specifics of this study and my interpretation of the results -- you can find my thoughts on this study in particular and on screening in general here. What I do want to reiterate is that 85% of all lung cancer is caused by a single exposure: smoking. And guess what? The same behavioral strategies that can help people stop overeating can be deployed towards smoking cessation. Yet, instead of spending 85% of all expenditures on smoking cessation efforts, we prefer to allocate it to early detection. My point is that we need both, but the balance has to be informed by pragmatism, not the marginal magical thinking.
And so it goes that the Pareto principle is bleeding into our healthcare policy decisions -- this is the steep price of the marginal magical thinking. What will it take to get the blinders off and face up to the idea that some intervention points are just more impactful than others? Marginal panaceas will improve our lives, but only at the margins. And without being addressed, the big elephants in the room are likely to stampede us.
SoMe in medicine: It's about communication, stupid!
My generation of doctors was almost proud of its paternalistic overbearing know-it-all archetype, with the my-way-or-the-highway attitude to patient care. Even today there are inter-specialty fights in medicine that demonstrate these entrenched and seemingly fundamental, albeit willfully exaggerated, differences of opinion and clinical approach. It used to be, and still is to an extent, a badge of honor for an internist to disagree with a surgeon, for a pulmonologist to recommend a course of action diametrically opposed to that suggested by the infectious diseases specialist, and for everyone to disparage neurologists (apologies to my neuro friends). The extent of the discussion with patients as modeled by some of my senior colleagues was to say "You have this, and I am giving you this prescription, and see you in 2 months." And even today, I have observed the best of doctors still respond to a cogent "why?" question from a patient with a "because this is how we do it" answer.
My peers' lack of communication skills is the stuff of urban legends. Yet here we are at what seems like a pivotal moment for so many aspects of medicine -- science, healthcare system, communication technologies -- where effectively communicating outside the profession is a make-or-break proposition. Along these lines, in this BBC documentary Sir Paul Nurse, the head of the Royal Society, examines the societal forces that are coalescing to bring "Science Under Attack." The unifying message that comes out of his inquiry is that other less informed parties with political agendas are co-opting the discussion. Yet there is a distinct lack of the antidote of countervailing communication by scientists in terms that are understandable to the lay public. Nurse's battle cry is that scientists need to do a better job communicating their craft themselves, and not just to each other.
In some ways the prevailing elitism of medicine in the 20th century set the stage for the backlash we are experiencing today. The erosion of trust in the profession, commodification and consequent devaluation of medicine, while multifactorial at their root, could no doubt have been mitigated with better communication. Yet, great communicators rarely choose medicine as the path.
And this brings me to the contentious topic of the role of social media in medicine. For many of the early adopters, the question is no longer "should we", but "how best to." But my sense is, that physicians engaging in social media are still a minority. I am not even sure what proportion of MDs are amenable to communication via e-mail with their patients, though these data may be out there. So, for what seems to me as the majority of MDs who are not sold on e-mail, Twitter, Facebook, blogging or Quora, the value must not be that obvious. This makes me wonder if there are certain unifying characteristics of these docs, one being lack of perceived value of communication outside the profession across all media, including in-person contact.
I am friends with many docs on Twitter and in the blogosphere. The vast majority of them have shown themselves to be patient-centric, knowledgeable and collaborative, the kind of people I would not hesitate to send a loved one to. Yet, this is a skewed sample born out of a selection bias. These are the people who are interested and confident in their ability to communicate outside medicine. These are the people to whom medicine is a humanistic pursuit, where communities of patients and doctors strengthen the discussion of how to transform our system and the patient encounter. My guess is, and this is purely unscientific, that many of those who are skeptical of social media are also skeptical of communication itself, or just do not see the value of it in the equation of providing good patient care within the crushing time constraints of today's healthcare.
So my point is this: before social media tools can be expected to diffuse broadly into the medical community, the value of all communication needs to become clear to physicians in general. At this moment of increasing societal skepticism of science and of usefulness and integrity of the medical profession, against the backdrop of healthcare changes and increasingly unfiltered media noise, willingness and skills to communicate clearly may be as useful to today's doctors as a stethoscope. Once communication becomes the backbone of all medicine, tweets and blog posts are sure to start flowing freely from the fingers of physicians everywhere. And that will be good for the patients, the science and the healthcare system.
My peers' lack of communication skills is the stuff of urban legends. Yet here we are at what seems like a pivotal moment for so many aspects of medicine -- science, healthcare system, communication technologies -- where effectively communicating outside the profession is a make-or-break proposition. Along these lines, in this BBC documentary Sir Paul Nurse, the head of the Royal Society, examines the societal forces that are coalescing to bring "Science Under Attack." The unifying message that comes out of his inquiry is that other less informed parties with political agendas are co-opting the discussion. Yet there is a distinct lack of the antidote of countervailing communication by scientists in terms that are understandable to the lay public. Nurse's battle cry is that scientists need to do a better job communicating their craft themselves, and not just to each other.
In some ways the prevailing elitism of medicine in the 20th century set the stage for the backlash we are experiencing today. The erosion of trust in the profession, commodification and consequent devaluation of medicine, while multifactorial at their root, could no doubt have been mitigated with better communication. Yet, great communicators rarely choose medicine as the path.
And this brings me to the contentious topic of the role of social media in medicine. For many of the early adopters, the question is no longer "should we", but "how best to." But my sense is, that physicians engaging in social media are still a minority. I am not even sure what proportion of MDs are amenable to communication via e-mail with their patients, though these data may be out there. So, for what seems to me as the majority of MDs who are not sold on e-mail, Twitter, Facebook, blogging or Quora, the value must not be that obvious. This makes me wonder if there are certain unifying characteristics of these docs, one being lack of perceived value of communication outside the profession across all media, including in-person contact.
I am friends with many docs on Twitter and in the blogosphere. The vast majority of them have shown themselves to be patient-centric, knowledgeable and collaborative, the kind of people I would not hesitate to send a loved one to. Yet, this is a skewed sample born out of a selection bias. These are the people who are interested and confident in their ability to communicate outside medicine. These are the people to whom medicine is a humanistic pursuit, where communities of patients and doctors strengthen the discussion of how to transform our system and the patient encounter. My guess is, and this is purely unscientific, that many of those who are skeptical of social media are also skeptical of communication itself, or just do not see the value of it in the equation of providing good patient care within the crushing time constraints of today's healthcare.
So my point is this: before social media tools can be expected to diffuse broadly into the medical community, the value of all communication needs to become clear to physicians in general. At this moment of increasing societal skepticism of science and of usefulness and integrity of the medical profession, against the backdrop of healthcare changes and increasingly unfiltered media noise, willingness and skills to communicate clearly may be as useful to today's doctors as a stethoscope. Once communication becomes the backbone of all medicine, tweets and blog posts are sure to start flowing freely from the fingers of physicians everywhere. And that will be good for the patients, the science and the healthcare system.
Wednesday, 26 January 2011
Webinar survey results
Last week I posted a survey link to gauge interest in and potential content for a webinar on how to review medical literature critically. I had a great response, and wanted to share the data with you.
The web page got 302 hits, resulting in 82 survey responses. This is a 27% rate of response, which certainly sets the results up to be biased and non-generalizable. But what the heck? I was looking to hear from people with some interest in this, not all-comers. So, here are the questions and the aggregated answers.
Q1: "I am thinking about creating a webinar based on some of the posts I have done on how to review medical literature. Would this be of interest to you?
R1: 82 people responded, of whom 81 (99%) answered "yes".
Q2: Are you a healthcare professional/researcher, an e-patient, or just an innocent bystander?
R2: 82 responses, 60 (72%) healthcare professionals/researchers, 5 (6%) e-patients, 17 (21%) innocent bystanders
Q3: Why do you feel the need to understand how to review medical literature
R3: This was a free text field, and I got 73 responses. Of these, many had to do with gaining a better understanding of the subject in order to help others (patients, clients, trainees) learn how to read and understand medical literature.
Q4: This question was only for those who responded "yes" to being a healthcare professional/researcher: Do you engage in journal peer review as a reviewer?
R2: Of 60 responses, only 7 (12%) were "yes".
Q5: Similar to Q4, this question was for only those who responded "yes" to Q4: Have you had formal training on how to be an effective peer reviewer?
R2: All 7 responded, of whom only 2 (29%) had formal training through a journal or a professional society, The remaining 5 (71%) have gained pertinent knowledge through reading about it. None of the responders got any reviewing courses during their medical training. Although the sample size is small, the responses are revealing and go along with my experience.
Q6: This question was targeted to only those responders who identified themselves as e-patients: How technical do you want the webinar information to get?
R2: All 5 e-patients answered this question, of whom 2 were comfortable with some degree of technicality, while the remaining 3 were comfortable with a greater degree of it.
Q7: This question was for all responders who expressed interest in having a webinar: Would you want one session or multiple sessions?
R2: Of the 80 responders, 21 (26%) felt that 1 session would suffice, 40 (50%) would be amenable to up to 3 sessions, and 11 (14%) would do up to 5 sessions. The remaining 8 (10%) of the responders chose "other", where their replies ranged from "no clue" to "as many as you see fit" to "let's start an ongoing discussion."
Q8: This was for those who would prefer a single session: How long should the session be?
R2: Of the 20 responses, 10 (50%) indicated 1 hour, while the majority of the rest indicated 2 hours.
Q9: If you are a part of an institution, do you think this would be of interest to your institution?
R9: 70 people responded, with 39 (56%) saying "yes" and 31 (44%) saying "no".
Q10: This was for those responding "yes" to Q9: What type of an institution are you a part of?
R10: All 39 people responded, and there was a range of institutions from medical schools to hospitals to government organizations to academic libraries. What was interesting here was that none of the "yes" responses to Q9 came from anyone in Biopharma or a professional organization or a patient advocacy organization. This I found surprising.
Overall, I am very pleased with the response. I am grateful to Janice McCallum (@janicemccallum on Twitter) for spreading the word to a lestserv of medical librarians. It certainly looks like there is enough interest in a webinar, and now I have to figure out how to execute one. If anyone has ideas, please, let me know in comments here or via e-mail.
Thanks again to all who took the time to respond!
The web page got 302 hits, resulting in 82 survey responses. This is a 27% rate of response, which certainly sets the results up to be biased and non-generalizable. But what the heck? I was looking to hear from people with some interest in this, not all-comers. So, here are the questions and the aggregated answers.
Q1: "I am thinking about creating a webinar based on some of the posts I have done on how to review medical literature. Would this be of interest to you?
R1: 82 people responded, of whom 81 (99%) answered "yes".
Q2: Are you a healthcare professional/researcher, an e-patient, or just an innocent bystander?
R2: 82 responses, 60 (72%) healthcare professionals/researchers, 5 (6%) e-patients, 17 (21%) innocent bystanders
Q3: Why do you feel the need to understand how to review medical literature
R3: This was a free text field, and I got 73 responses. Of these, many had to do with gaining a better understanding of the subject in order to help others (patients, clients, trainees) learn how to read and understand medical literature.
Q4: This question was only for those who responded "yes" to being a healthcare professional/researcher: Do you engage in journal peer review as a reviewer?
R2: Of 60 responses, only 7 (12%) were "yes".
Q5: Similar to Q4, this question was for only those who responded "yes" to Q4: Have you had formal training on how to be an effective peer reviewer?
R2: All 7 responded, of whom only 2 (29%) had formal training through a journal or a professional society, The remaining 5 (71%) have gained pertinent knowledge through reading about it. None of the responders got any reviewing courses during their medical training. Although the sample size is small, the responses are revealing and go along with my experience.
Q6: This question was targeted to only those responders who identified themselves as e-patients: How technical do you want the webinar information to get?
R2: All 5 e-patients answered this question, of whom 2 were comfortable with some degree of technicality, while the remaining 3 were comfortable with a greater degree of it.
Q7: This question was for all responders who expressed interest in having a webinar: Would you want one session or multiple sessions?
R2: Of the 80 responders, 21 (26%) felt that 1 session would suffice, 40 (50%) would be amenable to up to 3 sessions, and 11 (14%) would do up to 5 sessions. The remaining 8 (10%) of the responders chose "other", where their replies ranged from "no clue" to "as many as you see fit" to "let's start an ongoing discussion."
Q8: This was for those who would prefer a single session: How long should the session be?
R2: Of the 20 responses, 10 (50%) indicated 1 hour, while the majority of the rest indicated 2 hours.
Q9: If you are a part of an institution, do you think this would be of interest to your institution?
R9: 70 people responded, with 39 (56%) saying "yes" and 31 (44%) saying "no".
Q10: This was for those responding "yes" to Q9: What type of an institution are you a part of?
R10: All 39 people responded, and there was a range of institutions from medical schools to hospitals to government organizations to academic libraries. What was interesting here was that none of the "yes" responses to Q9 came from anyone in Biopharma or a professional organization or a patient advocacy organization. This I found surprising.
Overall, I am very pleased with the response. I am grateful to Janice McCallum (@janicemccallum on Twitter) for spreading the word to a lestserv of medical librarians. It certainly looks like there is enough interest in a webinar, and now I have to figure out how to execute one. If anyone has ideas, please, let me know in comments here or via e-mail.
Thanks again to all who took the time to respond!
Tuesday, 25 January 2011
Mirror neurons and the need for slow medicine
How long does it take for a silence to become uncomfortable? 5 seconds? 20 seconds? A minute? Students of education are taught to give a child roughly 20 seconds to answer a question posed to him. How long do teachers actually give? About 5 seconds, if that. Now sit there and count out 20 Mississippis and see what an astonishingly long time it seems. Why, what if a web page takes that long to load on your browser? This becomes a major technological tragedy for most of us. The point is that 20 seconds is a longer time than we appreciate.
Now, let's talk about empathy. Yes, empathy. This seeming non-sequitur has a solid connection. How do we like to experience empathy? Silent attentive listening is a great example of empathic engagement. When we talk with out friends about emotionally charged topics, we do not want them to respond with "yeah, yeah", and move on rapidly to the next topic, do we? So, empathy takes time and engagement. And when 20 seconds of silence seems like a long time, imagine it in a doctor's office, following a hard revelation or an emotional response by the patient. Can you? Are you counting the Mississippis?
Well, it is no wonder that doctors miss opportunities to express empathy to their patients. In a study from Canada, where oncologists were recorded during patient encounters, these doctors seized fewer than 1 in 4 opportunities to respond to their patients with empathy; the other 3 chances they squandered on discussing clinical information. And this is a pity, as is rightfully acknowledged by the investigator quoted in the article. His conjecture for why docs miss these opportunities to be empathic has to do with their apparently erroneous idea that it takes too much time, and his guidance is the following:
So, if the docs' intuition is correct, and empathy does mean non-detachment and time (after all 20 seconds represents 3% of a 10-minute appointment), how does the medical profession go about relishing and leveraging the other 3 opportunities for empathy instead of throwing them away? I agree with the point of the article that medical students should be taught empathic communication. At the same time, we learn by example, and if harried mentors continue to skirt these issues in the office because they are running two hours behind schedule already, the students will get the point loud and clear. The bigger issue is the incredible shrinking appointment, which is not only likely driving up healthcare costs and the frequency and intensity of testing, with its attendant adverse events, but is eroding the opportunity for a meaningful therapeutic relationship. After all, if the doctor herself provides a therapeutic benefit, is this not of utmost importance?
In short, this is another argument for slow medicine, an argument that should not be weakened by the detachment reasoning. My guess is that it is our biologic imperative as humans to exercise our mirror neurons avidly and often, and being forced to blunt their firing may be yet another path to demoralization. And is the medical profession not already demoralized enough?
Now, let's talk about empathy. Yes, empathy. This seeming non-sequitur has a solid connection. How do we like to experience empathy? Silent attentive listening is a great example of empathic engagement. When we talk with out friends about emotionally charged topics, we do not want them to respond with "yeah, yeah", and move on rapidly to the next topic, do we? So, empathy takes time and engagement. And when 20 seconds of silence seems like a long time, imagine it in a doctor's office, following a hard revelation or an emotional response by the patient. Can you? Are you counting the Mississippis?
Well, it is no wonder that doctors miss opportunities to express empathy to their patients. In a study from Canada, where oncologists were recorded during patient encounters, these doctors seized fewer than 1 in 4 opportunities to respond to their patients with empathy; the other 3 chances they squandered on discussing clinical information. And this is a pity, as is rightfully acknowledged by the investigator quoted in the article. His conjecture for why docs miss these opportunities to be empathic has to do with their apparently erroneous idea that it takes too much time, and his guidance is the following:
Showing empathy does not mean a doctor has to feel what his or her patient is feeling, Buckman says. Rather, it means acknowledging patients’ fears and other emotions.
“It is perfectly OK for the doctor to remain detached, but it is not OK to talk detached,” he says. “Acknowledging what a patient is feeling is not the same as feeling it yourself.”Well, I have to respectfully disagree. Here is the meaning of the word "empathy" from the trusted Merriam-Webster dictionary:
And in fact, looking to brain science to guide us on how we are wired to accomplish this, we realize that by definition empathy implies non-detachment, and, in fact, involves feeling what the other is feeling. Empathy is mediated by the so-called mirror neurons, residing in the cingulate gyrus of the brain. The great neurobiologist VS Ramachandran thinks that the discovery of these neurons is to the study of human behavior what the discovery of DNA was to biology. It has been said that mirror neurons help "dissolve the 'self vs. other' barrier." It is these neurons that make us feel others' pain, literally and figuratively. So, putting ourselves in the other person's shoes and "feeling what the patient is feeling" is truly the sine qua non of empathy.2: the action of understanding, being aware of, being sensitive to, and vicariously experiencing the feelings, thoughts, and experience of another of either the past or present without having the feelings, thoughts, and experience fully communicated in an objectively explicitmanner; also : the capacity for this
So, if the docs' intuition is correct, and empathy does mean non-detachment and time (after all 20 seconds represents 3% of a 10-minute appointment), how does the medical profession go about relishing and leveraging the other 3 opportunities for empathy instead of throwing them away? I agree with the point of the article that medical students should be taught empathic communication. At the same time, we learn by example, and if harried mentors continue to skirt these issues in the office because they are running two hours behind schedule already, the students will get the point loud and clear. The bigger issue is the incredible shrinking appointment, which is not only likely driving up healthcare costs and the frequency and intensity of testing, with its attendant adverse events, but is eroding the opportunity for a meaningful therapeutic relationship. After all, if the doctor herself provides a therapeutic benefit, is this not of utmost importance?
In short, this is another argument for slow medicine, an argument that should not be weakened by the detachment reasoning. My guess is that it is our biologic imperative as humans to exercise our mirror neurons avidly and often, and being forced to blunt their firing may be yet another path to demoralization. And is the medical profession not already demoralized enough?
Sunday, 23 January 2011
Top 5 this week
#5: Do private ICU rooms really reduce HAIs?
#4: Data mining: It's about research efficiency.
#3: To guideline or not to guideline, that is the ques...
#2: Reviewing medical literature, part 1: The study qu...
#1: A webinar survey -- Please, take this brief survey to help me gauge
interest in and content for a possible webinar on how to read and review
medical literature.
Thanks for visiting and reading!
#4: Data mining: It's about research efficiency.
#3: To guideline or not to guideline, that is the ques...
#2: Reviewing medical literature, part 1: The study qu...
#1: A webinar survey -- Please, take this brief survey to help me gauge
interest in and content for a possible webinar on how to read and review
medical literature.
Thanks for visiting and reading!
Friday, 21 January 2011
A webinar survey
Hi, folks,
I am conducting a survey to see how much interest there may be in a webinar on reviewing medical literature. This should take no more than 10 minutes of your time and would be enormously helpful to me to a). gauge interest and b). create appropriate content.
Thank you so much for doing this!
To get to the survey, click on this url: http://qtrial.qualtrics.com/SE/?SID=SV_bKEvpW0cEjYW5da
I am conducting a survey to see how much interest there may be in a webinar on reviewing medical literature. This should take no more than 10 minutes of your time and would be enormously helpful to me to a). gauge interest and b). create appropriate content.
Thank you so much for doing this!
To get to the survey, click on this url: http://qtrial.qualtrics.com/SE/?SID=SV_bKEvpW0cEjYW5da
Thursday, 20 January 2011
To guideline or not to guideline, that is the question in... pneumonia?
January 20, 2011
confounding by indication, EBM, guideline, hcap, methods, ventilator-associated pneumonia
No comments
Addendum 1/20/11, 1:27 PM
I want to add something to this, since I have been reflecting on the data more. It turns out that about 3/4 of all patients had an organism isolated felt to be causative of their pneumonia. Among these patients, over 80% in each group received empiric treatment that covered the pathogen. This means that 4 out of 5 patients in both groups received appropriate antibiotic coverage. What the authors skimmed over briefly is to talk about de-escalation. De-escalation is the guideline recommended strategy which entails reducing the spectrum of treatment after culture results become available to only those antibiotics that cover what has grown out. So, if, say, a patient is being empirically treated for Pseudomonas aeruginosa with double coverage, and the culture grows our MRSA and no Pseudomonas, the two anti-pseudomonal drugs should be stopped immediately. The investigators state that they did apply a de-escalation protocol, and that by day 3 50% and by day 5 75% were essentially de-escalated. The fact that they state this in the Discussion section makes me think that this was inserted in response to a reviewer. It is a pity that they did not include de-escalation in their stratified analysis, as it may be at least somewhat explanatory for the findings.
I always felt that there was something intangible and intuitive about my assessments of the critically ill for whom I cared. I could not always explain why I thought one particular patient was more ill than the next, but there was that little something that I must have noticed out of the corner of my eye, and if I tried too hard to focus on it, it would disappear like a puff of smoke. Yet, docs make these pre-conscious assessments all the time. And though these hints drive treatment choices, they are distinctly difficult to quantify scientifically.
A new paper that was just published in The Lancet Infectious Diseases online is a great illustration of what happens when our analyses fail to account for these intuitions. The phenomenon is referred to as "confounding by indication", and it is the perennial plague of observational clinical research. Just to summarize, the study was an observational study of guideline implementation for the treatment of healthcare-associated pneumonia among ICU patients. The central guideline was that for the choice of empiric antibiotics selection. The initial choice of antibiotics, even before the definitive results of cultures are available, is based on the clinician's best guess at what organism(s) may be causing the pneumonia. Among these severely ill patients, the risk of having a bug that is resistant to many antibiotics is higher than for patients who come from the community with pneumonia, and this propensity drives the recommendation for a broader antibiotic coverage for these cases. It has been shown by us and many others that missing this initial opportunity to cover the bug(s) adequately subjects patients to a doubling or even trebling of the risk of death, regardless of whether the coverage is broadened later to include the culprit organism(s).
Back to the study. The four academic medical center that participated in it enrolled 303 eligible patients, of whom 129 were treated with antibiotic combinations that comported with the guideline recommendations (guideline compliant treatment) and 174 received other combinations that did not fit the guideline recommendations (guideline non-compliant). To their surprise, the investigators discovered that 28-day survival was actually higher in the non-compliant group than in the compliant one. And even after doing a great job of adjusting for many potential factors that made the groups different, this paradoxical disparity persisted, with an overall near-doubling in the hazard of death at 28 days in the compliant as opposed to the noncompliant group. Now, this is a fine how-do-you-do! So, does this mean that the guideline is actually killing people by advocating broader coverage? Well, not so fast.
First, I have to acknowledge that I may be engaging in rescue bias right now. Having said this, taking biological plausibility into account, the findings are very likely explained by confounding by indication. Namely, the docs who choose, say, dual rather than single therapy against gram-negative bacteria may be pre-consciously incorporating some intangible patient data into their choices, data that are not well represented by either laboratory values or disease severity scoring systems. I know this is a bit "soft" and maybe even "touchy-feely", but ask any doc, and s/he will confirm this phenomenon.
On the other hand, to be fair and balanced, I do have to agree that there may be other explanations. These include the possibility that our guideline recommendations, never really prospectively validated, may be wrong. Perhaps there is something about the untoward effects of these broad spectrum regimens that is at play. Maybe it is as simple as the "no free lunch" principle, and that even in the situation of covering appropriately broadly, introducing additional drugs increases not only their benefits, but also the risks associated with them. Finally, I have to acknowledge the possibility that we just have no clue what any of this means because our understanding of how antibiotics work in the setting of these types of pneumonia is flawed.
Now, let's put all of this in the context of our multiple discussions about data and knowledge on this web site. Several factors suggest that my initial explanation is correct. The bulk of the evidence points to the fact that skimpy early coverage increases the risk of death. Also, over a century of understanding and the durability of the germ theory imply that antibiotics are important in treating serious bacterial infections. So, the pre-test probability of the validity of the finding in the paper is pretty low. This is not to say that the study should not inject caution and self-examination into how we treat severe pneumonia; it absolutely should! This is also a place where we definitely need well designed interventional studies to confirm (or debunk) what we think we know to be true. In the meantime, as we often intone on this blog, let us not throw the baby out with the bath water.
Disclosure: I have done a lot of work in this area, so I have a potential intellectual COI with the study. Also, at least some of my research has been funded by the manufacturers of some of the antibiotics included in the guidelines.
I want to add something to this, since I have been reflecting on the data more. It turns out that about 3/4 of all patients had an organism isolated felt to be causative of their pneumonia. Among these patients, over 80% in each group received empiric treatment that covered the pathogen. This means that 4 out of 5 patients in both groups received appropriate antibiotic coverage. What the authors skimmed over briefly is to talk about de-escalation. De-escalation is the guideline recommended strategy which entails reducing the spectrum of treatment after culture results become available to only those antibiotics that cover what has grown out. So, if, say, a patient is being empirically treated for Pseudomonas aeruginosa with double coverage, and the culture grows our MRSA and no Pseudomonas, the two anti-pseudomonal drugs should be stopped immediately. The investigators state that they did apply a de-escalation protocol, and that by day 3 50% and by day 5 75% were essentially de-escalated. The fact that they state this in the Discussion section makes me think that this was inserted in response to a reviewer. It is a pity that they did not include de-escalation in their stratified analysis, as it may be at least somewhat explanatory for the findings.
I always felt that there was something intangible and intuitive about my assessments of the critically ill for whom I cared. I could not always explain why I thought one particular patient was more ill than the next, but there was that little something that I must have noticed out of the corner of my eye, and if I tried too hard to focus on it, it would disappear like a puff of smoke. Yet, docs make these pre-conscious assessments all the time. And though these hints drive treatment choices, they are distinctly difficult to quantify scientifically.
A new paper that was just published in The Lancet Infectious Diseases online is a great illustration of what happens when our analyses fail to account for these intuitions. The phenomenon is referred to as "confounding by indication", and it is the perennial plague of observational clinical research. Just to summarize, the study was an observational study of guideline implementation for the treatment of healthcare-associated pneumonia among ICU patients. The central guideline was that for the choice of empiric antibiotics selection. The initial choice of antibiotics, even before the definitive results of cultures are available, is based on the clinician's best guess at what organism(s) may be causing the pneumonia. Among these severely ill patients, the risk of having a bug that is resistant to many antibiotics is higher than for patients who come from the community with pneumonia, and this propensity drives the recommendation for a broader antibiotic coverage for these cases. It has been shown by us and many others that missing this initial opportunity to cover the bug(s) adequately subjects patients to a doubling or even trebling of the risk of death, regardless of whether the coverage is broadened later to include the culprit organism(s).
Back to the study. The four academic medical center that participated in it enrolled 303 eligible patients, of whom 129 were treated with antibiotic combinations that comported with the guideline recommendations (guideline compliant treatment) and 174 received other combinations that did not fit the guideline recommendations (guideline non-compliant). To their surprise, the investigators discovered that 28-day survival was actually higher in the non-compliant group than in the compliant one. And even after doing a great job of adjusting for many potential factors that made the groups different, this paradoxical disparity persisted, with an overall near-doubling in the hazard of death at 28 days in the compliant as opposed to the noncompliant group. Now, this is a fine how-do-you-do! So, does this mean that the guideline is actually killing people by advocating broader coverage? Well, not so fast.
First, I have to acknowledge that I may be engaging in rescue bias right now. Having said this, taking biological plausibility into account, the findings are very likely explained by confounding by indication. Namely, the docs who choose, say, dual rather than single therapy against gram-negative bacteria may be pre-consciously incorporating some intangible patient data into their choices, data that are not well represented by either laboratory values or disease severity scoring systems. I know this is a bit "soft" and maybe even "touchy-feely", but ask any doc, and s/he will confirm this phenomenon.
On the other hand, to be fair and balanced, I do have to agree that there may be other explanations. These include the possibility that our guideline recommendations, never really prospectively validated, may be wrong. Perhaps there is something about the untoward effects of these broad spectrum regimens that is at play. Maybe it is as simple as the "no free lunch" principle, and that even in the situation of covering appropriately broadly, introducing additional drugs increases not only their benefits, but also the risks associated with them. Finally, I have to acknowledge the possibility that we just have no clue what any of this means because our understanding of how antibiotics work in the setting of these types of pneumonia is flawed.
Now, let's put all of this in the context of our multiple discussions about data and knowledge on this web site. Several factors suggest that my initial explanation is correct. The bulk of the evidence points to the fact that skimpy early coverage increases the risk of death. Also, over a century of understanding and the durability of the germ theory imply that antibiotics are important in treating serious bacterial infections. So, the pre-test probability of the validity of the finding in the paper is pretty low. This is not to say that the study should not inject caution and self-examination into how we treat severe pneumonia; it absolutely should! This is also a place where we definitely need well designed interventional studies to confirm (or debunk) what we think we know to be true. In the meantime, as we often intone on this blog, let us not throw the baby out with the bath water.
Disclosure: I have done a lot of work in this area, so I have a potential intellectual COI with the study. Also, at least some of my research has been funded by the manufacturers of some of the antibiotics included in the guidelines.








