Saturday, 2 April 2011

Invalidated Results Watch, Ivan?

My friend Ivan Oransky runs a highly successful blog called Retraction Watch; if you have not yet discovered it, you should! In it he and his colleague Adam Marcus document (with shocking regularity) retractions of scientific papers. While most of the studies are from the bench setting, some are in the clinical arena. One of the questions they have raised is what should happen with citations of these retracted studies by other researchers? How do we deal with this proliferation of oftentimes fraudulent and occasionally simply mistaken data?

A more subtle but no less difficult conundrum arises when papers cited are recognized to be of poor quality, yet they are used to develop defense for one's theses. The latest case in point comes from the paper I discussed at length yesterday, describing the success of the Keystone VAP prevention initiative. And even though I am very critical of the data, I do not mean to single out these particular researchers. In fact, because I am intimately familiar with the literature in this area, I can judge what is being cited. I have seen similar transgressions from other authors, and I am sure that they are ubiquitous. But let me be specific.

In the Methods section on page 306, the investigators lay out the rationale for their approach (bundles) by stating that the "ventilator care bundle has been an effective strategy to reduce VAP..." As supporting evidence they cite references #16-19. Well, it just so happens that these are the references that yours truly had included in her systematic review of the VAP bundle studies, and the conclusions of that review are largely summarized here. I hope that you will forgive me for citing myself again:
A systematic approach to understanding this research revealed multiple shortcomings. First, since all of the papers reported positive results and none reported negative ones, there is a potential for publication bias. For example, a recent story in a non-peer-reviewed trade publication questioned the effectiveness of bundle implementation in a trauma ICU, where the VAP rate actually increased directionally from 10 cases per 1,000 MV days in the period before to 11.9 cases per 1,000 MV days in the period after implementation of the bundle (24). This was in contradistinction to the medical ICU in the same institution, which achieved a reduction from 7.8 to 2.0 cases per 1,000 MV days with the same intervention (24). Since the results did not appear in a peer-reviewed form, it is difficult to judge the quality or significance of these data; however, the report does highlight the need for further investigation, particularly focusing on groups at heightened risk for VAP, such as trauma and neurological critically ill (25).             
Second, each of the four reported studies suffers from a great potential for selection bias, which was likely present in the way VAP was diagnosed. Since all of the studies were naturalistic and none was blinded, and since all of the participants were aware of the overarching purpose of the intervention, the diagnostic accuracy of VAP may have been different before as compared to after the intervention. This concern is heightened by the fact that only one study reports employing the same team approach to VAP identification in the two periods compared (23). In other studies, although all used the CDC-NNIS VAP definition, there was either no reporting of or heterogeneity in the personnel and methods of applying these definitions. Given the likely pressure to show measurable improvement to the management, it is possible that VAP classification suffered from a bias. 
Third, although interventional in nature, naturalistic quality improvement studies can suffer from confounding much in the same way that observational epidemiologic studies do. Since none of the studies addressed issues related to case mix, seasonal variations, secular trends in VAP, and since in each of the studies adjunct measures were employed to prevent VAP, there is a strong possibility that some or all of these factors, if examined, would alter the strength of the association between the bundle intervention and VAP development. Additional components that may have played a role in the success of any intervention are the size and academic affiliation of the hospital. In a study of interventions aimed at reducing the risk of CRBSI, Pronovost et al. found that smaller institutions had a greater magnitude of success with the intervention than their larger counterparts (26). Similarly, in a study looking at an educational program to reduce the risk of VAP, investigators found that community hospital staff were less likely to complete the educational module than the staff at an academic institution; in turn, the rate of VAP was correlated with the completion of the educational program (27). Finally, although two of the studies included in this review represent data from over 20 ICUs each (20, 22), the generalizability of the findings in each remains in question. For example, the study by Unahalekhaka and colleagues was performed in the institutions in Thailand, where patient mix and the systems of care for the critically ill may differ dramatically from those in the US and other countries in the developed world (22). On the other hand, while the study by Resar and coworkers represents a cross section of institutions within the US and Canada, no descriptions are given of the particular ICUs with respect to the structure and size of their institutions, patient mix or ICU care model (e.g., open vs. closed; intensivists present vs. intensivists absent, etc.) (20). This aggregate presentation of the results gives one little room to judge what settings may benefit most and least from the described interventions. The third study includes data from only two small ICUs in two community institutions in the US (21), while the remaining study represents a single ICU in a community hospital where ICU patients are not cared for by an intensivist (23).  Since it is acknowledged that a dedicated intensivist model leads to improved ICU outcomes (28, 29), the latter study has limited usefulness to institutions that have a more rigorous ICU care model.
OK, you say, maybe the investigators did not buy into my questions about the validity of the "findings." Maybe not, but evidence suggests otherwise. In the Discussion section on page 311 they actually say
While the bundle has been published as an effective strategy for VAP prevention and is advocated by national organizations, there is significant concern about its internal validity.
And guess what they cite? Yup, you guessed it, the paper excerpted above. So, to me it feels like they are trying to have it both ways -- the evidence FOR implementing the bundle is the same evidence AGAINST its internal validity. Much like Bertrand Russell, I am not that great at dealing with paradoxes. Will this contradiction persist in our psyche, or will sense prevail? Perhaps Ivan and Adam need to start a new blog: Invalidated Results Watch. Oh? Did you say that peer review is supposed to be the answer to this? Right.  
    

Friday, 1 April 2011

Another swing at the windmill of VAP

Sorry, folks, but I have been so swamped with work that I have been unable to produce anything cogent here. I see today as a gift day, as my plans to travel to SHEA were foiled by mother nature's sense of humor. So, here I am trying to catch up on some reading and writing before the next big thing. To be sure, I have not been wasting time, but have completed some rather interesting analyses and ruminations, which, if I am lucky, I will be able to share with you in a few weeks.

Anyhow, I am finally taking a very close look at the much touted Keystone VAP prevention study. I have written quite a bit about VAP prevention here, and my diatribes about the value proposition of "evidence" in this area are well known and tiresome to my reader by now. Yet, I must dissect the most recent installment in this fallacy-laden field, where random chance occurrences and willful reclassifications are deemed causal of dramatic performance improvements.

So, the paper. Here is the link to the abstract, and if you subscribe to the journal, you can read the whole study. But fear not, I will describe it to you in detail.

In its design it was quite similar to the central line-associated blood stream infection prevention study published in the New England Journal in 2006, and similarly the sample frame included Keystone ICUs in Michigan. Now, recall that the reason this demonstration project happened in Michigan is because of their astronomical healthcare-associated infection (HAI) rates. Just to digress briefly, I am sure you have all heard of MRSA; but have you heard of VRSA? VRSA stands for vancomycin-resistant Staphylococcus aureus, MRSA's even more troubling cousin, vancomycin being a drug that MRSA is susceptible to. Now, thankfully, VRSA has not yet emerged as an endemic phenomenon, but of the handful of cases of this virtually untreatable scourge that has been reported, Michigan has had plurality of them. So, you get the picture: Michigan is an outlier (and not in the desirable direction) when it comes to HAIs.

Why is it important to remember Michigan's outlier status? Because of the deceptively simple yet devilishly confounding concept of regression to the mean. The idea is that in an outlier situation, at least some of the effect is due to random luck. Therefore, if the performance of an extreme outlier is measured twice, the second time it will be closer to the population mean just by pure luck alone. But I do not want to get too deeply into this somewhat muddy concept right now -- I will reserve a longer discussion of it for another post. For now I would like to focus on some of the more tangible aspects of the study. As usual, two or three features of the study design reduce substantially the likelihood that the causal inference is correct.

First feature is the training period. Prior to the implementation of the protocol, which by the way consisted of the famous VAP bundle, which we have discussed on this blog ad nauseam, there was intensive educational training of the personnel on a "culture of change", as well as the proper definitions of the interventions and outcomes. It is at this time that the "trained hospital infection prevention personnel" were intimately focused on the definition of VAP that they were using. And even though the protocol states that the surveillance definition of VAP would not change throughout the study period, what are the chances that this intensified education and emphasis did not alter at least some of the classification practices?

Skeptical? Good. Here is another piece of evidence supporting my stance. A study from Michael Klompas from Harvard examined inte-rater variability in the assessment of VAP looking at the same surveillance definition applied in the Keystone (and many other) study. Here is what he wrote:
Three infection control personnel assessing 50 patients for VAP disagreed on 38% of patients and reported an almost 2-fold variation in the total number of patients with VAP. Agreement was similarly limited for component criteria of the CDC VAP definition (radiographic infiltrates, fever, abnormal leukocyte count, purulent sputum, and worsening gas exchange) as well as on final determination of whether VAP was present or absent.
And here is his conclusion:
High interobserver variability in the determination of VAP renders meaningful comparison of VAP rates between institutions and within a single institution with multiple observers questionable. More objective measures of ventilator-associated complication rates are needed to facilitate benchmarking and quality improvement efforts. 
Yet, the Keystone team writes this in their Methods section:
Using infection preventionists minimized the potential for diagnosis bias because they are trained to conduct surveillance for VAP and other healthcare-associated infections by using standardized definitions and methods provided by the CDC in its National Healthcare Safety Network (NHSN).
Really? Am I cynical to invoke circular reasoning here? Have I convinced you yet that CAP diagnosis is a moving target? And as such it can be moved by cognitive biases, such as the one introduced by the pre-implementation training of study personnel? No? OK, consider this additional piece from the Keystone study. The investigators state that "teams were instructed to submit at least 3 months of baseline VAP data." What they do not state is whether this was a retrospective collection or a prospective one, and this matters a little. First, retrospective reporting in this case would be a lot more representative of what has been, since these rates of VAP are already recorded for posterity and cannot presumably be altered. On the other hand, if the reporting is prospective, I can still conceive of ways to introduce a bias into this baseline measure. Imagine, if you will, that you are employed by a hospital that is under scrutiny for a particular transgression, and that you know the hospital will look bad if you do not demonstrate improvement following a very popular and "common-sense" intervention. Might you be a tad more liberal with identifying these transgressive episodes in your baseline period that after the intervention has been instituted? This is a subtle, yet all too real conflict of interest, which, as we know so well, can introduce a substantial bias into any study. Still don's believe me? OK, come to my office after school and we will discuss. In the meantime, let's move on.

The next nugget is in the graph in Figure 1, where VAP trends over the pre-specified time periods are plotted (you can find the identical graph in this presentation on slide #20). Look at the mean, rather than the median line. (The reason I want you to look at the mean is that the median is zero, and therefore not credible. Additionally, if we want to assess the overall impact of the intervention, we need to be embracing the outliers, which the median ignores). What is tremendously interesting to me is that there is a precipitous drop in VAP during the period called "intervention", followed by much smaller fluctuations around the new mean across the subsequent time periods. This to me confirms the high probability of reclassification (and Hawthorne effect), rather than an actual improvement in VAP rates, as the cause of the drop.

Another piece of data makes me think that it was not the bundle that "did it." Figure 2 in the paper depicts the rates of compliance with all 5 of the bundle components in the corresponding time periods. Again, here as in the VAP rates graph, the greatest jump in adherence to all 5 strategies is observed in the intervention period. However, there is still a substantial linear increase in this metric between the intervention period and through to 25-27 months period. Yet, looking back at the VAP data, no such robust commensurate reduction is observed. While this is somewhat circumstantial, it makes me that much more wary of trusting this study.

So, does this study add anything to our understanding of what bundles do for VAP prevention? I would say not, and it actually muddies the waters. What would have been helpful to see is whether any of the downstream outcomes, such as antibiotics administration, time on the ventilator and length of stay were impacted. Without impacting these outcomes, our efforts are Quixotic, merely swinging at windmills, mistaking them for a real threat.


       

          

Friday, 4 March 2011

Is this double dipping? A new bipartisan House bill on oncology reimbursements

Here is another gem from the House of Representatives: a bipartisan bill to increase Medicare reimbursements to community oncology practices. While at first glance this seems like a reasonable idea, this detail is puzzling:
The so-called "prompt pay" legislation excludes certain discounts extended to wholesalers when calculating Medicare reimbursements and is strongly supported by oncologists.
Confused? Met too. Here is how I understand it. Many community oncology practices have set up infusion clinics, where they administer intravenous chemotherapy on site to their patients. To stock these infusion centers they deal with drug manufacturers and distributors to purchase the drugs at wholesale prices. The bigger the buy, the bigger the manufacturer discount. To the best of my knowledge these discounts are proprietary information, guarded like state secrets. Yet despite these discounts, the clinics charge Medicare a premium for the drugs themselves as well as for the service of administering. The way this legislation looks to me is that it will completely eliminate any reduction in reimbursement related to these discounts. Double dipping, anyone?

Now, I have many friends who are oncologists, and this is really not a slur against them. But these infusion clinics have always represented a cash cow for these practices. And who would not want to have a steady source of income to maintain a robust practice and have some money left over for a life? Again, this is not an indictment of community oncology practices. If, however, one takes an external perspective, this bill becomes something of an anathema to improving efficiency of healthcare delivery. If the reimbursement rates for administering these already exorbitantly expensive drugs improve further, will it not become even more difficult for an oncologist to tread the fine line of the conflict of interest between treatment only when it is in the patient's best interest and treatment for income optimization? Again, I want to point out that I am not singling out oncologists, as it is a part of the human condition to rationalize our selfish decisions by putting them in an altruistic light. And given the amount of uncertainty about who might respond to these drugs, it is easy to convince oneself that a trial of a therapy may be a reasonable idea, with the reimbursements providing a nudge in that direction.

A couple of quotes from the sponsors of the legislation are also worth reprinting:

"On any legislation today, you have to find a way to pay for it. And like any legislation, that's an issue with this one," Whitefield said. 

"But to be truthful, because of the oncologist groups and patient groups and others, we think that there may be some provisions in the healthcare bill that passed last year that we may be able to utilize some of those funds for this. All of it's about healthcare, and if we can convince people that this is more important than the others then we can do it."     
On any legislation today? You mean it has not always been like this? I guess we have all gotten used to credit as a life style, and now it is time to pay the piper.

Now, what about this: "Because of the oncologist groups and patients groups and others..."? Are they saying what I think they are saying? That because groups are likely to benefit are saying that this is vital, it is in fact vital? I also have to wonder who those "others" are. Hmmm, I wonder...

And this: "All of it's about healthcare, and if we can convince people that this is more important than others then we can do this." OK, so the statement is so grammatically abominable and non-sensical that I can interpret it any way I like. And it seems to me that they are implying that increasing these already hefty reimbursements is more important than stuff like paying for prenatal care and immunizations to the poor? And other essential services to the Medicare population? Well, if this is not the a poster child for why we need to be articulating the value of healthcare, I do not know what is.  

Thursday, 3 March 2011

Why easy is not always good

My mother-in-law is a typesetter. She will not read a book unless it is not only appealing in its content, but also pleasing to the eye. When I was in medical school, she did quite a bit of work for medical textbook publishers. Comparing books typeset by her to what I was grinding through on a daily (and nightly basis) incensed her: unwieldy tables appearing three pages away from the corresponding text, small letters crammed to capacity onto oversized pages, few illustrations -- all baffling, annoying (and easily fixable) transgressions against readability. Yet, like all budding docs of all generations, I plowed through these morasses of knowledge without giving its readability much thought -- this was just what you did to get to your goal.

Yesterday I was listening to a program where the author Amy Chua was interviewed about her (ahem) embattled autobiography Battle Hymn of the Tiger Mother. Ms. Chua, though evenly humored throughout the interview, was on the defensive nearly the entire time, explaining how the intent of her opus has been grossly misunderstood by the public, thanks to attacks by critics on her parenting style. And granted, looking at the book as a parenting manual through the prism of our Western parenting norms is a bit disturbing. Yet putting its events in a culturally appropriate context, as well as looking at the content as a narrative rather than a guide, leads to completely different conclusions.

Why am I bringing up Amy Chua's interview after talking about my conquest of the unreadable? Well, it seems that ease is what we have come to expect from everything. What I mean by this is that not only do we expect easily readable texts, but we also expect people to present themselves in such a way as to make it easy for us to like them. Why else change your appearance through life-threatening eating disorders and grueling surgeries, get coached on how to make friends and influence people, and comment on how unlikable some of our female politicians are? Is this not a triumph of form over substance?

Amy Chua clearly bucks this trend in her book and is paying the price. But what worries me is that we are all paying a price. By creating another false dichotomy of "she is nice" or "he is nasty", we have eschewed a more realistic view of our human foibles. We are all nice sometimes and nasty at others. Yet this dichotomy has proven supremely fruitful to our political discourse, where for 30 years this new reality has been taking root. And it has born fruit, so that now people who do not hold similar opinions to ours are summarily dismissed as "nasty" or idiotic, and we are satisfied to surround ourselves with "nice" like-minded sycophants. How primitive it renders our political and social interactions!

Ms. Chua's immigrant parents' philosophy resonated with my upbringing. Coming from lands of uncertainty and deprivation, as immigrants, our parents subscribed to Maslow's pyramid and taught us that economic security trumped everything else. This is why only certain career choices were acceptable, while others were relegated to the back burner of a hobby. These choices were not about ease, but about doing what we were taught was the right thing. As John Adams said:
I must study Politicks and War that my sons may have liberty to study Mathematicks and Philosophy. My sons ought to study Mathematicks and Philosophy, Geography, natural History, Naval Architecture, navigation, Commerce, and Agriculture, in order to give their Children a right to study Painting, Poetry, Musick, Architecture, Statuary, Tapestry, and Porcelaine.
We all set priorities, and some of them may not be easy. I myself still read books even if they are not all that well presented; my priorities are content and writing style, though, to be sure, I do not frown upon the beauty of the visual form. I even enjoy characters who in, their multidimensionality, are a challenge to like. And I have learned in the rest of my life to enjoy people who do not necessarily hold easy or quick appeal for me, yet in the long run prove to add unimaginable richness to my life. Nietzsche coined the famous quote "What does not break you will make you stronger." In all aspects of our lives, while, based on Nietzsche's statement, adversity is a sufficient but not necessary road to strength, pushing ourselves a little bit out of our stuporous ease may prove to be one timely remedy.

The value of a test

Reading this vintage paper on C diff from the Archives of Pediatric and Adolescent Medicine, I came upon this irresistible conclusion:

Priceless!

Quality or value? A measure for the 21st century

Fascinating, how in the same week two giants of evidence-based medicine have given such divergent views on the future of quality improvement. Here (free subscription required), Donald Berwick, the CMS administrator and founder and former head of the Institute for Healthcare Improvement, emphasizes the need for quality as the strategy for success in our healthcare system. But here, one of the fathers of EBM, Muir Gray, states that quality is so 20th century, and we need instead to shine the light on value. So, who is right?

Well, let's define the terms. The Merriam-Webster dictionary defines quality as "the degree of excellence." The same source tells us that value is "a fair return or equivalent in goods, services or money for something exchanged." To me "value" is a holistic measure of cost for quality, painting a fuller picture of the investment vis-a-vis the returns on this investment. What do I mean by that?

Simply put, the idea behind value is to establish what is a reasonable amount to pay for a unit of quality. Let's take my used 1999 VW Passat as an example. If my mechanic tells me that it needs to have some hoses replaced, and it will cost me under $100, and the car will run perfectly, I will consider that to be a good value. However, if my transmission has fallen out in the middle of Brookline Ave. in Boston (really happened to me once, many years ago and with a different car), and it will cost me $5,000 to fix, I may say that the value proposition is just not there, particularly given that the car itself is worth much less than $5,000. Given that my budget is not unlimited, I have to make trade-off decisions about where to put my money, so I may instead spend the money on another used Passat that has good prospects.    

But in medicine, we routinely avoid thinking about value. There seems to be an overall impression that if it out there on the market, and especially if it is new, it is good and I am worth all of it. This impression is further enabled by the fact that CMS has no statutory power to make decisions based on value of interventions -- they are legislatively mandated to turn a blind eye to the costs. Does this make sense? How toothless is our comparative effectiveness effort likely to be if it has to ignore half of the story?

Let us now look at my favorite sticky wicket, ventilator-associated pneumonia, or VAP. Now, the IHI bundle aimed at eliminating VAP consists of 5 points of intervention: 1). semi-recumbent positioning, 2). daily screen for readiness to get off mechanical ventilation, 3). daily sedation vacation, 4). prophylaxis against GI bleeding, and 5). prevention of clots. As I have mentioned before elsewhere, adherence of 95% to all these measures is deemed compliance and may be ultimately used as a quality measure by payers to determine levels of reimbursement. And while each of these interventions is basically "motherhood and apple pie", applying them blindly and in toto to 95% of intubated patients may be a strategy for disaster. But what is even clearer is that, in order to implement this and all of the other quality improvement strategies, systems need to be put in place that will safeguard against failing to implement these quality measures. The time and resource expenditures needed to institute and maintain these systems, which have not been described in great enough detail as far as I am concerned, have never been quantified. So, what we are left with is a bunch of interventions that, while looking OK individually in clinical trials (until you really start looking at them critically), are likely providing small, if any, gains in quality at the margins, whose investment-return equation has not even been disclosed, let alone balanced. And because budgets are necessarily limited, as are clinicians' time and cognitive capacities, we need to select a sensible menu of interventions from this practically unlimited feast.

This is the quality conundrum, a clear case of chasing our tails to achieve perfection at the expense of good enough. And while no one in their right mind will argue with the language of improved quality in healthcare, I do think that Muir Gray and his camp are on to something that has been a long time coming. At this time of shrinking budgets, competing priorities and tightening resources, does it not make sense to look at value as a package deal, rather than merely at quality in isolation from its context? Instead of being bombarded by ever-increasing volume of quality measures coming from many directions, would it not be more sensible to prioritize these interventions based on the value that they bring rather than merely on their projected outcomes benefits, so frequently estimated based on data that have very little applicability to the real world? Let's start asking the question: how much quality and at what price? Without paying attention to this critical balance, we will not only bankrupt the system, but also worsen outcomes paradoxically, as we continue to overwhelm clinicians with infinite minutia that may or may not be generating helpful outcomes.

So, in my book, Muir Gray: score; Berwick: keep trying.            

Sunday, 27 February 2011

Friday, 25 February 2011

Guidelines: What really constitutes level I evidence?

There has been some interesting buzz in the blogosphere about where evidence-based guideline recommendations come from, and I wanted to add a little fuel to that fire today.

As you know, I think a lot about the nature of evidence, about the "science" in clinical science, and about pneumonia, specifically ventilator-associated pneumonia or VAP. Last week I wrote here and here about a specific recommended intervention to prevent VAP consisting of semi-recumbent, as opposed to supine, positioning. This recommendation, one of 21 maneuvers aimed at modifiable risk factors for VAP, had level I evidence behind it. Given my recent deconstruction of this level I evidence, consisting of a single unblinded RCT in a single academic urban center in Spain, and given that we already know that level I data represent a very small proportion of all the evidence behind guideline recommendations, I got curious about this level I stuff. How is level I really defined? Is there a lot of room for subjective judgment? So, I went to the source.

In its HAP/VAP guideline, the ATS and IDSA committee define the levels of evidence in the following way:
Level I (high)
Level II (moderate) 








Level III (low)
     Evidence comes from well conducted, randomized controlled trials


Evidence comes from well designed, controlled trials without randomization (including cohort, patient series, and case-control studies). Level II studies also include any large case series in which systematic analysis of disease patterns and/or microbial etiology was conducted, as well as reports of new therapies that were not collected in a randomized fashion

Evidence comes from case studies and expert opinion. In some instances therapy recommendations come from antibiotic susceptibility data without clinical observations
So, well conducted, randomized controlled trials. But what does "well conducted" mean? Seems to me that one person's well conducted may be another person's garbage. Well, I went to the text of the document for clarification:
The grading system for our evidence-based recommendations was previously used for the updated ATS Community-acquired Pneumonia (CAP) statement, and the definitions of high-level (Level I), moderate-level (Level II), and low-level (Level III) evidence are summarized in Table 1 (8). 
OK, then. We have to go to reference #8, or the CAP guideline to get to the bottom of the definition. And here is what that document states:
Therefore, in grading the evidence supporting our recommendations, we used the following scale, similar to the approach used in the recently updated Canadian CAP statement (46): Level I evidence comes from well-conducted randomized controlled trials; Level II evidence comes from well-designed, controlled trials without randomization (including cohort, patient series, and case control studies); Level III evidence comes from case studies and expertopinion. Level II studies included any large case series in which systematic analysis of disease patterns and/or microbial etiology was conducted, as well as reports of new therapies that were not collected in a randomized fashion. In some instances therapy recommendations come from antibiotic susceptibility data, without clinical observations, and these constitute Level III recommendations.
Again, we are faced with the nebulous "well-conducted" descriptor with no further defining guidance on how to discern this quality. I resigned myself to going to the next source citation, #46 above, the Canadian CAP statement:
We applied a hierarchical evaluation of the strength of evidence modified from the Canadian Task Force on the Periodic Health Examination [4]. Well-conducted randomized, controlled trials constitute strong or level I evidence; well-designed controlled trials without randomization (including cohort and case-control studies) constitute level II or fair evidence; and expert opinion, case studies, and before-and-after studies are level III (weak) evidence. Throughout these guidelines, ratings appear as roman numerals in parentheses after each recommendation.
Another "well-conducted" construct, another reference, another wild goose chase. The reference #4 above clarified the definition for me thus:
OK, so, now we have "at least one properly randomized controlled trial." So, having gotten to the origin of this broken telephone game, it looks like proper randomization trumps all other markers for a well-done trial. The price of such neglect is giving up generalizability, confirmation, appropriate analyses, and many other important properties that need to be evaluated before stamping the intervention with a seal of approval. 

And this is just one guideline for one syndrome. The bigger point that I wanted to illustrate is that, even though we now know that only 14% of all IDSA guideline recommendations have so-called level I evidence behind them, what is dubious is the value and validity of assigning this highest level of evidence to these recommendations, given the room for subjectivity and misclassification. So, what does all of this mean? Well, for me it means no foreseeable shortage of fodder for blogging. But for our healthcare policy and our public's health? Big doo-doo.

Thursday, 24 February 2011

New treatments: What benefits at what costs

Yesterday brought quite a bit of press coverage to a small biotechnology company in Cambridge called Vertex. All this attention was spurned by their gene therapy trial results in cystic fibrosis. The treatment, aimed at a genetic mutation present in about 4% of all CF sufferers, was able to improve the volume that a patient can force out of his lungs in 1 second by over 10%, from about 65% to 75%. Matthew Herper of Forbes on his blog, while being duly impressed by the results, also cautioned that the annual price tag for this medicine is likely to reach $250,000 per patient. So, what does all of this mean in the context of our ongoing national discussion about the value of therapies? Well, let's break things down a bit.

First, let's talk about CF. This is a genetic disorder that essentially makes mucus very sticky. Among its many effects, in its most familiar manifestation this mucus plugs up the airways making it difficult to breathe and predisposing the person to frequent and serious lung infections. When I was a resident back in the early '90s, I remember a devastating case of a young man in his late teens with CF whom we all knew so well from his frequent admissions for exacerbations. Though he was pretty high on the lung transplant list, he ended up succumbing to a devastating pneumonia in our ICU, leaving behind a devoted sister who had been fortunate enough to benefit from a transplant several years earlier. This was a typical course in those days: a brief life punctuated by frequent exacerbations, hospitalizations, antibiotics, gastrointestinal complications, and early death in the second or at best third decade of life with very little hope of procreation. Over the last 20 years things have changed dramatically in the treatment of CF: fewer exacerbations, much lengthened life expectancy and a good chance of having children. Yet we cannot attribute most of these changes to dramatic new breakthrough therapies. To be sure, while there have been tweaks to how we give antibiotics and how pancreatic enzymes are administered to replace the digestive enzymes that the pancreas in CF is unable to produce, most of the progress can be attributed to the increased attention to detail and the advent of almost ruthless care coordination at specialty centers. As a Fellow in the '90s I participated in a clinic where CF patients were transitioning from care by pediatric Pulmonologists to that by adult doctors. The CF specialist running this clinic did not only know all of his patients and their family members by names, but was available 24/7 to them and to his staff for consultation. This is the kind of dedication and vigilance necessary to improve the outcomes in CF.

Now, let's talk about the lesion addressed in the Vertex trial. The type of chronic lung disease caused by CF is called "obstructive." Simply put, it makes exhaling the air in the lungs difficult to do. On lung testing one manifestation of obstruction is the amount of air one is able to force out of his lungs in the first second of the effort, and this is called the FEV1, or forced expiratory volume in 1 second. Another important measure of the degree of obstruction is the amount of air that this volume expired in 1 second represents as a proportion of all of the air in the lungs that can be expired, known as the FVC or forced vital capacity. We say that if the FEV1/FVC ratio is under 75%, then obstruction is present. The size of FEV1 helps us understand how bad the obstruction is.

With this as a background, the primary outcome in many obstructive lung disease trials is the improvement in the FEV1. In the specific trial discussed, the average starting FEV1 in the intervention group was about 65%, which falls in the mild-to-moderate category of obstruction. What this means in terms of symptoms can vary widely. The 10% absolute improvement seen in the intervention group resulted in the average FEV1 of about 75% after treatment, definitely representing fairly mild obstruction (generally FEV1 over 80% is considered to be in the normal range). And this truly is impressive. However, equally interesting is the information that is not in the press coverage, largely based on press releases and sound bites from company executives, since the peer reviewed study is not available at this point. We are not told, but led to assume that, the control group started out on average with a similar deficit in lung function. We are informed that the treatment patients were 55% less likely than placebo patients to have an exacerbation of their disease, yet we do not know what the absolute numbers are; that is we are not told what proportion in each group had an exacerbation, how frequently or how severely. So, this 55%, in the absence of context, while an attention grabber, is not a substantive number. Herper does tell us that there was a remarkable difference in the weight gain (a desirable outcome in the CF population), on average 6.8 lb in the treatment vs. 0.9 lb in the placebo group. This is truly impressive, though it would be even more so if I knew that the trial was double blind, a piece of information I did not notice in any of the reports. Some of the reports have also alluded to symptomatic improvement in shortness of breath, though nowhere did I see this quantified.

The most important piece of data, however, is conspicuously absent from all the stories. What is the proportion of patients who responded to therapy? Why is this important? Well, we know that far from everyone responds to every treatment that they ostensibly qualify for; this is referred to as the heterogeneous treatment effect, or HTE. It is very likely that the 10% improvement in the FEV1 represents at once an inflated estimate referent to those non- or under-responders and a muted one for those patients with a terrific response. The question of a minimal clinically significant change in the FEV1 has haunted the lung trials community for a long time now. Yet, without setting some threshold for a minimum FEV1 improvement that correlates with a meaningful improvement in symptoms, one cannot quantify how well the drug works and hence articulate its value. This is crucial when trying to justify the ostensibly exorbitant price tag anticipated for this drug. How many patients will we need to treat in order to have one of them respond meaningfully with an improvement not just in a laboratory number, but also in their lives? If this targeted drug produces a desirable response even in 50% of all patients with the specific mutation it targets, then it means that we need to spend $500,000 annually to obtain a meaningful improvement in symptoms in one CF patient. But what if it only works this way in 20%? Then we will need to treat 5 patients with this drug to obtain 1 meaningful response at the price of $1.25 million annually. This becomes a bit more daunting, particularly given that the costs will have to be covered through some kind of public or pooled funds and given that this is one of many therapies in the pipeline likely to come with a similar conundrum.

I am not implying that improving a single life is not worth $1.25 million annually. In fact, it may well be a bargain. My point is that these are the serious discussions we need to have as a society, so that when the time comes to make these choices, the discussion will not be subverted by a few loud voices sensationalizing "death panel" slogans. Manufacturers need to know that they should disclose full data, not just selective tidbits that highlight benefits only, but also those difficult pieces of information that shed light on their costs. On our part, we need to understand the gargantuan effort and resources these companies expend to tame these elusive wild therapies that hold so much more promise in the abstract than they end up embodying.

We tread a fine line here. Information and how we assimilate it are the next frontier for cogent decision making. We need to get educated about this now because this train is leaving the station regardless of how we feel about it.                            

Tuesday, 15 February 2011

The rose-colored glasses of early trial termination

The other day I did a post on semi-recumbent positioning to prevent VAP. The point I wanted to make was that an already existing quality measure for a condition that is well on its way to becoming a CMS "never event" is based on one unreplicated single-center small unblinded randomized controlled trial that was terminated early for efficacy. In my post I cited several issues with the study that question its validity. Today I want to touch upon the issue of early termination, which in and of itself is problematic.

What is early termination? It is just that: stopping the trial before enrolling the pre-planned number of subjects. First, it is important to be explicit in the planning phases about how many subjects will need to be enrolled. This is known as the power calculation and is based on the anticipated effect size and the uncertainty in this effect. Termination can happen for efficacy (the intervention works so splendidly that it becomes unethical not to offer it to everyone), safety (the intervention is so dangerous that it becomes unethical to offer it to anyone) or for other reasons (e.g., the recruitment is taking too long, etc.).

Who makes the decision to terminate early and how is the decision made? Well, under the best of circumstances, there is a Data Safety Monitoring Board, a body that is specifically in place to look at the data at certain points in the recruitment process and look for certain pre-specified differences between groups. This DSMB is fire-walled from both the investigators and the patients. The interim looks at the data  should be pre-specified by the protocol also, as the number of these looks actually influences the initial power calculation, since the more you look, the more differences you are likely to find by chance alone.

So, without going into too much detail on these interim looks, understand that they are not to be taken lightly, and their conditions and reporting require full transparency. To their credit, the semi-recumbent position investigators reported their plan for one interim analysis upon reaching 50% enrollment. Neither the Methods section nor the Acknowledgements, however, specify who was the analyst and the decision-maker. Most likely it was the investigators themselves that ended up taking the look and deciding on the subsequent course of action. And this itself is not that methodologically clean.

Now, let's talk about one problem early termination. This gargantuan effort led by the team from McMaster in Canada and published last year in JAMA sheds the needed light on what had been suspected before: early termination leads to inflated effect estimates. The sheer massiveness of the work done is mind boggling -- over 2,500 studies were reviewed! The investigators elegantly paired meta-analyses of truncated RCTs with meta-analyses of matched but nontruncated ones, and compared the magnitude of the inter-group differences between the two categories of RCTs. Here is one interesting tidbit (particularly for my friend @ivanoransky):
Compared with matching nontruncated RCTs, truncated RCTs were more likely to be published in high-impact journals (30% vs 68%, P<.001).
But here is what should really grab the reader:

Of 63 comparisons, the ratio of RRs was equal to or less than 1.0 in 55 (87%); the weighted average ratio of RRs was 0.71 (95% CI, 0.65-0.77; P <.001)(FIGURE2). In 39 of 63 comparisons (62%), the pooled estimates for nontruncated RCTs were not statistically significant. Comparison of the truncated RCTs with all RCTs (including the truncated RCTs) demonstrated a weighted average ratio of RRs of 0.85; in 16 of 63 comparisons (25%), the pooled estimate failed to demonstrate a significant effect. [Emphasis mine]
The authors went on to conclude the following:

In this empirical study including 91 truncated RCTs and 424 matching nontruncated RCTs addressing 63 questions, we found that truncated RCTs provide biased estimates of effects on the outcome that precipitated early stopping. On average, the ratio of RRs in the truncated RCTs and matching nontruncated RCTs was 0.71. This implies that, for instance, if the RR from the nontruncated RCTs was 0.8 (a 20% relative risk reduction), the RR from the truncated RCTs would be on average approximately 0.57 (a 43% relative risk reduction, more than double the estimate of benefit). Nontruncated RCTs with no evidence of benefit—ie, with an RR of 1.0—would on average be associated with a 29% relative risk reduction in truncated RCTs addressing the same question.

So, what does this mean? It means that truncated RCTs do indeed tend to inflate the effect size substantially and to show differences by chance alone where none exists.

This is concerning in general, and specifically for our example of the semi-recumbent positioning study. Let us do some calculations to see just how this effect inflation would play out in the said study. Recall that microbiologically confirmed pneumonia occurred in 2 of 39 (5%) semi-recumbent cases and in 11 of 47 (23%) supine cases. The investigators calculated the adjusted odds ratio of VAP in the supine compared to semi-recumbent to be 6.8 (95% CI 1.7 - 26.7). This, as I mentioned before is an inflated estimate as odds ratios tend to be with frequent events. Furthermore, I obviously cannot do the adjusted calculation, as I would need the primary patient data for this. What we need is the relative reduction in VAP due to the intervention being investigated anyway, which is the reciprocal of what we have. So, I can derive the unadjusted relative risk thusly: (2/39)/(11/47) = 0.22. Now, if the RCT truncation alone reduces this risk by 29%, then if the trial had been allowed to go to completion, this relative risk would have been ~0.3. In this range, the difference does not seem all that impressive. But as all of the threats to validity we discussed in the original post begin to chisel mercilessly away at this risk reduction, the 29% inflation becomes a proportionally bigger deal.

Well, that does it.  

Monday, 14 February 2011

Redefining compassion: "A spiritual technology"

This is a TEDxUN talk by Krista Tippett. It is fantastic!
If you are in a rush, just go to around minute 10:00 or so. But really the whole talk is well worh considering.

Friday, 11 February 2011

CMS never events: Evidence of smoke in mirrors?

Let me tell you a fascinating story. In 1999, I was still fresh out of my Pulmonary and Critical Care Fellowship, struggling for breath in the vortex of private practice, when a cute little paper appeared in the Lancet from a great group of researchers in Spain, describing a study performed in one large academic urban medical center's two ICUs: one respiratory and one medical. Its modest aim was to see if semi-recumbent (partly sitting up) compared to supine (lying flat on the back) positioning could reduce the incidence of that bane of the ICU, ventilator-associated pneumonia (VAP). The study was a well done randomized controlled trial, and the investigators even went so far as to calculate the power (the number needed to enroll in order to detect a pre-determined magnitude of effect [in this case an ambitious 50% reduction in clinically suspected VAP]), and this number was 182 based on the assumption of a 40% VAP prevalence in the control (supine) group. The primary endpoint was the prevalence (percentage of all mechanically ventilated [MV] patients developing) and the secondary the incidence density (number of cases among all MV patients spread over all the cumulative days of MV [patient-days of MV]) of clinically suspected VAP, based on the CDC criteria, while microbiologically confirmed VAP (also rigorously defined) served as the secondary endpoint.

Here is what they found. The study was stopped early due to efficacy (this means that the intervention was so superior to the control in reaching the endpoint that it was deemed unethical after the interim look to continue the study), enrolling only 86 patients, 39 in the intervention and 47 in the control groups. And here are the results for the primary and secondary outcomes:

So, this is great! No matter how you slice it, VAP is reduced substantially; there is a microbiologically confirmed prevalence reduction of nearly 6-fold (this is unadjusted for potential differences between groups; and there were differences!). Well, you know what's coming next. That's right, the "not so fast" warning. Let's examine the numbers in context.

First of all, if we look at the evidence-based guideline on HCAP, HAP and VAP from the ATS and IDSA, the prevalence of VAP is generally between 5 and 15%; in the current study the control group exceeds 20%. Now, for the incidence density, for years now the CDC has been keeping and reporting these numbers in the US, and the rate in patients comparable to the ones in the study should be around 2-4 cases per 1,000 MV days. In this study, no matter how you slice it, clinically or microbiologically, the incidence density is exceedingly high, more in line with some of the ex-US numbers reported in other studies. So, they started high and ended high, albeit with a substantial reduction.

Second of all, there is a wonderful flow chart in the paper that shows the enrollment algorithm. One small detail has always been somewhat obscure to me: the 4 patients in the semi-recumbent group that were excluded from analysis due to reintubation (this means that they were taken off MV, but had to go back on it within a day or two), which was deemed a protocol violation. Now, you might think that 4 patients is a pretty small number to worry about. But look at the total number of patients in the group: 39. If the excluded 4 all had microbiologically confirmed VAP, that would bring our prevalence from 5% to 14% (6 out of 43). This would certainly be a less than 6-fold reduction in VAP.

Thirdly, and this I think is critical, the study was not blinded. In other words, the people who took care of the patients knew the group assignment. So what, you ask. Well remember that VAP is a pretty difficult, elusive and unclear diagnosis. So, let us pretend that I am a doc who is also an investigator on the study, and I am really invested in showing how marvelous semi-recumbent positioning is for VAP prevention. I am likely to have a much lower threshold for suspecting and then diagnosing VAP in the comparator group than in my pet intervention group. And this is not an indictment of anyone's judgment or integrity; it is just how our brains are wired.

Next, there were indeed important differences between groups in their baseline risk factors for VAP. For example, more patients in the control (38%) than in the intervention (26%) group were on MV for a week or longer, the single most important risk factor for developing VAP. Likewise, the baseline severity of illness was higher in the control than the intervention group. To be sure, the authors did statistical analyses to adjust these differences away, and still found an adjusted odds ratio of VAP among the supine group to be 6.8, with the 95% confidence interval between 1.7 and 26.7. This is generally taken to mean that, on average, the risk of VAP increases nearly 7-fold for supine position as opposed to semi-recumbent, and if the trial was repeated 100 times, 95 of those times this estimate would fall between a 1.7 and a 26.7-fold increase. OK, so we can accept this as a possible viable strategy, right?

But wait, there is more. Remember what we said about the odds ratio? When the event happens in more than 10% of the sample, the odds ratio vastly overestimates the risk of this event. 28.4% anyone?

Now, let's put it all together. A single center study from a Spanish academic hospital, among respiratory and medical ICU patients, with a minuscule sample size, yet halted early for efficacy, an exceedingly high baseline rate of VAP, a substantial number of patients excluded for a nebulous reason, unblinded and therefore prone to biased diagnosis, reporting an inflated reduction in VAP development in the intervention group. It would be very easy to write this off as a flawed study (like all studies tend to be in one way or another) in need of confirmatory evidence, if it were not so critical in the current punitive environment of quality improvement. (By the way, to the best of my knowledge, there is no study that replicates these results). The ATS/IDSA guideline includes semi-recumbent positioning as a level I (highest possible level of evidence) recommendation for VAP prevention, and it is one of the elements of the MV bundle, as promoted by the Institute for Healthcare Improvement, which demands 95% compliance with all 5 elements of the bundle in order to get the "compliant" designation. And even this is not the crux of the matter. The diabolical detail here is that CMS is creeping up on making VAP into one of their magical "never" events, and the efforts by hospitals will most assuredly be including this intervention. So, ICU nurses are already expected to fall in step with this deceptively simple yet not-so-easily executable practice.

And this is what is under the hood of just one simple level I recommendation by two reputable professional organizations in their evidence-based guidelines. One shudders to think...              

Wednesday, 9 February 2011

Evidence and profit: An unhealthy alliance

My JAMA Commentary came out this week, and I am getting e-mail about it. It seems to have resonated with many docs who feel that the research enterprise is broken and its output fails them at the office. But what I want to do is tie a few ideas together, ideas that I have been exploring on this blog and elsewhere, ideas that may hold the key to our devastating healthcare safety problem.

The last four decades can be viewed as a nexus between the growth of evidence-based medicine (EBM) on the one hand, and the unbridled proliferation of the biopharmaceutical industry and its technologies. The result has been rapid development, maximization of profit, and a juggernaut of poorly thought-out and completely uncoordinated research geared initially at regulatory approval and subsequently to market growth. It is not that the clinical research has been of poor quality, no. It is that our research tools are primitive and allow us to see only slivers of reality. And these slivers are prone to many of our cognitive biases to boot. So, the drive to produce evidence and the drive to grow business colluded to bring us to where we are today: inundated with evidence of unclear validity, unbalanced with regard to where the biggest difference to public health can be made. Yet we are constantly poked and prodded by the eager bureaucracy to do better at implementing this evidence, while the system continues to perform in a devastatingly suboptimal fashion, causing more deaths every year than strokes.

A byproduct of this technological and financial race has been the rapid escalation of healthcare spending, with the consequent drive to contain it. The containment measures have, of course, had the "unintended consequence" of increased patient volume for providers and of the incredible shrinking appointment, all just to make a living. The end-result for clinicians and patients is the relentless pressure of time and the straight jacket of "evidence-based" interventions in the name of quality improvement. And in this mad race against the clock and demoralization, very few have had the opportunity to think rationally and holistically about the root causes of our status quo. The reality is that we are now madly spinning our wheels at the margins, getting bogged down in infinitesimal details and losing the forest for the trees (pardon all of the metaphor mixing). Our evidence-based quality improvement efforts, while commendable, are like trying to plug holes in a ship's hull with bandainds: costly and overall making little if any difference.

But if we step back and stop squinting, we can see the big picture: stagnated and outdated research enterprise still rewarding spending over substance, embattled clinicians trying to stay afloat, and a $2.5 trillion healthcare gorilla feeding the economy at the expense of human lives. Will technology fix this mess? Not by itself, no. Will more "evidence" be the answer? No, not if we continue to generate it as usual. Is throwing more money at the HHS the solution? I doubt it. A radical change of course is in order. Take profit out of evidence generation, or at least blunt its influence (this will reduce the clutter of marginal, hair-splitting technologies occupying clinicians' collective consciousness), develop new tools for better patient care rather than for maximizing the bottom line, give clinicians more time to think about their patients' needs rather than about how to maintain enough income to pay for the overhead, these are some of the obvious yet challenging solutions to the current crisis. Challenging because there needs to be political will to implement them. And because we are currently so invested in the path we are on that it is difficult and perhaps impossible to stray without losing face. But what is the alternative?

Tuesday, 8 February 2011

Medical decision making: More signal less noise, please!

It's official, I'm a country bumpkin! Driving in Boston last week I was distracted, annoyed, made anxious and confused by the constant traffic, billboards and signs. Even highway markings confused me, particularly one indicating a detour to Storrow Drive East, which never materialized. Despite the fact that I know the geography of Boston like the back of my hand, I nearly went down the wrong streets multiple times, including driving the wrong way on some one-way roads. Yes, I am now the menace I used to save my prize driving language for in my younger days.

But it seems that over the years of my living away, there has been a sharp increase in the information thrown at me from all directions, accompanied by a decline in places to rest my gaze without suffering the perseveration of conscious processing. And while the value of this information is at best questionable, the sum total of this overstimulation is clearly confusion, wrong road choices and possibly a reduction in the safety of my driving. This whole experience reminded me of Thomas Goetz's distaste for how medical results are reported. If you have not seen him preach about it, you really should. Here is his excellent TED talk on the subject.


It is ironic that during this overwhelming city visit I also had the chance to speak to a doctor about "routine" preoperative testing and its value. Before surgery, it is recommended that a patient get a screening evaluation. Yet the components of this evaluation vary widely, and may include blood work, urinalysis, electrocardiogram, a chest X-ray and the like. Although evidence suggests that most of the points of this evaluation are useless at best, many institutions continue to order a shotgun panel of preoperative testing for everyone. This one-size-fit-all medicine results in reams of useless and distracting information, a high frequency of abnormal findings of questionable significance, a potential for harm, worry and needless healthcare spending. In my particular conversation I asked the anesthesiologist what the pre-test probability for someone with my characteristics was for a useful chest X-ray result, for example, and whether the fancy electronic medical record used by the hospital could help her determine this. While the answer to the former question was "probably exceedingly low", the answer to the latter was a definitive "no." So, given some elementary thinking, it became clear that a patient like me should not in fact be subjected to a chest X-ray, since any pathology found on one would likely represent a false positive finding, which would nevertheless require potentially invasive follow-up. And guess what? By focusing on the particular individual in the office, rather than all comers, we could have gone through the entire menu of the possible preoperative tests "routinely" ordered and eliminated most if not all of them. But my bet is that not all patients, not even all e-patients, either know or are able to initiate this type of a critical discussion. And yet what tests to obtain, if any, should always be a thoughtful and individualized decision. To approach testing in any other way is to risk generating noise, distraction and harm.

And this brings me back to Thomas Goetz's idea of redesigning how test results are reported. I love his idea. But to me what needs to happen before making the data patient-friendly, is making the decision-making provider-friendly. So, great idea, Mr. Goetz, but let us move it upstream, to the office, where the decision to get chest X-rays, cholesterols and urinalyses is made, and help the doctor visualize her patient's risk for a disease being present, the characteristics of the test about to be ordered, the probability of a positive test result, and all the downstream probabilities that stem from this testing, so as to put a positive test result in the context of the individual's risk for having the disease. Because getting the results of tests that perhaps should never have been obtained in the first place is following the GIGO principle. It is generating noise, distraction and detours going wrong way down one-way roads. And when applied to medicine, these are definitely unwelcome metaphors.