Showing posts with label quality. Show all posts
Showing posts with label quality. Show all posts

Wednesday, 15 February 2012

Big changes in the world of VAP?

As you may or may not be aware, the four main professional societies in the US that include large critical care constituencies, AACN, Chest, SCCM and ATS, have created something called a Critical Care Societies Collaborative (CCSC). Its purpose is essentially to give the critical care community a voice in shaping public policy. And, as you can imagine, one of the current-day issues they are tackling is performance measures.

Now, on this web site we have spent a lot of virtual ink talking about quality metrics, particularly where ventilator-associated pneumonia, or VAP, is concerned. Well, I am happy to report that finally, a strong voice in Critical Care medicine is in agreement with what we have been saying: the VAP bundle needs to go! In fact, here is the sum of the recommendations made to the National Quality Forum in a letter dated February 9, 2012, from the CCSC leaders about VAP (emphasis mine):
The Task Force felt that the VAP “care bundle” minimizes the importance of each individual component measure and neglects the fact that many elements of the existing VAP bundle are known to have important effects outside of VAP reduction, including improved patient survival. The task force also notes that one of the components of the VAP care bundle, stress ulcer prophylaxis, may actually increase the risk of VAP.51Therefore, the Task Force would like to make the following recommendations regarding measure gaps related to VAP:(1) Dissolve the VAP care bundle and instead develop a new group of quality measures related to general evidencebased practices for patients requiring mechanical ventilation (described above.) These potential measure gaps would include care processes known to reduce morbidity and mortality in patients who are ventilated.
(2) Develop measures using the VAPspecific measure gaps supported by recent guidelines.52,53 These may include measures for the following evidencedbased practices:
• Orotracheal rather than nasotracheal intubation to prevent VAP54;
• Subglottic secretion drainage to prevent VAP55;
• Elevating the head of bed to 45 degrees to prevent VAP56;
• Oral antiseptic administration to prevent VAP57;
• When empiric antibiotics are used to treat VAP, initial treatment based on qualitative endotracheal aspirates rather than quantitative bronchoscopic aspirates58; and
• No more than an 8day course of antibiotics as treatment for uncomplicated VAP.59
All of these VAP prevention strategies are supported by randomizedcontrolled trials. However, not all have favorable costbenefit profiles, and all have significant barriers, which may make widespread adoption unfeasible. Although we list them all here, we note that all may not be good quality measures.
So, here it is -- the recommendation. But will it be followed? When I was at SCCM, I heard a presentation that talked about some new metrics being developed by the CCSC in collaboration with the CDC, which will likely replace VAP as the focus of mechanical ventilation complications. I am in the process of learning more about these developments even as we speak, and will update my readers on what I learn. Suffice it to say, change is coming to the world of VAP. And it's about time. 

Tuesday, 22 November 2011

Lessons from Xigris

I have been wanting to write for a while about the demise of Xigris, but work and other commitments have stalled my progress. But it is time.

Here is my disclosure: I have received research funding from BioCritica, a daughter company of Eli Lilly, the manufacturer of Xigris. I also happen to know well and hold in high esteem the depth of knowledge and integrity of several colleagues who worked on Xigris internally at Lilly.

But on to the story. Xigris has had a short and bumpy life. When the PROWESS study, the Phase III Xigris trial, was first published in the NEJM in 2001 [1], it was the first therapy to succeed in sepsis, reducing mortality by 6% from about 31% to about 25%, yielding the number needed to treat of 16. This was huge, as so many trials to date had failed, and no progress had been made in sepsis management for years. These data opened the door to the FDA approval, despite a hung advisory committee, where equal numbers of members voted for and against approval. The controversy centered on concerns for bleeding complications, as well as some protocol changes during the trial and a switch in the manufacturing process. The latter concern was allayed by the Agency's detailed analysis and the finding of equivalence. There was a signal in a subgroup analysis that the drug might have been most effective among the most ill patients with a high probability of death, but not in their less ill counterparts. And despite the fact that the pivotal trial was not specifically performed in these patients, the approval for use specified just such a population.

So, despite the controversy, the drug was approved, though several post-marketing commitment studies were mandated. ENHANCE [2, 3] was an international study whose findings broadly confirmed the safety and efficacy of the drug, while the ADDRESS study [4], done in patients at low risk for death, was terminated early for lack of efficacy.

It seemed that PROWESS ushered in an era of positive results in sepsis. Shortly after its publication, other studies on the use of early goal-directed therapy [5], low-dose steroids [6] and tight glucose control [7] appeared in high impact journals, and the years of failure in sepsis management seemed to be over.  

In the meantime, and amid further controversy [8], Lilly supported the creation of the Values, Ethics and Rationing in Critical Care (VERICC) Task Force [9, 10], in addition to giving funding for the international Surviving Sepsis Campaign (SSC), which has resulted in the evidence-based practice guideline for sepsis management [11, 12] and an implementation program for the sepsis bundles, jointly sponsored by the SSC and the Institute for Healthcare Improvement [13]. The latter 2-year program enrolled over 15,000 patients world-wide, and achieved a doubling of bundle compliance from 18% to 36% with a concurrent drop in adjusted mortality of 5%. Because of several methodological issues and the lack of transparency about what it took to implement the bundle, it has never been clear to me a). whether there was causality between the bundle and mortality, and b). whether this effort was cost-effective.

But that aside, Xigris continued to stir up controversy, and there were still safety concerns. Some very well done observational studies, however, continued to confirm its effectiveness and safety in the real world setting [14]. Yet the final trial, PROWESS-SHOCK (done because of fears of an increase in bleeding complications), where patients in septic shock received Xigris as a part of their early management, brought doom. It was this study, whose preliminary results appeared in the press release from October 25, 2011, that prompted Lilly to pull the drug off the world market, since no difference in the 28-day mortality was detected between placebo and Xigris arms. Ironically, the preliminary reports indicate that no excess bleeding was noted in the treatment arm.

So, after roughly 10 years and millions of dollars, Xigris disappeared. But what can we learn from its story? There are many lessons that we should carry away, some about the way we do research, some about marketing practices, but all of them are about the need for a higher level of conversation and partnership. The biggest elephant in this room is whether a manufacturer should be allowed to fund guideline development. It is a complicated issue, particularly given our native proneness to cognitive bias, but in my opinion yes. This certainly cannot be done in a quid pro quo way. Perhaps this is naïve but should it not simply be a question of good data? And why wouldn't a manufacturer give money for the development of sensible guidelines without strings attached when the data are good?

Unfortunately, to me, Xigris is the poster child for how broken our research enterprise is, as I have discussed in this JAMA commentary [15]. Until all stake holders start talking to each other and arriving at common, useful and achievable goals, this is a story that will repeat itself again and again. The fact that regulatory trials, with all of their expensive and flashy internal validity, concern themselves only with statistical issues and care nothing about what happens in the real world is a travesty on many levels. The fact that it costs nearly $1 billion to bring a drug to market means that only big Pharma can bankroll such a gamble, and in return must demand big profits. The fact that this $1 billion fails to bring us studies that help clinicians and policy makers understand fully how to optimize the use of a drug once it is on the market is inexcusable. What we need is more intellectually honest discussions leading to novel pragmatic ways to answer the relevant questions in a timely manner and without bankrupting the system.

So, does the obvious financial interest mean that manufacturers should stay out of these discussions? I happen to think that they need a prominent place at the table. I actually think that the current fiasco is largely the result of too little interaction and too little cross-pollination of ideas: when we all sit around the table a nod in agreement, there is little progress. Deeper and novel understanding is built on disagreement and debate. Therefore, to leave the manufacturers out would invite further irrelevance. The bottom line is that we are all conflicted, and, according to the editors of PLoS, non-financial conflicts of interest, though more subtle and difficult to discern, may present an even bigger threat to much of what we do [16]. Elbowing out a party with an obvious conflict may have the unintended consequence of leaving some of the more insidiously conflicted others to run the show. And although we can argue whether profit is the healthiest driver for performance in healthcare, the reality is that our entire healthcare "system" is built around profit-making. Therefore it is disingenuous to single out one player over others.

On the positive side, the halo effect around Xigris brought a ton of attention to sepsis and its management. As Wes Ely conjectured in this piece, our improved understanding of sepsis (largely due to all the attention Xigris brought to it, in my opinion), is probably what rendered the drug useless in PROWESS-SHOCK. So, after all the hype, the noise and the hoopla, what is left is a company less one drug and hundreds of millions of dollars, and a disease area with a whole lot of what amounted to public health investment, with a vastly improved understanding of the disease state. How much is this benefit worth?

References

[1] Bernard GR, Vincent JL, Laterre PF, et al: Efficacy and safety of recombinant human activated protein C for severe sepsis. N Engl J Med 2001; 344:699–709
[2] Bernard GR, Margolis BD, Shanies HM, et al. Extended Evaluation of Recombinant Human Activated Protein C United States Trial (ENHANCE US). A Single-Arm, Phase 3B, Multicenter Study of Drotrecogin Alfa (Activated) in Severe Sepsis. Chest 2004;125:2206-16
[3] Vincent JL, Bernard GR, Beale R et al.
Drotrecogin alfa (activated) treatment in severe sepsis from the global open-label trial ENHANCE: further evidence for survival and safety and implications for early treatment. Crit Care Med 2005;33: 2266-77
[4] Abraham E, Laterre P-F, Garg R, et al. Drotrecogin Alfa (Activated) for Adults with Severe Sepsis and a Low Risk of Death. New Engl J Med 2005;353:1332-1341
[5] Rivers E, Nguyen B, Havstad S, et al. Early goal-directed therapy in the treatment of severe sepsis and septic shock. N Engl J Med 2001;345:1368-1377
[6] Annane D, Seville B, Charpentier C, et al. Effect of treatment with low doses of hydrocortisone and fludrocortisone on mortality in patients with septic shock. JAMA 2002;288:862-871
[7] van den Berghe G, Wouters P, Weekers F, et al. Intensive insulin therapy in the critically ill patients.  N Engl J Med 2001;345:1359-1367 
      [8] Eichacker PQ, Natanson C, Danner RL. Surviving Sepsis – Practice Guidelines, Marketing Campaigns and Eli Lilly. N Engl J Med 2006;355:1640-2 
      [9] Sinuff T, Kahnamui K, Cook DJ, et al. Rationing critical care beds: A systematic review. Crit Care Med 2004;32:1588-97
      [10] Truog RD, Brock DW, Cook DJ, et al. Rationing in the intensive care unit. Crit Care Med 2006;34:958-63
      [11] Dellinger RP, Carlet JM, Masur H, et al: Surviving Sepsis Campaign guidelines for management of severe sepsis and septic shock. 2004;32:858-73 
      [12] Dellinger RP, Levy MM, Carlet JM, et al: Surviving Sepsis Campaign: International guidelines for management of severe sepsis and septic shock: 2008. Crit Care Med 2008;36:296-327. Erratum in Crit Care Med 2008;36:1394-96
      [13] Levy MM, Dellinger RP, Townsend SR, et al. The Surviving Sepsis Campaign: Results of an international guideline-based performance improvement program targeting severe sepsis. Crit Care Med 2010;38:367-74
      [14] Lindenauer PK, Rothberg MB, Nathanson BH, et al. Activated protein C and hospital mortality in septic shock: A propensity-matched analysis. Crit Care Med 2010;38:1101-7
      [15] Zilberberg MD. The clinical research enterprise: Time for a course change? JAMA 2011;305:604-5 
      [16] The PLoS Medicine Editors (2008) Making Sense of Non-Financial Competing Interests. PLoS Med 5(9): e199. doi:10.1371/journal.pmed.0050199


Friday, 29 July 2011

Quality measures: Process, outcome, or both?

In the last week I wrote about our quality improvement, or QI, efforts in healthcare. And although there is a burgeoning field representing itself as the "science" of QI, I question much of its scientific validity. As always, VAP is my poster child for these discussions, where neither the definition of the condition itself nor its prevention efforts are subject to much scientific scrutiny. This makes VAP have a surreal, ghost-like quality: now you see it, now you don't. And this alone makes it difficult to assess prevention efforts. Much as in the heated mammography debate, where passionate anecdote prevails, the sanctity of the QI rubric blunts the usual critical approach to the data.

So, the central point that I made in this post was essentially to devalue the VAP eradication efforts as not grounded in solid scientific evidence. What has occurred to me, however, is that this position may be in fact at odds with a realization I blogged about here and here, wherein I agreed with Dan Arieli's suggestion that outcomes in the real world, where they are influenced by so much randomness, are not the thing to reward. It would be much more rational to reward best efforts at best results, thus the process rather than the outcome. So, here is the apparent contradiction: On the one hand I agree that outcomes may be too unpredictable, being that they are influenced by too many factors that are not in our control, yet I am also advocating that we start measuring such outcomes as antibiotic use associated with VAP and its reduction. What gives?

Well, on the one hand, I am OK with contradiction; life is full of instances where we have to hold conflicting information and feelings together. But as a scientist it is my predisposition to analyze (which literally means splitting into smaller, more manageable chunks), so I have given this ostensible paradox more thought. What I came up with is that measuring process is the right thing to do, but only under very specific conditions. Avedis Donabedian, who is considered the father of quality science, introduced the triad of structure-process-outcome as the backbone of quality science. This relationship certainly lends validity to the "process" metrics as surrogates for "outcome." But the condition that has to be met is that there be an actual correlation between the process and the said outcome. If there is no such solid correlation, then we are simply going through the motions, doing a rain dance to cause rain.

So, what I have said about VAP prevention in particular is that we are nowhere near being able to say that the recommended processes correlate with any changes in meaningful clinical outcomes. And because the data on these interventions are so weak, throwing massive resources behind implementing them is irrational and resembles religious fervor more than scientific pragmatism.

It is entirely understandable that we would jump on this bandwagon so rapidly, given the magnitude of harm in our healthcare system combined with the need to reign in the healthcare spending. But there is a more subtle point to be made here too. It relates to the fertile soil of our American psyche, where doing something is always perceived as better than thinking about our course of action, which is frequently referred to with contempt as "doing nothing." In the end, this crisis response mentality is good in a crisis, but potentially detrimental in the long term: we are unlikely to be altering meaningful outcomes, and we are spending billions of dollars on interventions lacking evidence.

So, I stand behind both of my assertions and maintain that they are not mutually exclusive. Yes, outcomes are subject to much randomness; yes, processes known to alter these outcomes are the sensible measures of our efforts to improve quality; and yes, these processes need first to be rigorously validated for their impact on the outcomes in question. Anything short of this pathway is not just a waste of our collective resources, but a manipulation of the public trust. And that is as far from the intent of science as it can get.

Wednesday, 27 July 2011

Health surveys: Run the other way!!!

Sometimes when I get an unsolicited call about answering survey questions, I feel a karmic obligation to participate; after all, if everyone said no to everything, I would not have any data to analyze. So, for this very reason, I just got off the phone with a poor young woman conducting a survey who called me randomly. The survey had to do with healthcare delivery, and she had no idea what she was getting herself into. First, I queried her whom the survey was for. She proceeded to tell me that she did not have that information specifically, but gave me a general idea of who the customers tend to be. Then she launched into the survey questions.

Now, I realize that they all have to ask the same questions the same way in order not to bias the data. But man, who writes these questions? "What would you say is the reputation of the cardiac surgery program at thus-and-such a hospital in your area: a). good locally, 2). good locally and state-wide, 3). good locally, statewide and regionally, 4). good locally, statewide, regionally and nationally, 5). good locally, statewide, regionally, nationally and internationally, or 6). not good at all?" Well, what the heck do you mean by "reputation"? You mean what is the gossip about Dr. Smith in my community? Or do you mean what kind of care they provide in terms of timeliness, evidence, shared decision making, post-operative complications, what? Then came "if you or your family member needed a cardiac procedure, how comfortable would you be going to this facility? 1). very comfortable, 2). somewhat comfortable, 3). somewhat uncomfortable, and 4). not at all comfortable?" How the heck should I know? I have not researched all the local facilities, I have not checked on their outcomes, I have not interviewed all of their cardiac surgical teams (yes, including anesthesia), I do not know what their infection control track records are, and, most importantly, how willing they are to treat me as an individual rather than a source of income. And then, for every hospital she mentioned (and there were quite a few), she went through the same litany of meaningless questions.

And then she asked me if I am familiar with some of the well-known quality-rating organizations. And she included US News and World Report Hospital Ratings! And I don't even believe the CMS got it anywhere near right!!! Oy! What do the answers to these questions from someone who is not steeped in the data mean anyway? If researchers and providers have not arrived at the appropriate metrics for quality, how meaningful are the lay public's opinions on these matters?

And finally, a group of questions that let the cat out of the bag as to the purpose of the survey. She told me a story first, of a large regional medical center in the area building a new multi-million dollar state-of-the-art cardiac care facility. Sexy new equipment, individual patient rooms, targeted and individually-tailored treatment plans, all the buzzwords of the brave new world of medicine. And then she asks me would I be comfortable going to this facility. What am I supposed to say? I have no idea! How do I tell this poor child that the questions are written in an absurd way and smack of marketing? How do I explain to her that this facility will probably need to recoup their capital investment, and, therefore, has a conflict of interest when it comes to caring for me? How do I teach her that this is the problem with American medicine, this very over-reliance on reputations and expertise to tell us to over-indulge in interventionism at the expense of our health and budgets?

Anyway, I will not belabor this further. My advice to survey fielders: If you want to market to the gullible, go ahead and call people randomly and ask your market-building question. And if a person tells you she is a physician and a health services researcher to boot, run, don't walk, the other way.

Thursday, 3 March 2011

Quality or value? A measure for the 21st century

Fascinating, how in the same week two giants of evidence-based medicine have given such divergent views on the future of quality improvement. Here (free subscription required), Donald Berwick, the CMS administrator and founder and former head of the Institute for Healthcare Improvement, emphasizes the need for quality as the strategy for success in our healthcare system. But here, one of the fathers of EBM, Muir Gray, states that quality is so 20th century, and we need instead to shine the light on value. So, who is right?

Well, let's define the terms. The Merriam-Webster dictionary defines quality as "the degree of excellence." The same source tells us that value is "a fair return or equivalent in goods, services or money for something exchanged." To me "value" is a holistic measure of cost for quality, painting a fuller picture of the investment vis-a-vis the returns on this investment. What do I mean by that?

Simply put, the idea behind value is to establish what is a reasonable amount to pay for a unit of quality. Let's take my used 1999 VW Passat as an example. If my mechanic tells me that it needs to have some hoses replaced, and it will cost me under $100, and the car will run perfectly, I will consider that to be a good value. However, if my transmission has fallen out in the middle of Brookline Ave. in Boston (really happened to me once, many years ago and with a different car), and it will cost me $5,000 to fix, I may say that the value proposition is just not there, particularly given that the car itself is worth much less than $5,000. Given that my budget is not unlimited, I have to make trade-off decisions about where to put my money, so I may instead spend the money on another used Passat that has good prospects.    

But in medicine, we routinely avoid thinking about value. There seems to be an overall impression that if it out there on the market, and especially if it is new, it is good and I am worth all of it. This impression is further enabled by the fact that CMS has no statutory power to make decisions based on value of interventions -- they are legislatively mandated to turn a blind eye to the costs. Does this make sense? How toothless is our comparative effectiveness effort likely to be if it has to ignore half of the story?

Let us now look at my favorite sticky wicket, ventilator-associated pneumonia, or VAP. Now, the IHI bundle aimed at eliminating VAP consists of 5 points of intervention: 1). semi-recumbent positioning, 2). daily screen for readiness to get off mechanical ventilation, 3). daily sedation vacation, 4). prophylaxis against GI bleeding, and 5). prevention of clots. As I have mentioned before elsewhere, adherence of 95% to all these measures is deemed compliance and may be ultimately used as a quality measure by payers to determine levels of reimbursement. And while each of these interventions is basically "motherhood and apple pie", applying them blindly and in toto to 95% of intubated patients may be a strategy for disaster. But what is even clearer is that, in order to implement this and all of the other quality improvement strategies, systems need to be put in place that will safeguard against failing to implement these quality measures. The time and resource expenditures needed to institute and maintain these systems, which have not been described in great enough detail as far as I am concerned, have never been quantified. So, what we are left with is a bunch of interventions that, while looking OK individually in clinical trials (until you really start looking at them critically), are likely providing small, if any, gains in quality at the margins, whose investment-return equation has not even been disclosed, let alone balanced. And because budgets are necessarily limited, as are clinicians' time and cognitive capacities, we need to select a sensible menu of interventions from this practically unlimited feast.

This is the quality conundrum, a clear case of chasing our tails to achieve perfection at the expense of good enough. And while no one in their right mind will argue with the language of improved quality in healthcare, I do think that Muir Gray and his camp are on to something that has been a long time coming. At this time of shrinking budgets, competing priorities and tightening resources, does it not make sense to look at value as a package deal, rather than merely at quality in isolation from its context? Instead of being bombarded by ever-increasing volume of quality measures coming from many directions, would it not be more sensible to prioritize these interventions based on the value that they bring rather than merely on their projected outcomes benefits, so frequently estimated based on data that have very little applicability to the real world? Let's start asking the question: how much quality and at what price? Without paying attention to this critical balance, we will not only bankrupt the system, but also worsen outcomes paradoxically, as we continue to overwhelm clinicians with infinite minutia that may or may not be generating helpful outcomes.

So, in my book, Muir Gray: score; Berwick: keep trying.            

Friday, 25 February 2011

Guidelines: What really constitutes level I evidence?

There has been some interesting buzz in the blogosphere about where evidence-based guideline recommendations come from, and I wanted to add a little fuel to that fire today.

As you know, I think a lot about the nature of evidence, about the "science" in clinical science, and about pneumonia, specifically ventilator-associated pneumonia or VAP. Last week I wrote here and here about a specific recommended intervention to prevent VAP consisting of semi-recumbent, as opposed to supine, positioning. This recommendation, one of 21 maneuvers aimed at modifiable risk factors for VAP, had level I evidence behind it. Given my recent deconstruction of this level I evidence, consisting of a single unblinded RCT in a single academic urban center in Spain, and given that we already know that level I data represent a very small proportion of all the evidence behind guideline recommendations, I got curious about this level I stuff. How is level I really defined? Is there a lot of room for subjective judgment? So, I went to the source.

In its HAP/VAP guideline, the ATS and IDSA committee define the levels of evidence in the following way:
Level I (high)
Level II (moderate) 








Level III (low)
     Evidence comes from well conducted, randomized controlled trials


Evidence comes from well designed, controlled trials without randomization (including cohort, patient series, and case-control studies). Level II studies also include any large case series in which systematic analysis of disease patterns and/or microbial etiology was conducted, as well as reports of new therapies that were not collected in a randomized fashion

Evidence comes from case studies and expert opinion. In some instances therapy recommendations come from antibiotic susceptibility data without clinical observations
So, well conducted, randomized controlled trials. But what does "well conducted" mean? Seems to me that one person's well conducted may be another person's garbage. Well, I went to the text of the document for clarification:
The grading system for our evidence-based recommendations was previously used for the updated ATS Community-acquired Pneumonia (CAP) statement, and the definitions of high-level (Level I), moderate-level (Level II), and low-level (Level III) evidence are summarized in Table 1 (8). 
OK, then. We have to go to reference #8, or the CAP guideline to get to the bottom of the definition. And here is what that document states:
Therefore, in grading the evidence supporting our recommendations, we used the following scale, similar to the approach used in the recently updated Canadian CAP statement (46): Level I evidence comes from well-conducted randomized controlled trials; Level II evidence comes from well-designed, controlled trials without randomization (including cohort, patient series, and case control studies); Level III evidence comes from case studies and expertopinion. Level II studies included any large case series in which systematic analysis of disease patterns and/or microbial etiology was conducted, as well as reports of new therapies that were not collected in a randomized fashion. In some instances therapy recommendations come from antibiotic susceptibility data, without clinical observations, and these constitute Level III recommendations.
Again, we are faced with the nebulous "well-conducted" descriptor with no further defining guidance on how to discern this quality. I resigned myself to going to the next source citation, #46 above, the Canadian CAP statement:
We applied a hierarchical evaluation of the strength of evidence modified from the Canadian Task Force on the Periodic Health Examination [4]. Well-conducted randomized, controlled trials constitute strong or level I evidence; well-designed controlled trials without randomization (including cohort and case-control studies) constitute level II or fair evidence; and expert opinion, case studies, and before-and-after studies are level III (weak) evidence. Throughout these guidelines, ratings appear as roman numerals in parentheses after each recommendation.
Another "well-conducted" construct, another reference, another wild goose chase. The reference #4 above clarified the definition for me thus:
OK, so, now we have "at least one properly randomized controlled trial." So, having gotten to the origin of this broken telephone game, it looks like proper randomization trumps all other markers for a well-done trial. The price of such neglect is giving up generalizability, confirmation, appropriate analyses, and many other important properties that need to be evaluated before stamping the intervention with a seal of approval. 

And this is just one guideline for one syndrome. The bigger point that I wanted to illustrate is that, even though we now know that only 14% of all IDSA guideline recommendations have so-called level I evidence behind them, what is dubious is the value and validity of assigning this highest level of evidence to these recommendations, given the room for subjectivity and misclassification. So, what does all of this mean? Well, for me it means no foreseeable shortage of fodder for blogging. But for our healthcare policy and our public's health? Big doo-doo.

Tuesday, 15 February 2011

The rose-colored glasses of early trial termination

The other day I did a post on semi-recumbent positioning to prevent VAP. The point I wanted to make was that an already existing quality measure for a condition that is well on its way to becoming a CMS "never event" is based on one unreplicated single-center small unblinded randomized controlled trial that was terminated early for efficacy. In my post I cited several issues with the study that question its validity. Today I want to touch upon the issue of early termination, which in and of itself is problematic.

What is early termination? It is just that: stopping the trial before enrolling the pre-planned number of subjects. First, it is important to be explicit in the planning phases about how many subjects will need to be enrolled. This is known as the power calculation and is based on the anticipated effect size and the uncertainty in this effect. Termination can happen for efficacy (the intervention works so splendidly that it becomes unethical not to offer it to everyone), safety (the intervention is so dangerous that it becomes unethical to offer it to anyone) or for other reasons (e.g., the recruitment is taking too long, etc.).

Who makes the decision to terminate early and how is the decision made? Well, under the best of circumstances, there is a Data Safety Monitoring Board, a body that is specifically in place to look at the data at certain points in the recruitment process and look for certain pre-specified differences between groups. This DSMB is fire-walled from both the investigators and the patients. The interim looks at the data  should be pre-specified by the protocol also, as the number of these looks actually influences the initial power calculation, since the more you look, the more differences you are likely to find by chance alone.

So, without going into too much detail on these interim looks, understand that they are not to be taken lightly, and their conditions and reporting require full transparency. To their credit, the semi-recumbent position investigators reported their plan for one interim analysis upon reaching 50% enrollment. Neither the Methods section nor the Acknowledgements, however, specify who was the analyst and the decision-maker. Most likely it was the investigators themselves that ended up taking the look and deciding on the subsequent course of action. And this itself is not that methodologically clean.

Now, let's talk about one problem early termination. This gargantuan effort led by the team from McMaster in Canada and published last year in JAMA sheds the needed light on what had been suspected before: early termination leads to inflated effect estimates. The sheer massiveness of the work done is mind boggling -- over 2,500 studies were reviewed! The investigators elegantly paired meta-analyses of truncated RCTs with meta-analyses of matched but nontruncated ones, and compared the magnitude of the inter-group differences between the two categories of RCTs. Here is one interesting tidbit (particularly for my friend @ivanoransky):
Compared with matching nontruncated RCTs, truncated RCTs were more likely to be published in high-impact journals (30% vs 68%, P<.001).
But here is what should really grab the reader:

Of 63 comparisons, the ratio of RRs was equal to or less than 1.0 in 55 (87%); the weighted average ratio of RRs was 0.71 (95% CI, 0.65-0.77; P <.001)(FIGURE2). In 39 of 63 comparisons (62%), the pooled estimates for nontruncated RCTs were not statistically significant. Comparison of the truncated RCTs with all RCTs (including the truncated RCTs) demonstrated a weighted average ratio of RRs of 0.85; in 16 of 63 comparisons (25%), the pooled estimate failed to demonstrate a significant effect. [Emphasis mine]
The authors went on to conclude the following:

In this empirical study including 91 truncated RCTs and 424 matching nontruncated RCTs addressing 63 questions, we found that truncated RCTs provide biased estimates of effects on the outcome that precipitated early stopping. On average, the ratio of RRs in the truncated RCTs and matching nontruncated RCTs was 0.71. This implies that, for instance, if the RR from the nontruncated RCTs was 0.8 (a 20% relative risk reduction), the RR from the truncated RCTs would be on average approximately 0.57 (a 43% relative risk reduction, more than double the estimate of benefit). Nontruncated RCTs with no evidence of benefit—ie, with an RR of 1.0—would on average be associated with a 29% relative risk reduction in truncated RCTs addressing the same question.

So, what does this mean? It means that truncated RCTs do indeed tend to inflate the effect size substantially and to show differences by chance alone where none exists.

This is concerning in general, and specifically for our example of the semi-recumbent positioning study. Let us do some calculations to see just how this effect inflation would play out in the said study. Recall that microbiologically confirmed pneumonia occurred in 2 of 39 (5%) semi-recumbent cases and in 11 of 47 (23%) supine cases. The investigators calculated the adjusted odds ratio of VAP in the supine compared to semi-recumbent to be 6.8 (95% CI 1.7 - 26.7). This, as I mentioned before is an inflated estimate as odds ratios tend to be with frequent events. Furthermore, I obviously cannot do the adjusted calculation, as I would need the primary patient data for this. What we need is the relative reduction in VAP due to the intervention being investigated anyway, which is the reciprocal of what we have. So, I can derive the unadjusted relative risk thusly: (2/39)/(11/47) = 0.22. Now, if the RCT truncation alone reduces this risk by 29%, then if the trial had been allowed to go to completion, this relative risk would have been ~0.3. In this range, the difference does not seem all that impressive. But as all of the threats to validity we discussed in the original post begin to chisel mercilessly away at this risk reduction, the 29% inflation becomes a proportionally bigger deal.

Well, that does it.  

Wednesday, 9 February 2011

Evidence and profit: An unhealthy alliance

My JAMA Commentary came out this week, and I am getting e-mail about it. It seems to have resonated with many docs who feel that the research enterprise is broken and its output fails them at the office. But what I want to do is tie a few ideas together, ideas that I have been exploring on this blog and elsewhere, ideas that may hold the key to our devastating healthcare safety problem.

The last four decades can be viewed as a nexus between the growth of evidence-based medicine (EBM) on the one hand, and the unbridled proliferation of the biopharmaceutical industry and its technologies. The result has been rapid development, maximization of profit, and a juggernaut of poorly thought-out and completely uncoordinated research geared initially at regulatory approval and subsequently to market growth. It is not that the clinical research has been of poor quality, no. It is that our research tools are primitive and allow us to see only slivers of reality. And these slivers are prone to many of our cognitive biases to boot. So, the drive to produce evidence and the drive to grow business colluded to bring us to where we are today: inundated with evidence of unclear validity, unbalanced with regard to where the biggest difference to public health can be made. Yet we are constantly poked and prodded by the eager bureaucracy to do better at implementing this evidence, while the system continues to perform in a devastatingly suboptimal fashion, causing more deaths every year than strokes.

A byproduct of this technological and financial race has been the rapid escalation of healthcare spending, with the consequent drive to contain it. The containment measures have, of course, had the "unintended consequence" of increased patient volume for providers and of the incredible shrinking appointment, all just to make a living. The end-result for clinicians and patients is the relentless pressure of time and the straight jacket of "evidence-based" interventions in the name of quality improvement. And in this mad race against the clock and demoralization, very few have had the opportunity to think rationally and holistically about the root causes of our status quo. The reality is that we are now madly spinning our wheels at the margins, getting bogged down in infinitesimal details and losing the forest for the trees (pardon all of the metaphor mixing). Our evidence-based quality improvement efforts, while commendable, are like trying to plug holes in a ship's hull with bandainds: costly and overall making little if any difference.

But if we step back and stop squinting, we can see the big picture: stagnated and outdated research enterprise still rewarding spending over substance, embattled clinicians trying to stay afloat, and a $2.5 trillion healthcare gorilla feeding the economy at the expense of human lives. Will technology fix this mess? Not by itself, no. Will more "evidence" be the answer? No, not if we continue to generate it as usual. Is throwing more money at the HHS the solution? I doubt it. A radical change of course is in order. Take profit out of evidence generation, or at least blunt its influence (this will reduce the clutter of marginal, hair-splitting technologies occupying clinicians' collective consciousness), develop new tools for better patient care rather than for maximizing the bottom line, give clinicians more time to think about their patients' needs rather than about how to maintain enough income to pay for the overhead, these are some of the obvious yet challenging solutions to the current crisis. Challenging because there needs to be political will to implement them. And because we are currently so invested in the path we are on that it is difficult and perhaps impossible to stray without losing face. But what is the alternative?

Thursday, 9 December 2010

1,000 lives per day or 45 lives every hour

In the wake of the recent studies confirming our suspicions that we are no better off today than a decade ago as far as the safety of our healthcare system is concerned, I have been doing a lot of thinking and writing about this issue. The other day I blogged about the fact that there are no simple solutions, yet we must pursue change. Today, this e-mail from 350.org really stopped me in my tracks:
Dear friends,
Climate negotiations can seem quite abstract sometimes.

I'm here in Cancún, Mexico, where UN delegates from around the world spend hours debating details of complex regulations.  Sometimes it seems that everyone has forgotten a crucial fact: the climate is changing much faster than these negotiations are moving. 

Meanwhile, out in the real world, climate impacts are all too visible. Since the negotations began 10 days ago, climate disasters have struck all over the world: flooding in Australia, Venezuela, the Balkans, Columbia, India; wildfires in Israel, Lebanon, Tibet; freak winter storms in Europe and the United States. These events have been devastating--hundreds are dead, and hundreds of thousands have been affected.
To put it in the context of our healthcare system, the unnecessary mortalities and morbidities are happening faster than our quality improvements are moving! In other words, if there are approximately 400,000 avoidable deaths annually attributable to healthcare encounters, this means that every day we delay implementing a viable solution we lose over 1,000 lives per day or about 45 lives every hour or 1 life every 1 and 1/2 minutes! In the time that it took me to write this post, 20 patients have lost their lives unnecessarily. Are any of them your loved ones?

All these lives come with stories, all these lives are loved by someone, and all these lives cannot just be written off as sacrificial lambs in the name of a growing bureaucracy that cannot move the meter. We can wring our collective hands and say that we wish we knew how to stop this gushing bleed. Yet, we continue to conduct business as usual, increasing revenues and testing and interventions and cognitive loads and questionable evidence. Ultimately, should eleven years of doing the same thing and getting the same woefully inadequate result encourage us to continue in the same direction, or should we just come to a full stop for a moment?

I realize that medicine cannot stop -- illness will not stop. But the lifestyle that feeds the gluttonous homicidal machine of healthcare can be altered. A combination of prevention, reduction of interventions of questionable effectiveness and safety, more time for doctors to think about their patients and make decisions together -- this is the path. It is not easy, but neither is losing a partner, a brother or a child to the very idol at whose altar we have come to worship and atone for all of our individual and societal bad choices. Today is the day. Who is with me?  

Monday, 6 December 2010

"Invisibility, inertia and income" and patient safety

Hat tip to @KentBottles for a link to this story

I spend a lot of time thinking about the quality and safety of our healthcare system, as well as our efforts to improve it. I have written a lot about it here in this blog and in some of my peer-reviewed publications. You, my reader, have surely sensed my frustration with the fact that we have been unable to put any kind of a dent in the killing that goes on within our hospitals and other healthcare encounter locations. So, it is always with much interest and appreciation that I learn that I am not alone, and that others have had it with the criminal lack of the sense of urgency to stop this medical holocaust. For this reason, I was really happy to read Michael Millenson's post on the Health Affairs Blog titled "Why We Still Kill Patients: Invisibility, Inertia and Income". I was very curious to see how he structured his argument to boil it down to these three I's, since I think that sexy slogans and memorable triplets are the way to go. So, here is how his arguments went.

First, establish the problem. And indeed, we have been killing around 100,000 people annually since the late 1970s (and probably since before then, as you actually have to look in order to find), which amounts to the total 20-year toll of 2.5 million unnecessary deaths due to healthcare in the US. This is truly appalling. And this is just up through the 1999 IoM report! Here is what I was thinking: And if we take into account not just the killing fields of the hospital, but all of life's interfaces with healthcare, we arrive at an even more frightening 400,000 deaths annually, as known back in 2000. Multiply this by 10, and now we really are talking about a killing machine of holocaust proportions! And I completely agree with Millenson that the fact that we continue to say "more research needed" and other pablum like that is utterly and completely irresponsible. However, is this really an invisible problem? The author makes a good argument for how we minimize these numbers by failing to add them up:

I laid out those numbers in a March, 2003 Health Affairs article that challenged the profession to break a silence of deed — failing to take corrective actions — and a silence of word — failing to discuss openly the consequences of that failure. This pervasive silence, I wrote:
continually distorts the public policy debate [and] gives individuals and institutions that must undergo difficult changes a license to postpone them. Most seriously of all, it allows tens of thousands of preventable patient deaths and injuries to continue to accumulate while the industry only gradually starts to fix a problem that is both long-standing and urgent.
Nearly eight years later, medical professionals now talk freely about the existence of error and loudly about the need for combating it, but silence about the extent of professional inaction and its causes remains the norm. You can see it in this latest study, which decries the continuing “patient-safety epidemic” while failing to do next what any public health professional would instinctually do: tally up the toll. Instead, we get dry language about the IOM’s goal of a 50 percent error reduction over five years not being met.
Let’s fill in the blanks: If this unchecked “epidemic” were influenza and not iatrogenesis, then from 1999 to date it would have killed the equivalent of every man, woman and child in the cities of Raleigh (this study took place in North Carolina) and Washington, D.C. Does a disaster of that magnitude really suggest that “further study” and a “refocusing of resources” are what’s needed?
I guess this makes sense -- adding up the numbers is pretty startling, yet we are reluctant to do so. At the same time I hesitate to call this "invisible", since as you saw in a paragraph above, I just multiplied by 10! Yet I am willing to concede the first "I" to Millenson, since I do see the power in these startling numbers.


On the to the next "I", inertia. I agree with Millenson generally, and we actually know this, that physicians do not practice evidence-based medicine, and, even when it does, evidence takes decades to penetrate practice. And there is every reason to be upset that the medical profession has not rushed to adopt evidence-based prevention measures that Millenson talks about. But there is a greater subtlety here than meets the eye. True, the Kestone project is frequently held as an example of a simple evidence-based bundled intervention resulting in in a huge reduction in central line-associated blood stream infections. Indeed, this is a great success and everyone should be practicing the checklist instituted in the project by Peter Pronovost's group. What is less obvious and even less talked about is that the same approach of evidence-based bundled approach to prevention of ventilator-associated pneumonia (VAP) has also been piloted by the Keystone group, yet none of us has seen any data from that. All I have is rumors at this point, but they are not good. Why is this? Well, I have discussed this before here and here: VAP is a very tricky diagnosis in a very tricky population. This is not to say that we need not work as hard as we can to prevent it. It is just to clarify that we are not sure of the best ways to accomplish this. Is this in and of itself shameful? Well, yes, if you think that medicine is a precise science. But if you have been reading my blog long enough, you know this is not the case.


Millenson further sites his reading of the Joint Commission Journal, which has been documenting the progress within one large Catholic healthcare system, Ascension, in its efforts to reduce infections, falls and other common iatrogenic harms. By the system's account, they are now able to save over 2,000 lives annually with these measures. This is impressive. But is it trustworthy? Unfortunately, without reading the primary studies I cannot comment on the latter. However, I did publish a review of studies from this very journal on VAP prevention efforts, and here is what I found:
A systematic approach to understanding this research revealed multiple shortcomings. First, since all of the papers reported positive results and none reported negative ones, there is a potential for publication bias. For example, a recent story in a non-peer-reviewed trade publication questioned the effectiveness of bundle implementation in a trauma ICU, where the VAP rate actually increased directionally from 10 cases per 1,000 MV days in the period before to 11.9 cases per 1,000 MV days in the period after implementation of the bundle (24). This was in contradistinction to the medical ICU in the same institution, which achieved a reduction from 7.8 to 2.0 cases per 1,000 MV days with the same intervention (24). Since the results did not appear in a peer-reviewed form, it is difficult to judge the quality or significance of these data; however, the report does highlight the need for further investigation, particularly focusing on groups at heightened risk for VAP, such as trauma and neurological critically ill (25).             
Second, each of the four reported studies suffers from a great potential for selection bias, which was likely present in the way VAP was diagnosed. Since all of the studies were naturalistic and none was blinded, and since all of the participants were aware of the overarching purpose of the intervention, the diagnostic accuracy of VAP may have been different before as compared to after the intervention. This concern is heightened by the fact that only one study reports employing the same team approach to VAP identification in the two periods compared (23). In other studies, although all used the CDC-NNIS VAP definition, there was either no reporting of or heterogeneity in the personnel and methods of applying these definitions. Given the likely pressure to show measurable improvement to the management, it is possible that VAP classification suffered from a bias.
Third, although interventional in nature, naturalistic quality improvement studies can suffer from confounding much in the same way that observational epidemiologic studies do. Since none of the studies addressed issues related to case mix, seasonal variations, secular trends in VAP, and since in each of the studies adjunct measures were employed to prevent VAP, there is a strong possibility that some or all of these factors, if examined, would alter the strength of the association between the bundle intervention and VAP development. Additional components that may have played a role in the success of any intervention are the size and academic affiliation of the hospital. In a study of interventions aimed at reducing the risk of CRBSI, Pronovost et al. found that smaller institutions had a greater magnitude of success with the intervention than their larger counterparts (26). Similarly, in a study looking at an educational program to reduce the risk of VAP, investigators found that community hospital staff were less likely to complete the educational module than the staff at an academic institution; in turn, the rate of VAP was correlated with the completion of the educational program (27). Finally, although two of the studies included in this review represent data from over 20 ICUs each (20, 22), the generalizability of the findings in each remains in question. For example, the study by Unahalekhaka and colleagues was performed in the institutions in Thailand, where patient mix and the systems of care for the critically ill may differ dramatically from those in the US and other countries in the developed world (22). On the other hand, while the study by Resar and coworkers represents a cross section of institutions within the US and Canada, no descriptions are given of the particular ICUs with respect to the structure and size of their institutions, patient mix or ICU care model (e.g., open vs. closed; intensivists present vs. intensivists absent, etc.) (20). This aggregate presentation of the results gives one little room to judge what settings may benefit most and least from the described interventions. The third study includes data from only two small ICUs in two community institutions in the US (21), while the remaining study represents a single ICU in a community hospital where ICU patients are not cared for by an intensivist (23).  Since it is acknowledged that a dedicated intensivist model leads to improved ICU outcomes (28, 29), the latter study has limited usefulness to institutions that have a more rigorous ICU care model.           
So, not to toot my own horn here, and not expecting you to read the long-winded Discussion, suffice it to say that we found many methodologic errors in this body of research from the Joint Commission's own journal to invalidate potentially nearly all of the reported findings. My point is again to reiterate that unless you read each study with a critical eye and then put it into the larger context, do not believe someone else's cursory reference to the staggering improvements. I guess pertinent to our discussion, inertia, while present, is a more nuanced issue than we are led to believe.


And finally, income. I do agree that it is annoying that economic arguments are even necessary to promote a culture of prevention and safety. What I disagree with is that these economic fallacies of the C-suite impact in any way the implementation of the needed prevention systems. Most of the evidence-based preventions are pretty low tech. And although they do require teams and commitment and systems to implement broadly, small demonstrations at the level of individual clinicians are possible. Also, I shudder at the thought that a group of dedicated clinicians could not persuade a group of equally dedicated administrators to do the right thing, even at the risk of losing some revenue. 


Bottom line? While I like Millenson's sexy little "three I's of safety", I think the solutions, as is always the case when you start looking under the hood, are more complicated and nuanced. In a recent post I cited 5 potential solutions to our quality problem, and I will repeat them here:
1. Empower clinicians to provide only care that is likely to produce a benefit that outweighs risks, be they physical or emotional.
2. Reward the signal and not the noise. I wrote about this here andhere.
3. Reward clinicians with more time rather than money. Although I am not aware of any data to back up this hypothesis, my intuition is that slowing down the appointment may result not only in reduction of harm by cutting out unnecessary interventions, but also in overall lowering of healthcare expenditures. It is also sure to improve the crumbling therapeutic relationship.
4. We need to re-engineer our research enterprise for the most important stakeholder in healthcare: the clinician-patient dyad. We need to make the data that are currently manufactured and consumed for large scale policy decisions more friendly at the individual level. And as a corollary, we need to re-think how we help information diffuse into practice and adopt some of the methods of the social sciences.
5. Let's get back to the tried and true methods of public health, where an ounce of prevention continues to be worth a pound of cure. Yes, let's strive for reducing cancer mortality, but let us invest appropriately in stuffing that tobacco horse back into its barn -- getting people to stop smoking will reduce lung cancer mortality by 85% rather than 0.3%, and at a much lower cost with no complications or false positives. Same goes for our national nutrition and physical activity struggles. Our social policies must support these well-recognized and efficient population interventions.
No, they are not simple, they are not sexy, and most importantly they may be painful. Yet, what is the alternative? We must stop this massive bleeder before the American public starts thinking that the cure is worse than the disease.     

Sunday, 5 December 2010

Top 5 this week

#5: Could our application of EBM be unethical?
#4: Evidence of harm
#3: Our nation's shocking Lady Macbeth moment
#2: Why are we still paying tobacco executives to kill...

And the #1 post this week is... Healthcare quality: 5 ways to stop the insanity

Thank you all for stopping by, commenting and broadening my thinking with your contributions!