Thursday, 20 January 2011

HOW TO SURVIVE A HEART ATTACK WHEN ALONE

Since many people are alone when they suffer a heart attack, without help,the person whose heart is beating improperly and who begins to feel faint, has only about 10 seconds left before losing consciousness.

However, these victims can help themselves by coughing repeatedly and very vigorously. A deep breath should be taken before each cough, and the cough must be deep and prolonged, as when producing sputum from deep inside the chest.

A breath and a cough must be repeated about every two seconds without let-up until help arrives, or until the heart is felt to be beating normally again.

Deep breaths get oxygen into the lungs and coughing movements squeeze the heart and keep the blood circulating. The squeezing pressure on the heart also helps it regain normal rhythm. In this way, heart attack victims can get to a hospital. Tell as many other people as possible about this. It could save their lives!!

A cardiologist says If everyone who gets this mail sends it to 10 people, you can bet that we'll save at least one life.

BE A FRIEND AND PLEASE SEND THIS ARTICLE TO AS MANY FRIENDS ! AS POSSIBLE

Wednesday, 19 January 2011

Data mining: It's about research efficiency.

I have taken a little break from my reviewing literature series -- work has superseded all other pursuits for a little while. But I did want to do a brief post today, since this JAMA Commentary really intrigued me.

First thing that interested me was the authors. Now, I know who Benjamin Djulbegovic is -- you have to live under a rock as an outcomes researcher not to have heard of him. But who is Mia Djulbegovic? It is an unusual enough surname to make me think that she is somehow related to Benjamin. So, I queried the mighty Google, and it spat out 1,700 hits like nothing. But only one was useful in helping me identify this person, and that was a link to her paper in BMJ from 2010 on prostate cancer screening. On this paper (her only one listed on Medline so far), she is the first author, and her credentials are listed as "student", more specifically in the Department of Urology at the University of Florida College of Medicine in Gainesville, FL. The penultimate author on the paper is none other than Benjamin Djulbegovic, at the University of South Florida in Tampa, FL. So, I am surmising from this circumstantial evidence that Mia is Benjamin's kid who is either a college or a medical student. Why does this matter? Well, there seem to be so few papers in high impact journals that are authored by people without an advanced degree, let alone in the first position, that I am in awe of this young woman, now with two major journals to her name -- BMJ and JAMA. This is evidence that parental mentorship counts for a lot (assuming that I am correct about their relationship). But regardless, kudos to her!

Secondly, the title of the essay really grabbed me: what is the "principle of question propagation", and what does it have to do with comparative effectiveness research (CER) and data mining? Well, basically, the principle of question propagation is something we talk about here a lot: questions beget questions, and the further you go down any rabbit hole, the more detailed and smaller the questions become. This is the beauty and richness of science as well as what I have referred to as "unidirectional skepticism" of science, meaning that a lot of the time, building on existing concepts, we just continue down the same direction in a particular research pursuit. This is why Max Planck was right when he said
A new scientific truth does not triumph by convincing its opponents and making them see the light, but rather because its opponents eventually die, and a new generation grows up that is familiar with it.
So, yes, we build upon previous work, and continue our journey down a single rabbit hole our entire career. Though of course there are countless rabbit holes all being explored at the same time. It is really more of a fractal-like situation than a single linear progression. What is clear, as the authors of the Commentary point out, is that this results in the ever-escalating theoretical complexity of scientific concepts. What does this have to do with anything? This, the authors state, argues for continued use of theory driven hypothesis testing, given that medical knowledge will forever be incomplete. And this brings them to data mining.

Here is where I get a little confused and annoyed. They caution the powers that be from consigning all clinical research to data mining, at the expense of more rigorous studies to pursue hypothesis testing. They argue that mining data that already exist is limiting precisely because it is constrained by the scope of our current knowledge, and that we cannot use these data to generate new associations and new treatment paradigms. They further state that emerging knowledge will require updating these data sets with new data points, and this, according to the authors
...creates a paradox, which is particularly evident when searching for treatment effects insubgroups—one of the purported goals of the IT CER initiative. As new research generates new evidence of the importance for tailoring treatments to a given subpopulation of patients, the existing databases will need to be updated, in turn undermining the original purpose to discover new relationships via existing records.
Come agin? And then they say that "consequently, the data mining approach can never result in credible discoveries that will obviate the need for new data collection". Mmhm, and so? Is this the punch line? Well, OK, they also say that because of all this we will still need to do hypothesis testing research. Is this not self-evident?

I don't know about you, but I have never thought that retrospective data mining would be the only answer to our research needs. Rather, the way to view this type of research is as an opportunistic pursuit of information from massive repositories of existing data. We can look for details that are unavailable in the interventional literature, zoom in on the potentially important bits, and use this information to inform more focused (and therefore pragmatically more realistic) interventional studies.

Don't take me wrong, I am happy that the Djulbegovics published this Commentary. It is really designed more as an appeal to policy makers, who, in their perennial search for one-size-fit-all panaceas, may misinterpret our zeal for data mining as the singular answer to all our questions. No indeed, hypothesis testing will continue. But using these vast repositories of data should make us smarter and more efficient at asking the right questions and designing the appropriate studies to answer them. And then generate further questions. And then answer those. And then... Well, you get the picture.

Friday, 14 January 2011

Reviewing medical literature, part 4: Statistical analyses -- measures of central tendency

Well, we have come to the part of the series you have all been waiting for: discussion of statistics. What, you are not as excited about it as I am? Statistics are not your favorite part of the study? I am frankly shocked! But seriously, I think this is the part that most people, both lay public and professionals, find off-putting. But fear not, for we will deconstruct it all in simple terms here. Or obfuscate further, one or the other.

So, let's begin with a motto of mine: If you have good results, you do not need fancy statistics. This goes along with the ideas in math and science that truth and computational beauty go hand in hand. So, if you see something very fancy that you have never heard of, be on guard for less than important results. This, of course, is just a rule of thumb, and, as such, will have exceptions.

The general questions I like to ask about statistics are 1). Are the analyses appropriate to the study question(s), and 2). Are the analyses optimal to the study question(s). The first thing to establish is the integrity and completeness of the data. If the authors enrolled 365 subjects but were only able to analyze 200 of them, this is suspicious. So, you should be able to discern how complete the dataset was, and how many analyzable cases there were. A simple litmus test is that if more than 15% of the enrolled cases did not have complete data for analysis or dropped out of the study for other reasons, the study becomes suspect for a selection bias. The greater the proportion of dropouts, the greater the suspicion.

Once you have established that the set is fairly complete, move on to the actual analyses. Here, first thing is first: the authors need to describe their study group(s); hence, descriptive statistics. Usually this includes so-called "baseline characteristics", consisting of demographics (age, gender, race), comorbidities (heart failure, lung disease, etc.), and some measure of the primary condition in question (e.g., pneumonia severity index [PSI] in a study of patients with pneumonia). Other relevant characteristics may be reported as well, and this is dependent on the study question. As you can imagine, categorical variables (once again, these are variables that have categories, like gender or death) are expressed as proportions or percentages, while continuous ones (those that are on a continuum, like age) are represented by their measures of central tendency.

It is important to understand the latter well. There are three major measures of central tendency: mean, median and mode. The mean is the sum of all individual values of a particular variable divided by the number of values. So, mean age among a group of 10 subjects would be calculated by adding all 10 individual ages and then dividing by 10. The median is the value that occurs in the middle of a distribution. So, if there are 25 subjects with ages ranging from 5 to 65, the median value is the one that occurs in subject number 13 when subjects are arranged in ascending or descending order by age. The mode, a measure used least frequently in clinical studies, signifies, somewhat paradoxically, the value in a distribution that occurs most frequently.

So, let's focus on the mean and the median. The mean is a good representation of the central value in a normal distribution. Also referred to as a bell curve (yes, because of its shape), or a Gaussian distribution, in this type of a distribution there are roughly equal numbers of points to the left and to the right of the mean value. It looks like this (from wikimedia.org):
For a distribution like the one above it hardly matters which central value is reported, the mean or the median, as they are the same or very similar to one another. Alas, most descriptors of human physiology are not normally distributed, but are more likely to be skewed. Skewed means that there is a tail at one end of the curve or the other (figure from here):
For example, in my world of health economics, many values for such variables as length of stay and costs spread out to the right of the center, similar to the blue curve in the right panel of the above figure. In this type of a distribution the mean and the median values are not the same, and they tell you different things. While the median gives you an idea of the central tendency of the entire distribution, the mean will tell you the central tendency of the majority of the distribution that is tightly clustered at the end opposite the tail. For a distribution similar to the one in the right panel, the mean will underestimate the central measure.

To round out the discussion of central values, we need to say a few words about scatter around these values. Because they represent a population and not a single individual, measures of central tendency will have some variation around them that is specific to the population. For a mean value, this variation is usually represented by standard deviation (SD), though sometimes you will see a 95% confidence interval as the measure of the scatter. Variation around the median is usually expressed as the range of values falling into the central one-half of all the values in the distribution, discarding the 25% at each end, or the interquartile range (IQR 25, 75) around the median. These values represent the stability and precision of our estimates and are important to look for in studies.

We'll end this discussion here for the moment. In the next post we will tackle inter-group differences and  hypothesis testing.      

Thursday, 13 January 2011

Reviewing medical literature part 3 continued: threats to validity

As promised, today we talk about confounding and interaction.

A confounder is a factor related to both, the exposure and the outcome. Take for example the relationship between alcohol and head and neck cancer. While we know that heavy alcohol consumption is associated with a heightened risk of head and neck cancer, we also know that people who consume a lot of alcohol are also more likely to be smokers, and smoking in turn raises the risk of H&N CA. So, in this case smoking is a confounder of the relationship between alcohol consumption and the development of H&N CA. It is virtually impossible to get rid of all confounding completely in any study design, save for possibly in a well designed RCT, where randomization presumably assures equal distribution of all characteristics; and even there you need an element of luck. In observational studies our only hope to deal with confounding is through statistical manipulation we call "adjustment", as it is virtually impossible to chase it away any other way. And in the end we still sigh and admit to the possibility of residual confounding. Nevertheless, going through the exercise is still necessary in order to get closer to the true association of the main exposure and the outcome of interest.

There are multiple ways of dealing with the confounding conundrum. The techniques used are matching, stratification, regression modeling, propensity scoring and instrumental variables. By far the most commonly used method is regression modeling. This is a rather complex computation that requires much forethought (in other words, "Professional driver on a closed circuit; don't try this at home"). The frustrating part is that, just because the investigators did the regression, does not mean that they did it right. Yet word limits for journal articles often preclude authors from giving enough detail on what they did. At the very least they should tell you what kind of a regression they ran and how they chose the terms that went into it. Regression modeling relies on all kinds of assumptions about the data, and it is my personal belief, though I have no solid evidence to prove it, that these assumptions are not always met.

And here are the specific commonly encountered types of regressions and when each should be used:
1. Linear regression. This is a computation used for outcomes that are continuous variables (i.e., variables represented by a continuum of numbers, like age, for example). This technique's main assumption is that the exposure and outcome are related to each other in a linear fashion. The resulting beta coefficient is the slope of this relationship if it is graphed.
2. Logistic regression. This is done when the outcome variable is categorical (i.e., one of two or more categories, like gender, for example, or death). The result of a logistic regression is an adjusted odds ratio (OR). It is interpreted as an increase or a decrease in the odds of the outcome occurring due to the presence of the main exposure. Thus, a OR of 0.66 means that there is a 34% reduction in the odds (used interchangeably with risk, though this is not quite accurate) of the outcome due to the presence of the exposure. Conversely, a OR of 1.34 means the opposite, or a 34% increase in the odds of the outcome if the exposure is present.
3. Cox proportional hazards. This is a common type of a model developed for a time to event, also known as "survival analysis" (even if not done for survival per se as the outcome). The resulting value is a hazard ratio (HR). For example, if we are talking about a healthcare-associated infection's impact on the risk of remaining in the hospital longer, a HR of, say, 1.8 means that a HAI increases the risk of being in the hospital by 80% at any time during the hospitalization. To me this tends to be the most problematic technique in terms of assumptions, as it requires that the risk of an even stays constant throughout the time frame of the analysis, and how often does this hold true? For this reason the investigators should be explicit about whether or not they tested for the assumption of proportional hazards and whether this was met.

Let's now touch upon the other techniques that help us to unravel confounding. Matching is just that: it is a process of matching subjects with the primary exposure to those without in a cohort study or subjects with the outcome to those without in a case-control study, based on certain characteristics, such as age, gender, comorbidities, disease severity, etc.; you get the picture. By its nature, matching reduces the amount of analyzable data, and thus reduces the power of the study. So, is is most efficiently applied in a case-control setting, where it actually improves the efficiency of enrollment.

Stratification is the next technique. The word "stratum" means "layer", and stratification refers to describing what happens to the layers of the population of interest with and without the confounding characteristic. In the above example of smoking confounding the alcohol and H&N CA relationship, stratifying the analyses by smoking (comparing the H&N CA rates among drinkers and non-drinkers in the smoking group separately from the non-smoking group) can divorce the impact of the main exposure from that of the confounder on the outcome. This method has some distinct intuitive appeal, though its cognitive effectiveness and efficiency dwindle the more strata we need to examine.

Propensity scoring is gaining popularity as an adjustment method in the medical literature. A propensity score is essentially a number, usually derived from a regression analysis, giving the propensity of each subject for a particular exposure. So, in terms of smoking, we can create a propensity score based on other common characteristics that predict smoking. Interestingly, some of these characteristics will be present also in people who are not smokers, yielding a similar propensity score in the absence of this exposure. Matching smokers to non-smokers based on the propensity score and examining their respective outcomes allows us to understand the independent impact of smoking on, say, the development of coronary artery disease. As in regression modeling, the devil is in the details. Some studies have indicated that most papers that employ propensity scoring as the adjustment method do not do this correctly. So, again, questions need to be asked and details of the technique elicited. There is just no shortcut to statistics.

Finally, a couple of words about instrumental variables. This method comes to us from econometrics. An instrumental variable is one that is related to the exposure but not the outcome. One of the most famous uses of this method was published by a fellow you may have heard of, Mark McClellan, where he looked at the proximity to a cardiac intervention center as the instrumental variable in the outcomes of acute coronary events. Essentially, he argued, the randomness of whether or not you are close to a center randomizes you to the type of treatment you get. Incidentally, in this study he showed that invasive interventions were responsible for a very small fraction of the long-term outcomes of heart attacks. I have not seen this method used that much in the literature I read or review, but am intrigued by its potential.

And now, to finish out this post, let's talk about interaction. "Interaction" is a term mostly used by statisticians to describe what epidemiologists call "effect modification" or "effect heterogeneity". It is just what the name implies: there may be certain secondary exposures that either potentiate or diminish the impact of the main exposure of interest on the outcome. Take the triad of smoking, asbestos and lung cancer. We know that the risk of lung cancer among smokers who are also exposed to asbestos is far higher than among those who have not been exposed to asbestos. Thus, asbestos modifies the effect of smoking on lung cancer. So, to analyze those smokers exposed to asbestos together with those who were not will result in an inaccurate measure of the association of smoking with lung cancer. More importantly, it will fail to recognize this very important potentiator of tobacco's carcinogenic activity. To deal with this, we need to be aware of the potentially interacting exposures, and either stratify our analyses based on the effect modifier or work the interaction term (usually constructed as a product of the two exposures, in out case smoking and asbestos) into the regression modeling. In my experience as a peer reviewer, interactions are rarely explored adequately. In fact, I am not even sure that some investigators understand the importance of recognizing this phenomenon. Yet, the entire idea of heterogeneous treatment effect (HTE) and our pathetic lack of understanding of its impact on our current bleak therapeutic landscape, is the result of this very lack of awareness. The future of medicine truly hinges on understanding interaction. Literally. Seriously. OK, at least in part.

In the next installment(s) of the series we will start tackling study analyses. Thanks for sticking with me.        

Wednesday, 12 January 2011

Reviewing medical literature, part 3: Threats to validity

You have heard this a thousand times: no study is perfect. But what does this mean? In order to be explicit about why a certain study is not perfect, we need to be able to name the flaws. And let's face it: some studies are so flawed that there is no reason to bother with them, either as a reviewer or as an end-user of the information. But again, we need to identify these nails before we can hammer them into a study's coffin. It is the authors' responsibility to include a Limitations paragraph somewhere in the Discussion section, in which they lay out all of the threats to validity and offer educated guesses as to the importance of these threats and how they may be impacting the findings. I personally will not accept a paper that does not present a coherent Limitations paragraph. However, reviewers are not always, as, shall we say, hard assed about this as I am, and that is when the reader is on her own. Let us be clear: even if the Limitations paragraph is included, the authors do not always do a complete job (and this probably includes me, as I do not always think of all the possible limitations of my work). So, as in everything, caveat emptor! Let us start to become educated consumers.

There are four major threats to validity that fit into two broad categories. They are:
A. Internal validity
  1. Bias
  2. Confounding/interaction
  3. Mismeasurement or misclassification
B. External validity
  4. Generalizability
Internal validity refers to whether the study is examining what it purports to be examining, while external validity, synonymous with generalizability, gives us an idea about how broadly the results are applicable. Let us define and delve into each threat more deeply.

Bias is defined as "any systematic error in the design, conduct or analysis of a study that results in a mistaken estimate of an exposure's effect on the risk of disease" (the reference for this is Schlesselman JJ, as cited in Gordis L, Epidemiology, 3rd edition, page 238). I think of bias as something that artificially makes the exposure and the outcome either occur together or apart more frequently than they should. For example, the INTERPHONE study has been criticized for its biased design, in that it defined exposure as at least one cellular phone call every week. Now enrolling such light users can really result in such a small exposure as not to be able to detect any increase in adverse events. This is an example of a selection bias, by far the most common form that bias takes. Another example of a frequent bias is encountered in retrospective case-control studies where people are asked to recall distant exposures. Take for example middle-aged women with breast cancer who are asked to recall their diets when they were in college. Now, ask the same of similar women without breast cancer. What you are likely to get is the effect, absent in women without cancer, of seeking an explanation for the cancer that expresses itself in a bias in what women with cancer recall eating in their youth. So, a bias in the design can make the association seem either stronger or weaker than it is in reality.

I want to skip over confounding and interaction at the moment, as these threats deserve a post of their own, which is forthcoming. Suffice it to say here that a confounder is a factor related to both, the exposure and the outcome. An interaction is also referred to as effect modification or effect heterogeneity. This means that there may be population characteristics that alter the response to the exposure of interest. Confounders and effect modifiers are probably the trickiest concepts to grasp. So, stay tuned for a discussion of those.

For now, let us move on to measurement error and misclassification. Measurement error, resulting in misclassification, can happen at any step of the way: it can be in the primary exposure, a confounder, or the outcome of interest. I run into this problem all the time in my research. Since I rely on administrative coding for a lot of the data that I use, I am virtually certain that the codes routinely misclassify some of the exposures and confounders that I deal with. Take Clostridium difficile as an example. There is an ICD-9 code to identify it in administrative databases. However, we know from multiple studies that it is not all that sensitive or all that specific; it is merely good enough, particularly for making observations over time. But even for laboratory values there is a certain potential for measurement error, though we seem to think that lab results are sacred and immune to mistakes. And need I say more about other types of medical testing? Anyhow, the possibility of error and misclassification is ubiquitous. What needs to be determined by the investigator and the reader alike is the probability of that error. If the probability is high, one needs to understand whether it is a systematic error (for example, a coder always more likely than not to include C. diff as a diagnosis) or a random one (a coder is just as likely to include as not to include a C diff diagnosis). And while a systematic error may result in either a stronger or a weaker association between the exposure and the outcome, a random, or non-differential, misclassification will virtually always reduce the strength of this association.

And finally, generalizability is a concept that helps the reader understand what population the results may be applicable to. In other words, will the data be applied strictly to the population represented in the study? If so, is it because there are biological reasons to think that the results would be different in a different population? And if so, is it simply the magnitude of the association that can be expected to be different or is it possible that even the direction could change? In other words, could something found to be beneficial in one population be either less beneficial or even more harmful in another? The last question is the reason that we perseverate on this idea of generalizability. Typically, a regulatory RCT is much less likely to give us adequate generalizability than a well designed cohort study, for example.

Well, these are the threats to validity in a nutshell. In the next post we will explore much more fully the concepts of confounding and interaction and how to deal with them either at the study design or study analysis stage.            

Tuesday, 11 January 2011

Do private ICU rooms really reduce HAIs?

We have known for quite some time now that the patient's environment in a hospital matters to his/her outcomes. The concept of biophilia was applied by Roger Ulrich back in the 1980s to surgical patients in a series of experiments. Famously, this work showed that looking out your hospital room's window on a bunch trees is associated with better and less eventful post-operative recovery than staring at a brick wall, for example. We have also known for some time that some of the hospital-associated delirium can be mitigated by having the patient dwell in a room with a window and be exposed to the diurnal light changes.

Another, perhaps even more tangible outcome that can be modified by hospital design is the spread of hospital-acquired infections. This week a paper in the Archives of Internal Medicine from the group in Quebec, who brought us detailed reports of the devastating multihospital hypervirulent Clostridium difficile outbreak in the last decade, generally confirms the effectiveness of private ICU rooms in containing the spread of HAIs. There are some interesting details to point out.

For example, the intervention hospital appears to have had a higher proportion of medical patients than the control institution. Why is this important? Well, medical patients generally experience more chronic and therefore longer stays in the ICU. This gives them a greater opportunity for exposure to HAIs than their surgical counterparts. On the other hand, we know that VAP, for example, an infection very likely to be caused by one of the resistant organisms listed in Table 2 of the paper, happens much more frequently in trauma ICUs than medical ICUs.

Second, the unadjusted ICU length of stay shows some interesting results, depicted in the graph below:
So, while at the intervention hospital the raw ICU LOS has remained stable, at the comparator institution it has been slowly creeping up. Of course, the investigators adjusted for all kinds of factors that may influence this outcome, and showed that there may be a (marginal) reduction in the ICU LOS in association with the switch to private rooms. The authors note that the adjusted average ICU LOS fell by 10%, though under similar circumstances in other similar investigations there is a 95% chance that this would fall somewhere between 0% and 19% reduction. So, under the best of circumstances, if we get a 20% reduction in the 5-day ICU LOS, this translates to about 1 day. And given that transfer timing is more likely to be driven by the availability of ward beds than by the patient's clinical readiness, I question whether this is truly a staggering reduction. Additionally, if you read on, you will realize that there is very little reason to believe that this maximal reduction in ICU LOS is unlikely to be achieved by an average institution. In fact, even the 10% seen on average in this investigation may be a bar that is too high in other less well organized ICUs.  

It is important to remember a couple of things: 1). In some circumstances there is unlikely to be any reduction in the ICU LOS; 2). Since LOS is not a normally distributed function, the mean value underestimates the true measure of central tendency in this outcome (this is due to the typically long right tail present in this distribution); and 3). This investigation, though not strictly speaking experimental, was done at 2 academic institutions with highly organized infrastructure and what looks like closed model ICUs (a dedicated specialized team of critical care professionals caring for all ICU patients). For this reason, a similar intervention at a less stringently streamlined institution is unlikely to produce the same magnitude of results.

But the mere fact that the rates of exogenous transmission of pathogenic organisms were reduced is itself encouraging. At the same time, by focusing on carriage rates and not just clinical infections, the authors may be overstating the clinical significance of the observed reduction. Additionally, one of the issues that does not appear to have been addressed explicitly has to do with the availability of sinks: In the intervention unit there was a plethora of sinks, missing in the pre- period and also not available in the comparator hospital. Is it possible then that simply putting in more sinks would accomplish the same for a lot less money?

And this brings me to my next issue with the paper -- cost effectiveness. Now, according to the AHA annual survey of US hospitals, the average age of the physical plant is on the order of 10 years. Given the rapid pace of change in medicine, this may well signal a time for capital investments in plant improvements. And surely from the patient's and family's perspective, private rooms are preferable. However, one must ask the pesky question of the return on such an investment in this era of much needed fiscal restraint in medicine. If the same outcomes of reducing the spread of infectious organisms can be achieved with merely adding sinks, this may be a less drastic and more immediately feasible intervention well worth considering.                

Monday, 10 January 2011

Reviewing medical literature, part 2b: Study design continued

To synthesize what we have addressed so far with regard to reading medical literature critically:
1. Always identify the question addressed by the study first. The question will inform the study design.
2. Two broad categories of studies are observational and interventional.
3. Some observational designs, such as cross-sectional and ecological, are adequate only for hypothesis generation and NOT for hypothesis testing.
4. Hypothesis testing does not require an interventional study, but can be done in an appropriately designed observational study.

In the last post, where we addressed at length both cross-sectional and ecologic studies, we introduced the following scheme to help us navigate study designs:
Let's now round out our discussion of the observational studies and move on to the interventional ones.

Case-control studies are done when the outcome of interest is rare. These are typically retrospective studies, taking advantage of already existing data. By virtue of this they are quite cost-effective. Cases are defined by the presence of a particular outcome (e.g., bronchiectasis), and controls have to come from a similar underlying population. The exposures (e.g., chronic lung infection) are identified backwards, if you will. In all honesty, case-control studies are very tricky to design well, analyze well and interpret well. Furthermore, it has been my experience that many authors frequently confuse case-control with cohort designs. I cannot tell you how many times as a peer-reviewer I have had to point out to the authors that they have erroneously pegged their study as a case-control when in reality it was a cohort study. And in the interest of full disclosure, once, many years ago, an editor pointed out a similar error to me in one of my papers. The hallmark of case-control is that the selection criteria are the end of the line, or the presence of a particular outcome, and all other data are collected backwards from this point.

Cohort studies, on the other hand, are characterized by defining exposure(s) and examining outcomes occurring after these exposures. Similar to case-control design, retrospective studies are opportunistic in that they look at already collected data (e.g., administrative records, electronic medical records, microbiology data). So, although retrospective here means that we are using data collected in the past, the direction of the events of interest is forward. This is why they are named cohort studies, to evoke a vision of Caesar's army advancing on their enemy.

Some of the well known examples of prospective cohort studies are The Framingham Study, The Nurses Study, and many others. These are bulky and enormously expensive undertakings, going on over decades, addressing myriad hypotheses. But the returns can be pretty impressive -- just look at how much we have learned about coronary disease, its risk factors and modifiers from the Framingham cohort!

Although these observational designs have been used to study therapeutic interventions and their consequences, the HRT story is a vivid illustration of the potential pitfalls of these designs to answer such questions. Case-control and cohort studies are better left for answering questions about such risks as occupational, behavioral and environmental exposures. Caution is to be exercised when testing hypotheses about the outcomes of treatment -- these hypotheses are best generated in observational studies, but tested in interventional trials.

Which brings us to interventional designs, the most commonly encountered of which is a randomized controlled trial (RCT). I do not want to belabor this, as RCT has garnered its (un)fair share of attention. Suffice it to say that matters of efficacy (does a particular intervention work statistically better than the placebo) are best addressed with an RCT. One of the distinct shortcomings of this design is its narrow focus on very controlled events, frequently accompanied by examining surrogate (e.g., blood pressure control) rather than meaningful clinical (e.g., death from stroke) outcomes. This feature makes the results quite dubious when translated to the real world. In fact, it is well appreciated that we are prone to see much less spectacular results in everyday practice. What happens in the real world is termed "effectiveness", and, though ideally also addressed via an RCT, is, pragmatically speaking, less amenable to this design. You may see mention of pragmatic clinical trials of effectiveness, but again they are pragmatic in name only, being impossibly labor- and resource-intensive.

Just a few words about before-and after studies, as this is the design pervasive in quality literature. You may recall the Keystone project in Michigan, which put checklists and Peter Pronovost on the map. The most publicized portion of the project was aimed at eradication of central line-associated blood stream infections (CLABSI) (you will find a detailed description in this reference, Pronovost et al. N Engl J Med 2006;355:2725-32). The exposure was a comprehensive evidence-based intervention bundle geared ultimately at building a "culture of safety" in the ICU. The authors call this a cohort design, but the deliberate nature of the intervention arguably puts it into an interventional trial category. Regardless of what we call it, the "before" refers to measurement of CLABSI rates prior to the intervention, while the "after", of course, is following it. There are many issues with this type of a design, ranging from confounding to Hawthorne effect, and I hope to address these in later posts. For now, just be aware that this is a design that you will encounter a lot if you read quality and safety literature.

I will not say much about the cross-over design, as it is fairly self-explanatory and is relatively infrequently used. Suffice it to say that subjects can serve as their own controls in that they get to experience both the experimental treatment and the comparator in tandem. This is also fraught with many methodologic issues, which we will be touching upon in future posts.

The broad category of "Other" in the above schema is basically a wastebasket for me to put designs that are not amenable to being categorized as observational or interventional. Cost effectiveness studies frequently fall into this category, as do decision and Markov models.

Let's stop here for now. In the next post we will start to address threats to study validity. I welcome your questions and comments -- they will help me to optimize this series' usefulness.                

Friday, 7 January 2011

Reviewing medical literature, part 2a: Study design

It is true that the study question should inform the study design. I am sure you are aware of the broadest categorization of study design -- observational vs. interventional. When I read a study, after identifying the research question I go through a simple 4-step exercise:
1. I look for what the authors say their study design is. This should be pretty easily accessible early in the Methods section of the paper, though that is not always the case. If it is available,
2. I mentally judge whether or not it is feasible to derive an answer to the posed question using the current study design. For example, I spend a lot of time thinking about issues of therapeutic effectiveness and cost-effectiveness, and a randomized controlled trial exploring efficacy of a therapy cannot adequately answer the effectiveness questions.
If the design of the study appears appropriate,
3. I structure my reading of the paper in such a way as to verify that the stated design is, in fact, the actual design. If it is, then I move on to evaluate other components of the paper. If it is not what the authors say,
4. I assign my own understanding to the actual design at hand an go through the same mental list as above with the current understanding in mind.

Here is a scheme that I often use to categorize study designs:
As already mentioned, the first broad division is between observational studies and interventional trials. An anecdote from my course this past semester illustrates that this is not always a straight-forward distinction to make. In my class we were looking at this sub-study of the Women's Health Initiative (WHI), that pesky undertaking that sank the post-menopausal hormone replacement enterprise. The data for the study were derived from the 3 randomized controlled trials (RCT) of HRT, diet and calcium and vitamin D, as well as from the observational component of the WHI. So, is it observational or interventional? The answer to this is confusing to the point of pulling the wool over even experienced clinicians' eyes, as became obvious in my class. To answer the question, we need to go back to definitions of "interventional" and "observational". To qualify as an interventional, a study needs to have the intervention be a deliberate part of the study design. A common example of this type of a study is the randomized controlled trial, the sine qua non of drug evaluation and approval process. Here the drug is administered as a part of the study, not as a background of regular treatment. In contradistinction, an observational study is just that: an opportunistic observation of what is happening to a group of people under ordinary circumstances. Here no specific treatment is predetermined by the study design. Given that the above study looked at multivitamin supplementation as the main exposure, despite its utilization of the data from RCTs, the study was observational. So, the moral of this tale is to be vigilant and examine the design carefully and thoroughly.

We often hear that observational designs are well suited to hypothesis generation only. Well, this is both true and false. Some studies actually can test hypotheses, while others are relegated to generation only. For example, cross-sectional and ecological studies are well suited to generating hypotheses to be tested by another design. To take a recent controversy as an example, the debunked link between vaccinations and autism initially gained steam from the observation that as the vaccination rates were rising, so was the incidence of autism. The type of a study that shows two events changing at the group/population level either in the same or in the opposite direction is called "ecologic". Similar types of studies gave rise to the vitamin D and cancer association hypothesis, showing geographic variation in cancer rates based on the availability of sun exposure. But, as demonstrated well by the vaccine-autism debacle, running with the links from ecological studies is dangerous, as they are prone to a so-called "ecological fallacy". It occurs when, despite the finding in groups of a linked change of the two factors under investigation, there is absolutely no connection between them at the individual level. So, don't let anyone tell you that they tested an hypothesis in an ecological study!

Similarly in cross-sectional studies an hypothesis cannot be tested, and, therefore, causation cannot be "proven". This is due to the fundamental property of "a snapshot in time" that defines a cross sectional study. Since all events (with few minor exceptions) happen at the same time, it is not possible to assign causation to the exposure-outcome couplet. These studies can merely help us think of further questions to test.

So, to connect the design back to the question, if a study purports to "explore a link between exposure X and outcome Y", either an ecologic or a cross-sectional design is OK. On the other hand, if you see one of these designs used to "test the hypothesis that exposure X causes outcome Y", run the other way screaming.

We will stop here for now, and in the next post will continue our discussion of study designs. Not sure yet if we can finish it in one more post, or if it will require multiple postings. Start praying to the goddess of conciseness now!

    

Reviewing medical literature, part 1: The study question

Let's start at the beginning. Why do we do research and write papers? No, not just to get famous, tenured or funded. The fundamental task of science is to answer questions. The big questions of all time get broken down into infinitesimally small chunks that can be answered with experimental or observational scientific methods. These answers integrated together provide the model for life as we understand it.

Clearly, the question is the most important part of the equation, and this is why in my semester-long graduate epidemiology course on the evaluative sciences we spend fully the first four to five weeks talking about how to develop a valid and answerable question. The cornerstone of this validity is its importance. Hence, the first question that we pose is: Is the study question important?

This is a bit of a loaded question, though. Important to whom? How is "important" defined? This is somewhat subjective, yet needs to be scrutinized nevertheless. In the context of an individual patient, the question may become: Is the study question important to me? So, importance is dependent on perspective. Nevertheless, there are questions upon whose importance we can all agree. For example, the importance of the question of whether our current fast-food life style promotes obesity and diabetes is hard to dispute.

Regardless of how we feel about the importance of the question, we must first identify the said research question. At least some of the time you will be able to find it in the primary paper, buried in the last paragraph of the Introduction section. Most of the questions we ask relate to etiologic relationships ("etiology" is medicalese for "causation"). Now, you have heard many times that an observational study cannot answer a causal question. Yet, why do we bother with the time, energy and money needed to run observational studies? Without getting too much into the weeds, philosophers of science tell us that no single study design can give us unequivocal evidence of causality. We can merely come close to it. What does this mean in practical terms? It means that, although most observational studies are still interested in causality rather than a mere association, we have to be more circumspect in how we interpret the results from such studies than from interventional ones. But I am jumping ahead.

Once we have identified and established the importance of the question, we need to evaluate its quality. A question of high quality is 1). clear, 2). specific, and 3). answerable. The question that I posed above regarding fast food and obesity possesses none of these characteristics. It is too broad and open to interpretation. If I were really posing a question in this vein, I would choose a single well defined exposure (consuming 3 cans of soda per day) influencing a single outcome (10% body weight gain) over a specific period of time (over 30 weeks). While this is a much narrower question that the one I proposed above, it is only by answering bundles of such narrow questions and putting the information together that we can arrive at the big picture.

A general principle that I like to teach to my student is the PICO or PECOT model (I did not come up with it, but am its avid user). In PICO, P=population, I=intervention or exposure, C=comparator, and O=outcome. The PECOT model is an adaptation of the PICO for observations over time, resulting in P=population, E=exposure, C=comparator, O=outcome, T=time. These models can help not only pose the question, but to unravel the often mysterious and far from transparent intent of the investigators.

Once you have identified the question and dealt with its importance, you are ready to move on to the next step: evaluating the study design as it relates to the question at hand. We will discuss this in the next post.

Series launch: Critical review of medical literature

Today I am launching a series of posts on how to read medical literature critically. The series should provide a solid foundation for this task and dove-tail nicely with some of the more dense methods themes that occur on this blog. Who should read the series? Everyone. Although the current model of dissemination of medical information relies on a layer of translators (journalists and clinicians), it is my belief that every educated patient must at the very least understand how these interpreters of medical knowledge (should) examine it to arrive at the information imparted to the public. At the same time, both journalists and clinicians may benefit from this refresher. Finally, my own pet project is to get to a better place with peer reviews -- you know how variable the quality of those can be from my previous posts. So, I particularly encourage new peer reviewers for clinical journals to read this series.  

First, a conflict of interest statement. What comes first -- the chicken or the egg? What comes first -- expertise in something or a company hiring you to develop a product? Well, in my case I would like to think that it was the expertise that came first and that Pfizer asked me to develop this content based on what I know, not on the fact that they funded the effort. At any rate, this is my disclaimer: I developed this presentation about three years ago with (modest) funding from Pfizer, and they had it on a web site intended for physician access. Does this mere fact invalidate what I have to say? I don't think so, but you be the judge.

Roughly, the series will examine how to evaluate the following components of any study:
1. Study question
2. Study design
3. Study analyses
4. Study results
5. Study reporting
6. Study conclusions
I am not trying to give you a comprehensive course on how all of this is done, but merely make the reader aware of what entails a critical review of a paper.

Look for the first installment of the series shortly.

Thursday, 6 January 2011

National Healthcare Expenditures, 2009 (In pictures)

Well, it's that time of the year again: CMS has given us the accounting of our National Healthcare Expenditures (NHE) in a paper published in Health Affairs. I am sure you have already heard that the spending only went up by 4% this year over last, an all-time low.

At the same time, we have achieved the highest ever NHE as a proportion of the GDP (17.6%) and as expenditures per capita ($8,086). But the GDP proportion is a somewhat deceptive number on the one hand, as the GDP has suffered a substantial drop from its 2008 value of $14.4 trillion to $14.1 trillion in 2009. On the other hand, this implies that healthcare is eating into the rest of our expenditures on life. At the same time the per capita expenditures have continued their relentless rise.

Let us look at the components of the NHE individually and see what they can tell us.



As usual, the bulk of the expenditures went to personal health care (85%). Public health got a measly 3% of the total NHE, and this continues to be one of our gravest misappropriations. You may recall that about a year ago I did a post where I cited some startling statistics about some broad categories of causes of premature death in the US. Access to medical care accounted for a measly 10% of those, and the rest were attributable to behavior, genetics, environment and social factors. So, while, by inference, fixing medicine may impact 10% of these premature deaths, in reality 97% of the entire NHE goes to medicine rather than to potentially more impactful public health interventions. And the real travesty is that, despite these astronomical expenditures, we are still losing 1,000 lives per day to our broken healthcare system.

Looking a bit more closely at the "personal health" category, we see that, just as in years past, hospital costs and professional services comprise the bulk of this spending.

The "professional services" category, 81% of which is physician and other clinical services, is a bit murky. Yet, without too many leaps of faith we can say that if this expenditure buys us better preventive care, it may be a cost-effective area. At the same time we know that we can make this area a lot more efficient by streamlining and realigning incentives to promote better health rather than more care. Hospital expenditures, on the other hand, are a juggernaut that without a doubt requires containing. It is very likely that exchanging our inflated personal healthcare budgets for well placed public health funding along with reimbursement reform and improved end-of-life decisions, could substantially alter this category of spending.

One final data point that interested me was the breakdown of what are considered investments in the healthcare system. This broadly includes government-funded research and allocations for structures and equipment. Now, I am not sure what "structures and equipment" means, so, if any of my readers know, please, enlighten me. I do know what "research" means, however, and am rather disappointed about this breakdown. What I do not understand is, given that structures and equipment should have some kind of a half-life and not be replaced annually, how it is that this budget also grows consistently year-over-year at a steady rate? Would love to get more details on this.

To be sure, the total research expenditure of $45 billion is nothing to sneeze at. The big question is, however, are we spending it on the right research. I am not at all sure that the answer is yes, given that we still struggle with the same issues at the bedside that we have been struggling with for over a decade. But more on this later.  

Wednesday, 5 January 2011

Radium, dopamine and innovation: Name your poison

Reading Deborah Blum's "The Poisoner's Handbook" is an intellectual treat. Although non-fiction, it paints in understated sepia tones the crevices of New York City at the dawn of the Industrial Revolution, where bootlegged booze and poisons were fare of the day, homicides went unpunished and the corrupt coroner system basked in the glow of its own willful ignorance and political approval. That is until Charles Norris and Alexander Gettler, two single-minded and tireless men, brought science into the lagging American medical jurisprudence and created the now burgeoning field of forensic medicine.

The chapter on radium in particular sparked my interest. Blum describes in vivid detail the well-known misadventure of the "radium girls", a label given to young women in a watch factory in Orange, NJ, in the early 1920s. She sets up the story with the fascinating background of radium discovery by the Curies and Marie Curie's penchant for carrying a "pet" bottle of radium in her skirt pocket, exhibiting its breathtaking beauty in a circus-like fashion (she died a horrible death from aplastic anemia induced by radiation exposure). Once its tumor shrinking properties became known, it did not take long for entrepreneurs, backed by the medical establishment, to create and sell all kinds of tonics and pills containing radium to the clueless public searching for the fountain of youth. The tragic tale of the radium girls, who, because of occupational ingestion of radium used for painting numbers on the faces of watches (according to Blum's account, the girls were encouraged to lick the paint brushes to make them pointy), and playful applications of this glow-in-the-dark paint on their lips and faces, developed debilitating jaw necrosis and other bony complications and early deaths, delivered a dose of sobriety to the public and policy makers about this new health panacea. Even the gifts of radium to Marie Curie were now delivered in a thick lead shield to contain its homicidal particles.

The story of radium raised all sorts of questions for me. When the element was first discovered, even the scientists could not conceive of its deadly health effects on human tissues. And for this reason there was no caution exercised in its use. What I puzzle over, as you may have guessed from many previous posts, is how we can balance our adoption of new glittering technologies, about which we do not have complete information, and keeping a modicum of caution about their currently unknown potentially adverse effects. I particularly wonder about this in the context of how our brains are wired and of our prevailing concerns for the economy even at the expense of our health.

Humans are seekers. I recently read Jonah Lehrer's "How We Decide", and it made me appreciate just how susceptible we are to the pleasurable effects of dopamine, and how craving its effects drives us to perform irrational acts that will soothe our neurons in a bath of dopamine bubbles. Addiction, the ultimate seeking-and-never-finding behavior, is, at least in part, mediated by dopamine. Does this addiction fuel our drive for innovation as well? And does it also make us throw caution to the wind when a desirable new object, like, say, glow-in-the-dark radium or a smart phone, is within grasp?

On the same side of this equation is the corporate voice, thundering in the background about the importance of innovation, injecting doubt about the potential for untoward effects and invoking the reigning rhetoric of Queen Economy as the ultimate justification. You don't believe me? Just look at the tobacco history, rife with denials, manipulation and lies. And this is exactly what our consumer brain wants to hear. So we paint caution as unscientific alarm and walk away from it, shaking our heads, filled with self-righteousness.

This balance that I am describing is once again the baby and bath water problem. We encounter it in every aspect of our modern lifestyle: the environment and the threat of climate change; the healthcare system with its record technology spending without commensurate results in health; our food system and obesity and superbug epidemics; the galloping pace of technological development, far outpacing our cognitive abilities to incorporate these technologies sensibly into our lives. Simply put, the question becomes, how do we harness innovation without demanding corpses (literally and figuratively) as proof of its potential untoward effects?

The first step is clearly understanding our history, and for this read Blum's book -- you won't be sorry! Next we need awareness of how our brains operate and how these biological principles set well known traps in our reasoning. Using metacognition to understand these pitfalls in thinking may at least put us on a smarter course walking this fine line. Finally, as I have advocated before, we need to stop shouting at each other and start listening. Perhaps we are not so drastically different in our views as the press and politicians will have us believe. After all, we are all susceptible to the same poisons. And dopamine.                

Tuesday, 4 January 2011

Guest post: How our brains are wired to advance science

We have a treat today. Today I am featuring a guest post from my brilliant 17-year-old niece Katherine Dana. She is currently applying to colleges, and this is one of her brief essays. Kathy is interested in animal communication specifically, but, as you can see, also spends a lot of time thinking about science in general. And oddly, she seems to be contemplating similar themes to the ones we address here. 
While it is hard for me to stop waxing poetic about how proud I am of her, I will now cut myself short, so that you can enjoy her lucid commentary.

By Katherine E. Dana

Marcel Proust once wrote, "The real voyage of discovery consists not in seeking new landscapes, but in having new eyes." Thus goes the song of science, humanity's great unifier. Science is not merely the means for collecting random information—it is the means through which we make sense of our world. It is messier than mathematics, less exact. And yet in some ways, it is this very inexactitude that gives science its potency, and allows it to cut to the very heart of nature's chaotic randomness. It works by taking the givens of nature and churning out elegant guesses, which predict as effectively as they describe.

One quality that distinguishes mind from machine is that leap of thought that psychologists term "heuristics"—mental shortcuts, expressly designed to help us connect the dots without having to consciously traverse the spaces between. This is our organic advantage.

While today's machines, no matter how complex, are restricted to lengthy algorithms, we may leap from branch to branch. Nowhere in human endeavors is this cognitive edge more apparent than in the combined efforts of humans seeking to find new truth. For before we can know, we must question; and this is where insight is most crucial. It is not enough to investigate the familiar. We must find the courage to ask uncomfortable questions, and be willing to uproot even our most cherished beliefs, all in the name of a deeper understanding.