Download as pdf
Download as pdf
You are on page 1of 9
Corres ReaeNis bo Bor tooth aa Misunderstanding Analysis of Covariance Gregory A. Miller University of Minos, Champaign Jean P. Chapman University of Wiseonsin—Madison Despite mumerous technical eatmens in any vanes, analysis of covariance (ANCOVA) semaine & widely misused approach to dealing with substantive group differences on potential covariates, paic= layin psychopathology esearch. Published aces reach unfounded coaclusons, and some sass texts neglect the itu. The problem with ANCOVA in such cases i reviewed In many cases, there is no means of achieving te superficially appealing gos! of “coreting” or “controlling for" real group Aitferences on a potential covariate In hopes of curing misse of ANCOVA nd promoting spr priate use «nontechnical discussion is provided, emphasizing a substantive confound rarely articulated In textbooks and other general presentations. o complement the mathematical eviqus already available Some atemstives ae discussed for contexts in which ANCOVA is inappropriate or questionable In research comparing groups of participants, classical experi- ‘meatal design (Campbell & Stanley, 1963) relies, whenever pos sible, on random assignment of participants to groups. Observed, differences between such groups, prior to experimental treatments, fare due to chance rather than being meaningfully related to the ‘10up variable. In contast, when preexisting groups are studied, observed pretreatment differences may reflect some meaningful, substantive differences that are atteibutable to group membership. It has been noted that, “Even in the absence of a true treatment effect, the outcome scores in the treatment groups are likely to differ substantially because of initial selection differences, AS a result selection differences area threat to validity. ..." (Reichardt & Bormann, 1994, p. 442) In this article, we discuss why attempts to control statistically for such differences ar, in general, inappropriate. For example, consider a data set consisting of age as a potential covariate grade in school as the grouping variable, and basketball performance as the dependent variable. An analysis of covariance (ANCOVA) sight be run in hopes of asking whether 3rd and 4th graders would differ in performance were they not different in age. This might, seem t0 be a reasonable question, in that one could ask whether some maturational change at that age makes a nonlinear contibu- tion to basketball ability. However, infact t makes no sense to ask hhow 3rd graders would do if they were 4th graders. They are, Gregory A. Miler, Departments of Pyehology and Psychiatry and Beckman Isitae, University of Minos, Champaign: Jean . Chapman, Deparment of Psychology. University of Wisconsi~ Maison. ‘The writing ofthis article was supported by Nationa Institute of Mental Health Grant MH39628, We thank Loren J. Chapman, Lawrence J. Hubert Brandy G. aac, Jack B. Nise, and E, Keolan Taitano for comments on an calier dat (Corespondence concerning this atcle should be adresed to Gregory A. Mille, Deparment of Psychology, University of Mlinos, 603 East Daniel Stet, Champaign, lois 61420, oro Jean P. Chapmaa, Depart ment of Prychology, University of Wisconsin, 1202 West Jonson Stel, Madison, Wisconsin $3706. Electronic mail may be seat to gamiller 2uie ey oF jhapmad @ acs. wise.ed, inherently, not 4th graders, and ANCOVA cannot “control for" that fact. Age is so intimately associated with grade in schoo! that, removal of variance in basketball ability associated with age would remove considerable (perhaps nearly all) variance in basketball, ability associated with grade. The results of the ANCOVA would, bbe meaningless. Asa complement to this problem of the covariate removing too much of the independent variable of interes, a problem can arse when preexisting groups differ systematically on ‘more than the covariate. The covariate will leave those differences intact, thus biasing the estimate of the teatment effect (Reichardt {& Bormann, 1994), which has been called specification error Nonrandom Group Assignment When group membership is determined nonrandomly, there is \ypically ao thorough basis for determining whether a given pre tweatment difference reflects random error or true group differ- cence. This uncertainty complicates interpretation of apparent teat- ment effects because itis impossible to distinguish a main effect of treatment from an interaction between the effects ofthe treatment and the pretreatment difference, and from meaningful overlap (wariance shared) between te treatment and the pretreatment char- ‘cteritic, The confound due to an undetected interaction is widely understood—the pooled regression is inappropriate as a basis for “correcting” for the covariate. However, the ther confound, due to undetected or misinterpreted overlap of pretreatment difference and grouping factor, is commonly neglected, although it was noted inthe first modem treatment of ANCOVA (Cochran, 1957) and. ‘occasionally more recently (eg, Porter & Raudenbush, 1987), The problem of preexisting group differences arises very com: monly in psychopathology research because random assignment .0 ‘group (diagnostic category in the typical experimental design in psychopathology but called “treatment” in much ofthe statistical iterator) is routinely infeasible andfor unethical. I is thus ex tremely tempting for psychopathologsts to seek to use analytic methods in an attempt to avoid the interpretative problems that, arise when groups differ pretreatment. Ii unfortunate that, inthe general case, no such analytic method is available, nor can one be SPECIAL SECTION: ANALYSIS OF COVARIANCE 41 ‘developed (a point that will be explained below). This logical fact, has proven difficult for psychopathologiss to aceept, perhaps ‘because ofthe burden that preexisting group differences place on interpretation of experimental results. Despite numerous technical treatments in the Inerature (eg Chapman & Chapman, 1973; Cochran, 1957; Elashoft, 1969; Fleiss & Tanur, 1973; Huitera, 1980; Jin, 1992; Lord, 1967, 1969; Maxwell & Delaney, 1990: ‘Maxwell, Delaney, & Manheimer, 1985; Porter & Raudenbush, 1987; Wainer, 1991; Wildt & Ahtola, 1978) and more accessible statemens (eg, Neale & Oltmanns, 1980; Side & Turpin, 1980), together making an overwhelming case against inappropriate at tempts to “control for” such group differences, they remain com ‘mon inthe research literature and, if anything, even more common in research grant applications Given the continuing popularity of inappropriate uses or interpretations of ANCOVA, the present nice offers a relatively nontechnical critique, in hopes of helping to popilarize the correct use of ANCOVA and helping researchers to avoid its more common abuses. Understanding Analysis of Covariance ANCOVA is part of the ANOVA (analysis of variance) tra tion, ANCOVA was developed to improve the power ofthe test of, the independent variable, not to “control” for anything. Itis helpful haere to place ANOVA and ANCOVA in the more general frame- ‘work of multiple regression and correlation (MRC), understood within the general linear model. In fact, at least one popular ANOVA package, BMDP2V (Dixon, 1992), actually computes is ‘analyses using multiple regression rather than using the equations typically presented in textbooks on ANOVA. Some sources place ANCOVA in the context of MRC, with the covariate in ANCOVA, Understood as a regression predictor entered before what are, in ANOVA, the main effets and interactions (eg. Harris, Bisbeo, & Evans, 1971: for extended treatments of this viewpoint, see Cohen, & Cohen, 1983, and Judd, McClelland, & Smith, 1996). More ‘commonly. other sources have argued that ANCOVA is not strictly ‘equivalent to regressing the dependent variable on the covariate, then doing an ANOVA on the residuals (eg. Elashoff, 1969; Maxwell & Delaney, 1990; Maxwell etal. 1985; Porter & Rat Most of his repeated in Cohen and Coben (1983, 9.425), with the ‘esauim coeluson das “4 MILLER AND CHAPMAN cians analyzing a hyposhetical dataset. The data are for boys and girls atthe beginning and end of an academic year, The boys, a & ‘group, stat and end the year weighing more than the gies. and neither groups’ average weight changes over time, AN issue is, whether diet affected boys and girls differentially during the School year. The statistician who favored an ANCOVA (incor recily 80, inthe opinion of Lord and numerous subsequent authors {quoted above) used inital weight as a covariate and concluded: one selects on the basi fini weight a subgroup of boys and a suberoup of gris ving tential frequency disibutons of ita weight. the subgroup of boys is going to gaia substanaly more uring the year than the subproup of gis. (Lod, 1967, p. 308) Lord's pro-ANCOVA statistician concluded from this that there is, ‘a meaningful differential effect of diet on boys and gils. However, the two selected subgroups are not representative of the larger ‘soups of boys and girls. In effect, the two levels of the gender effect no longer represent the samples of boys and girls in the coviginal analysis. What this statistician observed is merely an example of the principle of regression toward the mean, By se- Jeoting subsets of boys and girls matched on weight the statistician will have selected boys weighing less than the boys’ mean and girls weighing more than the girs’ mean. Regression towand the ‘mean, when the participants are reassessed, would be expected solely usa result of te lack ofa perfect coreation between initia ‘and final weight, Thus, coraparison of the subset of boys gaining ‘weight and the subset of girls losing weight is not evidence for systematic differential elfects of diet on boys aad girls Lord's (1967) example is compelling, in part because he con- structed a case in which itis known that the groups do not change lover time, so they cannot change differentially over time, as a function of diet or any other factor. We suggest that the more ‘general point in this example is that, because gender and weight, ate confounded (correlated) to begin with, there is no statistical ‘means 1 unconfound them in these examples—here, to study the Potential differential effect of diet on the two groups. ‘As another illustration, consider a dataset in which two groups are older men and younger women, and gender is of interest as an independent variable, Grp. Using age as a covariate does indeed remove age variance. The problem is that because age and gender ‘are correlated in this data set, removing variance associated with Cov will also remove some (shared) variance due to Grp. Within ‘this dats set, there is no way to determine what values of DV men ‘younger than those tested or women older than those tested would have provided. Far from “controlling for" age, the ANCOVA will, systematically distort the gender variable. As in our presentation of Lord's Paradox above, Grp will not be a valid measure of the construct of gender. Again, this is seen in the right panel of Figure 1, when area 2 and area $ are removed from Grp. ‘Consider a data set consisting of childrens’ age, height, and weight, If we conduct an ANCOVA in which height isthe covari- fate, age is the grouping variable, and weight is the dependent variable, we are attempting to ask whether younger and older children would difer in weight if they did not happen to differ in height. Ifthe groups indeed do not differ on the covariate, this ‘question can be asked. Bur if there is something about the construct ‘of age in childhood that inherently involves differences in height, the question makes no sense, because then age with height par- tialed out would no longer be age. There is no way to “equate” ‘older and younger children on height, because growth is an inher= ent (not chance or noise) differentiation of the two groups. ‘As noted above, it would be entirely reasonable, given such a dataset, co explore the relationships among age, height, and weight, (Cohen & Cohen, 1983). The point is that statistical contol, in the sense of cleanly removing the effect of Cov. is not what one Would, be able to accomplish with ANCOVA. Cohen and Cohen (1983), provided the following extreme example: “Consider the fact that the difference in mean height between the mountains of the Himalayan and Catskill ranges, adjusting for differences in atmo- spheric pressure is ero" (p. 425), the point being that one has not in any sense “equated” the two mountain ranges by using atmo: spheric pressure as a covariate Invalid ANCOVA in Psychopathology Research (One can readily imagine additional examples in psychopathol- ‘gy. If likelihood of diagnosis of schizophrenia rises with age, it ‘would not be appropriate to try to “contol for” age differences, ‘between two samples to evaluate the relationship between propor- tion of each sample receiving a diagnosis and some dependent variable such as current employment status. Age would be sys- tematically related tothe defining characteristic ofthe groups, so removing Variance associated with age would, in effect, corupt the grouping variable itself Ina real-world example, Deldin (1996) wanted to investigate whether @ brain-wave measure known as P300 is reduced in, ‘depression. This question is complicated by the fact of consider able comorbidity between anxiety and depression, noted above ‘Two groups of participants, diagnosed as depressed and nonde- pressed, completed a self-report anxiety measure. Anxiety score ‘might be used as a covariate in an ANCOVA with diagnosis asthe independent variable and P300 as the dependent variable, The hhope in such an analysis would be “to contr” anxiety and thus be able 10 observe the relationship between pure depression (not confounded with anxiety) and P300. However, this would make no sense on clinical, substantive grounds (fortunately, Deldin did not conduct an ANCOVA), The literature does not generally view anxiety as an independent, confounding variable in depression but, as intimately related to depression and perhaps, in part, indistin- sguishable from it (¢g., Helle, Etienne, & Miller, 1995; Watson et a, 1995), Statistical methods cannot remove the “effect” of anx- jety from depression if conceptually they are overlapping ‘As a final example, consider the effects of using gender as a covariate in a comparison of schizophrenic and depressed groups, diagnoses in which itis well established that men and women, respectively, are overrepresented (Kessler etal, 1994; Walker & Lewine, 1993). Removing gender variance could systematically alter the apparent nature of and relationships between the diagnos tic groups If, further, the dependent variable were performance on, ‘task on which women are believed to be superior (i.e. if Cov correlates with both Grp and DV), such an ANCOVA would corrupt both independent and dependent variables, providing a meaningless F test. ‘Misuse or misinterpretaion of ANCOVA is so widespread in the psychopathology literature, in our experience, that itis difficult, to cite examples without stepping on toes. The report by Rosvold, Mirsky, Sarason, Bransome, and Beck (1956) may be used without SPECIAL SECTION: ANALYSIS OF COVARIANCE 45 ‘causing offense, because it was published in a respected journal before the problem was recognized (Cochran, 1957) and because it isa deservedly famous article for other reasons, having recently been republished as a landmark Journal of NIH Research, 1997) ‘Among other important contributions, this paper introduced the still widely used Continuous Performance Test (CPT) to the neu ropsychology and psychopathology literatures. The design in- cluded several groups of child and adult brain-damaged and con- trol samples, with X and AX denoting two types of tral. Since there was a sgniticant age dtference between the Child sub- ‘roups and 4 significant IQ diflerence between the Adalt so. ‘roups the differences between the pied subgroups means on X and AX in these ovo groups were evaluated by means ofan analysis ‘of covariance. (Rosveld eta, 1956, p. 346) Unfortunately, 1Q would be very likely to be meaningfully related to brain damage, so using 1Q 2s a covariate would disrupt any comparison of brain-damaged and control groups’ performance: 1Q differences would almost certainly be part of group differences, jn braindamage status. AS a consequence, removing variance associated with 1Q would alter the diagnostic group variable sub- stantively, and Grp,, would not merely be a de-noised surrogate {or Grp. Age as a covariate presents the same potential problem, ‘but itis not as likely that age is substantively related to brain damage, so ANCOVA might be viable. (The legitimacy of ANCOVA in the face of group differences on the covariate if the differences arose by chance is discussed below.) This range of examples serves to convey the substantive (nat ‘metely mathematical) problem that arises when groups differ ‘meaningfully on the covariate, ANCOVA does indeed “remove” the variance due to the Cov, but it does not successfully “control” for Cow if Cov is a meaningful part of Grp. On the contrary, and unfortunately for well-intentioned psychopathology researchers, ANCOVA removes meaningful variance from Grp, leaving an ‘undercharacterized, vestigial Grp,., with an uncertain relationship to the construct that Grp represented. The relationship of such a Gy 10 DV of DV, is often not interpretable. In. general ANCOVA is appropriate when the groups do not differ on the covariate—when inclusion of the covariate serves merely to re- ‘move noise variance unrelated tothe grouping variable (left panel of Figure 1), Possibly Valid ANCOVA When Groups Differ on the Covariate Beyond this basic point, there are some gray areas regarding the use and misuse of ANCOVA. Bock (1975) and Wainer (1991) noted that the appropriateness of an ANCOVA depends not only ‘on meeting statistical assumptions but also on the nature of the {question posed. Heckman (1989, p, 166) stated, “A decision about the appropriate statistical procedure requires information outside of statistics.” Feiss and Tanur (1973) concluded that i s not the analysis but, in part, “the phrasing of the hypotheses and of the inferences which are usually invali (p. S18) in the misuse of ANCOVA, suggesting that if properly framed the analysis can be appropriate (see Fleiss & Tanur, 1973, pp. 518-521, for further «iscussion). Cochran (1957), Evans and Anastasio (1968), Max- well and Delaney (1990), and Porter and Raudenbush (1987) also discussed narrow purposes or conditions under which ANCOVA, ‘might be interpretable despite nonrandont assignment ‘There can thus be a range of views on how to interpret certain cases in which groups differ on Cov. Even in the case of random assignment to groups, nontrivial differences on Cov will occasion- ally arise by chance. In the right panel of Figure 1, the issue is, ‘whether area 2 and area 5 are substantively important parts of Grp. If not, then after their removal Grp... would stil be a valid measure of the Group construct. How is an investigator to know when a Type T error has ‘occurred in a test of group differences on Cov? A moderate position might be that, if the investigator has good reason to believe that group differences on Cov truly arose by chance, ANCOVA is appropriate (Maxwell & Delaney, 1990). The ratio- nale for this view is that ANCOVA would only be removing noise variance from Grp, not anything substantive about Grp. Ofcourse, this rationale turns on the strength of the assumption that the differences on Cov di, infact, arise by chance. In other words, the investigator must be convinced (and must convince readers) that, ‘among the populations from which the groups are sampled, there is no relationship with Coy and thatthe available samples are sometimes nevertheless a good representation of those popula- tions. Inthe case of random assignment, the basis for this judgment is quite strong, assuming good execution of the group assignment ‘process, In the case of nonrandom assignment—almost invariably ‘the case in studies of psychopathology —this is often a difficult, judgment to defend, However, investigators should be fee to consider it and to make their case. Overall and Woodward (197) went a step beyond the arising randomly criterion and argued that what matters is whether Grp ‘could have caused the group differences on Cov. If not, in theie view, then ANCOVA may be legitimate even if the group differ- ‘ences on Cov are substantive. Similarly, Wildt and Ahtola (1978) suggested that ANCOVA might be acceptable if the investigator is certain that the Grp could not have affected Cov. We (see also ‘passages quoted above from Elashoff, 1969; Maxwell & Delaney, 1990; and Porter & Raudenbush, 1987) do not find this argument compelling, in general. It will stil happen that, because Cov will be entered into the regression before Grp, Cov will in effect get credit for any relationship of theit shated variance that is also shared with DV. Thus, not only will Grp, be a much-diminished representation of the construct that Grp measures (a conceptual problem) but also the relationship of Grp, and DV,,. may be underestimated (a statistical problem). Cohen and Cohen (1983) ‘offered another objection, in that itis often impossible to deter- ‘mine causal relationships between Grp and Cov. Overall and ‘Woodward offered a more specific analysis, however, which is sound: Inthe absence of random assignment, one might be able 10 render ANCOVA legitimate given an experimentally appropriate assignment to group on the buss of scores on the covariate. Their argument turns on particular patterns of relationships among Cov, Grp, and DY. In effect, the point is that the legitimacy of ANCOVA depends on a careful examination of these relation- ships. Unfortunatly, this suggestion obviously would not apply ia studies of preexisting groups typical in psychopathology. ‘One other gray area to consider isthe context of exploration of 4 data set, in contrast t0 the more commonly explicit goal of hhypothesis testing. As Huitema (1980, p. 108) outlined, it may be useful to conduct a series of ANCOVAS, trying different variables 46 (MILLER AND CHAPMAN asthe covariate, as a means of understanding the patterns of shared variance in a dataset. He hastened to add that, when groups differ fon the covariate, “it would be reasonable to speculate that the ‘weatment effect on the dependent variable is mediated by the ‘covariate. It covldn’t be concluded, ...” In effec, Huitema was, escribing the strategy of Cohen and Cohen (1983) in using a variety of hierarchical egression analyses fo evaluate the relation: ships among variables. Procedures such as those discussed by Baron and Kenny (1986) might then be used to confirm media- tional relationships Beyond Statistical Concerns We have emphasized that the problem with ANCOVA in the face of group differences on the covariate is as much a substantive and interpretative issue as it is a mathematical issue, I ean be noted that the problem of misuse of ANCOVA putatively to “coniol for” differences on the covariate is not confined (0 aca deme, An article in prominent weekly political magazine sited: (On the average, stants in predominantly White distiets did much beter on reading tests than those in predominantly Black dsr. The reavon? Not poverty, parent educa, of cass size. The White disc were, of couse, generally beter off, the parents of pupils ‘wore more highly edvate, and so og. But when these diferences were held constant sane the difference that made a iference ‘vas the quality ofthe teaching af as measored by a language skils exam. (Themstim, 1991, p. 22; emphasis add) It is apparent from this example that the public-policy stakes associated with the misuse or misinerpretaion of ANCOVA can bbe quite high. As we have argued on substantive grounds, com- plementing previous technical discussions, cis simply not posible {or such differences to be “held constant statistically” as this article claims was done, unless we were to suppose that teachers were assigned to While and Black school districts randomly or that, differences in poverty, parental education, and class size existed ‘only by chance, Surly, teacher recruitment (and subsequent per ormance) is in part drven by differences in district poverty, parent education, class size, ete, All ofthese variables are correlated, and they must be understood as inherently confounded if one adopts, the superficially appealing but inappropriate goal of “controlling foe” some of them. Ukimately, “quality of teaching staf” is nota variable substantively lft in the analysis by the time all of those ‘covariates that are surely meaningfully correlated with it have been partaled out. Only a highly residualized variable given the same hame is being tested. One could not conclude anything about the original variable, “quality of teaching staff" At best, one could speculate, as Huitema (1980) suggested. Such nuances are often lost in public debate and, in our experience, in the psychopathol- ogy literature ‘Some Alternatives to ANCOVA. analysis of covariance cannot tell us how groups would differ if they did not differ on the covariate. What then should the investigator do instead of analysis of covariance? Max~ well and Delaney (1990) discussed the analysis of gain scores and the blocking of subjects onthe covariate, noting limitations in both strategies lative to ANCOVA (see also Reichardt & Bormann, 1994). Cohen and Cohen (1983) and Harris etal (1971) recom ‘mended incorporating the covariate into the analysis-—no longer conceived as a covariate but as another substantive variable. Both of these are sound, relatively familiar stategies. Rosenbaum and Rubin (1984; Rosenbaum, 1995: Rubin, in press) discussed the less widely Known concept of propensity score, the conditional probability of assignment to a particular group or treatment given 4 set of observed covariates. They noted that there are circum stances under which propensity score analysis can balance the ‘group on the covariates, [t cannot addres unobserved differences, in covariates, although there are methods of determining the extent, of bias arising from unobserved covariates (Rosenbaum, 1984). [Less specific bu potentially fruitful i the aggregation (pests via formal meta-analysis) of a large number of individually compro- ‘mised studies, none of which may have random assignment but collectively presenta strong inferential case (Rosenbaum, 1987), For example, no study of tung cancer has undertaken random assignment to chronic smoking and nonsmoking groups, but a strong case for causal factors in the face of potentially confounded covariates (Guch as hypertension, on which smokers and non- smokers might vary and on which smoking can have causal ef- fecs) has been made on the basis of a varity of different studies. Importantly for psychopathology research, Rosenbaum’s sugges tion can be generalized beyond traditionally defined control ‘soups. For example, ina study comparing symptoms of psychosis, ‘observed in schizophrenia and bipolar patients, there ste advan tages to recruiting multiple samples, known to be different, within each diagnosis, Placement of group means on the regression lin. Fleiss and ‘Tanur (1973) suggested another alternative to ANCOVA. They reasoned from the principle that a problem in using analysis of ‘covariance forthe comparison of intact groups is thatthe regres sion line within groups cannot legitimately be used for between: ‘groups comparisons. They inferred from ths that the investigator should consider the performance of many differing groups of Params, examining tow diferent groups lice the reae- ‘As a simple example to ilustrate their method, an investigator ‘might wish to test the hypothesis tha schizophrenic patients show 4 greater concreteness of proverb interpretation than is accounted for by their lower Verbal 1Q soore on the Wechsler Adult Intel ‘gence Scale (WAIS), (Schizophrenic individuals’ thinking. has ‘often been characterized as more concrete and less abstract than nonpatiens’. Judging that a hammer and a screwdriver are both tools is more abstract than that they are both made of metal.) The investigator recognizes that she or he cannot simply match sub- groups of schizophrenic and control participants on WAIS 10 and then compare the matched subgroups on concreteness because such a comparison would be contaminated by regression toward the mean (ee discussion of Lord's Paradox above). Feiss and ‘Tunur's (1973, p. $22) solution would be to extend the study of IQ and concreteness to “many different kinds of subjects (normals, rneurties, depressives, etc), all having in common the fact thet, they are not schizophrenic.” Then, ater obtaining mean scores on IQ and on WAIS Verbal 1Q for each group, the investigator finds ‘transformation of the scores such that @ nearly linear relation can be fitted to all pairs of group means. In our example a straight line ‘would be fitted o the elation of mean group concreteness score t0 SPECIAL SECTION: ANALYSIS OF COVARIANCE 47 ‘mean group 1Q score, Then the means ofthe schizophrenic group are examined in relation to the fitted line. Ifthe schizophrenic patients score more deviantly on concreteness than predicted by the multigroup regression line, one infers that their concreteness is, ‘unusual ‘This solution is problematic. It is, of course, rather impractical {o test many groups in most such studies. More importantly, the ‘method may be flawed in some eases. If as Fleiss and Tanur suggested, only wansformation data rather than orginal data ft a straight line, the meaning ofthat linearity is unclear. Some trans formations might alter the data considerably, and the transforma tion is selected to fit all of the pairs of group means to a straight line except those of the schizophrenics. This procedure appears ‘vulnerable to random error that could foster the predicted result, that the means of the schizophrenic group, but not the other groups, do not fit that line. Tt is quite possible that, if one included the schizophrenic group among those for which a straight line is, fied, but excluded some other group, such as depressed patients, ‘the means for that excluded group would fail to fll on the line, xs a procedural artifact. ‘Regressing the dependent variable on the independent variable Jor control participants of a wide range of performance. M. B. Miller, Chapman, and Chapman (1993) suggested supplementing, the “normal” control group with additional persons from a kind of group that isnot grossly pathological but is known to be impaired (0m Variable that might be a confound in the results. They studied the prior preparatory interval (PPD) effect in schizophrenia (where- in reaction time depends on the length and predictability of the imerttal interval) and wished to determine whether schizophrenic patients’ greater overall slowness might be a confound. Accont- ingly, the investigators added elderly individuals to their normal control group because the elderly tend to demonstrate overall, slowness on reaction time. They found a very similar regression slope of the PPI effect score as a function of overall slowness for elderly as for younger normal participants. Accordingly, they computed a combined regression line. Using the slope and inter- ‘cept ofthat line, they computed residualized scores forthe schizo- ‘phrenic patients’ PPI effect. To apply this method to the example of concreteness of proverb interpretation in schizophrenia, one might add borderline retarded individuals to the control group, compute the regression of con creteness score on 1Q score for this expanded group of noapatents, an then use the regression line to compute residualized scores of ‘concreteness forthe schizophrenic patients. A finding of deviantly high residualized concreteness scores for the schizophrenic pa tients would indicate that their concreteness is greater than ex pected by the standards of a group tht includes borderine retarded persons, ‘This method has several limitations. First itis based on the hhypothesis that the regression line for the original normal group and that forthe additional impaired nonschizophrenie group have the same slope and intercept. If either the slopes oF intercepts should differ much, the combining of the two groups to compute joint regression line would appear inappropriate. In addition, the range of scores on the predictor variable (IQ in the example) mast bbe as great forthe combined control group as forthe schizophrenic patients. This is erucially important forthe residualized scores to be meaningl ‘A third problem is thatthe answer yielded by the study does not precisely match the original question. The investigators’ hypeth- esis was probably not concerned with a comparison of the con= ‘reteness of schizophrenia with that of borderline retardation. In ‘our view, however, this kind of answer is often beter than the availabe alternatives, Substantive Questions Future work may produce more general or more satisfactory ‘means to address the question psychopathologists usually attempt, to answer with ANCOVA, a question for which the technical literature shows ANCOVA to be inappropriate: How would the ‘soups differ on DV if they did not differ on Cov? However, we hhave argued that fundamental logical problems with such a ques- tion preclude its being meaningful Rather than pursue such a question, we believe that psychopa- ‘thologists would do well to frame questions that arch experimen- tul design can address. For example, the high comorbidity berween depression and anxiety need not be seen as a bartier to research, ‘The comorbidity suggests that depression and anxiety are not entirely distinet concepts and not phenomena that should be sep arated or can be represented in a fully factorial design Instead, one ‘an articulate a specific relationship between them and then design ‘an appropriate study. An appealing research design might include depressed individuals varying in level of anxiety and/or anxious individuals varying in level of depression. A nonadditive model of the characteristics and dynamics of comorbid depression and anx- Jety may best fit the resulting data, Similarly, variables that have often been viewed as confounds in research on schizophrenia may be incorporated into models as significent concepsusl players. For example, given the large liter ature on cognitive deficits in schizophrenia, reduced 1Q need not be viewed as something to “control for” but rather asa feature of the disorder in many cases, to be studied as one studies more traditional clinial symptoms. The broader lesson here is that, if ‘one’s questions cannot be served by one's methods, one can reconceptualize the questions to suit the best available methods Its our hope that the present discussion provides an accessible portrayal of the proper use of ANCOVA. Relevant technical lit ‘erature overwhelmingly condemns its uncritical use in the face of aroup differences on the covariate despite its continuing popularity jn such situations. Noprandom assignment to groups presents psychopathologists with design, statistical, and interpretative chal- lenges, and we have provided some comments on attempts to ‘address those challenges. Studies of psychopathology often in volve group comparisons without random assignment. investiga tors should consider ANCOVA and the alternatives briefly re- Viewed hereto improve their designs. A more fundamental point is that investigators should reexamine the issues they may have wished to address with ANCOVA and consider whether other issues are more fruitful theoretically and more achievable ‘methodologically. References Baron, R.M, Kenny, D_ A. (1986), The mosertor-mediaor variable distinction in soil psychological researc: Conceptual, stati, and statistical considerations. Jounal of Personality and Socal Psychol 293, $1, MTS-U182, 48 MILLER AND CHAPMAN Bock, D. (1975) Multivariate statistical methods in behavioral research "New York: McGraw-Hill Campbell, D. T., & Stanley, J. C. (1963). Experimental and quasis ‘iperimental designs for research, Chicago: Rand MeNally. (Chapman, LJ. & Chapman, J.P (1973). Disordered thought in schizo Phreria, New York: Applctn-Conury Cros CCochvan, W. G. (1957). Analysis of covariance: Is ratre and uses. Biomerics, #4, 261-281, Cohen, 1. & Coben, P (1975) Aplled multiple regressiowcoreation ‘aalisfor the behavioral sciences. Hillsdale, NJ Bbaum, Coben, J, Coben, P. (983). Applied muliple regressiwtorrelation ‘anaiyls for the behavioral sciences (2nd eo). Hillsdale, NJ: Ee ‘baum, Deldin, PJ. (1996) formation processing In major depression: The ERP “conection. Unpublished doctoral disertation, Unversity of Tints, Champaign, Dixon, W. J (1982), BMDP saisical software manual. Los Angeles University of California ress last, J.D. (1969), Analysis of covariance: A delicate instrument ‘American Educational Research Jour, 6, 383-401. Bans, §H, & Anastasio E. J. (1968), Misuse of analysis of covariance Wwhen teatment effect and covariate ae confounded. Psychological Bulletin, 69, 225-238, leis, J. L, & Tan, JM (1973). The analysis of covariance in poycho- pathology. In M. Hammer K. Salzinger, S,Suton (Eas), Prychopa- ‘thology: Contributions from the social, Behavioral, and bolgical ‘ences (pp. 508-527), New York: Wiley Harris, DR, Bisbe, C.T., & Evans, §. H. (1971, Further comments Misuse of analysis of covariance. Peychological Bulletin 75, 220-222. Heckman, 1.1989). Casi inference and nonrandom samples. Jour ‘of Educational Saisies, 14 159-168, Heller, W..Eiene, M. A. & Miler, G. A. (1995). Paes of percept "asymmetry in depression and anxiety: Implication for neuropsycholog- ical models of emotion and psychopathology. Journal of Abnormal Psychology, 104, 327-333. Holland, P. W., & Rubin, D.B. (1983), On Loris parax. In H, Waies 1&S. Messick (Eds), Principles of moder psschological measurement (op. 3-38) Hillsdale, NJ Evbaum Huitema, B. (1980). Analysis of covariance and alternatives. New York ily. in, P. (192). Toward a reconcetusization ofthe Law of Iti Value Peychologcal Bulletin, 111. 76-188 Judd, C. M, McClelland, G. H., & Smith, E.R, (1996). Testing treatment by covariate interactions when Weatment varies within subjets. Psycho logical Methods, 1, 365-378. Keller, Jy Niels JB, Bhargava, T, Deldin, P.. Geegen.J. A. Mills, (GA. & Heller, W. (2000), Newropsycholopical differentiation of de- ‘ression and airy. Journal of Abnormal Psychology. 109, 3-10, Kessler, RC., MeGonagle, KA. Zha,S.. Nelson, C. B. Hughes. M., TEsleman, 8, Witchen, H, & Keadle, K. S. (1994). Lifetime and |2:monih prevalence of DSM-II-R psychiatric disorders inthe United Sates Archies of General Prychiany, $1, 8-19. Lord, FM (1967). A paradox inthe interpretation of group comparisons. Poychological Bulletin, 68, 304-205 Lod, FM. (1969), Statistical adjustments when comparing prexisting ‘sr0ups, Prychological Bulletin, 72, 336-337, Maris, E. (1998). Covariance adjustment versus gain scores—revisite. Prychological Methods, 3, 309-327. Maxwell, S. E, & Delaney, H. D. (1990) Designing experiments and analyzing deta: A model comparison perspective. Belmont, CA: Wads- worth ‘Maxwell S. E, Delaney, H. D, & Manbeimer, J. M. (1985). ANOVA of tesiduls and ANCOVA: Corecting an illson by using madel com parsons and graphs. Journal of Educational Stansis, 10, 197-208 Miler, M.B., Chapman, L. 1 & Chapman .P (1993). Sloweess and he ‘preceding preparsory interval eect in schizophrenia. Journal of Ab normal Psychology, 102, 145-181, [Neale M, & Oltmanns,T.F. (1980) Schizophrenia, New York: Wile. ‘Oven, & Woodward J, (1977), Nonrandom asignnent and the analysis of covariance. Psychological Bulletin, 84, 588-598 Ponet, A.C. & Raudenbush, S. W. (1987) Analysis of covariance: Ws ‘model and use in piyebological research. Journal of Counseling Pry chology, 3, 383-392, Reichard CS. (1979). The staistical analysis of data from nonequivalent soup designs, In T, D, Cook & D. T. Campbell (Eds), Quast-exper- ‘mentation: Design and anaes ives for fed stings (pp. 147-208), Boxioa: Houghion Mifi, Reichardt, C.S. & Bormann, C. A. (1994), Using regression modesto ‘eimate program effecs. Ia J. S. Wholey, HL P. Hauy, & K. E. "Newoomer (Eds), Handbook of practical program evaluation (417 1455) San Francie: Josey Bas. Richards, I.E, (1980) The sata analysis of heat rae: A review ‘emphasizing infancy dats, Psychophysiology, 17, 183-166, Rogos, D. (1980), Comparing nonparalel zegresion lines. Psychological Bulletin, 88, 307-321 Rosenbaum, P. R. (1984). From association to causation in observational ‘dies: The role of suongly ignrable weatment asignment. Journal of the Amerlan Statistical Atsciarion, 79, 41-48. Rosenbaum, P. R. (1987). The role of a second contol group in an ‘bscrutional stady (with cincusson). Staten Sconce, 2. 292-316, Rosenbaum, P. R (1995). Observational studies. New York: Springer ‘Verlag. Rosenbaum, PR. & Rubin, D. (1984). Reducing bss in observational ‘dies sing subcassifiatioas on the propensity Score, Journal of the American Statistical Association, 79, 16-528 Rosvol, HE, Mik, A. F, Sarason, I, Brasome, ED, Ja & Beck, 1. H.(9S6) A continuous performance ts of bain damage. Journal of Consating Psychology, 20. 46-350. Rubi, D.B. (in press). Propensity score methods. In L. Bickman (Ed), ‘Contributions to research design: Donald Campbell's legacy (Vol. 2) “Thousand Oaks, CA: Sage Side, DA. T., & Turpin, G. (1980), Measuremest, quantification, and analysis of cardiac activity. In 1. Matin & P. HL Venables (Eds), Techniques in psychophysiology (gp. 139-246). London: Wily. ‘Thernsom, A. (1991, December 16). Beyond the ple. New Republic Wainer, H. (1991), Adjusting for diferent base rates: Lord's Paradox ‘again Pexcholopical Bullen 109, 147-151 Walle, E, & Lewine. R, (1993). Sampling biases in stodies of pender and schizophrenia. Schcophvenia Bulletin 19, 1-7, ‘Watson, D., Weber, K, Asteneimar, JS. Clark, LA. Strauss M. E & MeCormick, R.A, (1095). Testing a tripnite model: Evaluating the ‘convergent and discriminant validity of amxety and depression symptom teales, Journal of Abnormal Psychology, 104, 3-18 Wild, A, Ry & Abola ©. T. (1978), Analysis af covariance. Bevery ils, CA: Sage Received November 3, 1998 Revision received September 7, 2000 ‘Accepted September 7, 2000

You might also like