Document 50gnZQO31mK1RBd00BokoKBzN

3Z/i. 2). / UNITED STATES ENVIRONMENTAL PROTECTION AGENCY Environmental Criteria and Assessment Office (MD-52) Research Triangle Park, North Carolina 27711 July 31, 1985 Dr. James Vare Dept. o Biostatistics Harvard School of Public Health 677 Huntington Ave. Boston, MA 02115 Dear Dr. Ware: I apologize for the long delay in getting back to you regarding the comments from Du Pont on the Corrigenda to the Lead Criteria Document and Joel Schwartz' response to them, but I have been away from the office for a week and a half on my honeymoon cruise to Mexico. Upon my return, I found the attached paper from Dick Landis that I had mentioned in my letter of July 1. If you would be so kind as to review this material, which was prepared specifically in response to the comments raised by Du Pont at the CASAC meeting on the Lead document, and integrate it into the opinion which you are preparing for us, we would be most grateful. I will call you on Wednesday, August 7, to inquire when we might be able to complete this task. If you should have any questions in the meantime, please feel free to call me (919/541-4163), Joel Schwartz (202/382-2782), or Dick Landis (313/764-5450). Thank you for your help in this matter. '.riavxd E. Weil, Ph.D. Project Manager TEH 0412349 C/IO- <2.0-g/.&) Z&.2). g> UNITED STATES ENVIRONMENTAL PROTECTION AGENCY Environmental Criteria and Assessment Office (MD-52) Research Triangle Park, North Carolina 27711 July 31, 1985 Dr. Richard Royall Dept, of Biostatistics Johns Hopkins University 615 N. Wolfe St. Baltimore, HD 21205 Dear Dr. Royall: I apologize for the long delay in getting back to you regarding the comments from Du Pont on the Corrigenda to the Lead Criteria Document and Joel Schwartz' response to them, but I have been away from the office for a week and a half on my honeymoon cruise to Mexico. Upon my return, I found the attached paper from Dick Landis that I had mentioned in my letter of July 1. If you would be so kind as to review this material, which was prepared specifically in response to the comments raised by Du Pont at the CASAC meeting on the Lead document, and integrate it into the opinion which you are preparing for us, we would be most grateful. 1 am also enclosing a copy of a memo that Joel Schwartz sent to Jim Ware upon request; hcpefully, it will be of use to you. I will call you on Wednesday, August 7, to inquire when we might be able to complete this task. If you should have any questions in the meantime, please feel free to call me (919/541-4163), Joel Schwartz (202/382-2782), or Dick Landis (313/764-5450). Thank you for your help in this matter. Sincerely, Project Manager A/o t e : Se e t mt &.I f o r . a t t a c c e n t s TEH 0412350 DUP050453402 N 30986.01 ZA-2>. 7 APPLICATIONS OF THE GENERALIZED MANTEL-HAENSZEL PROCEDURE TO THE ANALYSIS OF THE BLOOD PRESSURE/ BLOOD LEAD RELATIONSHIP IN THE NHANES n DATA J. Richard Landis Department of Biostatistics University of Michigan Ann Arbor, MI 48109 U.S.A. Presented to the EPA Relative to Analyses Described at the CASAC Meeting, May 13-14, 1985 Abstract Many research investigations involve the analysis of categorical data in which primary attention is directed at the relationship between two of the variables, while controlling for the effects of a set of covariables. Frequently, the two primary variables correspond to treatment and response and one or both are measured and/or reported on ordinal scales. The Generalized Mantel-Haenszel procedure, based on a randomization model, can be used to address this relationship under appropriate scores for the row and column variable. In contrast to this usual context, the row and column variables may be continuous, so that the scores are the actual measurements themselves. This use of the Generalized Mantel-Haenszel procedure is much less common, but does provide an alternative strategy for multiple regression models when the main focus is on the primary relationship averaged across the strata, rather than on a predictive model which includes effects for each of the strata. This methodology is applied to the analysis of the blood pressure/blood lead relationship in the NHANES II data as an alternative to multiple regression modelling. Adjustments are made for age and body mass index, the two main covariables in Hie variation of blood pressure, and the 64 sampling sites. 1.1 Introduction For many research situations the primary question of interest frequently involves the relationship between a set of s categories which correspond to distinct sub-populations (as defined in terms of pertinent independent variables) and a set of r categories which correspond to the response profiles associated with the specific dependent variables under study. Quite frequently, however, the distribution of the response profiles may also be influenced by the effects of a secondary set of t categories which correspond to distinct TEH 0412351 DUP050453403 N 30986.02 2 Generalized MH Analysis of Lead/DBP levels of relevant covariables such as investigators, hospitals, clinics, or pretreatment status. The resulting data obtained from such studies can be summarized in a set of t:(s x r) contingency tables 'which will be indexed by h = 1, 2, t. In this formulation, the basic hypothesis then can be expressed in terms of no partial association between the sub populations and the response profiles, after adjusting for the possible effects of the covariables. 1.2 Notation and Basie Model Let h = 1, 2, t index a set of (s x r) contingency tables which correspond to distinct levels of a covariable or combinations of several pertinent covariables. Let j = 1, 2, ..., s index a set of sub-populations such as treatment regimens which are to be compared with respect to a particular response variable for which the outcome categories are indexed by j = 1, 2, ..., r. Then let n^ = (n^j, nhlr.......nhsl' nhsP' wiiere n^y denotes the number of subjects in the sample who are jointly classified as belonging to the htk table, the i th sub-population and the jf/i response category. These frequency data, corresponding to the hth table, can be summarized as shown in Table 20.2, where N^. denotes the marginal total number of subjects classified as belonging to the i/ft sub population, denotes the marginal total number of subjects classified as belonging to the j th response category, and denotes the overall marginal total sample size in the hfft table. The basic hypothesis under investigation involves the relationship between the response variable and the sub-populations adjusted for the levels of the covariable set. Under the assumption that the marginal totals {N^} and {N^} are fixed(either by design or conditional distribution arguments), the overall null hypothesis of no partial association can be stated as H^: For each of the separate levels of the covariable set h = 1, 2, t, the response variable is distributed at random with respect to the sub-populations, July 15, 1985 Landis, J. R. Memo TEH 0 4 1 2 3 5 2 DUP050453404 Generalized MH Analysis of Lead/DBP Table 1 Observed Contingency Table for Level h of the Covariables pop'n 1 2 Response Variable Categories 12 nhll nhl2 nhl2 nh22 r nhlr nh2r 3 Total Nhl. Nh2- s Total nhsl Nh*l nhs2 Nh- 2 . nhsr N.h*r Nhs. Nh i.e., the data in the respective rows of the hth table can be regarded as a successive set of simple random samples of sizes from a fixed population corresponding to the marginal total distribution of the response variable {N^. j). (1) Under Hq in (1), it follows that the vector of observed frequencies, n^, has the multiple hypergeometric distribution n N. . ! II N, .! i=l hl*j=l Pr(nh|Ho) = sr N, 1 n n n, hi=lj=l *"> (2) as developed in Chapter 16 for a single s x r table. For each table indexed by h = 1, 2, t, let P^.. = CN^/N^) denote the marginal proportion of subjects classified as belonging to the ith sub-population(row) and let P^i< j = (N^VN^) denote the marginal proportion of subjects classified as belonging to the jth response category (column). These proportions, which are assumed to be fixed quantities, can be summarized in vector notation for h = 1, 2, t as = (P^ , Landis, J. R. Memo TEH 0412353 July 15, 1985 DUP050453405 4 Generalized MH Analysis of Lead/DBP Pj^ ) and p. Then by computing the first and second moments of the probability distribution in (2) as outlined in Chapter 16, it follows that the expected value of nh,, under HQ in (1) is mh,, = NhPhi*Phj` These expected values of the elements of can be expressed in vector notation as mh = E<nhlH0> (3) = NhtPh.^ where @ denotes left hand Kronecker product multiplication. Furthermore, the covariance between nh..i.j and n,hij under H0. in (1) is Nh<Nh - 1) where 5.a.. = 1, if i = i\ = 0, otherwise and 5J.J., = 1, if j = j\ = 0 otherwise. Thus, the covariance matrix of under in (1) can be expressed in matrix notation as Nf Var(nh|H0) = t DPh4 * "Ph* *Ph* " 1 [ DPh* ~Ph** Ph** ^ ' (Nh-D (4) where D_ and D,, are diagonal matrices with elements of the vectors P, _ and n** h** h~* * on the main diagonal. 1.3 Correlation Test Essentially, there are three different situations in which this randomization model structure can be used to test Hq . In order of increasing specificity, they can be summarized as i) nominal row variable, nominal column variable ii) nominal row variable, ordinal column variable Hi) ordinal row variable, ordinal column variable. Each of these cases are summarized in considerable detail in Landis et al (1978, 1979); however, only case iii) will be discussed further here. For situations where both the response categories j = 1, 2, r and the sub- TEH 0412354 July 15, 1985 Landis, J. R. Memo DUP050453406 Generalized MH Analysis of Lead/DBP 5 population categories i = 1, 2, ..., s are ordinally scaled with progressively larger intensities, it is often useful to test HQ in terms of composite mean scores which correspond to the respective products of an appropriate vector of response scores a^g, a^) and an appropriate vector of sub-population scores = (c^, ch2' chs^ ^0r ^ = ! <3* Specifically, let w^ = c^-a^. be the score assigned to the joint outcome corresponding to the ith sub-population and }th response category in the hth table which can be summarized in vector notation by letting = (w^j, ..., w^r, ..., Wjisr). Then let the composite mean score function for the hfA table be defined by Fh = Whnh/Nh- (5) Moreover, let the mean score on the row marginal proportions be denoted by s ch" J Ariftri* i=l (6) and let the mean score on the column marginal proportions be as defined in (20.18). Also, let the total variance of the sub-population variable in the overall population for the hth table with respect to be denoted by Sc,h = isx ^*hi ~ ch)2pbi* (7) and let the corresponding total variance of the response variable be given as in (20.20). In addition, let sr *ac,h ~ ah) (hi ch^ ^"hij^h (8) be the covariance between the response scores and the sub-population scores in the ht/t table. Then from (3) and (4) it follows that EW - "Vh (9) Landis, J. R. Memo TEH 0412355 July 15, 1985 DUP050453407 6 Generalized MH Analysis of Lead/DBP q2 q 2 a,h c,h Var(F,|HJ =. h (Nh-1) (10) Furthermore, from (5) and (8) it can be shown that rh-W = !W (ID Given this framework in (5) - (11), HQ in (1) can be tested in terms of a test statistic which is directed at alternatives pertaining to the association between the response variable the sub-population variable under the composite scores {w^}. In particular, let QMA,h [Fh-E(Ph|H0) ] MW s2a,,h sc,\h (12) = <Nh-R:c,h> 2 where denotes the squared Pearson correlation coefficient corresponding to the observed bivariate distribution for the {a^} and the {c^.} in the hth table. Under Hq in (1), 2 h asymptotically follows the x distribution with d.f. = 1. As defined in (12), Qma h essentially is a standard correlation analysis test statistic for HQ. As a result, an overall test statistic for the simultaneous test of the measures of association {F^} in (5) across all t tables is QMA = bl^MA.h- (13) If each of the respective for h = 1, 2, t is sufficiently large that all of the {QMA h} have approximate chi-square distributions, then has an approximate chi-square distribution with d.f. = t under Hn in (1). TEH 0412356 July 15, 1985 Landis, J. R. Memo DUPO 50453408 Generalized MH Analysis of Lead/DBP 7 Alternatively, if the overall sample size N is large but the individual sample sizes {N^} for many tables are small, then the statistic Qj^ in (13) is no longer useful for testing Hq in (1). Instead it is more appropriate to investigate Hq in terms of a generalized Cochran/Mantel-Haenszel/Birch strategy with respect to the measures of association {F^} in (5) based on the composite mean scores {w^}. For this purpose, let the weighted sum of the composite mean scores across the t tables be denoted be F a* l N,FU. h=l h h From (9) and (10) it follows that (14) E(f1H0) - l h= 1 . NXhSc2,h Var( FjH0) = J h-1 Nh-1 (15) (16) Consequently, the generalized Cochran/Mantel-Haenszel/Birch statistic for Hq in terms of composite mean scores with respect to {w^} is obtained as QRMA t fl a,h c,h I --------------h=l Nh-1 (17) Under Hq , Qjyy^ asymptotically has the chi-square distribution with d.f. -- 1. This statistic is directed at the extent to which there is a consistent positive (or negative) association between the response scores and the sub-population scores in the respective tables. Thus, it is directed at average partial association alternatives across the t tables between the response scores {a^} and the sub-population scores {c^}. In particular, Landis, J. R. Memo TEH 0412357 July 15, 1985 DUP050453409 8 Generalized MH Analysis of Lead/DBP for the case where ail the scores are either 0 or 1, (1?) simplifies to the Cochran/ Mantel-Haenszel/Birch statistic in (20.16) which corresponds to the collapsed set of (2 x 2) tables which are obtained by pooling the categories which are assigned "0" and those which are assigned "1" separately for the response categories and the sub-populations. Finally, if marginal rank or ridit-type scores are obtained From both the rows and columns of each table with midranks assigned for ties, the statistic in (17) is essentially equivalent to a partial Spearman Rank Correlation test; conditioning on the levels ofthe covariable set. Otherwise, the same types of comments could be made about power properties and central limit theory considerations for these statistics as were made for the multivariate test in Section 20.1 and the mean score test in Section 20.2, except that the concept of power is stronger relative to the association structure at which they are targeted and sample size requirements are minimal because only one degree of freedom is involved. 1.4 Continuous Variable Situation The test statistic in (17) simplifies considerably in situations for which the row and column variable are analyzed as continuous variables. In particular, let y^. denote the observed response variable value for the ifA individual in the hfA 1 stratum; x, . denote the observed predictorfrow) variable value for the ifA individual in the " 1 hfA stratum. Then the contribution to the numerator in the statistic (17) from the hfA stratum simplifies to (J*hi)(Eyhi) Fh - E(W " Iv b ---------- 3--' and the variance contribution to the denominator from the hfA stratum simplifies to July 15, 1985 TEH 0412358 Landis, J. R. Memo DUP050453410 Generalized MH Analysis of Lead/DBP 9 Var(Fh|H0) = As a result, the overall statistic for Hg, using the actual values for x and y, is based on weighted sums of the usual components of the regression coefficient across the strata. 1.5 Application to the NHANES II Data This generalized procedure can be used to test the significance of the linear relationship between blood lead(Pb) on the In scale and diastolic blood pressure(DBP), adjusting for the key predictive variables, age and body mass index, and the 64 sampling sites. An attractive feature of this methodology for the NHANES II design is that it provides adjustments for the blood pressure/blood lead test statistic for the potential effects of the sampling sites, without requiring a predictive model with 63 site parameters. Furthermore, this statistic averaged over the strata formed, by the cross-classification of age and body mass subgroups and sampling sites potentially is more robust than a multiple regression model with indicator variables for the sites. In order to adjust for age and body mass indexfquetlets), we need to create several ordinal variables that, in turn, are cross-classified to form strata. Although a large number of different possible partitions of these two risk factors could be considered, we will limit our attention to two alternatives for age and quetlets index. For notational convenience, let a denote age and q denote quetlets index as continuous variables. Then define the stratification subclasses to be Landis, J. R. Memo ' 1, A,, = 2 , * 3, a s 20 20 < a sS 40 a > 40 . TEH 0412359 July 15, 1985 DUP050453411 10 Generalized MH Analysis of Lead/DBP l, q 22 Qi 2, 3, 22 < q 26.5 q > 26.5 . ,1, q . 23 2 23 < q S 27 3, q> 27. These particular cutpoints for age and quetlets index were selected on the basis of meaningful substantive issues, as well as sample size considerations. To illustrate the effects of these two variables on diastolic blood pressure(DBP) and blood lead levels(lh scale), we can summarize the mean DBP by the subgroups. For example, the sample sizes, and the corresponding unweighted mean DBP and Blood Lead Levels for A2 by Q2 are displayed in Table 2. Table 2 Sample Size, Unweighted Mean Diastolic Blood FressureCDBP) and Unweighted Blood Lead Level(&i Scale) by Age and Quetlets Subgroups: Males, 12-74 years, NHANES 13, United States, Age (a 2) 1 1 1 2 2 2 3 3 3 Quetlets (Q2) 1 2 3 1 2 3 1 2 3 Sample Size 484 135 43 336 368 265 327 672 551 Mean DBP 70.9 75.6 84.2 75.4 79.6 86.6 79.9 83.4 87.6 Mean Blood Lead 2.55 2.54 2.67 2.78 2.77 2.75 2.77 2.76 2.74 The generalized Mantel-Haenszel statistic, q r ma > was computed for the relationship between blood lead level(ti scale) and diastolic blood pressure, adjusting for the possible effects due to sampling site, quetlets index and age, either singly or in July 15, 1985 Landis, J. R. Memo TEH 0412360 DUP050453412 Generalized MH Analysis of Lead/DBP 11 combination. The resulting statistics are summarized in Table 3. Table 3 Correlation Test Statistic for Diastolic Blood Pressure on Blood Lead Level(fii scale) Based on Randomization Model Adjusted for Stratification Variables: Males, 12-74 years, NHANES II, United States, 1976-1980 Strati fication None Sites Number of Strata with N^>2 1 64 Sample Size 3181 3181 Correlation Test Statistic 68.83 44.04 d.f. 1 1 Signif. Level <0.01 <0.01 Sites, Sites, Qg 192 3181 23.13 1 <0.01 192 3181 25.94 1 <0.01 Sites, Q^, A^ Sites, Q2, Aj Sites, Q^, Ag Sites, Q2, A2 355 3181 13.47 1 <0.01 368 3181 15.01 1 <0.01 463 3181 4.08 1 <0.05 478 3181' 4.62 1 <0.05 The primary objective in the analysis of these data was to investigate relationship between Pb level and DBP level, adjusting for other important sources of variation. As noted elsewhere, because of the strong multicollinearity between age, body mass and Pb levels on DBP, these adjustments are already adjusting for some of the lead effect on blood pressure. Moreover, because of the striking decline, both in Pb levels and blood pressure over time, adjustment for the sampling sites also reduces the potential for the Pb level variable to provide a significant incremental effect on this relationship. It is noteworthy that even in the multiple regression models with a large number of potential predictor variables, that quetlets index and age were the dominant sources of variation in DBP. Nevertheless, even with the severe adjustments by sampling sites and age and quetlets, the blood pressure/blood lead relationship is still statistically significant at the 5% Landis, J. R. Memo July IS, 1985 TEH Q412361 DUP050453413 12 Generalized MH Analysis of Lead/DBP level. In fact, the chi-square statistic of approximately 4 in the analysis with the most stringent adjustments corresponds to the square of the t statistic of 2 in the most conservative multiple regression models. References Landis, J.R., Heyman, E.R., and Koch, G.G. (1978), Average Partial Association in Three-Way Contingency Tables: A Review and Discussion, International Statistical Review 46, 237-254. Landis, J.R., Cooper, M.M., Kennedy, T. and Koch, G.G. (1979), A Computer Program for Testing Average Partial Association in Three-Way Contingency Tables(PARCAT), Computer Programs in Biomedicine 9, 223-246. July 15, 1985 TEH 0412362 Landis, J. R. Memo DUP050453414