Document 50gnZQO31mK1RBd00BokoKBzN
3Z/i. 2). /
UNITED STATES ENVIRONMENTAL PROTECTION AGENCY Environmental Criteria and Assessment Office (MD-52) Research Triangle Park, North Carolina 27711
July 31, 1985
Dr. James Vare Dept. o Biostatistics Harvard School of Public Health 677 Huntington Ave. Boston, MA 02115
Dear Dr. Ware: I apologize for the long delay in getting back to you regarding the
comments from Du Pont on the Corrigenda to the Lead Criteria Document and Joel Schwartz' response to them, but I have been away from the office for a week and a half on my honeymoon cruise to Mexico. Upon my return, I found the attached paper from Dick Landis that I had mentioned in my letter of July 1. If you would be so kind as to review this material, which was prepared specifically in response to the comments raised by Du Pont at the CASAC meeting on the Lead document, and integrate it into the opinion which you are preparing for us, we would be most grateful.
I will call you on Wednesday, August 7, to inquire when we might be able to complete this task. If you should have any questions in the meantime, please feel free to call me (919/541-4163), Joel Schwartz (202/382-2782), or Dick Landis (313/764-5450). Thank you for your help in this matter.
'.riavxd E. Weil, Ph.D. Project Manager
TEH 0412349
C/IO- <2.0-g/.&) Z&.2). g>
UNITED STATES ENVIRONMENTAL PROTECTION AGENCY Environmental Criteria and Assessment Office (MD-52) Research Triangle Park, North Carolina 27711
July 31, 1985
Dr. Richard Royall Dept, of Biostatistics Johns Hopkins University 615 N. Wolfe St. Baltimore, HD 21205
Dear Dr. Royall:
I apologize for the long delay in getting back to you regarding the comments from Du Pont on the Corrigenda to the Lead Criteria Document and Joel Schwartz' response to them, but I have been away from the office for a week and a half on my honeymoon cruise to Mexico. Upon my return, I found the attached paper from Dick Landis that I had mentioned in my letter of July 1. If you would be so kind as to review this material, which was prepared specifically in response to the comments raised by Du Pont at the CASAC meeting on the Lead document, and integrate it into the opinion which you are preparing for us, we would be most grateful.
1 am also enclosing a copy of a memo that Joel Schwartz sent to Jim Ware upon request; hcpefully, it will be of use to you.
I will call you on Wednesday, August 7, to inquire when we might be able to complete this task. If you should have any questions in the meantime, please feel free to call me (919/541-4163), Joel Schwartz (202/382-2782), or Dick Landis (313/764-5450). Thank you for your help in this matter.
Sincerely,
Project Manager
A/o t e : Se e t mt &.I f o r . a t t a c c e n t s
TEH 0412350
DUP050453402
N 30986.01
ZA-2>. 7
APPLICATIONS OF THE GENERALIZED MANTEL-HAENSZEL PROCEDURE TO THE ANALYSIS OF THE BLOOD PRESSURE/
BLOOD LEAD RELATIONSHIP IN THE NHANES n DATA
J. Richard Landis Department of Biostatistics
University of Michigan Ann Arbor, MI 48109 U.S.A.
Presented to the EPA Relative to Analyses Described at the CASAC Meeting, May 13-14, 1985
Abstract
Many research investigations involve the analysis of categorical data in which primary attention is directed at the relationship between two of the variables, while controlling for the effects of a set of covariables. Frequently, the two primary variables correspond to treatment and response and one or both are measured and/or reported on ordinal scales. The Generalized Mantel-Haenszel procedure, based on a randomization model, can be used to address this relationship under appropriate scores for the row and column variable. In contrast to this usual context, the row and column variables may be continuous, so that the scores are the actual measurements themselves. This use of the Generalized Mantel-Haenszel procedure is much less common, but does provide an alternative strategy for multiple regression models when the main focus is on the primary relationship averaged across the strata, rather than on a predictive model which includes effects for each of the strata.
This methodology is applied to the analysis of the blood pressure/blood lead relationship in the NHANES II data as an alternative to multiple regression modelling. Adjustments are made for age and body mass index, the two main covariables in Hie variation of blood pressure, and the 64 sampling sites.
1.1 Introduction For many research situations the primary question of interest frequently involves
the relationship between a set of s categories which correspond to distinct sub-populations
(as defined in terms of pertinent independent variables) and a set of r categories which
correspond to the response profiles associated with the specific dependent variables under
study. Quite frequently, however, the distribution of the response profiles may also be
influenced by the effects of a secondary set of t categories which correspond to distinct
TEH 0412351
DUP050453403
N 30986.02
2 Generalized MH Analysis of Lead/DBP
levels of relevant covariables such as investigators, hospitals, clinics, or pretreatment status.
The resulting data obtained from such studies can be summarized in a set of t:(s x r) contingency tables 'which will be indexed by h = 1, 2, t. In this formulation, the basic hypothesis then can be expressed in terms of no partial association between the sub populations and the response profiles, after adjusting for the possible effects of the covariables.
1.2 Notation and Basie Model
Let h = 1, 2, t index a set of (s x r) contingency tables which correspond to
distinct levels of a covariable or combinations of several pertinent covariables. Let j = 1,
2, ..., s index a set of sub-populations such as treatment regimens which are to be
compared with respect to a particular response variable for which the outcome categories
are indexed by j = 1, 2, ..., r. Then let n^ = (n^j, nhlr.......nhsl' nhsP' wiiere n^y denotes the number of subjects in the sample who are jointly classified as belonging to
the htk table, the i th sub-population and the jf/i response category. These frequency data,
corresponding to the hth table, can be summarized as shown in Table 20.2, where N^.
denotes the marginal total number of subjects classified as belonging to the i/ft sub
population,
denotes the marginal total number of subjects classified as belonging to
the j th response category, and denotes the overall marginal total sample size in the hfft
table.
The basic hypothesis under investigation involves the relationship between the
response variable and the sub-populations adjusted for the levels of the covariable set.
Under the assumption that the marginal totals {N^} and {N^} are fixed(either by design
or conditional distribution arguments), the overall null hypothesis of no partial
association can be stated as
H^: For each of the separate levels of the covariable set h = 1, 2, t, the
response variable is distributed at random with respect to the sub-populations,
July 15, 1985
Landis, J. R. Memo
TEH 0 4 1 2 3 5 2
DUP050453404
Generalized MH Analysis of Lead/DBP
Table 1 Observed Contingency Table for Level h of the Covariables
pop'n 1 2
Response Variable Categories 12
nhll nhl2
nhl2 nh22
r nhlr nh2r
3
Total Nhl. Nh2-
s Total
nhsl Nh*l
nhs2 Nh- 2
.
nhsr N.h*r
Nhs. Nh
i.e., the data in the respective rows of the hth table can be regarded as a
successive set of simple random samples of sizes
from a fixed
population corresponding to the marginal total distribution of the response
variable {N^. j).
(1)
Under Hq in (1), it follows that the vector of observed frequencies, n^, has the multiple
hypergeometric distribution
n N. . ! II N, .!
i=l hl*j=l
Pr(nh|Ho) =
sr
N, 1 n n n,
hi=lj=l *">
(2)
as developed in Chapter 16 for a single s x r table.
For each table indexed by h = 1, 2, t, let P^.. = CN^/N^) denote the marginal
proportion of subjects classified as belonging to the ith sub-population(row) and let
P^i< j = (N^VN^) denote the marginal proportion of subjects classified as belonging to the
jth response category (column). These proportions, which are assumed to be fixed
quantities, can be summarized in vector notation for h = 1, 2, t as
= (P^ ,
Landis, J. R. Memo
TEH 0412353 July 15, 1985
DUP050453405
4 Generalized MH Analysis of Lead/DBP
Pj^ ) and
p. Then by computing the first and second moments of
the probability distribution in (2) as outlined in Chapter 16, it follows that the expected
value of nh,, under HQ in (1) is mh,, = NhPhi*Phj` These expected values of the elements of can be expressed in vector notation as
mh = E<nhlH0>
(3)
= NhtPh.^ where @ denotes left hand Kronecker product multiplication. Furthermore, the covariance between nh..i.j and n,hij under H0. in (1) is
Nh<Nh - 1)
where 5.a.. = 1, if i = i\ = 0, otherwise and 5J.J., = 1, if j = j\ = 0 otherwise. Thus, the covariance matrix of under in (1) can be expressed in matrix notation as
Nf
Var(nh|H0) =
t DPh4 * "Ph* *Ph* " 1 [ DPh* ~Ph** Ph** ^ '
(Nh-D
(4)
where D_ and D,, are diagonal matrices with elements of the vectors P, _ and
n** h**
h~*
* on the main diagonal.
1.3 Correlation Test
Essentially, there are three different situations in which this randomization model structure can be used to test Hq . In order of increasing specificity, they can be summarized as
i) nominal row variable, nominal column variable ii) nominal row variable, ordinal column variable Hi) ordinal row variable, ordinal column variable.
Each of these cases are summarized in considerable detail in Landis et al (1978, 1979); however, only case iii) will be discussed further here.
For situations where both the response categories j = 1, 2, r and the sub-
TEH 0412354
July 15, 1985
Landis, J. R. Memo
DUP050453406
Generalized MH Analysis of Lead/DBP
5
population categories i = 1, 2, ..., s are ordinally scaled with progressively larger
intensities, it is often useful to test HQ in terms of composite mean scores which
correspond to the respective products of an appropriate vector of response scores
a^g, a^) and an appropriate vector of sub-population scores = (c^,
ch2' chs^ ^0r ^ =
! <3* Specifically, let w^ = c^-a^. be the score assigned to the
joint outcome corresponding to the ith sub-population and }th response category in the hth
table which can be summarized in vector notation by letting = (w^j, ..., w^r, ...,
Wjisr). Then let the composite mean score function for the hfA table be defined by
Fh = Whnh/Nh-
(5)
Moreover, let the mean score on the row marginal proportions be denoted by
s
ch" J Ariftri* i=l
(6)
and let the mean score on the column marginal proportions be as defined in (20.18). Also,
let the total variance of the sub-population variable in the overall population for the hth
table with respect to be denoted by
Sc,h = isx ^*hi ~ ch)2pbi*
(7)
and let the corresponding total variance of the response variable be given as in (20.20). In
addition, let
sr *ac,h ~
ah) (hi ch^ ^"hij^h
(8)
be the covariance between the response scores and the sub-population scores in the ht/t table. Then from (3) and (4) it follows that
EW - "Vh
(9)
Landis, J. R. Memo
TEH 0412355 July 15, 1985
DUP050453407
6 Generalized MH Analysis of Lead/DBP
q2 q 2 a,h c,h
Var(F,|HJ =. h (Nh-1)
(10)
Furthermore, from (5) and (8) it can be shown that
rh-W = !W
(ID
Given this framework in (5) - (11), HQ in (1) can be tested in terms of a test
statistic which is directed at alternatives pertaining to the association between the
response variable the sub-population variable under the composite scores {w^}. In
particular, let
QMA,h
[Fh-E(Ph|H0) ] MW
s2a,,h sc,\h
(12)
= <Nh-R:c,h>
2
where
denotes the squared Pearson correlation coefficient corresponding to the
observed bivariate distribution for the {a^} and the {c^.} in the hth table. Under Hq in (1), 2
h asymptotically follows the x distribution with d.f. = 1. As defined in (12),
Qma h essentially is a standard correlation analysis test statistic for HQ. As a result, an
overall test statistic for the simultaneous test of the measures of association {F^} in (5)
across all t tables is
QMA = bl^MA.h-
(13)
If each of the respective for h = 1, 2, t is sufficiently large that all of the {QMA h}
have approximate chi-square distributions, then
has an approximate chi-square
distribution with d.f. = t under Hn in (1).
TEH 0412356
July 15, 1985
Landis, J. R. Memo
DUPO 50453408
Generalized MH Analysis of Lead/DBP
7
Alternatively, if the overall sample size N is large but the individual sample sizes {N^} for many tables are small, then the statistic Qj^ in (13) is no longer useful for testing Hq in (1). Instead it is more appropriate to investigate Hq in terms of a generalized Cochran/Mantel-Haenszel/Birch strategy with respect to the measures of association {F^} in (5) based on the composite mean scores {w^}. For this purpose, let the weighted sum of the composite mean scores across the t tables be denoted be
F a* l N,FU. h=l h h
From (9) and (10) it follows that
(14)
E(f1H0) - l
h= 1
. NXhSc2,h
Var( FjH0) = J
h-1 Nh-1
(15) (16)
Consequently, the generalized Cochran/Mantel-Haenszel/Birch statistic for Hq in terms of composite mean scores with respect to {w^} is obtained as
QRMA
t fl a,h c,h I --------------h=l Nh-1
(17)
Under Hq , Qjyy^ asymptotically has the chi-square distribution with d.f. -- 1. This statistic is directed at the extent to which there is a consistent positive (or negative) association between the response scores and the sub-population scores in the respective tables. Thus, it is directed at average partial association alternatives across the t tables between the response scores {a^} and the sub-population scores {c^}. In particular,
Landis, J. R. Memo
TEH 0412357 July 15, 1985
DUP050453409
8 Generalized MH Analysis of Lead/DBP
for the case where ail the scores are either 0 or 1,
(1?) simplifies to the Cochran/
Mantel-Haenszel/Birch statistic in (20.16) which corresponds to the collapsed set of (2 x 2)
tables which are obtained by pooling the categories which are assigned "0" and those
which are assigned "1" separately for the response categories and the sub-populations.
Finally, if marginal rank or ridit-type scores are obtained From both the rows and
columns of each table with midranks assigned for ties, the statistic
in (17) is
essentially equivalent to a partial Spearman Rank Correlation test; conditioning on the
levels ofthe covariable set. Otherwise, the same types of comments could be made about
power properties and central limit theory considerations for these statistics as were made
for the multivariate test in Section 20.1 and the mean score test in Section 20.2, except
that the concept of power is stronger relative to the association structure at which they are
targeted and sample size requirements are minimal because only one degree of freedom is
involved.
1.4 Continuous Variable Situation
The test statistic in (17) simplifies considerably in situations for which the row and column variable are analyzed as continuous variables. In particular, let
y^. denote the observed response variable value for the ifA individual in the hfA 1 stratum;
x, . denote the observed predictorfrow) variable value for the ifA individual in the " 1 hfA stratum.
Then the contribution to the numerator in the statistic (17) from the hfA stratum simplifies to
(J*hi)(Eyhi) Fh - E(W " Iv b ---------- 3--'
and the variance contribution to the denominator from the hfA stratum simplifies to
July 15, 1985
TEH 0412358 Landis, J. R. Memo
DUP050453410
Generalized MH Analysis of Lead/DBP
9
Var(Fh|H0) =
As a result, the overall statistic
for Hg, using the actual values for x and y,
is based on weighted sums of the usual components of the regression coefficient across the
strata.
1.5 Application to the NHANES II Data This generalized procedure can be used to test the significance of the linear
relationship between blood lead(Pb) on the In scale and diastolic blood pressure(DBP), adjusting for the key predictive variables, age and body mass index, and the 64 sampling sites. An attractive feature of this methodology for the NHANES II design is that it provides adjustments for the blood pressure/blood lead test statistic for the potential effects of the sampling sites, without requiring a predictive model with 63 site parameters. Furthermore, this statistic averaged over the strata formed, by the cross-classification of age and body mass subgroups and sampling sites potentially is more robust than a multiple regression model with indicator variables for the sites.
In order to adjust for age and body mass indexfquetlets), we need to create several ordinal variables that, in turn, are cross-classified to form strata. Although a large number of different possible partitions of these two risk factors could be considered, we will limit our attention to two alternatives for age and quetlets index. For notational convenience, let a denote age and q denote quetlets index as continuous variables. Then define the stratification subclasses to be
Landis, J. R. Memo
' 1,
A,, = 2 , * 3,
a s 20
20 < a sS 40 a > 40 .
TEH 0412359 July 15, 1985
DUP050453411
10 Generalized MH Analysis of Lead/DBP
l, q 22
Qi
2, 3,
22 < q 26.5 q > 26.5 .
,1, q . 23
2 23 < q S 27 3, q> 27.
These particular cutpoints for age and quetlets index were selected on the basis of
meaningful substantive issues, as well as sample size considerations. To illustrate the
effects of these two variables on diastolic blood pressure(DBP) and blood lead levels(lh
scale), we can summarize the mean DBP by the subgroups. For example, the sample
sizes, and the corresponding unweighted mean DBP and Blood Lead Levels for A2 by Q2
are displayed in Table 2.
Table 2 Sample Size, Unweighted Mean Diastolic Blood FressureCDBP) and Unweighted Blood Lead Level(&i Scale) by Age and Quetlets Subgroups:
Males, 12-74 years, NHANES 13, United States,
Age (a 2)
1 1 1
2 2 2
3 3 3
Quetlets (Q2)
1 2 3
1 2 3
1 2 3
Sample Size
484 135 43
336 368 265
327 672 551
Mean DBP
70.9 75.6 84.2
75.4 79.6 86.6
79.9 83.4 87.6
Mean Blood Lead
2.55 2.54 2.67
2.78 2.77 2.75
2.77 2.76 2.74
The generalized Mantel-Haenszel statistic, q r ma > was computed for the relationship between blood lead level(ti scale) and diastolic blood pressure, adjusting for the possible effects due to sampling site, quetlets index and age, either singly or in
July 15, 1985
Landis, J. R. Memo TEH 0412360
DUP050453412
Generalized MH Analysis of Lead/DBP
11
combination. The resulting statistics are summarized in Table 3.
Table 3
Correlation Test Statistic for Diastolic Blood Pressure on Blood Lead Level(fii scale) Based on Randomization Model Adjusted for Stratification Variables: Males, 12-74 years, NHANES II, United States, 1976-1980
Strati fication
None
Sites
Number of Strata with N^>2
1
64
Sample Size
3181
3181
Correlation Test
Statistic
68.83
44.04
d.f. 1 1
Signif. Level
<0.01
<0.01
Sites, Sites, Qg
192
3181
23.13
1 <0.01
192
3181
25.94
1 <0.01
Sites, Q^, A^ Sites, Q2, Aj Sites, Q^, Ag Sites, Q2, A2
355
3181
13.47
1 <0.01
368
3181
15.01
1 <0.01
463 3181
4.08
1 <0.05
478
3181'
4.62
1 <0.05
The primary objective in the analysis of these data was to investigate relationship between Pb level and DBP level, adjusting for other important sources of variation. As noted elsewhere, because of the strong multicollinearity between age, body mass and Pb levels on DBP, these adjustments are already adjusting for some of the lead effect on blood pressure. Moreover, because of the striking decline, both in Pb levels and blood pressure over time, adjustment for the sampling sites also reduces the potential for the Pb level variable to provide a significant incremental effect on this relationship. It is noteworthy that even in the multiple regression models with a large number of potential predictor variables, that quetlets index and age were the dominant sources of variation in DBP. Nevertheless, even with the severe adjustments by sampling sites and age and quetlets, the blood pressure/blood lead relationship is still statistically significant at the 5%
Landis, J. R. Memo
July IS, 1985 TEH Q412361
DUP050453413
12 Generalized MH Analysis of Lead/DBP
level. In fact, the chi-square statistic of approximately 4 in the analysis with the most stringent adjustments corresponds to the square of the t statistic of 2 in the most conservative multiple regression models.
References
Landis, J.R., Heyman, E.R., and Koch, G.G. (1978), Average Partial Association in Three-Way Contingency Tables: A Review and Discussion, International Statistical Review 46, 237-254.
Landis, J.R., Cooper, M.M., Kennedy, T. and Koch, G.G. (1979), A Computer Program for Testing Average Partial Association in Three-Way Contingency Tables(PARCAT), Computer Programs in Biomedicine 9, 223-246.
July 15, 1985
TEH 0412362 Landis, J. R. Memo
DUP050453414