Document EdNnax9GoaavO9DeY7RzM6wbN

The MITRE Corporation Meirek Division 1820 Dailey Madison Boulevard McLean, Virginia 22102 WP- 81W00541 __ ______ ^ No. Voi. Series Rev. Supp. Coir. Subject: Statistical Analysis of Blood Lead Levels Among Children in Louisville, Kentucky To: L. Duncan Prom: J. Watson Contract No.: 68-01-5051 Sponsor: SPA Project No.: 167 2Y Dept: W53 Date: 30 June 1981 TEH 0531324 TABLE OF CONTENTS LIST OF FIGURES LIST OF TABLES 1.0 INTRODUCTION ' 1.1 Putpose and Scope 1.2 Major Conclusions 2.0 NATURE OF THE COMPUTER RUNS 2.1 General Characteristics of the Data Used 2.2 Data Structure for the Runs 2.3 Runs Ignoring Census Tract 2.4 Runs Using Census Tract Information 3.0 STATISTICAL RESULTS OF THE ANALYSIS 3.1 Analysis With No Use of Census Data 3.1.1 Correlation Results 3.1.2 Evidence of the Effect'of Other Factors 3.1.3 Linear Regression of In(GMBL) 3.2 Analysis of Data Using Census Tract Information 3.2.1 Objective of the Runs 3.2.2 Correlation Results 3.2.3 Regression Results REFERENCES APPENDIX gage iv ' iv 1-1 1-1 1-2 2-1 2-1 2-4 2-7 2-9 3-1 . 3-1 3-1 3-5 3-8 3-18 3-19 3-20 3-23 4-1 A-l iii TEH 0531325 DUP050032536 LIST OP FIGURES Figure Number i Map of Louisville Census Tracts 2 Major Census Block Map . Page 2-3 2-14 LIST OF TABLES Table Number 2-1 Illustration of Aggregated Quarterly Record Format (No Census Tract Info) 2-2 Variables Used in Regression - Corre lation Run/Sets Louisville, Kentucky, Data (Run Sets Arbitrarily Numbered) 2-3 Allocation of Census Tracts to Clusters and Major Blocks 3-1 Correlation Coefficients (Zero-Order) Between Geometric Mean Blood Level and Independent Variables and Between Gaso line Lead and Other Independent Variables .(Census Tract not Considered) 3-2 Partial Correlation of Selected Variables Compared with Zero-Order Coefficients Runs, Ignoring Census Tract, Louisville, KY,, Comprehensive Data 3-3 Regression Coefficients for Log Geometric Mean Blood-Lead Level, Ignoring Census Tract and Using Age Range Louisville, Kentucky, Comprehensive Data 3-4 Correlation Coefficients of LN(GMBL) and Gasoline Lead With Other Variables, Test Results Before and After 1977, Louisville, Kentucky Comprehensive Data 3-5 Correlation Coefficients of Blood Lead and Gasoline Lead with Other Variables, Using Census Tract Information, Louisville Ken tucky Comprehensive Data (Runs CBLK 1, CBLK 2, CBLK 3) 2-5 2-10 2-12 3-2 3-6 . 3-10 3-17 3-21 iv TEH 0531326 DUP050032537 > \ if <> ,.JT- LIST OF TABLES (Concluded) Table Number Page 3-6 Regression Results for Runs Using Census 3-24 Tract Information (CBLK1-CBLK4)> Louisville, Kentucky Data, Dependent Variable; LN(GMBL) v TEH 0531327 DUP050032538 > \ * TEH 0531328 DUP050032539 1.0 INTRODUCTION 1.1 Purpose and Scope This paper reports the results of statistical analysis performed by The MITRE Corporation on blood lead levels indicated in initial tests of children in Louisville, Ky,, during the period 1974 to 1979* The work represents a follow-up of that reported earlier in MITRE Working Paper WP-80W00466 and is based on more comprehensive test data provided to MITRE by Dr. Irwin Billick of the U.S. Department of Housing and Urban Development (HUD), The study was supported jointly by HUD and the Environmental Protection Agency. As in the preliminary study, the data were formatted by MITRE and then MITRE made computer runs at the MITRE Washington Computer Center to analyze the data. For these FORTRAN was used to a limited extent; chiefly, however, the Statistical Package for the Social Sciences (SPSS) was used (Nle, et al., 1975), A second line of analysis (being reported separately) investigated blood-lead levels in children repetitively tested over a period of up to 5 years. The objective of the analysis was to investigate the relationship of variations in geometric mean blood lead levels in the children tested to variations in other factors considered possibly associated. Of particularinterest was variation in the quantities of gasoline lead consumed in the Louisville area, this being of especial concern to EPA's Mobile Sources Enforcement Division, Other 1-1 TEH 0531329 DUP050032540 \ tf* r t * independent variables investigated in the analysis were race (black and white); residence (as indicated by census tract); age; and both time period and quarter ot season of the year during which the blood tests were made. Specifically, the following information was developed: a. A complete set of correlation and regression coefficients relating blood lead levels quantitatively to the independent variables investigated; b. Assessment from statistical evidence of the possible role that the various factors might play in variations observed in blood lead levels; c. Comparison of the statistical results for the Louisville < ' data with results previously obtained from less complete Louisville data and from very extensive data from New York City (Billick et al., 1979; MITRE 1980, 1980a). Major conclusions of the statistical analysis are summarised in the following section of this report. Section 2 explains the nature of the data and of the computer runs. The principal statistical results from which the conclusions were derived are then presented and discussed (Section 3). Other statistical data are given in the appendix. 1.2 Major Conclusions a. Variations in gasoline lead account for up to 45 percent of the observed variations in geometric mean blood lead levels in the Louisville children tested. This is by far the strongest interrelationship which any of the variables examined show with blood Lead but the correlation is less that that observed In New York City test data, b. As with the New York City data, the statistical results for Louisville are more consistent with the hypothesis that variation in gasoline lead levels principally explains 1-2 TEH 0531330 DUP050032541 variations in geometric mean blood lead level observed during a particular period of the year, rather than with a theory that some other seasonal factor largely accounts for blood lead variation. c. The complete set of variables incorporated into the analysis statistically explains at most about 55 percent of the variations observed in geometric mean blood lead level. This finding implies that compared to the regression model derived for New York City (MITRE, 1980), relatively poor results would be expected from trying to fit a regression equation to available data so as to project possible future blood lead levels as a function of gasoline lead and other relevant variables. d. A striking feature of this more comprehensive set of Louisville data was the fact that both gasoline lead and blood lead averages declined markedly after 19)7 (a year for which no usable test results were available). Statistical findings from the post-1977 data exclusively were consistent with an hypothesis that after this time gasoline lead exerted a strong influence on blood-lead variations, rather than that both levels had declined independently. e. Were obtained from the user of census : tract: data in that this additional information failed to provide any significant basis for interpreting the role which factors related to place of residence might have in accounting for blood-lead. Incorporation of residence (as indicated by groupings of contiguous census tracts) into the regression models did not increase the fraction of observed variations in geometric.mean blood-lead levels thereby explained. Only a few of the census tract groupings showed a significant statistical interrelationship with geometric mean blood levels; most correlated very weakly. These findings indicate that in general whatever factors besides gasoline lead are associated with blood--lead variations cannot be determined from place of residence within the available data base. A small number of census tracts were identified as associated with unknown factors related to blood-lead levels that tend to be significantly either higher or lower than averages for the metropolitan area of the whole, other things (i.e., age, race, and season of the year) being equal. 1-3 TEH 0531331 DUP050032542 2.0 NATURE OF THE COMPUTER RUNS 2.1 General Characteristics of the Data Used The statistical analysis reported here was based on computer runs made by MITRE at the Washington Computer Center. Dr. Billick of HUD provided a tape containing records of blood tests taken in Louisville* Ky., during the period 1974 through 1979. MITRE developed a computer program (using FORTRAN) to convert the basic data to a format convenient for subsequent organization into the files required for the SPSS programs. The FORTRAN program enabled the test data to be initially screened to eliminate incomplete results and/or those found to be in error on the basis of validation criteria deterministically applied (e.g., reported birth date not earlier than the test date; reported test results outside the possible range of values, etc.). In all, 12,913 records of blood-lead levels from the years 1974-76 and 1978-79* were found initially valid for analysis. Each record represented the initial test performed on a different child. (Repeated test results were handled separately In a study reported elsewhere.) This number of unique test results comprised an increase of nearly 502! over the body of similar test data used in the previous ^Results for 1977 were omitted because changes in selection of the sample for testing made the blood-lead data unsuitable for statistical, analysis with data from other years. 2-1 TEH 0531332 DUP050032543 analysis of Louisville data (MITRE, 1980a). About 70 percent of the children tested were black and nearly 30 percent white, while records of other and unknown races constituted less than 1 percent of the total:. Tests were heaviest in the years 1975, 1976 and 1978 (comprising over 70 percent of all records). Ninety-eight percent of the children tested were 6 years of age or younger, with age ranges 1 and A (i.e., children up to 12 months and children between 37 and 48 months respectively) making a combined total of 53 percent. Testing was heaviest in the first quarter of the years (36.9 percent for the months of January-March) and lightest in the third quarter (16,3 percent). Residence of each test subject as to census tract was indicated on each record (see map reproduced as Figure 1). Those census tracts within the inner city (i.e., towards the north and west up to the Ohio River) contained the heaviest concentration of test * results. Data within many of the tracts were too sparse to permit a comprehensive analysis by individual census tract. Therefore, two sets of runs were made, distinguished as to whether or not census tract was considered. In the first set of runs, this variable was ignored; all records were treated without regard to residence of the test subject. However, in. the second facet of analysis, location as indicated by tract number was taken into account (as is discussed in Section 2,4 below). 2-2 TEH 0531333 DUP050032544 moz TEH 0531334 DUP050032545 2,2 Data Structure for the Runs The unique test records were first aggregated into subgroups homogeneous as to year and quarter; race (black, or white, there being too few records of "other" race to permit analysis); and age range by 12 month period. As sex had been found in all previous analyses to be insignificant., it was ignored here. In the first set of runs, census tract was also ignored. Thus each subgroup record which was input to the first SPSS runs represented test results taken in a single quarter of a particular year for members of a specific race and within one 12-month age range. A single value for gasoline lead (as determined by average content per gallon and gallons of gasoline consumed for that quarter and measured in hundred millions of grams) was inserted into each subgroup record. The number of test results thus aggregated into a single subgroup record varied from 1 to over 200. For each subgroup the geometric mean of the blood lead level (GMBL) in micrograms of lead per 100 milliliters of blood (yg/100 ml) was computed and entered.^ Geometric mean was used because blood lead levels had previously been found (Billicit, et al., 1979) to approximate closely.the log-normal distribution. Table 2-1 illustrates the format of the aggregated Where a record consisted of only one test result, the "geometric mean blood lead level" obviously could not be obtained and the corresponding entry denoted blood lead level for the single individual. Only about 1,5 percent of the subgroup records consisted of as few as 2 test results. Records with fewer than 3 test results were not used in the analysis. 2-4 TEH 0531335 DUP050032546 TABLE 2-1 ILLUSTRATION OF AGGREGATED QUARTERLY RECORD FORMAT (NO CENSUS TRACT INFO) Race 1 1 * 1 l 2 2 Age Range 3 4 * 1 2 * 3 4 Test Year 76 76 78 78 .* * 75 75 Test Qtr. 2 2 * 1 1 * # 4 4 Gas Lead 0.737 0.737 '* * 0,434 0,434 0.701 0,701 GMBL 3.204 3.243 . 2.751 2,928 3.321 3,344 Nr in Aggregate Record 14 26 * 12 5 10 18 Census Tract -- -- " ILLUSTRATION OF AGGREGAtED QUARTERLY RECORD FORMAT,CENSUS TRACT INCLUDED l 1 76 2 0.737 3.224 2 1 2 76 2 0.737. 3.448 6 1 3 76 2 0.737 3.119 4 2 2 2 2-5 TEH 0531336 DUP050032547 records for the first set of runs. In addition to the basic variables shown there (which were available for use in all runs), additional values were provided for particular runs. Quarter of the year was handled as a "dummy variable." That is, a designated variable had its value set as either one. or zero depending on which 3-mon.th period of the year the record represented. Dummy variables represent a standard means by which each category of a nominal variable (such as quarter of the year) can be assigned arbitrary metric values of 0 and l and inserted into a regression equation where it represents the fraction of the sample within that category. Since, however, the number of degrees of freedom is always one less than the number of categories represented by dummy variables (i.e., the value of the kth category or fraction of the sample which it represents is completely determined by that of the first k-1), one of the dummies must, be excluded and "becomes a sort of reference point...referred to as the reference category" (Me, et al., 1975). In the runs, fourth quarter served as the reference category. Age range was also handled as a dummy variable in most of the runs. The reference category was provided by the range of 7 years and older. Other treatments of the age factor and additional dummy variables used in particular runs are explained in the following sections. 2-6 TEH 0531337 DUP050032548 Runs were made separately for the two races, white and black. The data were organized into sub-files according to race in order to facilitate implementation of the SPSS programs. The small number of "other" races precluded meaningful.analysis and the few records available were ignored. 2.3 Runs Ignoring Census Tract As noted, the first analysis grouped all records homogeneous as to time (i.e., year and quarter or "season" of the year), race, and age-range, irrespective of census tract. Aggregated records consist ing of fewer than 3 test results were excluded. The SPSS subprogram REGRESSION (Nie, et al., 1975) was used to calculate regression coefficients for ln(GMBL) as the dependent variable. This subprogram also provides as output mean values of all the variables and a complete matrix of correlation coefficients. Print-out of residuals from some of the regression runs was also provided for analysis, largely by eye. Higher-order coefficients between ln(GMBL) and selected variables were also calculated as is detailed in Section 3. Experimentation was made with various ways of representing age. It had been found in previous analysis that blood-lead levels tend to increase up to about age 3 and thereafter to decline, in one set of runs a variable termed "age deviation" was computed. This variable expressed (as a rational number) the square of the difference (in months) between the subgroup or aggregate mean age and age 40 months 2-7 TEH 0531338 DUP050032549 {taken to represent the approximate peak). Another set of runs used both mean age (in months) and its square in an attempt to represent the apparent curvilinear relationship between blood-lead level and age. Runs were also made using dummy variables to separate test results into twt> major time groupings. This approach took note of the observation that gasoline lead levels tended to decrease over time; the dummy variables were used as a means of representing the relative weight of test results before and after selected years. In one set of runs, the dummy variables separated results before and after the year 1977 (a year for which, as noted, no test results were usable). In another, the separation was between the years 1974 and 1975 on the one hand, and 1976 and following on the other hand. Another approach used with the time factor in one set of runs was to compute "quarterly difference" as a real value expressing the difference between 31 and the sequential number of the test quarter for the aggregated record; the number 31 was that assigned to a quarter during the year 1977, assumed in this set of runs as the breakpoint between relatively high and low gas lead levels. Tabulation of the correlation coefficients is shown in Section 3, which also discusses their interpretation. The complete set of correlation coefficients obtained can be seen from the listing of variables included in a run, as explained next. .. t 2-8 TEH 0531339 DUP050032550 Each separate grouping of independent "carrier" variables in a particular run represented a distinct "model" by means of which the dependent variable, In(GMBL), was expressed as a function of these variables. Different groupings were tested in an attempt to find the set of independent variables which best explained variation in In(GMBL). The distinct models tested in each arbitrarily-numbered set of runs is shown in Table 2-2, Here the column heading represents a specific set of tuns while the rows denote the factors included as independent variables. An "x" placed at the row-column intersection indicates that a particular factor (or set of related factors, such as age-group or season of the year) was used in a specific set of runs. Gasoline lead was included in all runs. For example, in Run Set #1, the independent variables were the age-groups and season of the year (represented by dummy variables), along with gasoline lead, whereas in Run Set #2 age-grouping was replaced by age deviation. Results of the regression runs are given and discussed in Section 3 below. 2.4 Runs Using Census Tract Information The second major area of analytic effort comprised an investigation of the role of different residential census tracts in statistically accounting for variation in ln(GMBL). Aggregation of individual records for these runs differed from that shown in Table 2-1 for the previous runs in that records were 2-9 TEH 0531340 DUP050032551 VARIABLES USED IN REGRESSION CORRELATION RUN/SETSLOUISVILLE, KENTUCKY, DATA (Run Sets A r b it r a r ily Numbered) X X assigned. In th e la t t e r , 1 Run Set (CBLK 4) te s t re co rd s were aggregated by each o f the 3 m ajor b lo cks w ith no fu r th e r d is t in c t io n among th e Census T ra c ts com prising each b lo c k . In both run s e ts using m ajor b lo c k s , 1 , 2 and 3 te s t re co rd s aggregated fo r o th e r census tra c ts were included. ^ xiao CM e xiao 0 P *uo00C3O) z xiao T xiao <D *3 KHpoUd Pd00)0 pto (fl 3CO oyd _<&p 30aud> CpP 5O5 Cp0O) ///da//' ///////// ip*HyCQHO '<OyU P 33 h- -ypOss vO 235 <0 MpC0dO CP<O5) d ^r ro 3 CM r*4 Py d P H CO 3d0) o ro y<0 yHcn *dH dp > XXX XXX XXX XXX XXX XX X X XX X -n4 sr-r4o O <--S4O P r--4 Or P r-l CaO. XPycO u-< o3OP a) 3? o Cycd0wQ rHCHO O P rS CO *i-Cy33Q HOydOCCHOO P0) HaP3CO CO OCd03O) ou >-4 UX] <--oo1 CQ CO oC3dyO *p0n S pd d0 Hp rys pyCuO iaH3 2od5 3 sop xa XX XX X XX XX X XX X X xX X XX XX X XX X X y yu pd yu pd dA P r-i O -4 r- e' o CT\ er. P0 H Tcy30 r-4 dd .fi p H dd js P CO y so < dyc 23 .Ao3. p daH> o y 50 < H0 oC<O0 CO yCO -3 p cyO > CO yCO *-3 p <y0 > ay y CO y aw y CO u p y 4-J r--1 doH p dH > oaJ CQ p *H <0 >3 yu . y Q cr v? <3c* p p O' *3 d y P3 j cd & o r**4 3y Q pci r>> r> cr* r-4 y00 <U P CO y w (0 <o 25 \ y p 3 y y a! r-* rrH U CO .0 < H a 4 A A y p y s y wp X3 y P5 O y yX pO u O rH cO xO y pP y Or-j 43 pE p 00 y g d ri PP 0O 4-t d y y d p y c r-4 43 P Xdi P d p> w d >g% 5 3 3 sr "3 P d *r4 dy mp X *a CXQ d d O p y0 pd yp CQ p cy 3 PS y3 d dy yQ y 3 p JO y p w y yp r-4 0 r4 43 d P dy H P p y TyP d p p > tpH d to >g1 g *o y y p to 3 <5 4: t-4 CO d --4 cS( 2-10 TEH 0531341 DUP050032552 grouped by census tract as well as by race, age and test quarter. There were thus fewer Individual test results for each aggregate ' ? record but many more aggregated records in these runs. The number of census tracts was too great and the test results within any single tract were too few to permit distinguishing each tract individually in the analytic runs. Therefore, census tracts were grouped into clusters and into major blocks for the sets of runs in this area of analysis. Dummy variables were used to denote a particular cluster or major block. Demographic characteristics of the Louisville census tracts had been investigated by Dr. Billick and in particular by his associate, Prof. Robert Earickson, of the University of Maryland. The results showed similarities among various tracts - largely contiguous Sufficient to warrant forming several clusters comprising two or more census tracts. These clusters are shown on the map in Figure L, In all, 14 such clusters were defined. Test records aggregated for each census tract were in effect combined into a single grouping or cluster according to this definition and a dummy variable, CBL to CB13, inclusive* was assigned to denote that the results represented a particular cluster. The fourteenth cluster served as the reference category. Allocation of individual census tract to a particular cluster is listed in Table 2-3. Aggregated records for these 14 census blocks were selected from each subfile (black and white) and 2-11 DUP050032553 TABLE 2-3 ALLOCATION OF CENSUS TRACTS TO CLUSTERS AND MAJOR BLOCKS Cluster NO. Census Tracts Major Block Census Tracts 1 1,2 1 1-11,16,17,21,28 2 3,4 3 6,7,8 2 20,22-27,29-34, 47-52,57-62 4 10,11,12,17 5 13,15,16 6 18.26,27,33,34 7 19,20,24,25 3' 37,40,41,44,53-55, 63-74,80-83 8 21-23 9 36-59 10 41-44 11 51,52 12 59-62 13 64-70 14 82,83 2-12 TEH 0531343 DUP050032554 regression runs made. Those records aggregated by census tract (as illustrated in Table 2-2) for tracts NOT included in any of the clusters were also investigated for each racial subfile in a separate running of the SPSS program REGRESSION (Mie, et al., 1975). There was also evidence of some connection between year of blood testing and major area of the Louisville city represented, testing in the early years was heaviest in Western Louisville and in the northwest corner of the city at the bend of the Ohio River, as shown in Figure 2. three principal target areas were identified and the contiguous census tracts within each were treated as a major block;', two sets of runs were made with records from these tracts. In one, aggregated records from these tracts were designated by dummy variables for BLKl, BLK2, and BLK3. Other census tract represented the reference category, the three blocks represented some 45 percent of the records aggregated by census tract in the subfile of whites and nearly two thirds of the black records, the tracts and their grouping are also given in Table 2-3. In the other set of runs, individual test results were first grouped according to which of the major blocks they represented. Aggregated records were then made so that each represented all the tests for a given year and quarter, age group, and census block. .Gasoline lead value was inserted and the ln(GMBL) for the test results in the aggregate was computed. 2-13 TEH 0531344 DUP050032555 2-14 TEH 0531345 DUP050032556 Regression runs with calculation of correlation coefficients between every pair of variables included in the runs were made for these census tract clusters and major blocks. Again aggregated records with less than 3 test results were excluded. The variables within each set of runs are listed in Table 2--2. Coefficients of correlation and of regression obtained are presented and discussed in Section 3. 2-15 TEH 0531346 DUP050032557 3.0 STATISTICAL RESULTS OF THE ANALYSIS 3.1 Analysis With No Use of Census Data 3.1.1 Correlation Results The runs in which dummy variables' for age group and for season of the year were the only factors included with ln(GMBL) and gasoline lead (Run Set #1 in Table 2-1) are regarded as the basic Set. The runs afforded a basis for statistically examining the effect which three types of factors have on variations in ln(GMBL): age of the subjects tested; season or quarter of the year during which the tests were made; and gasoline lead. Results for the white and black racial groups* were .quite similar. The interrelationship between blood lead and gasoline lead was by far the highest statistical association of any found in this set of runs. In contrast with results from any other similar analyses, the interrelationship was very nearly the same for the two races; studies of previous data have always reflected a higher correlation between blood and gasoline lead for blacks than for whites (MITRE, 1980, 1980a). Interdependence of all variables analyzed with geometric mean blood lead levels and with gasoline lead is shown in Table 3-1 as The number of test subjects not classified as in one of these two racial groups was, as noted, too small to permit meaningful analysis. The total pf such tests was less than one percent of all results. 3-1 8 TEH 0531347 ^ DUP050032558 CORRELATION COEFFICIENTS (ZERO-ORDER) BETWEEN CEOMETRIC MEAN BLOOD LEVEL AND INDEPENDENT VARIAM.ES AND BETWEEN CASOI.IHE LEAD AND OTHER INDEPENDENT VARIABLES (CENSUS TRACT NUT CONSIDERED) L o u ltiv ille , Ky. , Campruhetiaive Data V V .?>! ` 5C IL* C : i* c as pp e ,p c o .p p sa i* t P*P-A <erP>s o Om P P&> irt sor Ppp cPqH 1N p oo e * 2WoQ io-/eI 01 Ad --p c n <n PO si oa P N --s M O - Or* P OP P J* c a --* r>* * '* * O I* o uaCat c4uaJ> O c# irv -- pp r P <s in- op P !* * i* t 3<m e` tn AJ _ ea c a-- CaJ 3 -4 O) p Po <4 a. C. S. -N u5 Q 2 u vo P w5 <uy w * a so < V 5<0 P3 3-2 S<fHn3 *g a3 Pc r -- -- m u <a a > -t 2 * V3 VI 2 a a a & a &o a0 2a JJ < < 9a5 5 * 50 d a a < zX 2a 2 p a a5 ssyw9i f)I* u 2 ea ea 5h 1 a TEH 0531348 DUP050032559 measured by the coefficient of correlation, , and the square of that value, r ^ (where the subscript "1" denotes blood lead and the j denotes the jth "carrier" or "independent variable"), the square of the coefficient of correlation, often called the "coefficient of determination," expresses the fraction of the variation in blood lead levels accounted for by variation in a specific independent variable. On this basis, changes in gasoline lead account for about 45 percent of the geometric mean blood lead variation observed in the sample for both races, i.e,, when the gasoline lead is the same, the variations in blood lead are 55 percent of what they are overall. By contrast, the role of other factors is insignificant, no other variable accounting for as much as 4 percent of blood lead variation. These correlation coefficients express interrelationships within the sample. To.estimate the corresponding values for the underlying populations, confidence intervals at the 5 percent level of significance were also calculated for coefficients of correlation between blood lead level and gasoline lead for each race, based on Fisher's z-transform (Guttman and Wilks, 1965; Kendall and Stuart, 1967). From these calculations, confidence intervals for the coefficients of determination were then derived. These intervals represent a range within which the coefficients for the underlying populations may be presumed to lie with a probability of 95 percent. 3-3 TEH 0531349 DUP050032560 On this basis it is seen that the variations in geometric means of blood lead levels accounted for by gasoline lead variation is very unlikely to be less than 30 percent and may be as much as 57 percent* Racial Group Lower Limit r r2 Upper Limit r r2 ' Black White .568 .323 .551 .304 .756 .572 ,752 .566 These results are somewhat different from those obtained fa the analysis of the first set of Louisville data. In particular, the more comprehensive test results available in the present analysis show a substantially higher correlation of ln(GMBL) with gasoline" lead for whites (.664 as compared with .498) and one slightly higher for blacks (.673 against .605), A somewhat surprising difference is that with the larger data base gasoline lead correlated much more Weakly with quarter of the year than in the first Louisville analysis (MITRE, 1980a). Blood lead also correlated somewhat less strongly with season in the present results. The implication of these facts in regard to seasonal effects on blood lead levels is discussed below. The New York City data also showed both blood lead and gasoline lead correlating more strongly with quartet of the year (MITRE, 1980). Obviously, correlation results do not establish a causal relationship. However, the results reported here are highly consistent with an hypothesis that changes in gasoline lead represent 3-4 < TEH 0531350 DUP050032561 a stronger Influence in determining increase or decrease in blood lead levels than do any of the other factors analyzed in the first set of runs, i.e., those associated with season of the year or age of children tested. It is also clear that these factors are associated to a lesser extent statistically with changes in blood lead level. Hence they may be presumed to exert some influence on blood lead, as discussed next. 3.1.2 Evidence of the Effect of Other Factors The statistical results are quite incompatible with any theory that the correlation between gasoline lead and blood lead levels Could reflect the accidental result of some other causal factor associated with age or season of the year. These conclusions are indicated by Table 3-2 for which the higher-order correlation coefficients (and the corresponding coefficients of determination) were calculated for .selected variables showing a high statistical interrelationship with changes in In(GMBL). These coefficients measure the extent of variation between two variables when other variables are in effect held constant (i.e., their effect is subtracted out). As can be seen in Table 3-2,. in general the statistical correlation between gasoline lead and blood lead is slightly higher under conditions such that these other factors do not vary than when they <io, In particular, removing the effect of ALL three of the other variables with which ln(GMBL) correlated other 3-5 TEH 0531351 DUP050032562 cV %= r* c - \ oS *n -- c m c*i O' O Q ` * ?r* c *0 o TABLE 3-2 PARTIAL CORRELATION OF SELECTED VARIABLES COMPARED WITH ZERO-ORDER COEFFICIENTS RONS IGNORING CENSUS TRACT, LOUISVILLE, K Y ., COMPREHENSIVE DATA ua c r - -- o * O CO P*. n0O0 r>>0 <h o\0v <02> tr*t CM *-3-- "0~'0-a3cC^'CNor-o1inCo'c-l?H 9t rt * 2 im la <r C 0 .0 O 0-0 O <4* 41 J* <* ! <N CO x *a- ^x *j s r> in rt n wt o f>> n O xr> com tr*- m xn rs*o* qvo 1" oocoooooo * < frt CM CM CM ,cn , V >. Ui> oCC ce zo zo zco zc5 zoc zco * <3u0 < o w .w Or n n u* 30 e 3 oi .01 u (0 2 S4 m C ao es 3 O' 41 m3 m os 0 CO w< -n .01 Jj O' 0 fl c CM 01 60 -9 e q 41 4J * W 0) m oi 00 *0 c n -o os 01 g o co 3 g Q fiO H CM 0 < $ 3 6 CM oeso 3 <0 cr* oar cr .a> *3 -c* <--n n 13 3 ,* H5 *0? *"3 "*3 *3 <3 3-6 3 c G C 01 o o uo o I 1 % *2 3 si a a) O <00) - <0 3 1- w 3 Or O' "3 o C'CM W rt V, *0 * "3 T3 8 4i c TEH 0531352 DUP050032563 than trivially (as measured by the zero-order correlation coefficient) results in a third-order correlation coefficient which is about two percentage points higher than the coefficient of zero order, the very small changes between the zero-order and higher-order coefficients for blood lead and gasoline lead are exactly what would be expected from the fact that gasoline lead correlates quite weakly with these other variables. Therefore, removing the effect of this interrelationship has little effect on the measured correlation between In(GMBL) and gasoline lead,* The correlation coefficients between ln(GMBL) and both second quarter and age group 2 also increase slightly when the effects of gas lead are removed as can be seen in Table 3-2. This result is as expected from the fact that gas lead is virtually uncorrelated with either of these other two variables, as shown by the zero-order coefficients which are .012 and .022.' The statistical evidence points to the role which some (unknown) factors associated with this particular age and season of the year have in accounting to a lesser extent for some changes in In(GMBL) independent of gasoline lead. *As Kendall and Stuart in their monumental work on statistics (1963, p. 334) write: "If we find that holding another variable fixed reduces the correlation between two variables, we infer that the interdependence arises irt part through the agency of that other variable,,.. Conversely, if the partial correlation is larger than the original correlation between the variables we infer that the other variable was obscuring the strong connection...." 3-7 * ' *!> . TEH 0531353 DUP050032564 Attempts to find one or more real-valued variables to represent age differences were not successful. As can be seen from Table 3-1, age duration showed a weak correlation with ln(GMBL). The correlation was negative, as expected; the farther away a test subject's age was from the age at which blood lead level manifested a peak, the lower ln(GMBL) was, other things equal. Mean age and mean age squared showed only a trivial correlation. As is discussed in the next section, these variables also contributed little when used in linear regression models. 3.1.3 linear Regression of ln(GMBL) j, 3.1.3.1 Regression Coefficients. Regression coefficients were calculated On the assumption of a linear relationship between the dependent variable (natural logarithm of GMBL) and the other so-called independent variables; specifically, that the observed value of In(GMBL)(y) can be closely approximated by the value (y) calculated from the following expression: y + BiXi + e i where x^ value of the ith independent variable. The constant factor, o, is presumed to account for those elements which affect variation in blood lead but which ate not included in the regression, analysis. The value of each element, Bi, termed a regression coefficient, expresses the number of units of variation in y expected for each unit variation in x ,, holding all 3-8 TEH 0531354 DUP050032565 other independent variables constant. For example, in blacks, the coefficient of regression for gasoline lead has a value of 0.921 (table 3-3) measured in hundred millions of grams. It would therefore be expected that when all other independent variables (age group and quarter of the year) are identical, a difference of 1 unit in gasoline lead would be reflected as a difference of about 0.92 units in In(GMBL); i.e., when gasoline lead increases by one unit with no change in the other independent variables the natural logarithm of the GMBL would rise by 0.92 and the geometric mean blood level would increase about 2.51 pg/100 ml. These values as calculated are shown in Table 3-3. They represent the regression coefficients for the sample population. The SPSS subprogram REGRESSION routinely calculates values for the F-test and these are also given in table 3-3, The F-statistic is used in testing the hypothesis that the sample comes from a population for which the regression coefficient is actually zero and the calculated non-zero value results from sampling fluctuations (or measurement errors). In other words, the F-statistie tests the hypothesis that the independent variable for which the regression coefficient is calculated contributes nothing to predicting the value of the dependent variable, here ln(GM3L). The F-statistic is calculated as a ratio and expresses the probability of obtaining a value larger than that calculated for a sample of given size Erora a population in which the value of the 3-9 TEH 0531355 DUP050032566 TABLE 3-3 REGRESSION COEFFICIENTS FOR LOG GEOMETRIC MEAN BLOOD-LEAD LEVEL, IGNORING CENSUS TRACI AND USING AGE RANGE LOUISVILLE, KENTUCKY, COMPREHENSIVE DATA Independent Variable Age Range 1 Age Range 2 Age Range 3 Age Range 4 Age Range 5 Age Range 6 Whites (Data Sets 125) Regression F-Ratio Coefficient Value -0.078 0.059 0.120 0.075 0.043 0.049 1.917 1.119 4.715 1.820 0.593 0,769 Blacks (Data Sets = 136) Regression F-Ratio Coefficient Value -0.098 0.016 0.094 0.061 0.010 0.031 4.836 0.138 4,439 1.874 0.051 0.485' Quarter 1 Quarter 2 Quarter 3 0.131 0.170 0.111 8.56! 14.756 6.653 0.083 0.142 0.069 4.962 15.637 3.757 Gas Lead Constant Terra 1.051 2.297 R2: 2 Adjusted R : Critical F Values: 5% level 2.5% level .555 .515 112.562 '-- 3.93 5.31 0.921 2.507 .565 .530 131.817 -- 3.917 5.145 3-10 TEH 0531356 DUP050032567 regression coefficient is actually zero. Hence the hypothesis of a zero value is rejected (at a given level pf significance corresponding to a set probability of error) if the calculated value ratio exceeds the critical or threshold value. These critical values, which have been compiled in standard tables (e.g.. Mood and Graybill, 1963) depend on both number of independent variables and number of data points (sample size) used in the regression calculations. The critical value for the F-ratio for each set of regression coefficients is also given in Table 3-3. With sample size and number of independent variables involved here, the hypothesis of a zero value for a population coefficient of regression can be rejected at the 5-percent level of significance for a calculated F-value exceeding about 3,9 for either race individually. As can be seen from Table 3-3, the hypothesis can on this basis be rejected ONLY for gas lead, age range 3, seasons of the year, and age range 1 for blacks. For the other variables, the hypothesis cannot be rejected with only 5 percent chance of error: these regression coefficients may in the underlying populations . actually be zero, in other words, virtually as good results would be obtained if age were ignored in the model. 3.1.3.2 Coefficient of Multiple Correlation. From the coefficients of correlation expressing the interrelationship between each pair of variables, it is possible to calculate a coefficient of 3-11 TEH 0531357 DUP050032568 multiple correlation, usually denoted as R. . The square o 2 this value, R , often termed the coefficient of multiple determination, measures the extent of variation remaining in the dependent variable when the values of the independent variables are held constant. Thus, for a value of R=0,S, R 53 0.25, indicating that only 25 percent of the variation in the dependent variable is accounted for by variations in the independent variable. Perfect correlation would of course be indicated by a value of R (and hence 2 R )=1, which is the situation when the dependent variable can be completely predicted from knowing the independent variables; i.e., for a given set of values for the independent variables, the dependent variable always has the same value (for example the relationship that exists between cost of a bag of groceries and cost of each item in the bag). Values of R and R2 obtained for blood lead level in each race as a function of all the variables considered in the analysis (gasoline lead, race, age range and quarter of the year) are given in Table 3-3. About 45 percent of the ln(GMSL) variation is statistically indicated as resulting from unknown factors not used in the regression analysis, ' These influences could reflect other features in the environment, anomalies in the sample data, or both.^ By contrast, in the New York City data as much as 80 percent of the blood lead variation was explained by the independent 3-12 TEH 0531358 DUP050032569 variables considered * On the basis of these indications, results of using the regression coefficients to predict for Louisville expected values of ln(GMBL) as a function of a given set of values for the independent variables would be much poorer than expected for New York City, in summary, the regression model developed by the Louisville analysis does not appear to be as useful for Louisville as does the corresponding New York model. In general, the regression results obtained from Run Set #1 were as good as or better than any as is discussed next. The use of age deviation in Run Set it2 proved of no consequence as indicated by the F-ratio value for the regression coefficient. Nor did mean age and age square (Run Set #6) result in a more useful regression equation; the coefficients for these two variables were minute, the F-ratio values very low as computed for blacks and slightly below the critical level for whites. Regression coefficients computed for quarter or season of the year differed very little between Run Sets #1 and it6. In Run Set it2, the coefficients computed for these vari ables were smaller and the F-ratio values lower than in Run Set #1, These results compare closely wih those obtained in the earlier Louisville analysis, where the R4 values were .590 for blacks and only .501 for whites. Although with the much larger (by an order of magnitude) sample size, the New York City data resulted in a much closer fit of the regression model, in Louisville increasing the sample about 50 percent resulted in no improvement. 3-13 TEH 0531359 DUP050032570 Complete sets of regression coefficients and accompanying F-ratio values are provided in tables in the Appendix. Introduction of the squared term for gas lead in Run Set #2, with the coefficient of correlation with gas level near unity, resulted in multicolllnearity in the model. No gains to offset this drawback were observed in the equation which used both gas lead variables. In particular, the difference in sign of the regression coefficients (positive for gas lead and negative for gas lead squared), which is counter intuitive, indicated poor results from the double rate assigned gasoline lead. 3.1.3.3 Runs Using a Variable to Reflect Time Trends. Run Sets 3, 4, and 5 were designed to examine the possible effect on blood-lead levels of long-term time trends. . `That is, it was conceivable that the measured geometric mean of blood lead might reflect a general tendency either to rise or to fall oyer the testing period (1974-1.979) for any of a variety of reasons; possible influences thus manifested could include changes over time in gasoline lead levels, in the population from which the samples were drawn, or in extent of exposure to lead from other sources. Accordingly, as noted in Section 2, three sets of runs were made, each of which used a different variable to distinguish test period: dummy variable reflecting whether the test occurred before or after 1976; dummy variable similarly distinguishing tests before and after 3-14 TEH 0531360 DUP050032571 1977; and "quarterly deviation" - an Interval variable representing the difference between the sequentially-numbered quarter in which the test occurred and the sequential number of. a mid-quarter in 1977 (as will be recalled, no test results from that year were used in the analysis) 3,1.3.4 Correlation Results. All three of the above variables correlated strongly with both gasoline lead and ln(GMBL) as can be seen from Table 3-1. The dummy variable for year earlier than 1977 showed higher correlation coefficients with both than did that for year prior to 1976; this result was consistent with the fact that quarterly deviation (measured from 1977) also correlated quite 'Strongly with both gasoline and blood lead. It was subsequently observed (in another connection) that all quarterly gasoline lead values prior to 1977 were well above 0.5 (measured in billions of grams), whereas after 1977 in no quarter did it exceed 0.493 and tended to drop off even further. Indeed its mean values were about 0.74 for the earlier period and 0,45 for the later (as shown in Run Set in discussed below). Overall mean values of ln(GMBL) also declined from around .3.2 before 1977 which indeed provided a dividing line in the context of the study. These correlation results were, however, all quite ambiguous as to the reason for the decline in blood lead levels and in particular as to the possible role of gasoline lead therein. As already noted, 3-15 TEH 0531361 DUP050032572 statistical evidence would in no event establish a causative relationship* But while consistent with an hypothesis that declining blood gasoline lead consumption over time was the reason for a drop in blood-lead levels, the evidence was also consistent with an opposite theory that both gasoline and blood lead dropped independent of one another during the period of testing. In either event, one would expect the strong 3-way correlation that was observed among variables representing time, gasoline lead, and blood lead; the first-order coefficient of blood lead with either of the others, controlling for the third, would be very much lowered because of he strong mutual interrelationship. Id attach this ambiguity. Run Set #7 was made in which test results were separated into those prior to 1977 and those comprising test years 1978 and 1979. The results were quite interesting and appear significant. Although the reduced sample size results in much wider confidence intervals for the coefficients so that the numerical values must be treated with some caution. Table 3-4 leaves little room for doubt as to the significance which variations in gasoline lead consumption have in accounting for variations in ln(GMBL) after 1977. The correlation coefficient of over 0.62 between gasoline lead and ln(CMBL) for whites is more than twice the value of any other correlation coefficient with In(GMBL); test results from blacks are very nearly as strong in this regard. Interestingly enough, in the 3-16 TEH 0531362 DUP050032573 TABLE 3 -4 O r> < as 3^ Ed S < fits 5'<a MM 3W z Mcn 5a3 *f J ca a* OSwvHn>J 't5> ZP pMtil oX^ o t-A* CS .3 V] cn u H SI S4 u5(bM5d-J,cuUihjj)* Ifa. 03 > OOCd M0<5 MVaJ So os o i* oIt r o r i smeCS:) --4J< u * 3: *0os* 0C1l o A O frr ot m > r> n * vMn x o fS 3 o r >UM S u: 3 js t 1". 34 O(3 US F-* w w 5 $ T1-* 31 -2 * i O 3<n g; S 'u .-* JCZl 05 O *"> c- mCM CO ooo s CM *H iM i-4 3 O r 0 & 3-17 <a9 3 -40<J3 <ag- *9 M Q (I MaQ) *g n (9A <H <M O 0o s > sI _ TEH 0531363 DUP050032574 pre-1977 data gasoline lead is far less closely associated statistically with variations in ln(GMBL). Indeed among both whites and blacks other factors - age and Season of the year - show a stronger interrelationship.. 3.2 Analysis of Data Using Census Tract Information As noted in Section 2, four sets of runs (designated arbitrarily CBliK 1 through CBLK4) were made with data aggregated on the basis of census tract. In the first three, each subgroup or aggregated record consisted of test results from a single census tract (and of course homogeneous as to race, age group, gasoline lead, and test quarter,). The aggregated records were then further designated as to whether or not the census tract belonged to a particular one of 14 clusters (runs CBLK1 and CBLK3) or one of 3 major census blocks (CBLK2). The clusters and blocks in Runs CBDC1 and CBLK3 were distinguished by the fact that the former selected 0N1.Y aggregated records from one of the 14 clusters, whereas the latter selected those from other Census tracts not included in the cluster set. In the fourth set of runs (CBLK4), individual test records were aggregated on the basis of each major census block without regard for individual tract designation. Thus each aggregated subgroup record was homogeneous as to race, test quarter, age group, particular census block (or no block), and gasoline lead consumption. 3-18 DUP050032575 the results of these groupings were that in the fourth run there were fewer total aggregate records than in the other three hut the number of individual test results in each aggregated grouping was larger. The computed Ln(GMBL) value for each could therefore be expected to represent more nearly the underlying population, with the effect of fluctuations due to small sample size smoothed out. All of these runs used the SPSS subprogram REGRESSION (Nie, et al., 1975), with printout of mean values for the variables and a complete matrix of correlation coefficients. 3.2.1 Objective of the Runs ), The objective of all these sets of runs was to investigate what statistical interrelationships(s) existed between variations in blood-lead levels and differences in residence of the subjects tested. Despite the strong correlation between ln(GMBL) and gasoline lead noted, it was clear (as discussed above) that other factors also accounted for the variation in blood lead. In particular, it was hoped that these unknown factors might be related to place of residence (e.g., condition of home as to lead pipes and lead paint, proximity to lead from industrial discharges, etc.) and that this situation would be reflected in the statistical results of the runs. Because the data in these runs represented different aggregated groupings, it would not be expected that the correlation coefficients of blood and gasoline lead with variables reflecting age and quarter 3-19 TEH 0531365 DUP050032576 or season of the year (used both in these runs and in the runs which ignored census tract) would have the same values. For explanatory purposes, ideal results would have included: correlation coefficients between blood and gasoline lead that were similar to those found in the noncensus tract runs; high correlation of blood lead with variables denoting specific census clusters or blocks; and 2 greater R values than were obtained in the concensus tract runs, thus reflecting additional statistical interrelationships with ln(GMBL) not found in these first runs. 3.2.2 Correlation Results From the above point of view, however, the results obtained from the use of census tract data were disappointing. Coefficients of correlation between both gasoline lead and In(GMBL) on the one hand and the other variables used on the other hand are given in the columns of Table 3-5. These tables also show R2 values. Regression coefficients are given in Table 3-6. As can be seen from Table 3-5 no coefficient of correlation between a census variable and 'lti(GMBL) exceeded an absolute value of .236 so that none of the former variables accounted statistically for as much as 6 percent of the blood-lead variation. Highest positive values were observed for census-tract clusters (CB) 5 and 7 and major block 2, among whites, and for CB 7 and major block 2 among blacks. There is considerable commonalty among the census tracts comprising these census clusters 3-20 TEH 0531366 DUP050032577 TABLE 3-5 CORRELATION COEFFICIENTS OF BLOOD LEAD AND GASOLINE LEAD WITH OTIIKR VARIABLES, USING CENSUS TRACT INFORMATION 1XM1ISVILLE, K Y ., COMPREHENSIVE DATA (RUNS CULK 1, CBIR 2 , CBLK 3) 3liVcom|mtable, owing to sm a ll sample s iz e *Duimy V a ria b le s represented a l l values In th is column except Cas Load. >-- c m oq co -a m -T s ci m s . e a tilt cm m m <r -- c-i MOO t1 -CM CM CO 00 -- <o oi n n n n 0 a ii i 0 *H CM CM H <-- QHO 11 j 3 j n vw m.m >9 ; < r - o nn c 0.0 . oo i r j* f O < CO n c m m CM Q 0 * * r < CO o*M o o-- v a NI0M.-C8M MT O O ; co c m m m m m ncH : o o o 0 ooo M* * * f 1* i ij i i i i r i i i ti i t it i i i i C CO 0 9 <o- -i mt tmn Cl - AO CM CO O N CO <9 ----C-- o - mmm m O CQ Ih ^sOOWMCS | I 1 CO I H-m0- l m c m in rv yo m i i i i mOO e0--e a o i1 o2 tnnrp.CO n n *j ni in in nn O 'O mI rti <*13 0rt1 ci m 8 m r* -- c m i iA c o c n n o j ( i i 0 -r- o O -0 Q 0 -- 0 O o c m * " p r i* i* r j lift i Ji lOOcO^M.OvCiD- OHMOn O-h> ^O*hH Pmc iii ill! ! i i I Io o ^o -- in P*| CMft O N <*3 f(\ 02 H H o c oc o o eoo II m p i N N sr n N <o N -I N N O O N (MPl oapocc n o o tI I I I 1 I I I 1 t I ii i i i i i i i t t i r> c m rvw c m so *>.-!/> c O'H O mc m r m COO O OO O 00 m cm m crvm ONOCI C> O i i p*iNcqn-p3'0om m .5c0s ST-* C-- c--*i CinO >0 0 CC .<0M1 --O' iCvM. moo m <r i C j m -k -rv .c m .e m m .m O PM -r tM CM O Ho *C o f**V\ 00 01-1-0 30 ICVCM T-NT CM Tv O 'O C -- -- cC O -- -- O mm O' N t 1 O CM I I -- CM mv--TaO* m--cvoh m-ocvMr om----p<--^d -- <M n w* -- CM m* CM m T /> *6 Cv oo P <5 -- c m m 3-21 TEH 0531367 DUP050032578 and major block 2 as can be seen from Table 2-3. The obvious inference is that some (unknown) conditions associated with residence in these census tracts are related slightly to an increase in blood-lead, other factors being equal. Analogously, highest negative values among both blacks and whites were found for coefficients of correlation between ln(GMBL) and census cluster #2; a relatively large negative value was also found for the correlation coefficient between cluster number 13 and ln(GMBL) for whites, there being too ..few black subjects tested from this cluster to permit computation of the coefficient. For no major block was there a significant negative correlation coefficient with In(GMBL). Indeed, the coefficients are so low as to imply that residence in one of the tracts comprising major blocks 1 and 3 has little if any significance in explaining blood-lead variations. Most disappointing of all, from the standpoint of interpreting blood-lead variations, is the fact that R2 values for all of these sets of runs were substantially lower than in the sets which ignored census tract information. This observation may be related to the fact that the coefficients of correlation between blood gasoline lead and In(GMBL) - while the highest of any observed in the runs - were significantly reduced over those obtained in the run sets previously discussed. It would be tempting to infer that among children in the selected census tracts factors other than gasoline lead played a 3-22 TEH 0531368 DUP050032579 somewhat greater role in accounting for blood-lead variations than among Louisville children generally* While the results of run sets CBUC1 and 2 are consistent with such an hypothesis, those of CBLK3 are not- It can be observed that for test subjects outside the tracts grouped into 14 clusters, aggregation on the basis of census tract results in lower values both for overall R2 and for the coefficient of correlation between gasoline lead and .ln(GMBL) than when data are not so aggregated (Run Sets 1 - 6). A more satisfactory explanation - in that it accounts for mote of the significant observations - would be that aggregation by census tract, which resulted in a smaller number of test results per aggregated record, somehow obscured the relationships manifested with ln(GMBL) computed for each aggregated record from a larger sample size. In any event, the correlation results of the last four Sets of runs provide little basis for attempting to explain additional casual factors in blood-lead variation. 3.2.3 Regression Results The R2 values obtained in all of the runs using data aggregated by census tract were so low as to indicate that the models would be virtually worthless for predictive purposes. However, the regression results are of some interest in other regards. The regression coefficients and the F-ratio values are listed in Table 3-6. 3-23 TEH 0531369 DUP050032580 ii 3M-I5552 ' S,2sIS SJ 2|2IS=ll4J23iiSliill15S5SI dodooo S o' n o d o o Ji 833ni2a S isS I! msma s, . """ ""illiiI3IIiI!IriiI] ill iSSV.^ummvmiisvm 'H ' "" . 4 *si s illl " iu S358.SS 33*S23 I S3 || ^dddddddd xydddddd dd 3323*315 3.5*22 d8 = S-~~||||3|||31|SS|S'M-- _ "*! 22121211212122................" iii -p-a^333333|333|js3-,S3j?=3*.3 3 SUB*SMSIM"I,,I,,II*SSJSJI (jeeeeooeo ^ ooo s is ssssaasas d! ""'"'s==S|j|2i|3]!|i|s',|ii , "U |.H55M^I22222I222225-22I2 " od'"......... .. dd *3 35S.Ss33333 5235^252, 5`fS d d d d d d d J>1 d d <r d o A o o dp d o n v -o -o l! II Jill " IS. 513=S2iS2iS 25SI223SS3S Si'S ddddddddd<f<i ddddddddddd 11 illllllliSlIlIllllliSlilllii iii! 3-24 TEH 0531370 DUP050032581 First of all, it will be noted that where the number of aggregated records was small -- as in Run Set CBLK1, which had only 133 data points (n) for whites, by far the smallest number in any of these runs - low F-ratio values were obtained for virtually all regression coefficients. The only exception is for gas lead, which everywhere produced a high F-ratio value, indicating again the significance of: this factor. The low F values, which are not surprising, suggest that a small sample size tends to be less representative in regard to the characteristics of interest here of Che underlying population than does a larger sample. It will be t noted in the results for whites in Run Set CBLK1 that the hypothesis of a zero population coefficient of regression can be rejected at the 5-percent level ONLY for Cluster #2, in addition to gas lead. For blacks, where the sample size is over 4 times as large (611 aggregated records was compared with 133), the number of significant regression coefficients is substantially greater. However, even here the hypothesis of a zero population coefficient cah be rejected for only about one half of the census clusters and other independent ot carrier variables. In other words, the regression models - poor as they are, as indicated by the low R2 values - would not be significantly degraded if ALL independent variables except gas lead and the dummy variables for cluster 2 had been omitted for whites. , 3-25 TEH 0531371 DUP050032582 However, nearly all of the dummy variables for age and season are significant to the model derived for whites in Hun Set CBLK 2 with a sample size an order of magnitude greater; this run set used aggregated records from census tracts other than those in the 14 blocks. With blacks, virtually as good results would have been obtained in Run Set CBLK1 if about half of the independent variables were ignored; the model needs to use -- in addition to gas lead - only the dummy variables for age group3; possibly for age group 4, the F-ratio value for which is marginal; season of the year; and census clusters 2, 7, and 8. It will also be noted that adequate test results were lacking for each race from some of the census tracts, reflecting a difference in population distribution. On the other hand, where thesample size is very large - as with Run Set CBLK3 (grouping data into3 major census blocks) where there were over 2000 aggregated records for each of the two races hypothesis of a zero popolation coefficient can be rejected for virtually all variables used in the models. Except for age group 2 among blacks and age groups 1 and 5 among whites, the results indicate that all variables used make a significant contribution in fitting the observations to the regression surface. Unfortunately, however, even here the R2 values are too low to support the model for predictive purpose. It is also of interest to note that the regression results In Run Set CBLK3 are consistent with the 3-26 TEH 0531372 DUP050032583 correlation coefficients calculated in supporting the hypothesis that children in the three major census blocks have slightly higher blood-lead levels than children elsewhere in the metropolitan area, other things being equal* The results in this regard are the same for both races. Further, it can be seen that the regression results in Run Set CBLKi are consistent with the correlation results in that they point to specific census clusters where residential conditions appear most interrelated with blood-lead levels. The negative regression coefficients and the F-value (above the 5 percent significance level) for cluster #2 with both races agree with the negative correlation coefficients obtained for this cluster with ln(GMBL). Children in this cluster tend to have somewhat lower blood-lead levels than children in the other clusters. Conversely, the high F-ratio values obtained for the positive regression coefficients in the results for blacks in clusters 7 and 8 imply higher blood-lead levels here than elsewhere. These are also the clusters for. which the highest positive correlation coefficients with ln(GMBL) were obtained. As to what residential conditions may induce these deviations, the statistical results of course provide no evidence. As a minimum, then, the regression results support the hypothesis that residence in the census tracts comprising the , clusters and the major blocks is somehow associated with higher .3-27 TEH 0531373 DUP050032584 4 blood-lead levels. They also point to a few particular census tracts (within specific clusters) where it would be of greatest potential value to examine residential conditions for factors that might be causally related to elevated levels of blood-lead, finally, the runs using census tract information are consistent with the evidence widely observed elsewhere in this study that variation in gasoline lead consumption is the factor most strongly associated with variations in Ln(GMBL). 3-28 TEH 0531374 DUP050032585 APPENDIX This appendix lists results of regression models developed in runs which were not of statistical significance for purposes of the study. They are given here for the sake of completeness. A-l TEH 0531375 DUP050032586 TABLE A-l REGRESSION COEFFICIENTS FOR LN(GMBL) BY RACE, NO CENSUS TRACT DATA, RUN SET #2, USING AGE DEVIATION Variable Age Deviation White Regression F-Ratio Coefficient Value -0.3xl0~4 2.921 Quarter 1 0.081 3,229 Quarter 2 0,110 6.044 Quarter 3 0,070 2.380 Gas Lead 3.675 28.584 Gas Lead Squared -2.022 14.785 Constant Term R2 Value 1.600 .568 -- Critical F-Valvie 3.9 # Aggregated Records 123 Black Regression F-Ratio Coefficient Value' -.9xl06 0.010 0.038 0.770 0.088 5.408 0.025 0.461 3,178 31.831 -1.728 16.424 1,871 -- ,562 3.9 136 A-2 TEH 0531376 DUP050032587 TABLE A-2 REGRESS IQS COEFFICIENTS FOR LN(GHBL) BY RACE, NO CENSUS TRACT DATA RUN SET #3, USING TEST QUARTER DEVIATION Variable White Regression Coefficient F-Ratio Value Age Group 1 -0.088 2.689 Age Group 2 0.055 1.103 Age Group 3 0.117 4,938 Age Group 4 -0.071 1.840 Age Group 5 0.393 0.559 Age Group 6 0.452 0.739 Test Quarter Deviation* -0.022 13.104 Quarter 1 Quarter 2 0.022 0.144 0.184 11,772 Quarter 3 0.112 7,435 Gas Lead 0.192 0.564 Constant Term R2 Value ' 2.828 .602 -- Critical F Value 3.9 ft Aggregated Records 123 Black Regression F-Ratio Coefficient Value rCOdM 0 1 5.651 0.010 0.057 0.090 4.427 0.057 1.785 0.006 0.022 0,027 0.405 -0.017 11.483 -0.008 0.119 0.070 0.268 2.914 0.032 11.459 4.085 1.672 -- .602 3.9 136 *Numeric deviation between sequential number of test quarter and 31 (a quarter in 1977 taken as dividing point). A-3 TEH 0531377 DUP050032588 TABLE A-3 REGRESSION COEFFICIENTS FOR LN(GMBL) BY RACE, NO CENSUS TRACT DATA, RUN SET 04, USING DUMMY VARIABLE TO DISTINGUISH TESTS BEFORE AND AFTER 1977 Variable White Regression Coefficient F-Ratio Value Age Group 1 -0.085 2.330 Age Croup 2 0.057 1.102 Age Group 3 0.119 4.748 Age Group 4 0.073 1,809 Age Group 5 0,041 0.575 Age Group 6 l Dummy Variable for Time 0,047 0.139 0,751 4.205 Quarter 1 0.101 4.778 Quarter 2 0.170 15,503 Quarter 3 0,114 7.182 Gas Lead 0.656 9.179 Constant Term R2 Value 2.471 .571 -- Critical F Value 3 .9 0 Aggregated Records 123 Black Regression F-Ratio Coefficient Value -0.098 5.136 0,015 0.120 0.093 . 4.619 0.060 1.939 0.010 0.049 0.030 0.494 0.147 7.449 0.048 0.138 0,067 0.516 2.684 1,561 15.660 3.728 9.492 .589 3 .9 136 1 Value of I denoted test before 1977, value of zero test after 1977. A-4 TEH 0531378 DUP050032589 TABLE A-4 REGRESSION COEFFICIENTS FOR LN(GMBL) BY RACE, NO CENSUS TRACT DATA, RUN SET #5, USING DUMMY VARIABLE TO ` DISTINGUISH TEST BEFORE AND AFTER 1976 Variable White Regression Coefficient F-Ratio Value Age Group 1 Age Group 2 -0.086 0.051 2.647 0.964 Age Group 3 Age Group 4 0.113 0.067 4.694 1.668 Age Group 5 0.035 0.457 Age Group 6 ^Dummy Variable for Time (1976) 0.041 0.170 0.624 15.926 Quarter 1 0.123 8.615 Quarter 2 0.182 19.420 Quarter 3 0.123 9.891 Gas Lead 0.746 38.239 Constant Term ?,Z Value 2.435 15.926 0.611 Critical F Value 3.9 # Aggregated Records 123 Black Regression F-Ratio Coefficient Value -0.104 5.974 0.007 0.027 0.088 0.055 4.265 1.670 0.004 0.009 0.025 0,346 0.128 13.305 0.070 3.914 0.148 18.595 0.084 5.907 0.690 48,172 2.614 -- 0,607 3.9 136 ^Value of 1 denoted test before 1976, value of zero test after 1976. A-5 TEH 0531379 DUP050032590 TABLE A-5 REGRESSION COEFFICIENTS FOR LN(GMBL) BY RACE, NO CENSUS TRACT DATA RUN SET If6, USING MEAN AGE AND AGE SQUARED Variable White Regression F--Ratio Coefficient Value Mean Age 0,004 3,588 Age Square Quarter 1 -0.0005 0.132 - 3.913 8.319 Quarter 2 Quarter 3 0.166 0,111 - 13,727 6.299 Gas Lead ,1.054 107.820 Constant Terra R2 Value 2.258 .517 -- Critical F-Value 3.9 If Aggregated Records 123 Black Regression F-Ratio Coefficient Value 0.002 -0,0001 1.393 0.907 0.080 4.313 0.141 14.195 0.067 3.198 0.913 119.425 2.486 -- .513 3 ,9 136 A-6 TEH 0531380 DUP050032591 REFERENCES Billick, Irwin H,,, Anita S. Curran, and Douglas R. Shier, 1979. Relation of Pediatric Blood Lead Levels to Lead in Gasoline. Submitted to Environmental Health Perspectives. Coerr, Stanton and Irwin H. Billick, 1979. The effect of lead in gasoline on lead in the blood of urban populations. Memorandum to David Hawkins, Assistant Administrator for Air and Hazardous Materials, U.S. Environmental Protection Agency. April 3, 1979. Dixon, Wilfrid J. and F.J. Massey, Jr., 1969, Introduction to Statistical Analysis. Third edition. McGraw-Hill. Guttman, Irwin and S.S. Wilks, 1965. Introducing Engineering Statistics. John Wiley & Sons, Inc., New York, NY. Kendall, Maurice G. and Alan Stuart, 1967. The Advanced Theory of Statistics. Vol. 2, Section Edition. Hafner Publishing Co. Mood, A.M. and F.A. Graybill, 1963. Introduction to the Theory of Statistics. Second edition. McGraw-Hill. Nie, Norman H., C.H. Hull, J.G. Jenkins, K. Steinbrenner, and D.H. Bent, 1975. 'Statistical Package for the Social Sciences. Second edition. McGraw-Hill. The MITRE Corporation, 1980a. Statistical Analysis of Blood Lead Levels in Relation to Gasoline Lead, Louisville, Ky. WP-80W00466. McLean, VA. June 2, 1981. The MITRE Corporation, 1980, Correlation and Regression Analysis of Blood Lead Levels. WP-80W00447. McLean, VA. December 1, 1980. 4-1 TEH 0531381 DUP050032592 D-12 S. W. GOUSE V-50 L. KUSHNER R. P. OUELLETTE L. W. THOMAS W-80 M. M. SCHOLL E. G. SHARP W-53 T. J. WRIGHT L. J. DUNCAN (A) W-58 G. ERSKINE J. WATSON (3) W-53 DISTRIBUTION LIST UP# 81W00541____ Metrek Library (2) Warehouse (Storage) (I) Records Resources (I) W-50 Data File (5) (W187) SPONSOR (6) i . AUTHORIZED FOR DISTRIBUTION Thomas yj. Wrj4ht Department Head, W-53 aftn Project Approval TEH 0531382 DUP050032593