Document NeXr9X7qkgkb6XOr2VD5g2mvy
The effect of analytical eriror on the interpretation of blood lead data obtained from population studies is discussed. Statistical analy sis of measured blood lead values reported in four representative population studies shows over 40 percent of the variability in blood lead measurements is due to analytical error. Because of this analytical error, a large percentage of the observed blood lead measurements which exceed a threshold limitare "false positives" in that the true blood lead values forthese individuals do not exceed the threshold. The EPA has suggested that the distribution of blood lead values could be used as a guide for determining the portion of a population at risk. "At risk" is defined as having blood lead values above a certain threshold. If this is done, the effect of analytical error which causes "false positives" should be considered properly to obtain an accurate estimate of the population at risk.
Effect of analytical variability on measurements of population blood lead levels
JAMES M. LUCAS, Ph.D., Consultant E.l. Du Pont de Nemours & Company, Jnc., Wilmington, DE 19898
introduction
The Environmental Protection Agency has suggested that the distribution of blood lead values could be used as a biological guide for determining the portion of a population at risk as defined by blood lead levels above a certain threshold.01 A similar guideline has been adopted by the Commission on European Communities.1201 In general,
these biological guidelines, based on the cumulative frequency distribution of blood lead values in a population, enable the determination of the portion of a population which exceeds a specified threshold blood lead value. Although this approach takes into account the inherent variability of true blood lead values of individuals in a population, it can be misleading if it fails to consider fully the impact of analytical errors on the measured blood lead values.
The size of the analytical error in blood lead determinations that can be encountered in a highly experienced laboratory is shown in a recent publication.14' In this study, a blood sample obtained from a single donor was divided into 35 separate samples. These samples were submitted along with other blood samples to the Kettering Laboratory over a period of nine months. The mean measured blood lead value was 19.13 fig/100 g with a standard deviation of 5.72. Values ranged from 12 to 42 /ug lead/100 g of blood. If the extreme value of 42 #tg/ 100 g is rejected as an outlier, the coefficient of variation is 21 percent. If it is included, as it wou ld be in a test where only a single sample is made on an individual, the coefficient of variation is above 25 percent. The Kettering Laboratory is accredited by the American Industrial Hygiene Association, has long experience with blood lead analysis, and has performed well in many interlaboratory comparisons. Larger analytic errors have been reported in other studies*51 so errors such as these are- not uncommon. While the presence of this analytical variability has been recognized,161 its effect on the biological guideline blood lead distribution and on the proportion of individuals who will exceed a
specified threshold has not been assessed. This paper uses a simple analysis of variance model to show the large effect that analytical errors can have on the blood lead distribution and on the proportion of individuals who will exceed commonly used thresholds.
In an effort to define the factors accounting for variation in blood lead levels in populations, analysis of blood lead data obtained from a number of different studies have been compared. For these studies, we have estimated the individual-to-individual variation and the analytical error variance components. These variance component estimates are then used in our calculations which estimate the percentage of individuals who will falsely or truly exceed selected threshold blood lead levels. Our calculations show that for typical populations the percentage of "false positives", which are measured values that exceed the threshold value because of analytical error, is large.
variations in measured blood lead values
The total variation in measured blood lead level for a population contains several components -- true blbod lead variation between individuals, within-person variability (over time) for an individual and a number of analytical error components which includes both sampling and measurement error. For most of this paper a simple separation of the total variability into two components, an analytical error component and the remainder (called the individual-to-individual component), will suffice. The effects of the individual-to-individual variation and of the analytical error are schematically shown in Figure 1.
In Figure 1-A, a hypotheticalbut reasonable distribution of blood lead results for a population is shown. As blood lead measurements are approximately lognormally distributed,0'5' the drawings represent the distribution ofthe logarithm of blood lead values. As drawn, this distribution includes both the variation in the true individual-to-
Copyright 1981, American Industrial Hygiene Association
88
Am. Im). Hyg. Assoc. J (42)
February. 1981
TEH 0531870
N33781
A Total Variability in Measured Blood Leads
Figure 1 - Effects of variability on measured blood lead levels.
individual blood lead levels for the population group and the variation due to analytical error. In Figure 1-B, the hypothetical distribution which only includes true individual-to-individual variability in blood lead levels is pictured. In Figure 1-C, the hypothetical distribution which includes only analytical error is shown. This corresponds to the case where a single individual is sampled and analyzed repeatedly for lead. As shown in Figures 1-B and 1-C, almost none of the population is above the threshold, yet when the two effects are combined as pictured in Figure 1-A, "false positives" occur.
estimates of blood lead variability
The percentage contribution of the individual-to-individual variance and of the analytical variance to the total measured variance was determined for blood lead values from four major population studies. These studies were chosen as representative of the best recent studies, using yenous blood samples, with repeat blood samples from individuals. A need for more research to determine all components of blood lead variability is indicated as there have been few studies in which repeat blood samples have been taken on individuals and no studies have been carried out to determine all components of individual-to-individual variability, and the components of analyticalerror. The four selected studies included a multi-city survey,'7' a longitudinal study,'*' a personal air monitor study,'" and a Southern California rural-urban study.`10) Results of the analyses of variance on the studies axe summarized in Table I. The analysis of variance breakdown used to obtain Table I
American Industrial Hygiene Association JOURNAL
(42) 2/81
is described in detail in Appendix A. The analytical error was over 46 percent of the total measured blood lead variability in the four studies.
In the multi-city study, blood samples were obtained from more than 2000 suburban housewives in 10 communities across the United States. Duplicate samples were obtained from about 10% ofthe subjects. This study also measured air lead levels using stationary samplers, thus simulating how an air quality standard for lead would be implemented. In this study the value of the individual-to-individual variance was the smallest of all studies reported; it accounted for only
TABLE I Summary of Analysis of Variance for Blood Lead Values from Four Studies
Component Standard Deviation
Individual-toIndividual
Analytical Error
Multi-City Survey'71
Longitudinal Study1'"
Personal Air Monitor StudyTM
Southern California Study""'
0.052 (20%)* 0.063 (12%) 0.097 (54%)
0.0f7(21%)
0.103 (80%) 0.174 (88%) 0.089 (46%)
0.169 (79%)
*% Total Variance
[ ]Component Standard Deviation Total Standard Deviation
X100
All analyses are carried out on logm blood lead levels.
TEH 0531871
89
' /< - DUP050032367
Incorrectly Classified Below Threshold
Correctly Classified Above Threshold
CD
Correctly Classified Below Threshold
Incorrectly Classified Above Threshold
AB
m
Measured Blood Lead
y - True Blood Lead Threshold Limit
m - Measured Blood Lead Threshold Limit Figure 2.- True blood lead vs. measured blood lead.
20% of the total variance even through the method variance was not excessively large. For a homogeneous group of subjects, such as suburban housewives, the variability between individuals will be small.
The longitudinal study Was a five-year study of more than 2000 workers from 23 locations around the U nited States. Urine samples were obtained from all subjects and blood lead samples were obtained from more than 200 individuals. On 217 individuals blood samples were obtained in each of the Five years. No trends in blood lead were observed over the period of the study. Without duplicate blood samples in a year, long-term changes in an individual's blood lead cannot be separated from the analytical error. Thus, for this study the analytical error is inflated by long term withinperson variability.
The personal air monitor study was conducted to determine the relationship between air lead levels and blood lead levels with the air lead being measured using personal air monitors. Five groups of thirty subjects each were monitored for a 2-4 week period. Two to eight blood samples were obtained from each subject and each sample Was analyzed in duplicate. In three of the groups, two blood samples were taken each week of the study. With this sampling procedure week-to-Week variability for an individual can be separated from analytical error. The analytical error Component shown for this study does not contain this within-person variability. This study shows the smallest analytical error of all the studies, in fact, the analytical error estimated in this study is smaller than the measurement error shown in the Kettering Study14' [where the measurement error is one of the components of analytical error as described in Appendix A]. To be conservative all our calculations showing the effect of analytical error use the variance components from this study.
so
^ ::|>r|fSS-. ?
In the Southern California study blood samples from two populations, an urban group and a rural group, were obtained. The urban group lived in a housing development close to a freeway (and a stationary air sampler which monitored air lead levels) while the rural group lived in an area of low lead exposure. While the study appeared to be well designed and conducted, there were often large differences in measured blood lead level between samples taken from the same individual even though these blood samples were taken only one day apart. For this study the analytical error was large.
estimation of "false positives"
When blood lead levels are measured on a typical human population, there are often individuals whose measured values are above a threshold value (e.g., 40 jug/100 mL). Repeat samples taken on these individuals usually have blood lead measurements at a much lower level -- below the threshold value. This can occur because of analytical error or because of within-person variability (where an individual is above the threshold at the time of measurement but is below the threshold by the time of the follow-up sample). In this manuscript, we evaluate the effect of analytical error. Because of analytical error, a large percentage of the measurements that exceed a threshold are "false positives" in that the individual's true blood lead level is not above the threshold.
The occurrence of "false positives" is schematically illustrated by a bivariate plot shown in Figure 2 where the measured blood lead value is plotted on the X axis and the true blood lead value is plotted on the Y axis. The "measured threshold limit" and "true threshold limit" divide the distribution into four quadrants - "A", "B", "C", and "D".
"A" represents the measured values Which are correctly classified as being below the threshold value since they lie below the measured and true threshold limits. The values in "D" are correctly classified as being above the threshold limits. The values in "B" and "C" are incorrectly classified. The values in "B" lie above the measured threshold but below the true threshold limit. The values in quadrant "B" will be incorrectly Classified as being above the threshold; these are "false positives". In a similar manner, the values in "C" will be incorrectly classified as being below the threshold. They are "false negatives".
A numerical estimate of the percentage of"false positives" can be made by using a bivariate distribution model of blood lead values.11" A summary of the results from such a calculation is given in Table II. The specific results in Table II were derived using the variance components from the personal air monitor study. The results from this study are conservative in that they show fewer "false positives" than the other studies because the data from this study had the smallest analytical errors. Similar results could be calculated for any study (or for other responses besides blood lead) when the individual-to-individual variance component and the analytical variance component can be determined. The description of the procedure used to obtain
Am. Ind. Hyg. Assoc. J (42)
February. 1981
TEH 0531872
DUP050032368
TABLE II Estimation of Falsa and True Exceedences of Bipod Lead Threshold Values
Population Geometric Mean bag/1.00 mL)
15 20 25 30 35 40
Percent*' of values observed above a threshold when the threshold is
30 1.11 35 0.26 40 0.06
9,06 3.25 1.11
27.38 13.36 6.06
50,00 30.56 17.14
50.00 32.98
50.00
Percent" of values which are "false positives" when the threshold is
30 1.05 6.77 12.76 11.80
35 0.26 286
8.93 13.07 11.80
40 0.06
1,0$
4.91 10.41 13.19 11.80
Percent1 of values actually above the threshold when the threshold is
30 35 40
0.10 0.01 0.00
3.50 0,62 0,10
20.76 6.63 1.78
50.00 24.54
9.92
50.00 27.63
50.00
'Percent in Categories B and D in Figure 2. "Percent in Category B in Figure 2. ` Percent in Categories C and D in Figure 2.
the results in Table II using the logarithm of the blood lead values is given in Appendix B.
Table II shows that the analytical error will ajmostalways cause some "false positives" when measuring blood lead levels in a population. When the population mean blood lead level is well below the threshold value, most of the values measured over the threshold will be falsely categorized. For example, in.a population with a true mean blood lead level of 25 /eg/100 mL (the upper end ofthe range for the average blood lead level in typical populations01), and a threshold value of 40 /tig/100 mL, Table II shows that 6.06 percent ofthe individuals will be measured to be above the threshold. However, 4.91 percent of the population (8.1 percent of these individuals) will be "false positives" as their true blood lead is below 40 ngj 100 mL.
The percentage of "false positives" increases as the population mean decreases for a given threshold value, although the absolute number of individuals classified above the threshold decreases. If a group has its mean near the threshold limit, an opposite problem occurs. In this group there can be many individuals whose measured blood lead is below the threshold limit while their true blood lead level is above the threshold. These would be "false negatives",
example of "false positives" To illustrate how analytical error can account for the difference in the number of individuals above a threshold between the original sample and the follow-up sample, we will use the data from a recent survey in Pittsburgh that measured blood lead levels in children.00 In this study a representative sample of households from the metropolitan Pittsburgh area was surveyed for levels of lead in paint and dust in homes. Following this survey, blood samples were requested from children in households having childrenaged 7 or younger. While not all households complied with the request, the complying households had higher lead content
in dust and paint than the noncomplying households. Four hundred fifty-six children were sampled. The average of the 456 blood lead values was 21,9 (median 21.0) and the standard deviation was 7.9. Fifteen children (3%) showed blood lead values above 40 on the first sample, A resample was requested for these children. Fourteen of these 15 children were found for a resample. Most children were
TABLE III Blood Lead Values for 15 Out of 456 Children in Pittsburgh Study Whose Initial Blood Lead Values
Were Greater Than 40 pg/1O0 mL
Child
Measured Blood Pb (pg/100 mL)
First Sample
Second Sample
Third Sample
Predicted Blood Pb ' (ug/100 mL)
1 2" 3 4 5 6 7 8 9 10 11 12" 13 14" 15 RMS'
45 73 42 43 45 41 46 42 47 59 42 43 40 60 41 16.1
26 73 76 35 27 38 28 34 38 35 28 19 Could not be located. 21 48 43 32
32 41 31 31 32 30 32
31 33 37 31 31 30 34 30
6,3
31 39 30 30
31 29 31 30
31 35 30 30 29 32 29
6.1
*First column - Regression Caused by analytical variability; Second column - Regression caused by analytical variability and withinperson variability.
"Not used in RMS calculation.
CRMS Root Mean Square Difference from second sample.
American Industria1 Hygiene Association JOURNAL
(42) 2/SI
TEH 0531873
DUP050032369
resampled within one month of the original sample. There was no attempt to change exposures between the original sample and the resample. In the resample, 2 of the 15 children remained above 40 while 12 dropped below. These results are shown in Table III. The very high value was obtained from an eight-year-old being treated for lead poisoning.
Using the regression equation shown in Appendix C, we can predict the blood lead levels for individuals in the Pittsburgh study whose initial blood lead was above40. The regression has two causes; the first is analytical error, and the second is within-person variability (as the blood lead level could be above the threshold at the time ofthe original measurement but below the threshold by the time a resample is taken). In Table I'll, we show the results of two regression calculations. One considers only the regression effect due to analytical error while the second gets slightly better predictions as it predicts the combined regression caused by both effects. The blood lead values predicted using the regression equations are much closer to the resampled blood lead values than are the initial sample values. The good agreement between the predicted and the resample blood lead levels shows the validity of the calculation of "false positive" percentages (Table II).
In summary, we have shown that the percentage of blood lead Values measured to be above a threshold will usually overestimate the percentage actually above the threshold. For typical populations, the percentage of "false positives" will be substantial. A numerical estimate of the extent of "false positives" has been given along with a predictor that is a better estimate of a confirmatory blood lead reading than is the original reading.
The existence of false positives depends only on the fact that the total observed variability is bigger than variability in the true values due to the presence of analytical error. The Analysis of Variance model used here is more general than the classification models used elsewhere05* as it shows how the classification probabilities change with changes in population mean level and with the threshold level. Thus, the approach illustrated here is appropriate for any quantitative variable.
acknowledgement
I would like to thank R.D. Snee for providing the original data for the personal air monitor study, for performing some of the analysis of variances described and for many helpful discussions.
references
1. Office of Research and Development: Air Quality Criteria (or Lead, EPA-600/8-77-017. U.S. Environmental Protection Agency, Washington, D.C. (1977).
2. Zielhuis. R.L.: Biological Quality Guide for Inorganic Lead. int. Arch. Arbeitsmed. 32:103-127 (1974).
3. Commission of the European Communities: Council Directive of March 19, 1977, on Biological Screening of the Population for Lead, Official Journal of the European Communities, Voi. 20, No. 4/05:10-17, Brussels (1977).
4. Lamer, S.: Blood Lead Analysis -- Precision and Stability. J. Occup, Med. 77:153-154 (1975).
5. Billick, I.H., A.S. Curran and D.R. Shier; Analysis of Pediatric Blood Lead Levels in New York City for 1970-1976. Environ. Health Perspect. 37:133-190 (1979).
6. Pierce, J.O., S.R. Koirtyohann, T.E. Clevenber and F.E. Lichte: the Determination of Lead in Blood. International Lead Zinc Research Organization, Inc., New York (1976).
7. Tepper, L.B. and L.S. Levin: A Survey of Air and Population Lead Levels in Selected American Communities. Environmental Quality and Safety, Supplement Voi. it -- Lead. pp. 152-196. George Thieme Verlag. Stuttgart, W. Germany (1975).
8. McLaughlin, M,, A.L. Linch and R.D. Snee: Longitudinal Studies of Lead Levels in a U.S. Population. Arch. Environ. Health 26:131-136 (1973).
9. Azar, A., K. Habib! and R.D. Snee: An Epidemiologic Approach to Community Air Pb Exposure Using Personal Air Samplers. Environmental Quality and Safety, Supplement Voi. it--Lead, pp, 254-270. George Thieme Verlag, Stuttgart, W. Germany (1975).
10. Johnson, D.E., B.J. Prevost and J.B, Tillery: Baseline Levels of Platinum and Palladium in Human tissues. EPA600/1-76-019, Environmental Health Effects Research Science Document, U.S. Environmental Protection Agency, Washington, D.C. (1976).
11. National Bureau of Standards: tables of the Bivariate Norma!Distribution Function and Related Functions. Applied Mathematics Series No. 50, U.S, Government Printing Office, Washington, D.C. (1959).
12. Urban, W.D.: Statistical Analyses of Biood Lead Levels of Children Surveyed in Pittsburgh. PA: AnalyticalMethodology and Summary Results. NBSIR 76-1024, Institute of Applied Technology, National Bureau of Standards. Washington, D.C. (1976).
13. Colton, T.: Statistics in Medicine. Little Brown and. Co., Boston, Mass. (1974).
14. Seheffe, ti.: The Analysis of Variance. John Wileyand Sons, New York (1959).
15. Snee, R.O. and P.E. Smith, Jr.: Statistical Analysis of Interlaboratory Studies. Am. Ind. Hyg. Assoc. J. 33:784-790 (1972).
16. Keppler, J.F., M.E. Maxfield, W.E. Moss, G. Tietjen and A.L. Linch: Interlaboratory Evaluation of the Reliability of Blood Pb Analyses. Am. Ind. Hyg. Assoc. J. 37:412-429 (1970).
APPENDIX A
Estimates of the total variance and the components of the total variance for blood lead values were obtained for four studies using a standard hierarchical analysis ofvariance,1"" As the distribution of blood lead value is approximately log normal,0'5' all results are developed on a lagio scale. The
92
variance components for theSe studies are discussed in this appendix.
Consider a study where:
g = number of groups
Ant. Ind. Hyg. Assoc. J (42)
TEH
February, 1981
0531874
''msBB/KKm DUP050032370
Source
TABU A-1 Analysis of Variance
Degrees of Number Freedom
Groups Person-tQ-Person Within-Person Sampling Measurement
9 a~i
rt g(n-1)
t gn(t-1) s gnt(s-l)
m gnts(m-1)
Variance Component
Ogroups Operson OwMhhvperson ffiampUng ^measurement
CT.linalytiest
*
^sampling +
^measurement
m
n = number of individuals (per group) t -- number of sampling times (per individual) r = number of samples (per sampling time) d =? number of determinations (per sample) a2 -- variance component
The data from such a study could be analyzed using the analysis of variance, shown in Table A-1. Even more extensive breakdowns could be considered. For example, the measurement variance component could be further broken down by considering the component parts of the measurement error, such as the operator-to-operator error and the instrument-to-instrument error. Data to make this more comprehensive breakdown is seldom available.
In the analysis of variances shown in Table A-1, the analytical error contains two components, a sampling component and a measurement component as shown by the following formula:
^analytical " Campling *1" Cmeasuicment/m
In most screening studies only a single measurement is made on each sample (m=l). When this occurs, the sampling variance component and the measurement variance
component cannot be estimated separately; only the total analytical error can be estimated. Our estimate of the effect of analytical error will assume that a single measurement is made on each sample, so for our calculations
0n,ly!ic! " Oumpling ^measurement
Generally,- enough samples are not taken to estimate all the variance components shown in Table A-1. For example, in a study in which each individual is sampled on two occasions and each sample is measured once, only two variance components can be estimated. These two components are:
1) <7parson 4" (k) OwElhin-rpenron
2) the S.Um of (1 k) 4" 4"Owilhin-pcoon Oninpltag Oatoaurement*
In this formula, the k value which multiplies the withinperson variance component is a number between 0 and 1. if the two samples from an individual are taken close together, k will approach one; however, if there is a long time between samples then k will approach zero.
In this latter case, if the second component is interpreted as analytical error, it will be inflated as it contains much of the within-person variability.
When only a single sample is taken from an individual, but duplicate determinations are made on the blood sample, crLuuremetu can be estimated. However, the total analytical error cannot be estimated as no estimate of sampling variability can he obtained, in this case the measurement error may be (falsely) reported as the total analytical error. This causes the precision of the study to be overestimated. To estimate the analytical variability, more than one sample must be taken from an individual. Table A-2 summarizes analyses of variance for blood lead values from four studies for which multiple samples were taken from individuals. In the longitudinal study'8' the repeat samples from an individual were taken a year apart. For this study, the within-person variability is combined with the analytical error causing an overestimate of the analytical variability.
TABLE A-2 Analysis of Variance for Blood Lead Values from Four Studies
Study
Component Standard Deviation''
Individual-toIndividual
Analytical Error
Total
Geometric Standard Deviation
Multi-City Survey"'
Longitudinal Study1*'
Personal Air Monitor Study1'1
Southern California Study1"1'
0.0521 (20)"
0.063 (12)
0.0972 (67) 0.09721 (54)
0.0871 (21) 0.0784" (18)
0.1030(80)
0.174 (SB)
0.0688 (33) 0.0889 (46)
0.1689(79) 0.1662(82)
01164
0.185
0.1191 0-1317
0.1900 0.1838
1.30
1.53
1.32 1.35
1.55 1.53
'Component Standard Deviation = (Variance Component)13 'Values in parentheses are % Total Variance ^
Component Standard Deviation ~|3 x 100
[ JTotal Standard Deviation
'Modified oloaiyticii obtained from duplicates (see Equation (1)). "Deleted six outlier blood lead values.
American Industrial Hygiene Association JOURNAL
(42) 2/81
TEH 0531875
93
DUP050032371
TABLE A-3 Personal Air IVIonitor Study Analysis of Variance Table
Source
Variance Degrees of Variance Component Freedom Component Estimates
Group Person Sampling Time Analytical
4 145 195 194
0.00893 0.00854 0.00090 0.00474
(7group Cfpersan <7withuv person
(7sampling + <7measuremnl/2
When single readings are made
Gtanalydcal <7sampling + Omeasuremenl -- (0.0889)' (7individual -- <7person + <7wKhrn-perfon = 0.00944 = (0.0972)2
Qtai^l = (Jimt.iduy! 1 Oanslyttol = (0.0972)" "T (0-0889)" = (Q.1317)
For the other three studies described, duplicate samples were taken much closer together (so k -- I). For example, in the Southern California study"01 which had the largest percentage analytical error, duplicate samples were taken one day apart.
In the four studies summarised in Table A-2, a component estimating group-to-group differences was also obtained as each of the studies was made on more than one group. This component which accounts for differences between groups was not reported as part of the individualto-individual variance.
In the personal air monitor studies,1" two blood samples were taken from most individuals during each week of the study. Thus, for this study it was possible to separate the week-to-week within-person variability from the analytical error. Table A-3 shows the detailed analysis of variance for the personal air monitor study.
In the personal air monitor study, duplicate determinations were carried out on each blood sample. Only the average of duplicate determinations was used in the analysis since the values of the individual determinations were not available. In the analysis, therefore, only the sum of oumpiint an(l Omusvroncnt/2 could be estimated. The following equation was used to estimate the analytical error which would be obtained if only single determinations were used:
Uuitlyltea) {.ingle determination) " Uanalytlce! (duplicate .determination) 7" Umcaauivinent/ 2 0)
For this study:
^analytical i,in*le determination! = 0.00474 4" (0.0795) / 2 -- (0.0889)"
The value of (0.0795)2 for aLaaurnt was obtained from an AIHA study."5'161 This modification was used to make the personal air monitor study consistent with the other studies which used only single determinations.
The analytical variability for the personal air-monitoring study is the smallest of all four studies considered. In fact, the analytical error for the personal air-monitor study is smaller than the sampling error found in the Kettering Lab study.*" Had the estimate offalse positives used the variance components from any of the other studies, the percentage of false positives would have been greater.
APPENDIX B
calculation of bivariate lognormal distribution probabilities The relationship between measured blood lead and true blood lead values can be described by a bivariate lognormal distribution (a bivariate normal distribution using log blood lead values).
Five parameters are required to describe this distribution:
Ut = the mean of the true blood lead values Um the mean of the measured blood lead values <j\ = the true variance in blood lead values aTM = the variance in measured blood lead values
p = the correlation between measured and true blood lead values
The means of the true and the measured blood lead values are considered to be the same; they art equal to the population mean. A bias in the measurement could easily be accounted for by adjusting the mean levels. The true blood lead variance ol is estimated toy o^m.n + oLhin-pcrTM which for the personal air monitor study is (0.0972)2. The variance in measured blood lead levels (oiL) is estimated by a?oui which is (0.1317)2. The correlation (p) is:
94
-- Covariance (XY) p [Variance (X) Variance (Y)], ;
- of _ g| P [ai X aT]i : Pm
Thus our estimate of p is
^ 0.0972 p 0.1317
0.738
The probabilities in Tables II and B-l were obtained by first calculating Zi and Zj where:
_ _ log (true threshold value) - log ut fh
^ log (measured threshold value) -- log um
<7m
Note that the measured threshold value and the true threshold value need not be equal.
Values of Zi and Z2 were used with normal probability tables to determine the probabilities in quadrants A + C, B + D, A + B, and C + D. (These refer to quadrants shown in Figure 2.) The probability in quadrant D Was obtained from tables of the bivariate normal distribution"11 using Zi, Z2
4m. Ind. Hyg. Assoc. J (42)
February. 1981
TEH 0031876
DUP050032372
Population .Geometric Mean Blood
Lead (Mg/too mL)
15 20 25 30
15 20 25 30 .35
15 20 25 30 35
15 20 25 30 35 40
15 20 25 30 35 40
15 20 25 30 35 40
TABLE B-1 Estimation of False and True Exceedences of Blood Lead Threshold Values* **
Threshold for Blood Lead (mIj /100 mL)
Meas. True
Bivariate Distribution * (percent)
A CD
% Values Observed
Above Measured
Threshold, (B+D)
% Positives Which Are
False
Positives, 1005 (B + O)
% True Exceedences of Threshold,
(C + D)
30
30 98.85
1.05 0.04
0.06
89.73
6:77 1.21
2.29
66.49 12.76 6.13 14.62
38.20 11.80 11.80 38,20
1.11 9.06
27.38 50.00
30 35 98.89 1.11 0.00 0.01
1.11
90.86 8.53 0.09 0.53
9.06
71.79 21.59 0.83 5.79
27.38
47,42 2804 2.58 21.96
50.00
26.43 23.58 4-13 45.87
69.44
35 35 99.74 026 0-00 0.00
0.26
96.54
2.88 0.23
0.39
3.25
84,44 893 2.21
4.42
13.36
62.39 13.07 7.05 17.49
30.56
38.20 11.80 11.80 38.20
50,00
30
40 98.89
1,11 0.00 0.00
1.11
90.94 8.96 0.00 0.09
9.06
72.54 25.68 0.08
1.71
27.38
49.63 40.44 0.37
9.56
50.00
29.72 42.74 0.84 26.70
69,44
15.94 34.04 1.19 48.82
82.86
35 40 99.74 0.26 0.00 0.00
0.26
96.74 3.17 0.02 0.08
3.25
.86.36 11.86 0.28
1.50
13.36
68.04 22.04 1.40 8.52
30.56
46.70 25.77 3.30 24.23
50.00
28.12 21.89 4-86 45.13
67.02
40 40 99.94 0,06 0.00 0.00
0.06
98.85
1.05 0.04
0.06
1.11
93.31
4.91 0.64
1.15
6.06
79.67 10.41
320
6.73
17.14
59.27 13.19 7.74 19.79
32.98
38.20 11.80 11.80 38.20
50.00
95 75 47 24
100 94 79 56 34
100 88 67 43 24
100 99 94 81 62 41
100 98 89 72 52 33
100 95 81 61 40 24
o.io 3.50 20.76 50.00
0.01 0.62 6.62 24.53 50.00
0.01 0.62 6.63 24.54 50.00
0.00 0.10 1.78 9.92 27.54 50.00
0.00 0.10 1.78 9.92 27.52 50.00
0.00 0.10 1.78 9,92 27.53 50.00
*A - Values correctly classified as below threshold. 6 - Values .incorrectly classified as above threshold. C - Values incorrectly classified as below threshold. D - Values correctly classified as above threshold. ** - % False Positives for the Measured Threshold.
and p. (Actually, a computer program that followed the above procedure was used, as interpolation would be required in the tables.) Once the probability in quadrant D is known, the probability in all quadrants can be calculated.
In Table B-1, the four classification probabilities of the
bivariate distribution are presented for different population mean and assumed threshold blood lead values. In addition, the last three columns present the percent ofvalues observed above a threshold, (B + D), the percent positives which are "false positives", (B/(B -I- D) X 100), and the percent exceedences of the threshold, (C + D).
APPENDIX C
prediction of blood load values for repeat blood samples
The value of a follow-ttp blood lead sample can be predicted using the following regression equation which estimates the statistical phenomena of "regression toward the mean".
American Industrial Hygiene tecciafion JOURNAL
(42)2/81
Individuals whose measured blood lead levels are above a threshold will either be those who have positive analytical errors or those who have high true blood lead levels at the time measurements are taken. When only these individuals are resampled, the resample will tend to be lower because
,, _______
TEH 0531877
DUP050032373
positive analytical errors will occ ur less often or because the individual's blood lead may have dropped in the time between the original sample and the resample. Using the estimated variance components from the personal air monitor study we will estimate the regression effect caused by analytical error and the regression effect caused by the combined effect of analytical error and within-person variability. The regression calculation enables one to make a good prediction of the resample blood lead value.
The regression equation has the form log Y = log u + (log X - log u)
(1)
where: Y = the predicted blood lead value for the repeat blood sample
X = the original blood lead reading
u = the geometric mean of the population (assumed to be 21 for the Pittsburgh study)
d = regression coefficient
For the personal air monitor study the regression coefficient is:
a -- 5es k b. -- *1.00854 _ 04o3
P oL,,
0.01734
for the regression equation that includes both the effect of analytical error and the effect of within-person variability and
- _ gp.*ron 4- .Pwithin-pcmon _ 0:00944 _ q 544
P~
oiu.1
~ 0.01734
for the regression equation that includes only the effect of analytical error
oLa = the total variability (for single samples) -- (0-1317)2
uiLson -- person-to-person variability = 0.00854 ~ the within-person variability == 0.00090
The variance component values from the personal air monitor study that were used as the analytical error could not be estimated from the Pittsburgh data as there were no duplicate samples.
9$
j
Am. Ini. Hyt. Assoc. J142)
February, 1981
TEH 0531878
38NMMI DUP050032374