hilum

Theory · The statistics of the test report · Agreement

The tool passes when it disagrees like a second analyst.

A counting tool is not validated by accuracy against a label. It is validated by reading the same material as trained analysts, without seeing their answers, and showing that it disagrees with them as much as they disagree with each other. This page gives the calculations of that study: Bland-Altman, Lin's coefficient, κ, the ICC, the non-inferiority criterion and how many plates it requires.

Go deeper · mathematics and references

The design: the same material, read blind

The agreement study has four parts. The material is plates that the tool never saw in training, covering the range of viability found in routine work, from about 10% to about 95%, and the high densities that are hard to count, with touching seeds. The reference is at least three analysts, each reading without seeing the other analysts' readings or the tool's. The individual readings are kept, because they are what measures how much trained people disagree.

There are three units of measurement: the count per image, the percentage of viable seeds per replicate (the number that goes into the test report) and the class of each seed. The tool's mode is also declared. In automatic mode nobody checks anything; in assisted mode the machine proposes and the person corrects. The two answer different questions, and the study needs both: the assisted mode is the one in use, and the automatic mode shows what the checking corrects.

The question is not whether the tool gets a true value right, which nobody knows. It is whether, if it takes the place of one of the analysts, it increases the disagreement. The sections that follow give the calculations, and the pair generator lets you vary each of them.

Bland and Altman: the difference against the mean

For each plate \(i\) there are two readings, \(y_i\) from the tool and \(x_i\) from the analyst. Bland and Altman [1] proposed looking at the difference against the mean of the two:

\[ d_i = y_i - x_i, \qquad m_i = \frac{x_i + y_i}{2}. \](1)

Definition

The bias is \(\bar d\), the mean of the differences. The 95% limits of agreement are \(\bar d \pm 1{,}96\, s_d\), with \(s_d\) the standard deviation of the differences. If the differences are approximately normal and do not depend on the size of the count, 95% of new plates will have a difference that lies between the two limits.

Why against the mean, and not against the analyst's reading? Write \(x = T + e_x\) and \(y = T + e_y\), with \(T\) the true value and errors \(e_x\) and \(e_y\) independent of \(T\) and of each other. Then

\[ \mathrm{Cov}(y - x,\; x) = -\mathrm{Var}(e_x), \qquad \mathrm{Cov}\Big(y - x,\; \frac{x+y}{2}\Big) = \frac{\mathrm{Var}(e_y) - \mathrm{Var}(e_x)}{2}. \](2)

Against \(x\), the difference always has a downward trend, even with no defect at all in the tool. Against the mean, the trend vanishes when both methods have the same error variance, and appears only when one of the methods is truly noisier.

The bias and the limits are estimates and have confidence intervals. The interval for the bias is \(\bar d \pm t\, s_d/\sqrt{n}\), with Student's \(t\) and \(n-1\) degrees of freedom. The interval for each limit uses the approximate variance of Bland and Altman [1], revisited in the 1999 paper [2]:

\[ \mathrm{Var}(\bar d \pm 1{,}96\, s_d) \approx s_d^2\left(\frac{1}{n} + \frac{1{,}96^2}{2(n-1)}\right) \approx \frac{3\, s_d^2}{n} \](3)

The first term comes from \(\bar d\), and the second from \(s_d\), whose variance is about \(s_d^2/(2(n-1))\) for normal data. With \(n\) plates, each limit is known to within \(\pm t\sqrt{3/n}\, s_d\). In a simulation with normal differences, this interval contains the true limit 94.2% of the time with 30 plates and 94.7% with 100. This calculation determines the size of the study (see the How many plates section).

When the spread grows with the count

In counting, the error is rarely the same on a plate of 50 seeds and on one of 500. Dense plates have more touching seeds, and each pair that is not separated removes one seed from the count. On the Bland-Altman plot this shows up as a funnel: the cloud of differences fans out as the mean increases, and limits computed with a single \(s_d\) are too wide on small plates and too narrow on large ones.

The simplest way out is to replace the difference with the ratio, taking logarithms before subtracting. With \(d_i = \ln y_i - \ln x_i = \ln(y_i/x_i)\), the bias and the limits are computed as before and transformed back to the original scale by exponentiation [2]:

\[ \exp\big(\bar d \pm 1{,}96\, s_d\big) \](4)

The result reads as a ratio: in 95% of plates, the tool counts between, say, 0.96 and 1.02 times what the analyst counts. If the error is proportional to the count, the cloud on the log scale has constant width, and a single pair of limits applies to all plates. When the funnel is not proportional, for example an error that grows with the square root of the count, as in a Poisson count, Bland and Altman [2] also show how to model the spread of the differences as a function of the mean. In the pair generator, the count-dependence control creates the funnel, and the ratio scale removes it when the error is proportional.

Lin: agreement with the line y = x

Pearson's correlation coefficient measures how well the points line up on any line. A tool that always counts 10% more than the analyst has \(r = 1\) yet is wrong by 10% on every plate. And \(r\) grows with the range of the sample: with the same error, plates of 20 to 800 seeds give a larger \(r\) than plates of 180 to 220. For this reason Bland and Altman [1] rejected correlation as a measure of agreement.

Lin [3] measured the distance to the line \(y = x\), and not to just any line. The mean squared deviation between the readings decomposes into

\[ \mathrm{E}\big[(y-x)^2\big] = (\mu_y - \mu_x)^2 + \sigma_x^2 + \sigma_y^2 - 2\sigma_{xy}, \](5)

and the concordance correlation coefficient (CCC) compares this deviation with the value expected if \(x\) and \(y\) were independent, with the same means and variances.

Result: the CCC and its decomposition

\[ \rho_c = 1 - \frac{\mathrm{E}[(y-x)^2]}{(\mu_y-\mu_x)^2+\sigma_x^2+\sigma_y^2} = \frac{2\sigma_{xy}}{\sigma_x^2+\sigma_y^2+(\mu_y-\mu_x)^2} = \rho\, C_b \]
\[ C_b = \frac{2}{v + 1/v + u^2}, \qquad v = \frac{\sigma_x}{\sigma_y}, \qquad u = \frac{\mu_y - \mu_x}{\sqrt{\sigma_x\sigma_y}} \]

Proof. The first equality is the definition, and the second follows from (5). For the third, write \(\sigma_{xy} = \rho\,\sigma_x\sigma_y\) and divide numerator and denominator by \(\sigma_x\sigma_y\).

\(\rho\) is the precision, how much the points spread around the best-fit line, and \(C_b\), between 0 and 1, is the accuracy, how far that line departs from \(y = x\) by shift (\(u\)) or by scale (\(v\)). The CCC reaches 1 only if the points fall on \(y = x\). Lin estimated means, variances and covariance with denominator \(n\) and worked on the inverse hyperbolic tangent scale (Fisher's \(z\) transformation) for the interval. A bootstrap resampling of the plates gives an equivalent interval without a formula, and that is what the generator uses.

The CCC inherits a defect of \(r\): it also grows with the range of the counts. It complements the Bland-Altman limits and does not replace them.

κ: the class of each seed

In the tetrazolium reading, the question per seed is one of class: viable, non-viable or empty. The observed agreement \(p_o\), the fraction of seeds on which the two readings coincide, is misleading because part of it comes from chance. If 80% of the seeds are viable, two readers who guess "viable" at that proportion, without looking, agree on 68% of the seeds (\(0{,}8^2 + 0{,}2^2\)). Cohen [4] corrected for that part:

\[ \kappa = \frac{p_o - p_e}{1 - p_e}, \qquad p_e = \sum_k p_{k\cdot}\, p_{\cdot k}, \](6)

with \(p_{k\cdot}\) and \(p_{\cdot k}\) the fractions of class \(k\) in each reading. \(p_e\) is the agreement expected if the two readings were independent, with the observed marginals. \(\kappa = 1\) is perfect agreement and \(\kappa = 0\) is chance.

Hypothetical example for the calculation: 400 seeds read by the tool and by the consensus of the analysts.
consensus: viableconsensus: non-viabletotal
tool: viable30018318
tool: non-viable127082
total31288400

Here \(p_o = 370/400 = 0{,}925\), \(p_e = (318 \cdot 312 + 82 \cdot 88)/400^2 = 0{,}665\) and \(\kappa = 0{,}776\). The same table gives the sensitivity for non-viable seeds, \(s = 70/88 = 0{,}795\), and the specificity, \(e = 300/312 = 0{,}962\). These are the \(s\) and the \(e\) of the Rogan-Gladen correction, which is used to correct the tool's count.

With \(m\) analysts reading the same \(N\) seeds, Fleiss [5] generalized the calculation. For seed \(i\), with \(n_{ij}\) analysts placing it in class \(j\):

\[ P_i = \frac{\sum_j n_{ij}^2 - m}{m(m-1)}, \qquad p_j = \frac{1}{Nm}\sum_i n_{ij}, \qquad \kappa = \frac{\bar P - \sum_j p_j^2}{1 - \sum_j p_j^2}, \](7)

with \(\bar P\) the mean of the \(P_i\). \(P_i\) is the fraction of analyst pairs that agree on seed \(i\), and chance uses the marginals of all analysts together.

Two cautions. First, \(\kappa\) depends on the prevalence of the class. The same classifier, with \(s = 0{,}80\) and \(e = 0{,}98\), has \(\kappa = 0{,}82\) in a seed lot with 30% non-viable seeds and \(\kappa = 0{,}56\) in a lot with 2%, even though the raw agreement rises from 92.6% to 97.6%. Feinstein and Cicchetti [6] described this paradox, and for that reason \(\kappa\) is always reported alongside the confusion matrix and the sensitivity and specificity per class, with the non-viable class and the faint-staining boundary reported separately. The verbal bands of Landis and Koch [7], such as "substantial" for 0.61 to 0.80, are a convention of the authors, not a test.

Second, the seeds of one plate are not independent: a problem with lighting or focus affects all of them. The interval for \(\kappa\) must resample plates, not seeds, and thousands of seeds from a few plates count as only a few plates, by the same design effect as on the proportions page.

ICC for absolute agreement

For a continuous measurement read by \(k\) readers on the same \(n\) plates, the two-way model separates the plate, the reader and the remainder:

\[ x_{ij} = \mu + a_i + b_j + e_{ij}, \qquad \mathrm{Var}(a_i) = \sigma_a^2,\;\; \mathrm{Var}(b_j) = \sigma_b^2,\;\; \mathrm{Var}(e_{ij}) = \sigma_e^2. \](8)

\(a_i\) is the true difference between plates, \(b_j\) the systematic deviation of reader \(j\) (the tool is one of the readers) and \(e_{ij}\) the error. The ICC for absolute agreement is the fraction of the total variance that comes from the plates:

\[ \mathrm{ICC}(A,1) = \frac{\sigma_a^2}{\sigma_a^2 + \sigma_b^2 + \sigma_e^2}, \qquad \widehat{\mathrm{ICC}} = \frac{MS_R - MS_E}{MS_R + (k-1)MS_E + \frac{k}{n}(MS_C - MS_E)}. \](9)

This is the ICC(2,1) of Shrout and Fleiss [8], or ICC(A,1) in the notation of McGraw and Wong [9]. \(MS_R\) is the mean square of the plates, \(MS_C\) that of the readers and \(MS_E\) the residual, all from the two-way analysis of variance without replication. The consistency version, ICC(3,1) or ICC(C,1), removes the reader term from the denominator and is blind to a constant bias; for comparing methods, it is the wrong choice. There are ten forms of ICC, and published work must state which one was used [10].

Like the CCC, the ICC is a ratio of variances and grows with the variation among plates, for the same error. With two readers, the two coefficients are numerically very close, and the generator shows both side by side.

The criterion: no worse than a second analyst

If two trained analysts may disagree within a tolerance, the yardstick for the tool is that disagreement, not a single label. The non-inferiority criterion fixes a margin \(\delta\) before data collection and requires that the disagreement between tool and analyst exceed the disagreement between analysts by no more than \(\delta\).

One way to quantify the criterion uses the two differences on the same plates: \(D_{FA}\), tool minus analyst A, and \(D_{AA}\), analyst B minus analyst A. For each, the agreement range \(\mathcal{A} = |\bar d| + 1{,}96\, s_d\) is the distance between zero and the limit farther from it: on the worse side, only 2.5% of plates go beyond that distance. The ratio

\[ \theta = \frac{\mathcal{A}_{FA}}{\mathcal{A}_{AA}} \](10)

equals 1 when the tool, standing in for analyst B, disagrees with A exactly as B did.

Criterion

The tool is non-inferior to a second analyst if the upper limit of the 90% confidence interval for \(\theta\) (one-sided 95%) lies below \(1 + \delta\). It is inferior if the lower limit exceeds \(1 + \delta\). Between the two, the study is inconclusive: the answer calls for more plates.

The two differences share analyst A and are not independent. The interval for \(\theta\) comes from resampling the plates: plates are drawn with replacement and the two ranges are recomputed together, preserving the pairing. With three or more analysts, the calculation is repeated for each pair. The margin \(\delta\) has to come from outside the data, from what the standard tolerates between replicates or from what the laboratory considers a difference of no practical consequence. If it is fixed after the results are seen, it decides the result.

Instrument · pair generator, tool against analyst Simulation

−0,5%
1,25%
1,5%
0,5

The noise values are the standard deviation, as a percentage of the count, on a typical plate of 155 seeds. With dependence 0 the error has the same size in seeds on every plate; with 1 it grows in proportion to the count; 0.5 is the middle ground, as in a Poisson count.

plates in the study
margin δ
scale

–

Bland-Altman: tool minus analyst A

platesbiaslimits, with 95% CIlimits between analysts

Tool against analyst A dashed: y = x

Each plate has a true count drawn between 40 and 600 seeds. Analysts A and B make unbiased errors, with the chosen noise; the tool makes errors with the chosen bias and noise. The count dependence applies to all of them. The interval for \(\theta\) and that for the CCC come from 1,000 resamplings of the plates. Note that CCC, ICC and \(r\) stay near 1 under any reasonable adjustment of the controls, because the range of counts is wide; the limits say how many seeds separate the readings.

How many plates: the precision of the limits

By equation (3), the half-width of the 95% interval of each limit of agreement, in units of \(s_d\), depends only on \(n\): it equals \(t_{0{,}975;\,n-1}\sqrt{1/n + 1{,}96^2/(2(n-1))}\). For the bias, it is \(t/\sqrt{n}\). The table recomputes both for the most common study sizes, together with the short form \(t\sqrt{3/n}\).

Half-width of the 95% interval, in standard deviations of the differences (\(s_d\)).
plates, n\(t_{0{,}975;\,n-1}\)each limitshort form \(t\sqrt{3/n}\)bias
302,045±0,645±0,647±0,373
502,010±0,489±0,492±0,284
1001,984±0,340±0,344±0,198
2001,972±0,239±0,242±0,139
Exact calculationFigure 1. Half-width of the 95% interval of a limit of agreement and of the bias, in standard deviations of the differences, against the number of plates, on a log scale. The points mark 30, 50, 100 and 200 plates; the ring marks the number chosen in the generator.

With 100 plates, each limit is known to within about one third of \(s_d\); with 30, to almost two thirds. Going from 100 to 200 plates removes only a tenth of \(s_d\). The gain falls with the square root of \(n\), and the size of the study comes from the acceptable width for the limits, not from habit. Lu and colleagues [11] give a more complete calculation, which also fixes the probability that the interval of the limits falls within a maximum difference defined before the study.

For \(\kappa\) per seed, the number of seeds easily reaches the thousands. The real constraint is the number of plates and of independent sessions.

Repeatability and reproducibility among operators

Agreement with the analyst answers whether the tool reads like people do. Repeatability and reproducibility answer something else: how much the reading changes when the measurement is repeated. The typical design has several plates, three or more operators and two sessions per operator, days apart, in random order and without access to the earlier reading. The mixed model

\[ x_{ijk} = \mu + P_i + O_j + (PO)_{ij} + e_{ijk}, \](11)

with plate \(P_i\), operator \(O_j\), the plate × operator interaction and the session error \(e_{ijk}\), all random, separates two variances. The repeatability variance, \(\sigma_r^2 = \sigma_e^2\), is the variation of the same operator reading the same plate again. The reproducibility variance, \(\sigma_R^2 = \sigma_O^2 + \sigma_{PO}^2 + \sigma_e^2\), adds the variation between operators. The convention of ISO 5725 [12] turns each one into a limit: the difference between two independent readings falls below \(1{,}96\sqrt{2}\,\sigma \approx 2{,}8\,\sigma\) in 95% of cases, because the difference between two readings has standard deviation \(\sqrt{2}\,\sigma\).

In assisted mode the operator is part of the method: the same plate checked by two people can give different counts, and that difference is the reproducibility of the assisted method. It is also useful to measure reacquisition, with the same plate repositioned and scanned again three times: it measures the variation of the image, which the operator does not control. The same design applied to the manual reference method gives the variances against which the tool's variances are compared.

What the standard asks for

The ISTA method validation document [13] asks for repeatability within the laboratory, reproducibility between laboratories, and comparison with the existing method, when there is one. For a method validated in several laboratories, the minimum is 6 to 8 collaborating laboratories; with 2 to 6, it is a peer validation. Samples are sent coded, blind to whoever runs the test, and each laboratory receives at least two samples of each level, because it is the replicates that make it possible to compute repeatability and reproducibility. The levels cover the working range, with at least one low, one intermediate and one high. For uncertainty, the standard prefers the tolerance-table approach, the same one used in the Brazilian Rules for Seed Testing.

The point is that the standard does not ask for accuracy against a label. It asks for repeatability, reproducibility and uncertainty comparable to those of the reference method, measured with blind samples and replicates. The criterion of non-inferiority to a second analyst is the direct translation of this for a counting tool.

In SeedCounter

In SeedCounter

For an agreement study, use a single version of the app throughout data collection and record which one it was. Count each plate without seeing the analysts' readings, and declare the mode of each count. In assisted mode, the machine proposes the outline and the class and the person checks each seed, which is the mode the product advocates; automatic mode, without checking, shows what the checking corrects. Save the count and the class of each seed per plate: the count feeds the Bland-Altman plot, and the seed-by-seed class feeds the \(\kappa\) and the \(s\) and \(e\) of the Rogan-Gladen correction.

Where it fails

The Bland-Altman limits assume differences that are independent between plates and close to normally distributed. Plates from the same session or the same seed lot can have correlated errors, and the limits hold only for the range of counts and densities of the study. CCC and ICC grow with the range of the sample, and a study with only easy plates inflates both. \(\kappa\) drops with a rare class even though the reading does not get worse. The analysts' consensus is not the truth: where experts disagree, as with faint staining, no criterion separates tool error from legitimate doubt, and that boundary is reported separately. And non-inferiority depends on \(\delta\), which is a decision, not a statistic.

Data

This page does not use real analyst readings. The generator draws true counts and readings with the errors chosen in the controls, in your browser, and the numbers change with each draw. The sample size table and the examples of \(\kappa\) are exact computations; the script that reproduces them is in the site's source code.

References

  1. Bland J.M., Altman D.G. (1986). Statistical methods for assessing agreement between two methods of clinical measurement. The Lancet 327(8476), 307–310. doi:10.1016/S0140-6736(86)90837-8
  2. Bland J.M., Altman D.G. (1999). Measuring agreement in method comparison studies. Statistical Methods in Medical Research 8(2), 135–160. doi:10.1177/096228029900800204
  3. Lin L.I.-K. (1989). A concordance correlation coefficient to evaluate reproducibility. Biometrics 45(1), 255–268. doi:10.2307/2532051
  4. Cohen J. (1960). A coefficient of agreement for nominal scales. Educational and Psychological Measurement 20(1), 37–46. doi:10.1177/001316446002000104
  5. Fleiss J.L. (1971). Measuring nominal scale agreement among many raters. Psychological Bulletin 76(5), 378–382. doi:10.1037/h0031619
  6. Feinstein A.R., Cicchetti D.V. (1990). High agreement but low kappa: I. The problems of two paradoxes. Journal of Clinical Epidemiology 43(6), 543–549. doi:10.1016/0895-4356(90)90158-L
  7. Landis J.R., Koch G.G. (1977). The measurement of observer agreement for categorical data. Biometrics 33(1), 159–174. doi:10.2307/2529310
  8. Shrout P.E., Fleiss J.L. (1979). Intraclass correlations: uses in assessing rater reliability. Psychological Bulletin 86(2), 420–428. doi:10.1037/0033-2909.86.2.420
  9. McGraw K.O., Wong S.P. (1996). Forming inferences about some intraclass correlation coefficients. Psychological Methods 1(1), 30–46. doi:10.1037/1082-989X.1.1.30
  10. Koo T.K., Li M.Y. (2016). A guideline of selecting and reporting intraclass correlation coefficients for reliability research. Journal of Chiropractic Medicine 15(2), 155–163. doi:10.1016/j.jcm.2016.02.012
  11. Lu M.-J., Zhong W.-H., Liu Y.-X., Miao H.-Z., Li Y.-C., Ji M.-H. (2016). Sample size for assessing agreement between two methods of measurement by Bland-Altman method. The International Journal of Biostatistics 12(2), 20150039. doi:10.1515/ijb-2015-0039
  12. ISO 5725-6:1994. Accuracy (trueness and precision) of measurement methods and results, Part 6: Use in practice of accuracy values. International ISO standard.
  13. ISTA (2006). TCOM-P-01: ISTA Method Validation for Seed Testing, version 2.0. ISTA method validation standard. seedtest.org, Method Validation Programme

The machine proposes, the person checks, agreement measures.

SeedCounter counts and classifies seed by seed in the browser, and every seed is checked by the person doing the count.

Open SeedCounter