Theory · The statistics of the test report · Proportions
The margin of a percentage comes from the count.
The test report says 85% germination, but what the bench saw was 340 seedlings in 400 seeds. This page works out the margin from that count: the binomial, the Wald, Wilson and Clopper-Pearson intervals, what happens when replicates vary more than random sampling alone predicts, and how to correct the count of a classifier that gets some seeds wrong.
Go deeper · mathematics and references
The count is binomial
In a germination test with \(n\) seeds, each seed either germinates or does not. If all come from the same seed lot, with the same probability \(p\) of germinating, and one does not interfere with another, the number \(X\) of germinated seeds follows the binomial distribution. The same holds for viable seeds in the tetrazolium test, for damaged seeds in a quality class, or for any class counted seed by seed.
Definition
\(X \sim \mathrm{Bin}(n, p)\) when \(X\) counts the successes in \(n\) independent trials, each with success probability \(p\):
These two assumptions carry the rest of the page. The same \(p\) for all seeds fails when the seed lot is uneven. Independence fails when a fungus passes from one seed to its neighbor on the same plate. In both cases the replicates vary more than the binomial predicts, and the section on overdispersion measures by how much. The track page shows this random variation with 20,000 simulated test reports, in the section Chance.
The estimator and its variance
The natural estimator of \(p\) is the observed fraction, \(\hat p = X/n\). It is the maximum likelihood estimator and is unbiased.
Result
\(\mathrm{E}[\hat p] = p\) e \(\mathrm{Var}(\hat p) = p(1-p)/n\).
Proof. Write \(X = B_1 + \dots + B_n\), with \(B_i = 1\) if seed \(i\) germinated. Each \(B_i\) has mean \(p\) and variance \(\mathrm{E}[B_i^2] - p^2 = p - p^2 = p(1-p)\). With independence the variances add, \(\mathrm{Var}(X) = n\,p(1-p)\), and dividing \(X\) by \(n\) divides the variance by \(n^2\).
In percentage points (pp), this is the formula from the track, \(100\sqrt{p(1-p)/n}\). With 85% and 100 seeds it gives 3.6 pp; with 400 seeds, 1.8 pp. Quadrupling the seeds only halves the margin. And the standard error depends on \(p\), which is precisely what is not known: every interval has to decide what to put in place of \(p\) under the square root. That choice separates the methods that follow.
Wald, and where it breaks down
By the central limit theorem, \((\hat p - p)/\sqrt{p(1-p)/n}\) has a distribution close to the standard normal when \(n\) is large. The Wald interval replaces \(p\) with \(\hat p\) under the square root:
It is the textbook interval and the easiest to compute. It breaks down in three situations that come up every week in seed analysis. With no seed germinated, or with all of them, the square root is zero and the interval collapses to a point: 0 germinated in 50 would give \([0, 0]\), as if the absence were proven. Near 0 or 100%, with small \(n\), the interval extends outside \([0, 1]\). And even when it looks reasonable, it contains the true \(p\) in well under 95% of tests. With \(p = 5\%\) and 10 seeds, the 95% Wald interval contains the true value in 40.0% of tests, by the exact sum of equation (9).
The reason is that the square root uses \(\hat p\), which is poor precisely where the margin matters most. When \(\hat p\) comes out small by chance, the estimated margin also comes out small, and the interval ends up short and far from \(p\).
Wilson, by inverting the score test
Wilson [1] took the opposite path: instead of estimating the margin with \(\hat p\), he asked for which values of \(p\) the observed count would not be surprising. The score test of the hypothesis \(p = p_0\) uses the standard error under the null hypothesis, not the estimated one:
The interval is the set of \(p_0\) that the test does not reject, \(|Z(p_0)| \le z\). Square both sides and rearrange, and \((\hat p - p)^2 \le \tfrac{z^2}{n}\,p(1-p)\) becomes a quadratic inequality in \(p\):
Result: the Wilson interval
Proof. The parabola (5) opens upward, so the set is the interval between the roots. Set \(a = 1 + z^2/n\), \(b = -(2\hat p + z^2/n)\) and \(c = \hat p^{\,2}\). The center is then \(-b/(2a) = \tilde p\). The discriminant is \(b^2 - 4ac = \tfrac{4z^2}{n}\big[\hat p(1-\hat p) + \tfrac{z^2}{4n}\big]\), and half the distance between the roots, \(\sqrt{b^2-4ac}/(2a)\), is the margin given above.
The margin is the same formula as on the track, in the section Rare class. Three properties explain why it behaves well. The center \(\tilde p\) is pulled from \(\hat p\) toward 1/2, and the more so the smaller \(n\) is. The interval never leaves \([0, 1]\): at the extremes the parabola equals \(\hat p^{\,2} \ge 0\) and \((1-\hat p)^2 \ge 0\), and the vertex lies between them. And with \(x = 0\) the upper limit is \(z^2/(n + z^2)\), which for 50 seeds gives 7.1%, instead of the point that Wald returns. With 8 successes in 10, as on the track, the interval runs from 49.0% to 94.3%.
Clopper-Pearson and Agresti-Coull
Clopper and Pearson [2] invert the exact binomial test, with no normal approximation. The lower limit is the lowest \(p\) at which drawing \(x\) or more germinated seeds still has probability \(\alpha/2\); the upper limit, the highest \(p\) at which drawing \(x\) or fewer still has probability \(\alpha/2\):
The binomial tail is an incomplete beta function, \(P(X \ge x \mid p) = I_p(x,\, n-x+1)\), and for that reason the limits come out as quantiles of the beta distribution:
with \(p_L = 0\) when \(x = 0\) and \(p_U = 1\) when \(x = n\).
Result
The coverage of Clopper-Pearson is never below \(1-\alpha\), for any \(p\) and any \(n\).
Sketch. Fix \(p\). The interval excludes \(p\) from above only when \(P(X \le x \mid p) \lt \alpha/2\), and the counts \(x\) for which this happens form a lower tail with probability less than \(\alpha/2\). The same holds for the upper tail. Together, the two failure probabilities stay below \(\alpha\).
The price is the width. Because the binomial is discrete, the tails rarely add up to exactly \(\alpha\), the actual coverage is almost always above the nominal level, and the interval is wider than it needs to be. Agresti and Coull [3] called this conservative and proposed the opposite, an interval that is approximate but centered in the right place:
At 95%, \(z^2 = 3{,}84 \approx 4\): it is the Wald interval after adding two germinated and two non-germinated seeds to the count. The center is the same as Wilson's and the width is slightly larger. With 162 germinated out of 200, the four methods nearly coincide: Wald from 75.6% to 86.4%, Wilson from 75.0% to 85.8%, Agresti-Coull from 75.0% to 85.9% and Clopper-Pearson from 74.9% to 86.2%. The difference between them lies at the extremes and at small \(n\).
Coverage that oscillates
A 95% interval promises to contain \(p\) in 95% of tests. You can check the promise without any simulation. For fixed \(p\) and \(n\), the coverage is the sum of the binomial probabilities of the counts whose interval contains \(p\):
Because \(k\) takes only \(n+1\) values, \(C(p)\) is a sawtooth function: each time \(p\) crosses the endpoint of some interval, a whole term enters or leaves the sum. Brown, Cai and DasGupta [4] showed that for the Wald interval this sawtooth is deep and persistent. Coverage does not improve monotonically with \(n\), and lucky and unlucky sample sizes sit side by side. By the sum (9), with \(p = 20\%\) the 95% Wald interval covers 95.0% with 31 seeds and 89.0% with 32; with \(p = 0{,}5\%\), it covers 94.5% with 591 seeds and 79.2% with 592. The authors recommend the Wilson interval, or the Jeffreys interval (which we do not treat here), for small \(n\), and the Agresti-Coull interval for larger \(n\).
Instrument · actual coverage, from the exact binomial sum
| method | coverage at the selected p | mean width at the selected p | lowest coverage, p from 1% to 99% | mean coverage, p from 1% to 99% |
|---|
The dashed line is the promised level. Each curve is the sum (9) computed in your browser for hundreds of values of \(p\). Move a finger or the mouse over the chart to read the coverage of each method. Below 70% the curve leaves the frame. Try \(n = 30\) and \(n = 31\), or take \(p\) close to 1% or 99%.
Your count
What n is
In the formulas, \(n\) is the number of seeds, never the number of replicates. A test report with 4 replicates of 50 seeds and germination of 78%, 82%, 80% and 84% has 162 germinated seeds out of 200, and the Wilson interval with \(n = 200\) runs from 75.0% to 85.8%. If \(n\) were 4, the calculation would treat the report as 3 germinated seeds out of 4 and give 30.1% to 95.4%, an interval that says nothing about the seed lot.
The replicates serve another purpose: showing whether the seed lot varies more than chance sampling alone predicts. With them you can build an interval that does not depend on the binomial, from the mean and standard deviation of the percentages of the \(r\) replicates, with Student's \(t\) and \(r-1\) degrees of freedom. In the example, it runs from 76.9% to 85.1%. When the replicates vary as the binomial predicts, the two intervals are similar. When they vary more, the binomial interval is too short, and the next section measures by how much.
When the replicates vary more than the binomial predicts
In an uneven seed lot, each replicate comes from a part of the lot with slightly different germination. The simplest model draws the probability of each replicate from a beta distribution and then draws the seeds of the replicate with that probability. This is the beta-binomial, the same one as in the \(\rho\) control on the track, in the Chance section. Here, as there, \(n\) is the number of seeds in one replicate, and the test report has \(r\) replicates.
Definition
In replicate \(j\), \(P_j \sim \mathrm{Beta}(a, b)\), with mean \(p = a/(a+b)\), and \(X_j \mid P_j \sim \mathrm{Bin}(n, P_j)\). The number \(\rho = 1/(a+b+1)\) is the correlation between two seeds of the same replicate.
Result
\(\mathrm{E}[X_j] = np\) e \(\mathrm{Var}(X_j) = n\,p(1-p)\,\big[1 + (n-1)\rho\big]\).
Proof. For the beta distribution, \(\mathrm{Var}(P_j) = p(1-p)/(a+b+1) = p(1-p)\rho\) and \(\mathrm{E}[P_j(1-P_j)] = p(1-p) - \mathrm{Var}(P_j) = p(1-p)(1-\rho)\). By the law of total variance:
The factor \(1 + (n-1)\rho\) is the design effect. The variance of the percentage in the test report is multiplied by it, and the \(r\,n\) seeds are equivalent to \(n_{\mathrm{ef}}\) independent seeds. A small \(\rho\) can matter a great deal, because it is multiplied by \(n-1\): with \(\rho = 0{,}02\) and replicates of 100 seeds, the variance almost triples (\(1 + 99 \times 0{,}02 = 2{,}98\)), and 400 seeds are equivalent to 134.
The \(\rho\) is estimated from the replicates themselves. Pearson's statistic compares the variation among the replicate counts with the variation the binomial predicts:
Under the binomial, \(X^2\) approximately follows a chi-square distribution with \(r-1\) degrees of freedom, and \(\hat\varphi\) is close to 1. A \(\hat\varphi\) greater than 1 indicates overdispersion. Williams [5] used this same statistic to estimate the factor \(1 + (n-1)\rho\) in logistic models. For reporting, there are two honest options: multiply the binomial margin by \(\sqrt{\hat\varphi}\) (the quasi-binomial model) or use the \(t\) interval from the replicates in the previous section. To compare treatments, a generalized linear model with binomial response and this dispersion factor performs better than the arcsine transformation [6].
With 4 replicates, however, the test has little power. With 3 degrees of freedom, the 5% critical value of the chi-square is 7.81, and the test rejects the binomial only when \(\hat\varphi\) exceeds 2.6. Moderate overdispersion goes undetected. The effect shows up better when many plates are pooled, as in the corn data that follow.
With real seeds: corn hour by hour
Chen and colleagues [7] photographed 120 Petri dishes, each with 10 corn seeds, every hour in a germination test at 25 °C, and recorded the state of each seed. Here, a germinated seed is one with a root of at least 2 mm. The plates form six groups of 20 with different treatments, but the dataset does not say which group received which treatment. The groups are therefore called blocks 1 to 6 here, in plate order, with no treatment name.
Within a block, if the binomial holds, all the plates have the same \(p\), and the number of germinated seeds per plate at a fixed hour varies like a \(\mathrm{Bin}(10, \hat p)\). At 72 h all 120 plates were still being photographed. The table compares the variance observed among the 20 plates of each block with the binomial variance \(10\,\hat p(1-\hat p)\).
| block | plates | germinated | \(\hat p\) | variance among plates | binomial variance | \(\hat\varphi\) | \(X^2\) | p-value |
|---|---|---|---|---|---|---|---|---|
| 1 | 20 | 23 | 11,5% | 1,29 | 1,02 | 1,27 | 24,1 | 0,192 |
| 2 | 20 | 50 | 25,0% | 3,00 | 1,88 | 1,60 | 30,4 | 0,047 |
| 3 | 20 | 73 | 36,5% | 2,24 | 2,32 | 0,97 | 18,4 | 0,499 |
| 4 | 20 | 56 | 28,0% | 4,91 | 2,02 | 2,43 | 46,2 | 0,0005 |
| 5 | 20 | 30 | 15,0% | 1,53 | 1,27 | 1,20 | 22,7 | 0,249 |
| 6 | 20 | 1 | 0,5% | 0,05 | 0,05 | – | – | – |
| Sum of blocks 1, 2, 3, 4 and 5 | 1,49 | 141.9 (95 df) | 0,0013 | |||||
In blocks 1, 3 and 5, \(\hat\varphi\) lies between 0.97 and 1.27, within what chance explains (p-values from 0.19 to 0.50). Block 6 had a single germinated seed in 200 and has nothing to measure: with \(n\hat p\) near zero, the chi-square approximation does not hold, and the block is left out of the sum. Block 2 has a variance 1.6 times the binomial (p = 0.047). Block 4 has 2.4 times: two plates with 8 germinated seeds out of 10 next to others with 0 or 1, a pattern that the binomial with \(\hat p = 28\%\) almost never produces (p = 0.0005).
Summing the five blocks in which seeds germinated gives \(X^2 = 141{,}9\) with 95 degrees of freedom, p = 0.0013 and \(\hat\varphi = 1{,}49\), which corresponds to \(\hat\rho \approx 0{,}055\) with 10 seeds per plate. There is overdispersion, and it is concentrated in some of the blocks. For the test report of a whole block, the true margin is about \(\sqrt{1{,}49} \approx 1{,}22\) times the binomial one, and the 200 seeds are equivalent to about 134 independent seeds.
Instrument · real data, plate by plate
plates observed with that countplates expected from the binomial with the \(\hat p\) of the blockp-value below 0.05horizontal axis: germinated seeds on the plate, from 0 to 10
Real data Germinated seeds (root of at least 2 mm) per plate of 10 corn seeds, converted from the annotations published by Chen et al. (2024) [7], CC BY 4.0. The division into blocks follows plate order and is a hypothesis, because the dataset does not include a map from group to treatment. After 81 h some plates were no longer photographed and are left out of that hour. Blocks with fewer than 10 germinated seeds, summed over plates, do not enter the sum.
Correcting the count of an imperfect classifier
When a classifier decides the class, the count is subject to two types of error: non-viable seeds called viable and viable seeds called non-viable. With the notation of the track, \(s\) is the sensitivity, the fraction of non-viable seeds that the classifier gets right, and \(e\) the specificity, the fraction of viable seeds that it gets right. If \(p\) is the true fraction of non-viable seeds, the probability that a seed is counted as non-viable is
\(q\) is a straight line in \(p\), running from \(1-e\) in a seed lot with no non-viable seeds to \(s\) in a lot made only of non-viable seeds. Inverting the line gives the Rogan and Gladen estimator [8], with \(\hat q\) the fraction counted as non-viable in \(N\) seeds:
Result: Rogan-Gladen correction
For fixed \(s\) and \(e\), \(\hat p\) is a linear function of \(\hat q\): it is unbiased, and its variance is that of \(\hat q\) divided by the square of the slope \(s + e - 1\).
This is the factor \(1/(s+e-1)^2\) of the Classifier section of the track. In practice \(s\) and \(e\) also come from checking by hand, with \(n_s\) non-viable and \(n_e\) viable seeds checked, and they carry error. By the delta method, with \(J = s + e - 1\) and the derivatives \(\partial p/\partial q = 1/J\), \(\partial p/\partial s = -p/J\) and \(\partial p/\partial e = (1-p)/J\):
With the starting values of the track (\(s = 0{,}90\), \(e = 0{,}95\), \(\hat q = 22\%\) in 200 seeds), \(\hat p = 20{,}0\%\), with an interval of 13.2% to 26.8% if \(s\) and \(e\) are treated as exact. If they came from a reference plate with 100 non-viable and 100 viable seeds, the standard error rises from 3.4 to 4.1 pp and the interval runs from 12.0% to 28.0%.
The formula gives a negative \(\hat p\) when \(\hat q \lt 1-e\), that is, when the classifier flagged fewer non-viable seeds than false alarms alone would produce, and a value greater than 1 when \(\hat q \gt s\). With small \(p\) this happens by pure chance: with \(p = 2\%\), \(s = 0{,}90\), \(e = 0{,}95\) and 200 seeds, \(q = 6{,}7\%\), and the chance of \(\hat q\) falling below 5% is 13.2%. Truncating at zero hides what happened. The honest reading is that the count does not distinguish this seed lot from a lot with no non-viable seeds, and the interval starts at zero. If the value comes out far off, \(s\) and \(e\) have changed since they were measured and need to be measured again in the session.
In SeedCounter
In SeedCounter
Each seed in the test report is a seed counted and checked by a person, and the \(n\) in the calculations is the number of seeds, summed over all replicates. Count each replicate in a separate image and write down the percentages side by side: these are what make it possible to compute \(X^2\) and \(\hat\varphi\) and to know whether the binomial margin holds for your seed lot. To correct the count of a class proposed by the machine, check the class of each seed on a reference plate by hand at the start of the session. The fraction of non-viable seeds the machine got right is \(s\), that of viable seeds is \(e\), and the number of seeds of each class checked enters equation (13).
Where it fails
All the intervals on this page assume independent seeds. Touching seeds that contaminate one another, or a whole plate with different moisture, break this assumption in a way that no formula can detect from a single replicate. The beta-binomial summarizes unevenness in a single \(\rho\) for the seed lot, and \(X^2\) with few replicates, or with small \(n\hat p\), is a weak approximation. The Rogan-Gladen correction assumes that \(s\) and \(e\) are the same in the seeds of the test report and on the reference plate; if the reference plate comes from another seed lot or was imaged under different lighting, the correction applies that plate's error rates to this seed lot. And none of these intervals replaces the tolerances in the standard: for the official test report, the Brazilian Rules for Seed Testing apply.
Data
Chen C. et al. (2024). An RGB image dataset for seed germination prediction and vigor detection: maize. Frontiers in Plant Science 15, 1341335. doi:10.3389/fpls.2024.1341335. License CC BY 4.0. The counts per plate were converted from the annotations of the dataset; the script that generates the data and the table on this page is in the site code.
References
- Wilson E.B. (1927). Probable inference, the law of succession, and statistical inference. Journal of the American Statistical Association 22(158), 209–212. doi:10.1080/01621459.1927.10502953
- Clopper C.J., Pearson E.S. (1934). The use of confidence or fiducial limits illustrated in the case of the binomial. Biometrika 26(4), 404–413. doi:10.1093/biomet/26.4.404
- Agresti A., Coull B.A. (1998). Approximate is better than "exact" for interval estimation of binomial proportions. The American Statistician 52(2), 119–126. doi:10.1080/00031305.1998.10480550
- Brown L.D., Cai T.T., DasGupta A. (2001). Interval estimation for a binomial proportion. Statistical Science 16(2), 101–133. doi:10.1214/ss/1009213286
- Williams D.A. (1982). Extra-binomial variation in logistic linear models. Applied Statistics 31(2), 144–148. doi:10.2307/2347977
- Warton D.I., Hui F.K.C. (2011). The arcsine is asinine: the analysis of proportions in ecology. Ecology 92(1), 3–10. doi:10.1890/10-0340.1
- Chen C., Bai M., Wang T., Zhang W., Yu H., Pang T., Wu J., Li Z., Wang X. (2024). An RGB image dataset for seed germination prediction and vigor detection: maize. Frontiers in Plant Science 15, 1341335. doi:10.3389/fpls.2024.1341335
- Rogan W.J., Gladen B. (1978). Estimating prevalence from the results of a screening test. American Journal of Epidemiology 107(1), 71–76. doi:10.1093/oxfordjournals.aje.a112510
The margin starts at the n that you checked.
SeedCounter counts seed by seed in the browser, and a person checks each one. The margin comes from that number.