Theory · statistics
Every test report has a margin.
Counting seeds is sampling. Anyone who counts 400 seeds from a seed lot takes a sample of it, and another sample would give a different number. The gap between the two has a size, and it can be calculated before counting.
Understand · visual
One test report, four replicates of 100
germinateddid not germinateSimulation · the same seed lot at every draw
01 · Chance
How much a test report varies by chance
A seed lot with 85% germination does not deliver 85 seedlings in every replicate of 100. Each replicate is a draw, and the number that germinates follows the binomial distribution. The calculation below draws 20,000 test reports from the same seed lot and shows where the replicates fall and how far the highest is from the lowest, in percentage points (pp).
SD of one replicate= 100 · √(p(1 − p) / n) pp·heterogeneous seed lot,SD × √(1 + (n − 1)ρ)
One drawn test report
–
Difference between the highest and the lowest replicate in 20,000 test reports
test reportsthe 2.5% farthest out
What the number lets you say–
The standard decidesThe tolerance that applies to a test report is the one in the standard. Check the value in the tolerance table of the Brazilian Rules for Seed Testing, in the chapter for your test and for your number of replicates. This simulation shows what chance does. It does not decide whether a test report passes.
In SeedCounterCount each replicate in a separate image and write the percentages down side by side. The difference between the highest and the lowest is the number you take to the table.
02 · Classifier
Correcting the classifier's count
A classifier that separates viable from non-viable seeds makes errors in both directions. Sensitivity s is the fraction of non-viable seeds it gets right. Specificity e is the fraction of viable seeds it gets right. Counting what it proposes gives a biased percentage. The Rogan-Gladen correction removes the bias using s and e measured in the same session, at a cost in variance: each seed counts for less.
p = (q + e − 1) / (s + e − 1)·variance × 1 / (s + e − 1)²
It holds for any class. Replace non-viable with damaged, broken or stained.
–
Corrected percentage against counted percentage
correcteduncorrectedoutside the range of s and e
What the number lets you say–
In SeedCounterAt the start of the session, check by hand the proposed class of each seed on a reference plate. The fraction of non-viable seeds the machine got right is s. The fraction of viable seeds it got right is e.
A classifier with s = 0.45 and e = 0.89 multiplies the variance by 8.7. In that case, 400 seeds count as 46.
03 · Rare class
How many seeds of the rare class
Saying that the model gets 80% of the damaged seeds right is only meaningful when the checking includes enough damaged seeds. With 8 correct out of 10, the 95% interval runs from 49% to 94%. The Wilson interval gives the margin for each number of seeds in the class and behaves well with small n and with proportions near 0 or 100%. The calculation inverts the formula to find how many seeds the margin calls for.
margin = z / (1 + z²/n)· √(p(1 − p)/n + z²/(4n²))·z = 1.96
Margin against the number of seeds in the class, on a log scale
margin, Wilson 95%targetneededthe ones you have
What the number lets you say–
In SeedCounter–
04 · Drift
Control chart of the reference plate
A plate with a known proportion of the class, counted at the start of each session, shows whether the system has changed. Light, lens, camera and model change over time. As long as the measured percentage stays within three standard deviations of the binomial, the variation is consistent with chance. A point outside the limits calls for a review before the day's test report. The size of the plate determines what the chart can see: with few seeds the limits are wide, and a small drift passes without raising an alarm.
limits = p ± 3 · 100· √(p(1 − p) / n) pp
Percentage measured on the plate, session by session
measured3σ bandoutside the band
What the number lets you say–
In SeedCounterCount the reference plate at the start of each session and record the class percentage with the date. A point outside the band calls for a review of light, lens, camera or model before counting the day's samples.
05 · Budget
Error budget of the test report
The percentage in a test report carries error from five sources: the draw of the seeds, the heterogeneity of the seed lot, the reading method, the drift between sessions and the disagreement between analysts. The variances add up and the bias enters squared. Counting more seeds reduces only the error from the draw and the method noise. The rest does not depend on n, and only the protocol reduces it.
error² = 10⁴ · p(1 − p)/n· [k + (n − 1)ρ]+ bias² + drift²+ analyst²·in pp²
–
Where the error comes from, as a share of the squared error
falls with more seedsdoes not fall
Total error against the number of seeds, on a log scale
total errorfloor
What the number lets you say–
In SeedCounterEach term is measured in its own way. Drift comes from the control chart in section 04. Bias and method noise come from checking the reference plate by hand, as in section 02. Disagreement comes from two people checking the same plate without seeing each other's answers.
Limits
What this page does not decide
These calculations show what to expect from a seed lot, a plate and a method. What counts in an official test report is the Brazilian Rules for Seed Testing. Tolerances, number of replicates and sample size are given there, by test and by species.
The numbers come from simple models. The binomial assumes independent seeds. Heterogeneity enters as a beta-binomial, with a single ρ for the seed lot. The Rogan-Gladen correction assumes that s and e have not changed since they were measured. When the model does not describe your case, the number does not describe it either.
In section 01, 20,000 test reports are drawn, always with the same sequence of draws, so the number does not jump with each click. In 02, the interval uses the delta method, without continuity correction. In 03, it is the Wilson interval at 95%. In 04, the limits are the 3σ limits of the binomial. In 05, variances are added and the bias is squared. Everything runs in your browser.
Go deeper
The math, in full
The same math with formulas, proofs, real data and references.
Go deeper
Proportions and intervals
Wald, Wilson, Clopper-Pearson and Agresti-Coull, exact coverage, what n is, overdispersion in real corn data and the Rogan-Gladen correction.
Go deeper
Agreement with the analyst
Bland-Altman, Lin, Cohen's and Fleiss's κ, ICC and the non-inferiority criterion against a second analyst, with the sample size.
Count, check and keep the n.
SeedCounter counts seed by seed in the browser, and a person checks each one. The n these calculations ask for comes from there.