Theory · Image and measurement · Scale
From pixel to millimeter: scale and uncertainty.
Every millimeter measurement comes from a single number, the size of the pixel in the plane of the seeds. This page shows how to measure that number with a ruler, how large its error is, how the error reaches length and area, and where the calculation stops holding.
Go deeper · mathematics and references
The scale
A digital image is a grid. Each pixel covers a small square of the seed plane, and the side of that square is the scale of the image. It is not written in the pixels: it has to be measured.
Definition
The scale \(s\) is the length on the object that corresponds to one pixel, in µm/px. The inverse, \(k = 1000/s\), is the resolution in px/mm. When the resolution comes in dots per inch, \(k = \mathrm{dpi}/25{,}4\), because an inch is 25.4 mm.
A length measured in pixels becomes micrometers when multiplied by \(s\), and an area, when multiplied by \(s^2\):
The exponent 2 in the area recurs throughout the page: any relative error in the scale arrives doubled in the area. For reference, 300 dpi gives 11.8 px/mm, or 84.7 µm/px; 1,200 dpi, 47.2 px/mm and 21.2 µm/px; 4,800 dpi, 189.0 px/mm and 5.29 µm/px. The camera in the site's animations, at 40 px/mm, comes to 25 µm/px.
Two points on the ruler
A ruler photographed together with the seeds gives the scale as a simple quotient. Mark two points on ruler ticks that are \(L\) micrometers apart. If the clicks land at positions \(p_1\) and \(p_2\) in the image, the distance between them in pixels is \(d = \lVert p_2 - p_1 \rVert\), and
With 10 mm between the ticks and 400 px between the clicks, \(s = 10\,000/400 = 25\) µm/px.
How wrong the scale can be
Calculation (2) uses two measured numbers, and both carry error. The click lands a few pixels from the center of the tick, and the tick itself has some tolerance in width and position. The Guide to the Expression of Uncertainty in Measurement, the GUM, handles this case with a first-order approximation: for a quotient of independent quantities, the relative uncertainties add in quadrature [1].
Result · scale uncertainty
If \(d\) and \(L\) are independent, with standard uncertainties \(u(d)\) and \(u(L)\), then
Sketch. Take logarithms: \(\ln s = \ln L - \ln d\). Differentiating, \(\delta s/s = \delta L/L - \delta d/d\). With small, independent errors, the variance of the sum is the sum of the variances, and the minus sign disappears when squared. This is the GUM law of propagation applied to the function \(L/d\).
Where do \(u(d)\) and \(u(L)\) come from? If the error of each click has a standard deviation of \(\sigma_c\) pixels in each direction, only the error along the ruler changes the distance, and the two clicks add: \(u(d) = \sqrt{2}\,\sigma_c\). Likewise, if the position of each tick has deviation \(\sigma_m\), \(u(L) = \sqrt{2}\,\sigma_m\). With a click deviation of 1.5 px over a distance of 400 px, \(u(d)/d = 0{,}53\%\). Marking 20 mm instead of 10 mm halves this term, because \(d\) doubles and the click error does not change.
More readings, and the cost of estimating the error from only a few
With \(n\) independent readings of the scale, the mean has uncertainty \(u(s)/\sqrt{n}\). But the deviation of each reading is generally not known in advance: it is estimated from the readings themselves, and with few readings that estimate is poor. The correction is Student's t distribution [2]. The 95% interval for the mean scale is
where \(S\) is the standard deviation of the \(n\) readings. With two readings, \(t\) equals 12.71. A real example: two scans from the same scanner, with the ruler scanned alongside, gave 5.3648 and 5.3240 µm/px. The coefficient of variation is low, 0.54%, and even so the 95% interval of the mean is ±4.9%. With the same coefficient of variation, three readings give ±1.34%; five, ±0.67%; ten, ±0.39%. A small coefficient of variation with two readings does not yet mean the scale is good. It is the third reading that narrows the interval.
From scale to measurement
Applying the same rule to (1), with the uncertainty of the measurement in pixels next to the uncertainty of the scale:
The scale term carries over unchanged to the length and doubled to the area, because the derivative of \(\ln s^2\) is \(2\,\delta s/s\). And it is systematic: every seed in the image inherits the same scale error, and measuring more seeds does not reduce it. Only a better calibration does.
Instrument · calibration simulator
…
±2u of the mean, known deviationt interval that covers the true scalet interval that does not cover it
…
Each line of the plot is a session: \(n\) readings drawn with the chosen errors, their mean as a dot and the interval (4) as a bar. Bars that run outside the frame end in an arrow. The tick error is the deviation of the position of each ruler tick; 20 µm is an illustrative value. The seed is 5.0 × 2.3 mm, and its interval uses only the scale term of (5).
One pixel at the edge
The second term of (5) comes from the edge. The seed boundary almost never falls cleanly between two pixels, and exactly where it sits depends on the threshold, the focus and the light. A simple model: the whole edge shifts half a pixel outward or inward.
Result · half a pixel at the edge
For a seed of length \(L\) and width \(W\), in pixels, a shift of ±0.5 px along the whole edge changes the length and the width by ±1 px, and the area, to first order, by
Proof for the ellipse. The area is \(A = \pi a b\), with semiaxes \(a = L/2\) and \(b = W/2\). Adding 0.5 px to both, \(\Delta A/A \approx \Delta a/a + \Delta b/b = 0{,}5/(L/2) + 0{,}5/(W/2) = 1/L + 1/W\). For a circle, \(L = W\), and the relative error in the area is twice that in the diameter: the same doubling rule as for the scale.
In an elongated seed, the width dominates. An orchid seed of 0.9 × 0.24 mm at 1,200 dpi is 42 × 11 px, and ±1 px amounts to 2.4% of the length, 8.9% of the width and 11% of the area.
Two cautions. First, (6) is the systematic shift of the edge, which comes from the threshold, the focus or the light. The random quantization error, pixel by pixel along the edge, largely cancels out in the area count and matters far less. Second, pixel conventions also shift the measurement: measuring the length between the centers of the extreme pixels gives one pixel less than counting the pixels end to end. Both conventions exist, and the difference between them is exactly the ±1 px of this section.

Declared DPI and measured DPI
Every image file can store a DPI value, and that value is whatever the program that saved the file wrote there. Many cameras record 72 dpi on every photo, near or far from the table, and a crop or a resize can leave the old value in the file. The scanner comes closer, because the pixel pitch depends on the mechanics of the device, and even so the declared value needs checking.
| Scale | dpi | px/mm | µm/px | deviation |
|---|---|---|---|---|
| declared | 3.600 | 141,7 | 7,056 | reference |
| ruler, scan 1 | 4.735 | 186,40 | 5,365 | +31,5% |
| ruler, scan 2 | 4.771 | 187,83 | 5,324 | +32,5% |
Measuring with the declared value would make every seed look 32% longer and, by the square, have 73% to 76% more area. The two scans agreed with each other within 0.8%, and checking by hand against the same ruler gave 4,814 dpi. All three values are close to the nominal resolution of the device, 4,800 dpi: the declaration was wrong, not the ruler.
The ruler checks the pixel pitch, not the sharpness. A scanner can record more pixels than its optics resolve, by interpolation, and then the pitch is right but the fine detail of the edge is invented. The perimeter is the measurement that suffers most from this.
In large collections, the DPI changes from file to file. In a folder of orchid seed scans there were files at 200, 1,200, 2,400, 3,200 and 4,800 dpi, and using 4,800 for all of them would give a scale error by a factor of up to 24 on the 200 dpi files. A TIFF can also store several scans stacked, of different sizes, and anyone who reads only the first one loses the others. The scale is read per file and per frame.
The flatbed scanner is already a reading instrument in the seed laboratory, with the ruler alongside. In signal grass (Urochloa) seeds, the tetrazolium reading on the image scanned at 1,200 dpi was equivalent to the one done under the stereo microscope [12].
Ruler off the plane
An ordinary camera follows, to a good approximation, the pinhole model: a point at distance \(Z\) from the lens and at height \(X\) from the optical axis lands in the image at position \(x = fX/Z\), where \(f\) is the focal length [3]. The same object, farther away, looks smaller in the image. The ruler gives the scale of the seeds only if it is at the same distance from the lens.
Result · ruler off the plane
If the seeds are at distance \(Z_s\) from the lens and the ruler is \(\Delta\) closer, at \(Z_r = Z_s - \Delta\), the scale of the ruler applied to the seeds gives
Sketch. By projection, the side \(p\) of a sensor pixel corresponds to \(s = pZ/f\) in the plane at distance \(Z\). The true scale in the plane of the seeds is \(s_s = pZ_s/f\) and the scale measured on the ruler is \(s_r = pZ_r/f\). Measuring the seeds with \(s_r\) multiplies every length by \(s_r/s_s = Z_r/Z_s\).
With the camera 44 cm from the seeds and the ruler 2 mm closer to the lens, resting on the rim of a plate, the seeds come out 0.45% smaller in length and 0.91% in area. With the phone at 15 cm and the ruler 5 mm above, 3.3% and 6.6%. On a flatbed scanner, the ruler and the seeds touch the same glass, and the problem nearly vanishes. A ruler with one end raised by 5° also looks shorter in the image, by the factor \(\cos 5° = 0{,}9962\), that is, 0.38%.
Instrument · ruler off the plane
A negative height means the ruler is below the seeds, farther from the lens: then the seeds come out larger. The line has slope \(-1/Z_s\): the closer the camera, the costlier each millimeter of height difference.
The digital perimeter
The outline of a region of pixels is a chain: a sequence of edge pixels, each a neighbor of the next. In the Freeman 8-neighbor chain, each step goes to one of the eight surrounding pixels, straight or diagonal [4]. Adding the steps with weight 1 for the straight ones and \(\sqrt{2}\) for the diagonal ones seems the natural way to measure the perimeter. It overestimates, and the size of the excess has a formula.
Result · excess of the 8-neighbor chain
A line segment of length \(\ell\) in direction \(\theta\), with \(0 \le \theta \le 45°\), becomes a chain of length
and, with all directions equally likely, the mean ratio between the chain and the true length is
Proof. The line advances \(\ell\cos\theta\) horizontally and \(\ell\sin\theta\) vertically. Each diagonal step advances 1 in both directions and each straight step advances 1 horizontally only, so the chain has \(\ell\sin\theta\) diagonal steps and \(\ell(\cos\theta - \sin\theta)\) straight ones. The length is \(\ell(\cos\theta - \sin\theta) + \sqrt{2}\,\ell\sin\theta\), which is (8). By the symmetries of the grid, any direction repeats the value of some direction between 0° and 45°, and the mean over uniform \(\theta\) is the mean over that interval:
For a closed curve the same holds, piece by piece. If the orientation of the seed in the image is random, each piece of the edge has a uniform direction, and the expected length of the chain is the factor (9) times the perimeter. The calculation requires the edge to be nearly straight on the scale of a few pixels, which holds for a seed that is tens of pixels wide.
Checked on published data. In 1,904 rice grains from a dataset that publishes the descriptors of each grain [5], the ratio between the length of the 8-neighbor chain and the published perimeter was 1.054052, within 0.07% of the constant. The reciprocal of the constant, 0.948, is the correction factor Kulpa proposed for the perimeter of blobs in binary images [6]: multiplying the chain by it removes the mean bias.
Better estimators give different weights to the steps or look at several steps at once, and have been studied in depth [7] [8] [9] [10] [11]. For anyone who measures seeds, the practical lesson is a different one: perimeter and circularity, \(4\pi A/P^2\), depend on the estimator, and circularity inherits the error squared. A perfect circle measured by the 8-neighbor chain comes out with a circularity near \(1/1{,}0548^2 = 0{,}90\), not 1.
With real seeds
An orchid seed after tetrazolium, in a crop where it measures 175 px in length. The crop has no ruler, so here the scale is assumed: say the seed is 0.9 mm, a common size for an orchid seed. Shrinking the image by area-average downsampling, as a sensor with larger pixels would, lets us see the same seed at 300, 600, 1,200, 2,400 and 4,800 dpi.

300 dpi10 × 3 px

600 dpi20 × 6 px

1,200 dpi42 × 11 px

2,400 dpi85 × 22 px

4,800 dpi171 × 43 px
| Resolution | Pixels L × W | Area (px) | ±1 px on L | ±1 px on W | ±1 px on area | Measured L × W (mm) | Area (mm²) |
|---|---|---|---|---|---|---|---|
| 300 dpi | 10,0 × 3,0 | 19 | 10,0% | 33,3% | 43,3% | 0,851 × 0,254 | 0,136 |
| 600 dpi | 20,4 × 6,2 | 78 | 4,9% | 16,0% | 20,9% | 0,863 × 0,264 | 0,140 |
| 1,200 dpi | 42,4 × 11,2 | 293 | 2,4% | 8,9% | 11,3% | 0,898 × 0,238 | 0,131 |
| 2,400 dpi | 85,2 × 22,1 | 1.205 | 1,2% | 4,5% | 5,7% | 0,901 × 0,234 | 0,135 |
| 4,800 dpi | 170,7 × 43,3 | 4.832 | 0,6% | 2,3% | 2,9% | 0,903 × 0,229 | 0,135 |
At 300 dpi the seed is 3 px wide: ±1 px is a third of it, and the measured width comes out 11% above the measurement at 4,800 dpi. A width of 2 or 3 pixels is not a measurement. The length holds up better, because it has almost four times as many pixels. And for ±1 px to amount to less than 5% of the area of this seed, not even 2,400 dpi is enough: at 4,800, it falls to 2.9%.
The same effect appears in a whole scan. A sheet of orchid seeds scanned at 4,800 dpi was downsampled to 2,400 and 1,200 dpi, and the same 22 seeds were measured in the three versions.
| Ratio to 4,800 dpi | Length | Width | Area | Perimeter | Circularity |
|---|---|---|---|---|---|
| 2,400 dpi | 0,973 | 0,944 | 0,949 | 0,830 | 1,429 |
| 1,200 dpi | 0,905 | 0,853 | 0,850 | 0,665 | 1,977 |
All the measurements change with resolution, but perimeter and circularity change far more: at 1,200 dpi the perimeter has already dropped 33% and the circularity nearly doubled. The rule that follows is to store the resolution alongside the measurement, read the DPI per file, and compare perimeter and circularity only between images of the same resolution.
In SeedCounter
In SeedCounter
The scale comes from the ruler in the image itself. You mark two points on ruler ticks and say how many millimeters lie between them; the app computes the scale in µm/px, and each seed comes out in millimeters. You can mark the ruler more than once: the app uses the mean of the readings and shows the spread between them, the coefficient of variation. From it and the number of readings, calculation (4) tells you whether one more reading is worth it. The file's DPI enters as a declaration, and the ruler does the checking. The scale is stored with the measurements, along with the declared DPI and the measured DPI, for whoever redoes the calculation later. The scale is the decision of the person measuring: the machine proposes, the person checks.
Where it fails
Where it fails
The first-order calculation assumes small, independent errors. It does not catch systematic error: a cheap ruler with wrong graduation, the same pair of ticks marked again on every reading, or the scale read in one corner of the image and used in another. Against that, the only remedies are a checked ruler and readings at several points in the image.
A ruler off the plane of the seeds errs by the ratio of the distances, as in (7). With the camera close to the table, a few millimeters of height difference already cost several percent.
Lenses with barrel or pincushion distortion make the scale vary from the center to the edge of the image, and a camera tilted relative to the table makes the scale vary from one side to the other. A ruler read at the center is not valid at the corner.
The DPI of the file can be wrong, missing, or different between files in the same folder, and a TIFF can store several scans. With 2 or 3 pixels in width, the width is not a measurement, and perimeter and circularity can only be compared within the same resolution.
In a simplified outline, circularity has another bias besides that of the chain: in SeedCounter, the 48-sided outline overestimates circularity in seeds with low solidity, by 9.9% at the median. It has been measured, and the correction is the next step.
Data
Figures 2 and 5 and Table 2: "Sementes de Orquídeas" dataset v8, Roboflow Universe (universe.roboflow.com/sementes-de-orqudea/sementes-de-orquideas), CC BY 4.0 license; profile, downsampled versions and measurements computed for this page, with an assumed scale. Tables 1 and 3 and the 1.054052 ratio: measurements made during the development of SeedCounter, on scans with a ruler and on the rice dataset of reference 5.
References
- Joint Committee for Guides in Metrology (2008). JCGM 100:2008. Evaluation of measurement data: guide to the expression of uncertainty in measurement (GUM). doi:10.59161/JCGM100-2008E
- Student (1908). The probable error of a mean. Biometrika 6(1), 1–25. doi:10.2307/2331554
- Hartley R., Zisserman A. (2004). Multiple View Geometry in Computer Vision, 2nd ed. Cambridge University Press. doi:10.1017/CBO9780511811685
- Freeman H. (1961). On the encoding of arbitrary geometric configurations. IRE Transactions on Electronic Computers EC-10(2), 260–268. doi:10.1109/TEC.1961.5219197
- Cinar I., Koklu M. (2022). Identification of rice varieties using machine learning algorithms. Journal of Agricultural Sciences 28(2), 307–325. doi:10.15832/ankutbd.862482
- Kulpa Z. (1977). Area and perimeter measurement of blobs in discrete binary pictures. Computer Graphics and Image Processing 6(5), 434–451. doi:10.1016/S0146-664X(77)80021-X
- Proffitt D., Rosen D. (1979). Metrication errors and coding efficiency of chain-encoding schemes for the representation of lines and edges. Computer Graphics and Image Processing 10(4), 318–332. doi:10.1016/S0146-664X(79)80041-6
- Vossepoel A. M., Smeulders A. W. M. (1982). Vector code probability and metrication error in the representation of straight lines of finite length. Computer Graphics and Image Processing 20(4), 347–364. doi:10.1016/0146-664X(82)90057-0
- Dorst L., Smeulders A. W. M. (1987). Length estimators for digitized contours. Computer Vision, Graphics, and Image Processing 40(3), 311–333. doi:10.1016/S0734-189X(87)80145-7
- Koplowitz J., Bruckstein A. M. (1989). Design of perimeter estimators for digitized planar shapes. IEEE Transactions on Pattern Analysis and Machine Intelligence 11(6), 611–622. doi:10.1109/34.24795
- Barba J., Chan K. S., Gil J. (1992). Quantitative perimeter and area measurements of digital images. Microscopy Research and Technique 21(4), 300–314. doi:10.1002/jemt.1070210407
- Custódio C.C., Damasceno R.L., Machado Neto N.B. (2012). Imagens digitalizadas na interpretação do teste de tetrazólio em sementes de Brachiaria brizantha [Scanned images in the interpretation of the tetrazolium test in Brachiaria brizantha seeds]. Revista Brasileira de Sementes 34(2), 334–341. doi:10.1590/s0101-31222012000200020
Mark the ruler more than once on your image.
In SeedCounter, the scale comes from the ruler in the photo itself, and the spread of the readings tells you when to stop.