hilum

Theory · Color and class · Masking

When the line hides a class.

With two classes, the least-squares classifier finds the same direction as discriminant analysis. With three or more, it can erase an entire class, even one well separated from the others. This page proves why, shows the remedies and reports what was measured on real seeds.

Go deeper · mathematics and references

The indicator matrix

Suppose n seeds, each described by p numbers, such as length, width and mean color (L*, a*, b*). Stack the descriptions in a matrix X with n rows, with a first column of ones for the intercept. The class of each seed goes into the indicator matrix Y, with n rows and K columns: in the row of seed i, the column of its class equals 1 and the others equal 0.

Definition

The indicator matrix regression fits one affine function per class by least squares, \(\hat f(x)^{\top} = (1, x^{\top})\,\hat B\), with X of full rank and

\[ \hat B = \arg\min_{B}\ \lVert Y - XB \rVert_F^2 = (X^{\top}X)^{-1}X^{\top}Y, \](1)

and classifies a new seed by the rule of the largest fitted value:

\[ \hat G(x) = \arg\max_{k = 1, \dots, K}\ \hat f_k(x). \](2)

The problem splits into K ordinary regressions, one per column of Y. Each class gets a function that tries to equal 1 on its own seeds and 0 on the others, and the new seed goes to the class whose function gives the largest value [1]. It is the cheapest linear classifier there is: one decomposition of X and K products.

Lemma 1

If X contains the column of ones, \(\sum_{k=1}^{K} \hat f_k(x) = 1\) for every x.

Proof. Each row of Y sums to 1, so \(Y\mathbf 1_K = \mathbf 1_n\). The column of ones is the first column of X, \(\mathbf 1_n = Xe_1\), and regressing a vector that already lies in the column space of X returns the vector itself: \((X^{\top}X)^{-1}X^{\top}\mathbf 1_n = e_1\). Hence \(\hat B\mathbf 1_K = e_1\) and \((1, x^{\top})\hat B\mathbf 1_K = 1\). ∎

The sum is 1, but each \(\hat f_k\) can come out negative or exceed 1. They are not probabilities, and rule (2) does not know that [2]. The lemma comes up again in the proof below.

Masking

Put three classes on a line: light, medium and red seeds, for example, described by a single number x. Figure 1 does this with 60 simulated seeds per class, drawn around −4, 0 and 4. The classes barely overlap, and anyone would draw two boundaries, one at −2 and one at 2.

Line: basis (1, x) · error 33.3% 011/3 −4 0 4 x predicted Parabola: basis (1, x, x²) · error 2.2% 011/3 −4 0 4 x predicted
SimulationFigure 1. Three classes in one dimension, 60 seeds each, with standard deviation 1 around −4, 0 and 4. The curves are the values fitted to the three indicators; the strip below is the predicted class, and the three rows of ticks are the seeds of each class. On the left, with the basis (1, x), the line of the middle class stays at 1/3 and is never the largest: 33.3% error, the entire middle class. On the right, with x² in the basis, it becomes a parabola that wins at the center: 2.2% error. Linear discriminant analysis has an error of 2.8% on the same seeds.

The phenomenon has a name, masking, and is described with these same three classes in [1, §4.2]. For the seeds in the figure, the three fitted lines are exactly 1/3 at the same point, the mean of all seeds, and the line of the middle class is almost horizontal (slope −0.0003, versus −0.114 and +0.114 for the others). This is not a coincidence of the simulation.

The proof

Proposition

If the K classes have the same number of seeds and their centers \(\mu_1, \dots, \mu_K\) lie on a single line, rule (2) with a linear basis predicts only the two classes at the ends. The K − 2 middle classes never win; at most they tie, on a set of volume zero.

Proof. Write the fit in the centered basis: \(\hat f_k(x) = \bar y_k + \beta_k^{\top}(x - \bar x)\), where \(\bar x\) is the mean of all seeds and \(\bar y_k = n_k/n\) is the fraction of class k. The normal equations give

\[ \beta_k = S^{-1}\sum_{i=1}^{n}(x_i - \bar x)\,y_{ik} = n_k\,S^{-1}(\mu_k - \bar x), \qquad S = \sum_{i=1}^{n}(x_i - \bar x)(x_i - \bar x)^{\top}, \](3)

because \(y_{ik}\) equals 1 only on the seeds of class k. With classes of equal size, \(\bar y_k = 1/K\) for every k, and all the functions equal 1/K at the point \(\bar x\). If the centers lie on a line, \(\mu_k - \bar x = t_k\,u\) for a fixed vector u and numbers \(t_1 \lt t_2 \lt \dots \lt t_K\). Then

\[ \hat f_k(x) = \frac{1}{K} + \frac{n}{K}\,t_k\,s(x), \qquad s(x) = u^{\top}S^{-1}(x - \bar x). \](4)

Lemma 1 checks the computation: with classes of equal size, the mean of the centers is \(\bar x\), hence \(\sum_k t_k = 0\) and the functions (4) sum to 1. The number s(x) is the same for all classes. Where \(s(x) \gt 0\), the largest function is the one with the largest \(t_k\); where \(s(x) \lt 0\), the one with the smallest \(t_k\); where \(s(x) = 0\), all tie. Classes 2 to K − 1 are never the largest. ∎

Note what does not appear in the computation: the spread within the classes and the distance between them. Perfectly separated classes are masked in the same way. What decides is the geometry of the centers and the size of the classes. With different sizes, the functions no longer pass through the same point, and the middle class can win in a band if it is much larger than its neighbors. With the middle center off the line, the vector u is no longer shared and the argument fails. The instrument in the next section lets you vary both.

The remedies

More terms in the basis

With the basis (1, x, x²), each \(\hat f_k\) can be a parabola, and the one for the middle class can rise at the center, as in Figure 1. The general rule is harsh: K classes in a row can require terms up to degree K − 1, and in p dimensions this reaches the order of \(p^{K-1}\) terms [1, §4.2]. For three classes, the quadratic basis is enough.

Linear discriminant analysis

LDA assumes that, within each class, the descriptors follow a normal distribution with center \(\mu_k\) and the same covariance Σ for all classes. Bayes' rule becomes the largest of K affine scores,

\[ \delta_k(x) = x^{\top}\Sigma^{-1}\mu_k - \tfrac{1}{2}\,\mu_k^{\top}\Sigma^{-1}\mu_k + \log \pi_k, \](5)

with \(\pi_k\) the fraction of the class. The form is the same as in least squares, but nothing forces the scores to pass through the same point. In one dimension, with equal variances and fractions, the boundaries fall at the midpoints between neighboring centers, and each class gets its own interval [1, §4.3]. Fisher reached the discriminant direction by another route, without assuming normality: the projection w that maximizes the ratio of the variance between classes to the variance within them, \(w^{\top}S_B\,w \,/\, w^{\top}S_W\,w\) [3].

Logistic regression

Multinomial logistic regression models the probabilities directly,

\[ P(G = k \mid x) = \frac{\exp(\beta_{k0} + \beta_k^{\top}x)}{\sum_{l=1}^{K}\exp(\beta_{l0} + \beta_l^{\top}x)}, \](6)

and finds the coefficients by maximum likelihood [1, §4.4]. The decision regions are also given by the largest of K affine functions. The difference is in the fit: least squares also penalizes seeds that lie far from the boundary on the correct side, the predictions that are "too right", and this pushes the boundaries to the wrong place [2]. The likelihood hardly cares about them.

Instrument · three classes in the plane

3,5 0,0 ×1,0

Fitted values along the line y = 0

–accuracy on the drawn seeds
–accuracy on the middle class
–classes that never win

class 1, circles middle class, triangles class 3, squares. Each cloud has 60 simulated seeds with standard deviation 1 (the middle one, 60 times the chosen size), and the center of each cloud sits exactly where the controls say. The dashed line is the line y = 0, which passes through the centers of the end classes. With the middle class on the line and of the same size, the least-squares line never chooses it, however large the separation. Move it off the line, increase its size or switch methods and watch it come back. Simulation

Two classes

With K = 2, masking has nowhere to happen: there is no middle class. And there is a stronger result. With two classes, the least-squares coefficient vector is proportional to \(\hat\Sigma^{-1}(\hat\mu_2 - \hat\mu_1)\), the LDA direction; only the cutoff changes [1, exercise 4.2].

Figure 2 checks this with the 1,614 real orchid seeds from the Color and class page, described by the mean color (L*, a*, b*) inside the label polygon. The cosine between the least-squares direction and the LDA direction came out as 1.0000000000. The coefficients, in the order (L*, a*, b*), are −0.025, 0.108 and 0.013: a*, the red, carries almost all of the weight.

0,0 0,5 1,0 1,5 rule: largest fitted value (cutoff at 0.5) fitted value for the viable class, f̂ (L*, a*, b*)
Real dataFigure 2. Fitted value for the viable class, by least squares on (L*, a*, b*), in the 1,614 orchid seeds in the tetrazolium test. Solid cyan bars: viable label. Dashed magenta bars: non-viable label. The line marks the largest-value rule, which with two classes is a cut at 0.5.

The largest-value rule agrees with the label in 92.8% of the seeds. Leaving one crop out at a time and fitting on the other 37, least squares and LDA agree on 1,493 of the 1,614 (92.5%) and logistic regression on 1,501 (93.0%). With two classes the three methods tie, and the descriptor sets the ceiling. Using only a*, least squares reaches 1,504 (93.2%): mean L* and b* add nothing here.

SVD and conditioning

To solve (1), \(X^{\top}X\) is never formed. With the singular value decomposition \(X = U\Sigma V^{\top}\),

\[ \hat B = V\,\Sigma^{+}\,U^{\top}Y, \qquad \hat Y = X\hat B = UU^{\top}Y, \](7)

and the fitted values are the orthogonal projection of Y onto the column space of X. The reason is numerical. The condition number of the matrix of the normal equations is the square of that of X,

\[ \kappa(X^{\top}X) = \kappa(X)^2, \qquad \kappa(X) = \sigma_{\max}/\sigma_{\min}, \](8)

and the computation loses about \(\log_{10}\kappa\) digits of precision [4]. In the public rice dataset with 106 descriptors per grain [5], κ(X) is 101,526 and κ(XᵀX) exceeds 10¹⁰: forming the normal equations throws away about 10 of the 16 digits of double precision. In dry beans, with 16 descriptors [6], the figures are 2,230 versus 4.97 million. For the orchid seeds with raw (L*, a*, b*), they are 878 versus 7.7 × 10⁵; with centered and standardized columns, κ(X) drops to 3.4.

The quadratic expansion has a cost here. When x does not change sign, x and x² move almost together and the conditioning gets worse; centering before squaring helps. The instrument above solves each fit by QR, which, like the SVD, never forms \(X^{\top}X\). What the SVD says about the directions of the data itself is in PCA, SVD and the mean seed shape.

In 88 species

Masking is a textbook result with simulated classes. Does it appear with real seeds? In six datasets with 2 to 7 classes, among them rice [5], beans [6] and soybean with defects [9], least squares stayed within 3 percentage points of LDA and logistic regression, with no class erased. It did appear in a public dataset of 88 forage seed species, with 4,496 images [7], described by 12 descriptors measured on the images. The same dataset was cut to 2, 5, 10, 20, 40 and 88 species, with the same descriptors.

Accuracy 2 5 10 20 40 88 number of classes K (log scale) Classes that never win 2 5 10 20 40 88 number of classes K (log scale) 0,00 0,25 0,50 0,75 1,00 least squares logistic 25% 50% 0% 0% 0% 10% 28% 51%
Real dataFigure 3. Accuracy and classes that never win, by number of species K. Measurements by the project on the public 88-species dataset [7]; no image from the dataset is reproduced here.
Species KLeast squaresLogisticRatioNever predicted
20,9170,9430,9730%
50,7580,8260,9180%
100,5610,6710,8360%
200,4160,6060,68710%
400,2980,4740,62928%
880,1590,3410,46651%

Table 1. Accuracy by number of species and the ratio between the two methods. The last column is the fraction of species that least squares never gives as an answer.

If the descriptors lacked information, both methods would fall together. They do not: the ratio between them goes from 0.973 to 0.466, and at K = 88 half of the species are never the answer given by least squares. The loss comes from the method. Table 2 confirms this through the remedy: changing only the basis, without changing a single descriptor, closes almost all of the gap to logistic regression.

At K = 88ColumnsAccuracySpecies that never win
Least squares, linear basis130,14645 of 88
LDA120,268–
Least squares, quadratic basis910,3377 of 88
Logistic120,340–

Table 2. The remedies at K = 88. This table comes from another run of the same measurement, in which the linear basis gave 0.146 instead of the 0.159 of Table 1. The 91 columns are the 12 descriptors, their squares and pairwise products, plus the intercept.

Color undoes masking

In a dataset with images of 65 forage species, with 775 whole seeds, the same computation was done with two sets of descriptors. With only 7 shape descriptors, accuracy was 0.223 and 12 of the 65 species (18%) never won. With shape and color, 27 descriptors, accuracy was 0.487 and only 2 species (3%) were erased.

The proposition explains why this is not just "more descriptors help". Masking arises from centers lined up along the few directions that the descriptors offer. Shape descriptors are strongly correlated: area, perimeter and diameters measure nearly the same size, and in beans the 16 descriptors have an effective rank of 7, at 99% of the energy. The species line up along these few directions, and the middle ones vanish. Color adds directions that the outline lacks, such as hue and color spread, and takes classes off the line. It is the instrument's "move the middle class off the line" control, carried out with data.

One caveat remains. The prediction that the erased class is precisely the middle one could not be tested here. With shape and color only 2 species are erased, too few to compare positions.

Nor does shape settle the classes that the standard asks for. In rice [5], the 106 descriptors give 0.997 accuracy; color and texture alone, 0.994; shape alone, 0.955. In durum wheat photographed on a conveyor belt [8], with 4,970 objects in three classes, shape alone recognizes 93.8% of the vitreous kernels, 82.3% of the starchy ones and 52.6% of the foreign matter, which is precisely the purity class. In soybean with defects [9], five classes, shape reaches 0.358 and shape with color, 0.707. In wheat, adding color reduces the bias of the estimated purity from 6.41 to 2.15 percentage points above the true value, and the number of grains beyond which the bias exceeds the sampling variation rises from 31 to 274.

The descriptor ceiling

A public soybean dataset from the tetrazolium test [10] provides seed halves scanned at 1,200 dpi, classified by type of damage (mechanical, stink bug, moisture or none) and intensity, and publishes hundreds of global color and texture descriptors per image, without the images. The project measured five things on it.

  1. With up to 10 seeds per class, least squares, LDA, logistic regression and random forest tie, within 0.03. Plain least squares stops improving when the number of seeds gets close to the number of descriptors; shrinkage versions, such as ridge regression and regularized LDA, fix this.
  2. Plain accuracy is misleading: 79%, versus 77.3% for always guessing "no damage". Balanced accuracy is 0.66.
  3. Switching seed lots drops balanced accuracy from 0.66 to near 0.50, which is chance. Validating on randomly drawn seeds inflates the figure.
  4. Labeling a few seeds of the new seed lot helps little: the old seed lot helps only until about 5 seeds are labeled, and even then by only 0.03.
  5. The ceiling, near 0.67, does not rise with more seeds.

The last point is the one that matters. The tetrazolium test is read from the position and extent of the damage in the embryo [11], and a global descriptor, which sums the whole embryo, throws away where the damage is. No classifier recovers what the descriptor did not keep. On the Color and class page, Figure 4 shows the same limit with two embryos of equal mean color.

In SeedCounter

In SeedCounter

The class of each seed is proposed by the machine and checked by you, with the classes of your protocol. Each seed goes into the CSV with its shape measurements and mean color, one row per seed, and this table is what lets you redo the computations on this page: least squares, LDA or logistic regression, with the descriptors you choose. Before combining many classes in a linear classifier, check whether the middle ones appear in the predictions.

Where it fails

The proposition requires classes of equal size and centers exactly on a line; real data come close to this but do not meet it exactly, and real masking is partial. The numbers for 88 and 65 species come from descriptors measured by the project on public datasets, with one train and test split; another split changes the decimal places. The quadratic expansion grows fast: with 26 descriptors, the quadratic basis exceeds 370 columns, and with few seeds per class it memorizes instead of learning. And none of the methods here replaces checking: the agreement is with labels, which are also readings.

Data

Figure 1 and instrument: simulation, computed in the browser. Figure 2: dataset "Sementes de Orquídeas" v8, Roboflow Universe (universe.roboflow.com/sementes-de-orqudea/sementes-de-orquideas), license CC BY 4.0. Figure 3 and tables: measurements by the project on the public datasets cited in the references of the papers; no image from these datasets is reproduced.

References

  1. Hastie T., Tibshirani R., Friedman J. (2009). The Elements of Statistical Learning, 2nd ed., chapter 4. Springer. doi:10.1007/978-0-387-84858-7
  2. Bishop C.M. (2006). Pattern Recognition and Machine Learning, §4.1.3, 184–186. Springer.
  3. Fisher R.A. (1936). The use of multiple measurements in taxonomic problems. Annals of Eugenics 7(2), 179–188. doi:10.1111/j.1469-1809.1936.tb02137.x
  4. Trefethen L.N., Bau D. (1997). Numerical Linear Algebra. SIAM. doi:10.1137/1.9780898719574
  5. Koklu M., Cinar I., Taspinar Y.S. (2021). Classification of rice varieties with deep learning methods. Computers and Electronics in Agriculture 187, 106285. doi:10.1016/j.compag.2021.106285
  6. Koklu M., Ozkan I.A. (2020). Multiclass classification of dry beans using computer vision and machine learning techniques. Computers and Electronics in Agriculture 174, 105507. doi:10.1016/j.compag.2020.105507
  7. Yuan M., Lv N., Dong Y., Hu X., Lu F., Zhan K., Shen J., Wu X., Zhu L., Xie Y. (2024). A dataset for fine-grained seed recognition. Scientific Data 11, 344. doi:10.1038/s41597-024-03176-5
  8. Kaya E., Saritas İ. (2019). Towards a real-time sorting system: identification of vitreous durum wheat kernels using ANN based on their morphological, colour, wavelet and gaborlet features. Computers and Electronics in Agriculture 166, 105016. doi:10.1016/j.compag.2019.105016
  9. Lin W., Fu Y., Xu P., Liu S., Ma D., Jiang Z., Zang S., Yao H., Su Q. (2023). Soybean image dataset for classification. Data in Brief 48, 109300. doi:10.1016/j.dib.2023.109300
  10. Pereira D.F., Bugatti P.H., Lopes F.M., Souza A.L.S.M., Saito P.T.M. (2019). Contributing to agriculture by using soybean seed data from the tetrazolium test. Data in Brief 23, 103652. doi:10.1016/j.dib.2018.12.090
  11. Moore R.P. (1985). Handbook on Tetrazolium Testing. International Seed Testing Association, Zurich.

Classes you define, seeds you check.

Build the classes of your protocol in SeedCounter and use the per-seed table for the computations on this page.

Open SeedCounter