hilum

Data and credits

Where every seed on the site comes from.

The figures and screen recordings with real seeds use public datasets of images and measurements. Each block below says what the dataset is, where it appears on the site, where it came from, its license and how to cite it. The licenses were checked against the original sources on 8 October 2026. When the origin or license of a dataset still needs checking, the block says so. The animations are a different matter: they are drawn in code, and the end of the page explains the difference.

With an image on the site

Datasets that appear in a figure or screen recording

Photos, crops or outlines from these datasets are on the pages listed in each block.

Images and photos from the orchid research group (GPEOrq) and the seed research group (GPSEM)

Scans of orchid seeds after tetrazolium, made on the group's flatbed scanner, and photos of the lab and the team.

From the group, used with permission

Source: GPEOrq and GPSEM. To use these images, write to the contact address on the site.

Photos of the plants and seeds

Photos of orchids, grasses, soybeans, corn, rice, coffee and beans, and of the fun facts, to show each crop whole, from plant to seed. They are by third parties under open licenses, and each one appears with its author and license in the caption.

CC0, public domain, CC BY and CC BY-SA

Sources: Wikimedia Commons, iNaturalist and USDA PLANTS. The photos were downsized and converted to WebP, with no other changes.

Two images from Kew, for comparison

The capture bench and a slide read by both the person and the model, from the Millennium Seed Bank project that trains a model for the tetrazolium test in orchids. They appear only next to our own images, to compare the two approaches.

© RBG Kew, used with attribution

Source: "How machine learning can help us conserve orchid seeds", kew.org, 20 June 2025. The rights belong to the Royal Botanic Gardens, Kew.

Orchids in the tetrazolium test

Crops from tetrazolium plates of orchid seeds on blue paper, each seed with its outline and its class, viable or non-viable.

CC BY 4.0 license

Source: Roboflow Universe, version 8.

Citation: Sementes de Orquídeas, version 8. Roboflow Universe. universe.roboflow.com/sementes-de-orqudea/sementes-de-orquideas.

Composite sheet and 145 silhouettes

This is not a dataset but an assembly. Seeds from five datasets with one seed per photo (rice, coffee, corn, soybeans with defects and 88 species) were cropped and pasted into a single scene, with randomized position and rotation. The seeds were not photographed together. The quality-classes figure, with 23 seeds, is a scene of the same kind, using only the soybeans with defects. The 145 silhouettes used in the instruments, in the shape figures and in the Hilum symbol come from the same five sources.

CC BY-SA 4.0 license (assembly)

Sources: the Rice, Coffee, Corn, Soybeans with defects and 88 species blocks on this page, all with the license checked. Because coffee is CC BY-SA 4.0, the assembly is released under CC BY-SA 4.0.

Rice, five cultivars

Photos of rice grains of the cultivars Arborio, Basmati, Ipsala, Jasmine and Karacadag, one grain per photo, and a table of 106 shape and color descriptors measured on those photos. We checked SeedCounter's measurements against that table on 1,904 grains.

CC0 1.0 license

Source: the photos on Kaggle and the table on the author's datasets page, which asks for the citation below.

Citation: Koklu M., Cinar I., Taspinar Y. S. (2021). Classification of rice varieties with deep learning methods. Computers and Electronics in Agriculture 187, 106285. doi:10.1016/j.compag.2021.106285. Also: Cinar I., Koklu M. (2022). Identification of rice varieties using machine learning algorithms. Journal of Agricultural Sciences 28(2), 307–325. doi:10.15832/ankutbd.862482

Soybeans with defects

Photos of soybean seeds, one per photo, in five classes: intact, broken, immature, with a damaged seed coat, and spotted.

CC BY 4.0 license

Source: Mendeley Data, version 6. The Kaggle copy used here is a re-upload of this dataset.

Citation: Lin W., Fu Y., Xu P., Liu S., Ma D., Jiang Z., Zang S., Yao H., Su Q. (2023). Soybean image dataset for classification. Data in Brief 48, 109300. doi:10.1016/j.dib.2023.109300. The dataset is at doi:10.17632/v6vzvfszj6.6.

Soybean pods (YOLO POD)

Photos of soybean plants pulled up and laid on black cloth, with a box marked on each pod, made for counting pods. The outlines in the recording were generated inside the boxes and are not part of the original dataset.

CC BY 4.0 license on the AgML copy

Source: the AgML copy on Hugging Face, which assigns the CC BY 4.0 license. The original is on the Google Drive linked in the paper below, with no declared license.

Citation: Xiang S., Wang S., Xu M., Wang W., Liu W. (2023). YOLO POD: a fast and accurate multi-task model for dense Soybean Pod counting. Plant Methods 19, 8. doi:10.1186/s13007-023-00985-4

Three soybean cultivars

Scans of soybean seeds of the cultivars Anjasmoro, Dega 1 and Grobogan, with each seed cropped and the mask in the file itself.

CC BY 4.0 license

Source: Mendeley Data, version 3.

Citation: Syahraza M. A., Hanafiah D. S., Purnamasari F., Nurhasanah R. (2025). Image Dataset of Local Indonesian Soybean Seed Varieties (Anjasmoro, Grobogan, and DEGA-1). Mendeley Data, version 3. doi:10.17632/c733bjz4m3.3

Seeds of 88 species (LZUPSD)

Macro photos of seeds of 88 forage species, one seed per photo, on a black background. In the 145 silhouettes, they appear as “native”.

CC BY 4.0 license

Source: the images on figshare, under CC BY 4.0, described in the paper below.

Citation: Yuan M., Lv N., Dong Y., Hu X., Lu F., Zhan K., Shen J., Wu X., Zhu L., Xie Y. (2024). A dataset for fine-grained seed recognition. Scientific Data 11, 344. doi:10.1038/s41597-024-03176-5

Only in the calculations

Datasets used only as numbers, with no images

Measurements taken from these datasets appear in tables, charts and in the text. None of their images is reproduced on the site.

Germinating corn, hour by hour

Plates with 10 corn seeds photographed every hour during the germination test, with the state of each seed annotated. The site uses only the annotations, summed per plate: 120 plates, 1,200 seeds.

MIT license

Source: Kaggle, described in the paper below.

Citation: Chen C., Bai M., Wang T., Zhang W., Yu H., Pang T., Wu J., Li Z., Wang X. (2024). An RGB image dataset for seed germination prediction and vigor detection: maize. Frontiers in Plant Science 15, 1341335. doi:10.3389/fpls.2024.1341335

Dry beans, seven cultivars

Table of 16 size and shape descriptors measured on 13,611 dry bean grains of seven cultivars.

CC BY 4.0 license

Source: UCI Machine Learning Repository, dataset 602, doi:10.24432/C50S4B.

Citation: Koklu M., Ozkan I. A. (2020). Multiclass classification of dry beans using computer vision and machine learning techniques. Computers and Electronics in Agriculture 174, 105507. doi:10.1016/j.compag.2020.105507

Durum wheat on a conveyor belt

Durum wheat kernels filmed on a conveyor belt, in three classes (vitreous, starchy and foreign matter), with a spreadsheet of descriptors for each object.

CC0 1.0 license

Source: Kaggle, described in the paper below.

Citation: Kaya E., Saritas İ. (2019). Towards a real-time sorting system: identification of vitreous durum wheat kernels using ANN based on their morphological, colour, wavelet and gaborlet features. Computers and Electronics in Agriculture 166, 105016. doi:10.1016/j.compag.2019.105016

Soybean after tetrazolium

Global color and texture descriptors of soybean seed halves scanned after the tetrazolium test, classified by type and intensity of damage. The dataset publishes the descriptors, without the images.

License being verified

Source: the paper that publishes the descriptors.

Citation: Pereira D. F., Bugatti P. H., Lopes F. M., Souza A. L. S. M., Saito P. T. M. (2019). Contributing to agriculture by using soybean seed data from the tetrazolium test. Data in Brief 23, 103652. doi:10.1016/j.dib.2018.12.090

Seeds of 65 grass species

Photos of 775 whole seeds of 65 grass species, used to compare shape-only descriptors with shape-and-color descriptors.

Origin to be confirmedLicense being verified

Source: origin to be confirmed. The reference will be added here as soon as it has been checked.

The Theory pages also use numbers published in papers, such as the constants of the viability equation. These numbers are cited on the page itself, in the reference list.

Simulation or real data

What is drawn and what is measured

Every figure on the site carries a badge. It says whether the image comes from real seeds or was drawn in code.

Real data

Real seeds

A figure or recording made with one of the datasets on this page. The numbers come from SeedCounter's calculations on that data. The composite sheet, assembled from several photos, is explained above.

Simulation

Animations drawn in code

The animations on the site, on the home page, in Learn, in For labs, in Theory and on the blog, use model seeds drawn in code. They show the idea behind each measurement, not a measured result. The plates in the Test your eye challenge and the 3D bench in For labs are also simulations.

Demo

Examples that come with the app

The Soybean, Orchid TZ and Forage scenes inside SeedCounter are synthetic, with the correct answer known. The screen recordings with the twelve soybean seeds use that demo scene.

Real, coming soon

Dashed frame

Marks where a real photo or recording will go, with the code from the examples list or the recording script.

Is a credit missing, or is one wrong?

If you published one of these datasets and want a different form of citation, or you found an error on this page, write to us. We will correct the page.

Write to us