Summary
How are we improving soil biodiversity research?
This research project aimed to overcome key challenges in soil biodiversity research through development of ‘MESOSCAN’, a machine‑learning dataset and tool for automatically identifying soil biodiversity.
What is the challenge?
It focused on mesofauna (mites and springtails) which are particularly challenging to study due to the hidden nature of the soil as a habitat, the small size of these creatures (0.2 mm to 2 mm), how little we know about their ecology, and the evolving nature of their taxonomic classification. Additionally, there is a lack of taxonomy experts or identification keys in the UK, and identification can be time consuming, taking up to two days per sample in monitoring schemes that may be collecting up to 400 samples.
How does MESOSCAN help?
MESOSCAN helps enable the collection of robust, accurate mesofauna data that can ultimately be used to calculate biodiversity metrics for soil health assessments – working alongside the taxonomist allowing them to focus on identifying rarer and more unique specimens, while the MESOSCAN tool delivers a faster preliminary result.

Research Objectives
This project objectives were:
- Develop a machine learning–based biodiversity identification tool through a multidisciplinary approach integrating ecology, taxonomy, machine learning, and engineering.
- Build a robust training dataset by incorporating collected soil biodiversity samples into the machine learning workflow.
- Improve the efficiency of sample processing to enable larger, higher‑quality taxonomic resolution of biological datasets.

Findings and Recommendations
The two main challenges of the project have been the data collection of sufficiently high-quality images, and the development of automated classification using Machine Learning. To train such a system using a Deep Learning approach requires an extensive and well-curated set of training images. To create the best possible collection, we have constructed a macro-photography system mounted on a robotic arm that can scan many images at a very high magnification.
As with any new piece of scientific apparatus, it is critical to calibrate and monitor its performance. This is especially important in the case of a Machine Learning system, as it can be difficult for a human to interpret why a specimen was classified as such. Currently, we are performing this assessment by comparing the results of the systems output to the precise counts produced by an expert taxonomist.
Development of this machine learning-based identification tool will generate stronger baselines for soil fauna and enhance correlations with soil chemical and physical properties. It will also help advance understanding of soil food webs and ecosystem functioning.
Integration of additional biological data, e.g. molecular and trait data, could furthermore be integrated in future iterations to further expand capability and accuracy.
Latest Update
A dataset paper describing the methodology and linking to the data collected is currently under review and will be published in 2026.
Additionally, a follow up project has been commissioned for this work by Defra DNA CoE (2026-2029) using NCEA mesofauna data, it looks to “Enhancing indicator readiness of soil mesofauna DNA metabarcoding: pipeline benchmarking, taxonomic harmonisation and integration with morphology and ML-based monitoring for use in the production of soil health metrics”.
Funding & Partners
- Funded by the UK Government through Defra’s Natural Capital and Ecosystem Assessment programme (NCEA)
-
-