Machine learning is helping to speed up Antarctic climate research by automating much of the painstaking work of diatom identification, cutting analysis time from weeks to months, to a fraction of that time.
Buried in layers of marine sediment are the microscopic remains of organisms that once lived near the surface of the ocean. Among the most scientifically valuable are fossil diatoms, tiny algae enclosed in intricate glass-like shells. Although invisible to the naked eye, their remains can provide a detailed record of how the marine environment has changed over time.
For Antarctic research, diatoms are particularly useful because different species thrive under different environmental conditions. Some are associated with cold waters or sea ice, while others flourish where nutrients are abundant or where ocean upwelling brings deep, nutrient-rich water to the surface. As environmental conditions change, so does the composition of the diatom community. This means that scientists can study diatoms preserved in Antarctic marine sediment cores to reconstruct, for example, past wind strengths, sea-ice concentrations, and climate shifts.
But there’s a practical challenge to studying diatoms: there can be thousands of specimens in a single sample. Identification of diatoms has traditionally been performed by the human eye of a specialist using a microscope. Researchers must distinguish between species based on subtle differences in shape, size, and the intricate structures of their silica shells. Doing this consistently across hundreds of samples and thousands (or millions) of individual specimens is time-consuming and requires considerable specialist expertise.
Martin Tetard, Paleontology Data Scientist at Earth Sciences NZ, says “all of that really takes lots of time. Also, manually looking through a microscope at all the diatoms or other protists groups requires lots of expertise, because usually even the experts might not agree with each other”.

Automating Diatom Identification
Martin has been leading innovative ASIS-supported research to automate marine diatom identification from microscope images. He explains that “we are using artificial neural networks to basically train on our own datasets, in collaboration with taxonomic experts.”
Essentially, the artificial neural networks are trained to recognise the species in a sample and then the computer counts everything that it can find, measuring it automatically.
The process involves attaching a Recognition-assisted Camera for Automated Microscopy (RaCAM) to a microscope, where it automatically captures images of the samples being viewed. The Raspberry Pi computer inside the camera then processes these images by detecting and separating individual objects of interest. An AI model analyses each object and identifies what it is based on patterns it has learned from previously trained images.
This means the entire process, from taking the microscope images to processing and identifying the objects, can happen automatically within a few seconds, without needing a separate computer. This innovation has led to around 90% successful identification. Any unidentified samples can be further visually identified by experts and sorted accordingly.
This advance in the automated identification of diatoms builds on previous research by Martin Tetard and colleagues, led by Ross Marchant and published in the Journal of Micropalaeontology. This 2020 study presented a method for automatically classifying large numbers of foraminifera (a different microscopic marine animal) images using convolutional neural networks (CNNs) and comparing automatically calculated foraminifera abundances with manual counts. The results, achieving 90% accuracy, demonstrated that an image-based method could make the analysis of large foraminifera datasets more efficient and support automated palaeoceanographic reconstruction.

The resolution of the images used in training datasets must strike a balance between being large and accurate. A lower-resolution image is easier for the artificial neural networks to learn from and process, but there is a minimum image quality required to achieve the desired recognition accuracy. Identification also takes more processing time if there is greater variation for the neural networks to learn.
The potential applications extend beyond marine diatoms. Martin points to the broader possibilities for automated identification, noting that the same approach is already being explored across a range of microscopic organisms: “We applied this method to marine diatoms, but in other projects we are applying it to other plankton organisms. In another project we are applying it to continental microfossils, like phytoliths and pollen.”

With thanks to the Diatom Research Team: Martin Tetard, Christina Riesselman, Xavier Crosta, Olivia Truax and Nancy Bertler - for their contributions.

