Somewhere in the Pacific, an autonomous underwater vehicle is carefully navigating an 80-meter-deep reef slope as its cameras are constantly filming everything that comes into view. A tiny fish floats by, looking unimpressive from a distance. It is pale and somewhat translucent, and it is hovering close to a sponge-like structure. It is recorded by the car. In a matter of seconds, a computer vision model back on the research vessel analyzes the frame, identifies the fish as a possible match for a genus that hasn’t been documented in this ocean basin, and compares the fish’s morphology against a database of known species. The notification goes to a marine biologist. The AUV continues to travel.

In actuality, the term “AI oceanographer” refers to a collection of machine learning tools that have been absorbing the pattern-recognition problem at the heart of marine taxonomy and solving portions of it far more quickly than human experts working alone. The typical scientific procedure, which includes collecting specimens, transporting them to laboratories, having experts examine them, comparing them taxonomically, and publishing the results, was not designed to handle the volume of data produced by the ocean. Every voyage, the camera systems on research submersibles produce hundreds of hours of video. Acoustic data from large areas of open water is continuously recorded by hydrophone arrays. eDNA sampling extracts genetic material from water that has never come into contact with a human hand or a net. It’s not a bottleneck to process everything by hand. It cannot be done structurally.
Perhaps the most obvious application is computer vision. Video footage from ROV or AUV cameras can be scanned by models trained on annotated image databases of marine species, which draw from collections from organizations like the Smithsonian, MBARI, and several national natural history museums. These models can identify organisms with a speed and consistency that would require teams of taxonomic experts working in parallel to match. Novel species that don’t closely resemble anything in the database are tagged as questionable, which is the right result and one that seasoned researchers may investigate further. The algorithms aren’t perfect, and they are noticeably less reliable outside of their training distribution. What they excel in, they do at a scale that alters what is manageable.
The application that has garnered the most public attention is acoustic monitoring, in part because the results are more emotive than spreadsheet data. Individual whale cries inside hydrophone recordings can now be recognized in real time by neural networks trained on recordings of marine mammal vocalizations, differentiating species and occasionally individual individuals based on minute acoustic fingerprints. Fish are being included in this research more and more because the underwater acoustic environment is far more biologically rich and complicated than it was even twenty years ago, and AI techniques are starting to comprehensively define it for the first time.
The most effective tool for altering the access issue is environmental DNA. Conventional species surveys necessitate either physical specimen acquisition or direct observation, both of which are challenging in remote, deep, or structurally complicated habitats. All that is needed for eDNA sampling is a water sample. Skin cells, excrement, mucus, and minuscule pieces are just a few examples of the genetic material that organisms constantly release into the water around them. A sufficiently sensitive genomic study can identify species from these traces without ever seeing the animal. The sequencing and matching process has been significantly sped by genomic AI models, making eDNA surveys feasible at scales that were not possible when each sample needed weeks of laboratory examination. In a given water column, a survey that could have previously found thirty species can today find hundreds.
There is a significant enough discrepancy between predicted actual marine diversity and known marine species to be fairly depressing. There are thot to be between 700,000 and one million marine species, of which about 25% have formal descriptions. There are hundreds of thousands of undescribed specimens in natural history museum collections. These specimens were gathered by researchers who had the institutional means or time to finish official taxonomic descriptions. AI systems are starting to tackle that backlog head-on by analyzing specimen data and museum photos to produce initial identifications that may then be confirmed and formalized by human taxonomists. It doesn’t take the place of the knowledge needed to formally describe a new species. The pipeline is becoming much faster as a result.
