Human Tech Tree
Current research2024-2026 · Present (2015 – Oct 2026)

Life / Biology & Genetics

Animal-sound foundation models

NatureLM-audio (2024) and DolphinGemma (2025) are audio-language models for animal sounds; so far they give labels and predictions, not call meanings.

Open in the interactive tree →

These models are trained on animal recordings plus speech and music so that one network can label species, call types or life stages from a text prompt, or predict the next sound in a sequence. Almost all performance figures come from the developers' own benchmarks.

As of October 2026

NatureLM-audio (Earth Species Project, a nonprofit; preprint 11 November 2024, ICLR 2025) reports a new state of the art on its own BEANS-Zero benchmark, including zero-shot classification of unseen species, and has released weights and code. Google announced DolphinGemma on 14 April 2025, a model of about 400 million parameters trained on the Wild Dolphin Project's recordings of Atlantic spotted dolphins, collected since 1985; Google planned an open release, which the sources reviewed did not confirm. Earth Species Project says it is also studying crows, beluga whales and elephants with partner biologists. None of the sources reports a decoded call meaning.

Open steps

  • Common tests across species Medium AI leverageExtend benchmarks such as BEANS-Zero (call type, life stage, captioning, counting) to context and sequence tasks, so models are compared on more than species labels.
  • Generalise from few labels Medium AI leverageTest whether models trained on common species work on rare species and new sound types: NatureLM-audio reports zero-shot gains and Perch 2.0 transferred to marine tasks.
  • Peer-reviewed decode pilots Medium AI leverageEarth Species Project says it is studying crows, beluga whales and elephants with partner biologists; results need peer review and tests with the animals.
  • Real-time audio reply models Low AI leverageUse audio-in, audio-out models such as DolphinGemma to anticipate and generate sounds on a phone in the field; Google reported only early tests.

Where AI could help

Medium AI leverage. One model can label many species from a text prompt, but labels are not meanings and the benchmark scores come from the developers.

  • Label species, call types and life stages from a text prompt
  • Predict the next sound in a sequence to test for structure
  • Run on a phone in the field
  • Share open weights so small groups can adapt the model to their species

Shown so far

  • NatureLM-audio (Earth Species Project, ICLR 2025) reports state-of-the-art results on its own BEANS-Zero benchmark, including zero-shot classification of unseen species, and releases weights and code. source
  • Google's DolphinGemma (announced 14 April 2025) is a roughly 400-million-parameter audio-in, audio-out model trained on Wild Dolphin Project recordings and sized to run on field phones (company announcement). source

Prerequisites

Unlocks

Sources

More in Biology & Genetics · Present

All 66 points in Biology & Genetics →

Open in the interactive tree →