Species identity in image models
- Question
- When an image model is asked for a particular species, does the picture show that species, and how can that be checked automatically?
- Status
- Research prototype. Controlled experiments and a verification gate run; no service is deployed.
- Scope
- European animals, plants, fungi and lichens
Ask an image model for a gentian and it draws a plausible blue flower, not necessarily that species. For ecology, education and journalism, a picture of the wrong species is a factual error. I measure how well image models preserve the distinguishing features of the requested species, what improves that, and how to check an image automatically. No organisms are involved: it is about the accuracy of images.
What I built
- A taxonomic index of about 232,000 European taxa assembled from roughly a hundred open checklists, atlases and trait databases, in four layers (raw per-source claims, reconciled values, provenance, derived products). Each value retains its source.
- A merge engine that classifies disagreement between sources (agreement, complementary, contradiction, self-doubted) and keeps "unresolved" as a legitimate end state. Mirrors of one source do not count as corroboration.
- A structured description of each species' visible identifying features, shared by the prompt and the verification contract.
- A congener-controlled test protocol: scoring against same-genus relatives, shuffled-label and shuffled-description controls, paired seeds and exact tests.
- A staged verification gate: deterministic image checks, then a biology-pretrained retrieval model (BioCLIP, nearest-centroid scoring against real photographs), then a panel of distinct vision-language models checked against positive and negative controls.
- Training on licence-filtered photographs only, with a registry gate that blocks any fetch from an unregistered or refused source. A latent-diffusion trainer ported from TPU to a single GPU, with a capacity ladder from rank-8 LoRA to full fine-tune.
What I found
- Naming the species is not enough. Describing its diagnostic traits steers the image toward it, and the same words shuffled between species do not.
- Fine-tuning on a few look-alike species can destroy the species the model was not trained on, unless a supervision term limits the damage. Training added no identity gain for the species it saw.
- Several apparent failures were instrument failures: a conditioning channel carrying under one per cent of the signal, trigger tokens that were never expanded, and reference sets contaminated with herbarium sheets that three automatic filters had passed.
- Species name only 0.222 12/54
- Name + named traits 0.519 28/54
- Name + shuffled traits (control) 0.056 3/54
BioCLIP agreement with the requested species among nine look-alike gentians (chance is 1 in 9), 54 renders per prompt type with identical seeds, scored against 287 real photographs. A small pilot. The shuffled-trait arm tests whether diagnostic content matters beyond adding more words.