Land cover from aerial imagery, Switzerland
- Question
- Can a perception system say what it sees, how sure it is, and what it cannot tell, in a form a downstream system can check?
- Status
- Working prototype with an API and an inspector. Parcel-level rules are a draft.
- Data
- 10 cm aerial orthophotos and airborne LiDAR from national open data, Switzerland
Forest, grassland, shrub, bare rock and soil from high-resolution aerial imagery, built as a service whose output is structured. Every claim is a typed object with an epistemic status (observed, inferred, or derived from an external layer), a confidence, an effective resolution and a per-attribute observability score. The validator rejects claims that break evidence, resolution or schema rules, and supported classifications use fitted calibration cells. Where no calibrated estimate exists, the answer stays unknown.
What I built
- Interpretable features: structure-tensor orientation, spectral and periodicity statistics, canopy height from surface-minus-terrain models, and flow-routed height above nearest drainage.
- A closed ontology with per-class resolution floors, graceful fallback to a coarser class, and five observation states, so that the absence of a claim is never read as a claim of absence.
- Per-class, per-observability-band probability calibration, fitted on a systematic national survey grid, with out-of-fold reliability reporting and immutable artefacts bound to the ontology version.
- Labelled corpora assembled from open geodata and a hand-interpreted survey grid, across four datasets from three countries.
- A typed HTTP API with audit endpoints for features, ontology, rules and reliability curves, and a browser inspector for the evidence behind each value.
- Open-weight vision-language models benchmarked as baselines and labelling tools, constrained to propose a hypothesis with a named falsifying measurement.
- A separate, versioned rules layer for parcel-level determinations, so that perception never decides eligibility by itself. A draft with placeholder thresholds.
What I found
- Corpus breadth beat modelling: with the same features and model, the spread of held-out scores across random seeds fell from about 0.5 to about 0.01 as the number of sites grew from 14 to 500.
- A visual cue that ranked first was falsified on re-measurement, and in one case the effect reversed sign once region and slope were controlled.
- Transfer to imagery from other countries is poor, three transfers falling below chance, because a shared class name hides different label meanings.
- Giving terrain information to a vision-language model, as text or as an image panel, changed nothing. Where the prompt put the information mattered more than what it said.
- Bare soil, raw 0.263
- Bare soil, calibrated 0.007
- Dwarf shrub, raw 0.310
- Dwarf shrub, calibrated 0.003
Expected calibration error for two of the rarest classes, before and after per-class calibration. Lower is better. Raw scores on rare classes are badly overconfident.