Research
In my tax work, a source can be on the right subject and still be wrong for the question, because it belongs to another jurisdiction or has been superseded. I study how retrieval can recognise that. The same concern with evidence and limits runs through my other work: pseudonymising confidential legal documents in the browser before they reach a model provider, land-cover detection from aerial imagery that returns claims with their evidence, and knowledge graphs that keep their sources through automated content pipelines. I also fine-tune diffusion models and encoders, build a self-supervised model of market microstructure, and run interpretability experiments on open language models.
Evidence and retrieval
- Retrieval for law and tax How a retrieval system can tell that a source which looks relevant does not apply: wrong jurisdiction, superseded version, or a different concept.
- Knowledge graphs and pipelines A layered graph in which every sentence traces to a source, built by pipelines whose language-model steps are constrained and checked.
Privacy
- Confidential documents and cloud LLMs Using a cloud language model on a client file without the identifiers reaching the provider: in-browser pseudonymisation, my own multilingual NER model, an encrypted second check, and how to evaluate it.
Perception and generation
- Land cover from aerial imagery Forest and land-cover detection from aerial imagery that returns structured claims with their evidence and a stated confidence.
- Species identity in image models When an image model is asked for a particular species, does the picture show it? A taxonomic index with provenance, controlled tests, and an automatic check.
- Routes and maps Hiking routes synthesised from crowd GPS traces, and pedestrian routing on OpenStreetMap graphs.
Models and markets
- Models and tooling My own fine-tuned judges, encoders and adapters (0.6 to 9 billion parameters), and local benchmarking, routing and agent workflows.
- Markets A self-supervised model of trade and quote data tested against named volatility baselines, and an earlier QuantConnect strategy.
- Interpretability Logit-lens, ablation and coefficient-swap experiments on a small open language model.
Paper
- What doesn't match matters more Working paper on typed negative evidence for retrieval. All papers and articles, with DOIs, are on the publications page.