Donald Murre

Research

I build and study applied AI systems: retrieval for law and tax, privacy-preserving use of language models, knowledge graphs and content pipelines, fine-tuning and evaluation of models, quantitative finance, drug-safety data and interpretability. A second strand is nature: species identity in image models, land cover and cattle terraces from aerial imagery, and hiking routes from GPS traces. Each area below says what I built and the techniques behind it.

Evidence and retrieval

Retrieval for law and tax

Retrieval that recognises when a source looks relevant but does not apply.

The paper models jurisdiction, concept and version mismatches as typed negative evidence. Next to it I build the engineering side: a legal research system over millions of multilingual documents.

Techniques retrieval-augmented generation, hybrid retrieval, reciprocal rank fusion, typed negative evidence, calibrated retrieval, citation graphs, version-aware indexing, gradient-boosted rerankers, weak supervision, doc2query, shadow-mode rollout, randomised controlled trials

Paper: What doesn't match matters more.

Back to contents

Evidence and retrieval

Knowledge graphs and automated content pipelines

A knowledge graph where every sentence traces to a source, built by automated pipelines.

Techniques knowledge graph, claim-level provenance, schema-constrained LLM extraction, NLI cascades, content-addressed Parquet builds, DuckDB, PostGIS, pgvector, binary-quantised embeddings, shadow-mode validation

Back to contents

Privacy

Using cloud LLMs on confidential documents

Using cloud language models on confidential legal files without the identifiers reaching the provider.

A browser-based pseudonymisation gateway for confidential legal documents in five European languages. Detection runs on the device and the answer is restored locally. Every optional layer can only add protection.

  1. Document in the browser
  2. Local detection and pseudonymisation
  3. Optional encrypted check
  4. Protected text to the provider
  5. Answer restored locally

Techniques pseudonymisation, multilingual NER, blind (letterless) server models, BiLSTM and gradient-boosted ensembles, inversion attacks, ONNX Runtime Web, homomorphic encryption (BFV, SEAL on WebAssembly), exposure accounting, egress-level black-box evaluation, red teaming, re-identification attacks

Back to contents

Perception and generation

Land cover from aerial imagery, Switzerland

Land cover and cattle-terrace (erosion) detection from aerial imagery, with every claim carrying its evidence and confidence.

Forest, grassland, shrub, bare rock and soil from 10 cm aerial imagery and LiDAR in Switzerland, and the faint step-like terraces ("terracettes") that grazing cattle leave on alpine slopes. Each result is a typed claim with its evidence, an effective resolution and a confidence, and the answer stays unknown where no calibrated estimate exists.

Techniques remote sensing, orthophoto and LiDAR features, canopy height models, probability calibration (isotonic and Platt), ontology-constrained structured output, Pydantic and FastAPI, vision-language baselines, site-held-out cross-validation

Back to contents

Perception and generation

Species identity in image models

Checking that image models draw the species that was asked for.

A picture of the wrong species is a factual error, for European animals, plants, fungi and lichens alike. I measure how well image models keep the distinguishing features of the requested species, what improves that, and how to check an image automatically.

Does the prompt steer the species?
  • Species name only 0.222 12/54
  • Name + named traits 0.519 28/54
  • Name + shuffled traits (control) 0.056 3/54

BioCLIP agreement with the requested species among nine look-alike gentians (chance 1 in 9), 54 renders per prompt type, paired exact test p = 6e-8. A small pilot.

Techniques diffusion models (FLUX, Z-Image), LoRA, DreamBooth, textual inversion, ControlNet, BioCLIP, vision-language jurors, entity resolution, data provenance, paired exact tests

Back to contents

Perception and generation

Routes and maps

Recovering reliable hiking routes from many noisy GPS traces.

Techniques GPS trace consensus, OpenStreetMap, GraphHopper, OR-Tools orienteering, LiDAR line of sight, 3D flyover rendering

Back to contents

Health data

Drug-safety signals

Finding drug-safety signals in public adverse-event reports and checking them against drug labels.

A pharmacovigilance pipeline over the FDA Adverse Event Reporting System: twelve quarters of reports, linked to the history of drug labels and approvals, with statistics and language-model tools on top for an analyst.

Techniques pharmacovigilance, disproportionality analysis (PRR, ROR, IC, EBGM), Mantel-Haenszel, entity resolution, DuckDB, tool-using LLM agents, drug-label versioning

Back to contents

Models and markets

Models and tooling

Fine-tuned judges, classifiers and adapters, and the tooling to run open models locally.

A sentence-support judge (does the cited source support this sentence?) fine-tuned on Qwen models and a multilingual NLI model, on a test set held out by entity and labelled by model readers.

Techniques LoRA and QLoRA, PEFT, XLM-RoBERTa, NLI, adapters, llama.cpp, speculative decoding, MCP, multi-agent orchestration

Back to contents

Models and markets

Markets

A self-supervised model of trade and quote data, tested against named volatility baselines.

A causal transformer of about 20 million parameters on per-symbol dollar-volume states, read out frozen through one linear layer. Hypotheses and decision rules were recorded before runs, against named baselines, in two eras. A result only counts if it holds in both.

Techniques self-supervised learning, transformers (grouped-query attention, rotary embeddings, SwiGLU), market microstructure, HAR-RV and HARQ, time-series foundation models (Chronos-2, TimesFM, Toto), walk-forward retraining, LSTM-style sequence models in Keras, pre-set decision rules

Back to contents

Models and markets

Interpretability

Finding which internal components of a small language model carry a concept.

Techniques mechanistic interpretability, logit lens, lens vectors, ablation, coefficient-swap steering, Qwen3-4B

Back to contents

Paper

What doesn't match matters more: a paper on typed negative evidence for retrieval. All papers and articles, with DOIs, are on the publications page.

The code is not public. Write-ups and releases follow as they are ready.