Research
I build and study applied AI systems: retrieval for law and tax, privacy-preserving use of language models, knowledge graphs and content pipelines, fine-tuning and evaluation of models, quantitative finance, drug-safety data and interpretability. A second strand is nature: species identity in image models, land cover and cattle terraces from aerial imagery, and hiking routes from GPS traces. Each area below says what I built and the techniques behind it.
Evidence and retrieval
Retrieval for law and tax
Retrieval that recognises when a source looks relevant but does not apply.
The paper models jurisdiction, concept and version mismatches as typed negative evidence. Next to it I build the engineering side: a legal research system over millions of multilingual documents.
- Hybrid search (BM25 and dense vectors, rank fusion) that is version-aware and jurisdiction-aware, with a citation graph of dated effects.
- Corpus pipelines: about two dozen source builders, nightly refresh with versioned validity, reversible data repairs.
- Generated retrieval questions, document profiles from weak supervision, and a tariff-classification agent.
- Deterministic citation checks, pseudonymisation before any hosted model call, and randomised trials of index changes.
Techniques retrieval-augmented generation, hybrid retrieval, reciprocal rank fusion, typed negative evidence, calibrated retrieval, citation graphs, version-aware indexing, gradient-boosted rerankers, weak supervision, doc2query, shadow-mode rollout, randomised controlled trials
Paper: What doesn't match matters more.
Evidence and retrieval
Knowledge graphs and automated content pipelines
A knowledge graph where every sentence traces to a source, built by automated pipelines.
- Claims keep what the source said and how sure it was. Imprecise dates stay intervals, and conflicts become typed records instead of a silent winner.
- Language models only return ids and sentence numbers through a JSON schema, and plain code checks every number, name and citation against the source.
- Model judges are replaced by cheaper measured checks (entailment cascades, rules) that run in shadow mode first.
- Immutable, content-addressed builds on Postgres with PostGIS and pgvector and DuckDB, with binary-quantised embeddings.
Techniques knowledge graph, claim-level provenance, schema-constrained LLM extraction, NLI cascades, content-addressed Parquet builds, DuckDB, PostGIS, pgvector, binary-quantised embeddings, shadow-mode validation
Privacy
Using cloud LLMs on confidential documents
Using cloud language models on confidential legal files without the identifiers reaching the provider.
A browser-based pseudonymisation gateway for confidential legal documents in five European languages. Detection runs on the device and the answer is restored locally. Every optional layer can only add protection.
- Document in the browser
- Local detection and pseudonymisation
- Optional encrypted check
- Protected text to the provider
- Answer restored locally
- 87–91%identifier recall
- 95–98%precision, 5,940 documents
- A deterministic evidence engine with explained redact, review or keep decisions, plus my own multilingual NER model compressed for in-browser use.
- A blind server model: the client sends only per-token features (shape, length, rule and checksum flags, document-local equality groups, ids for certified safe words) and no letters. A server ensemble of BiLSTMs and gradient-boosted trees labels the identifiers, and a trained inversion attack prices what each feature group leaks.
- A server tagger that sees only protected text, and a homomorphic-encryption second check. Evaluation compares the masked text that leaves the device with open detectors on public benchmarks.
Techniques pseudonymisation, multilingual NER, blind (letterless) server models, BiLSTM and gradient-boosted ensembles, inversion attacks, ONNX Runtime Web, homomorphic encryption (BFV, SEAL on WebAssembly), exposure accounting, egress-level black-box evaluation, red teaming, re-identification attacks
Perception and generation
Land cover from aerial imagery, Switzerland
Land cover and cattle-terrace (erosion) detection from aerial imagery, with every claim carrying its evidence and confidence.
Forest, grassland, shrub, bare rock and soil from 10 cm aerial imagery and LiDAR in Switzerland, and the faint step-like terraces ("terracettes") that grazing cattle leave on alpine slopes. Each result is a typed claim with its evidence, an effective resolution and a confidence, and the answer stays unknown where no calibrated estimate exists.
- 0.94forest AUC
- 0.26 → 0.007bare-soil calibration error
- 0.51 → 0.01seed spread, 14 to 500 sites
- Interpretable texture and terrain features, a closed ontology with resolution floors, and per-class calibration fitted on a national survey grid.
- Cattle-terrace (terracette) detection for erosion assessment from 10 cm imagery and LiDAR relief, using periodicity and structure-tensor features at parcel scale.
Techniques remote sensing, orthophoto and LiDAR features, canopy height models, probability calibration (isotonic and Platt), ontology-constrained structured output, Pydantic and FastAPI, vision-language baselines, site-held-out cross-validation
Perception and generation
Species identity in image models
Checking that image models draw the species that was asked for.
A picture of the wrong species is a factual error, for European animals, plants, fungi and lichens alike. I measure how well image models keep the distinguishing features of the requested species, what improves that, and how to check an image automatically.
- A taxonomic index of about 232,000 European taxa from roughly a hundred open sources, where each value retains its source and disagreement is typed.
- A congener-controlled test protocol and a staged gate: deterministic checks, BioCLIP scoring, then vision-language models.
- Species name only 0.222 12/54
- Name + named traits 0.519 28/54
- Name + shuffled traits (control) 0.056 3/54
BioCLIP agreement with the requested species among nine look-alike gentians (chance 1 in 9), 54 renders per prompt type, paired exact test p = 6e-8. A small pilot.
Techniques diffusion models (FLUX, Z-Image), LoRA, DreamBooth, textual inversion, ControlNet, BioCLIP, vision-language jurors, entity resolution, data provenance, paired exact tests
Perception and generation
Routes and maps
Recovering reliable hiking routes from many noisy GPS traces.
- 2 mmedian to a surveyed line
- Canonical routes and variants from crowd GPS traces by overlap bundling and consensus centrelines, checked against OpenStreetMap.
- Pedestrian routing with one engine per country, a custom walking rule, orienteering-based tour composition and line-of-sight checks against surface models.
Techniques GPS trace consensus, OpenStreetMap, GraphHopper, OR-Tools orienteering, LiDAR line of sight, 3D flyover rendering
Health data
Drug-safety signals
Finding drug-safety signals in public adverse-event reports and checking them against drug labels.
A pharmacovigilance pipeline over the FDA Adverse Event Reporting System: twelve quarters of reports, linked to the history of drug labels and approvals, with statistics and language-model tools on top for an analyst.
- 4.2Mdistinct cases, 12 quarters
- 23Mdrug records
- 479klabel-section versions
- Signal detection: proportional reporting ratio, reporting odds ratio, chi-squared and Fisher tests, information component and empirical-Bayes scores, Mantel-Haenszel pooling across strata, temporal trends, and an interaction statistic for drug pairs.
- Causality evidence: dechallenge and rechallenge extraction and Bradford Hill scoring.
- Data cleaning: three layers of duplicate-report detection, case versioning, and drug-name normalisation.
- Signal-to-label gap: a date-aware check of whether a signal appeared before the matching warning was added to the label.
- Analyst layer: about thirty language-model tools and a set of reusable analysis recipes, with interactive visual components.
Techniques pharmacovigilance, disproportionality analysis (PRR, ROR, IC, EBGM), Mantel-Haenszel, entity resolution, DuckDB, tool-using LLM agents, drug-label versioning
Models and markets
Models and tooling
Fine-tuned judges, classifiers and adapters, and the tooling to run open models locally.
A sentence-support judge (does the cited source support this sentence?) fine-tuned on Qwen models and a multilingual NLI model, on a test set held out by entity and labelled by model readers.
- 0.72 → 0.869B judge AUC after LoRA
- 0.66 → 0.780.6B judge AUC
- Token classifiers for multilingual NER, an XLM-RoBERTa detector for machine-written text, and a small adapter into a frozen 4-billion-parameter model.
- Local serving and benchmarking on one GPU, cost-aware routing behind a privacy gate, and agent workflows with MCP servers.
Techniques LoRA and QLoRA, PEFT, XLM-RoBERTa, NLI, adapters, llama.cpp, speculative decoding, MCP, multi-agent orchestration
Models and markets
Markets
A self-supervised model of trade and quote data, tested against named volatility baselines.
A causal transformer of about 20 million parameters on per-symbol dollar-volume states, read out frozen through one linear layer. Hypotheses and decision rules were recorded before runs, against named baselines, in two eras. A result only counts if it holds in both.
- +0.07 to +0.10R² over best baseline
- Earlier, on QuantConnect: a long-only equity strategy ranking S&P 500 stocks by a crash-risk measure (down-to-up volatility, after Chen, Hong and Stein, 2001), with a scored daily exit.
- Earlier, on QuantConnect: a deep-learning strategy. Keras sequence models read 25-day windows of momentum, money-flow and crash-risk features (down-to-up volatility, negative skewness), classify the next move, and are retrained every month on a rolling window. Two models vote on each trade.
Techniques self-supervised learning, transformers (grouped-query attention, rotary embeddings, SwiGLU), market microstructure, HAR-RV and HARQ, time-series foundation models (Chronos-2, TimesFM, Toto), walk-forward retraining, LSTM-style sequence models in Keras, pre-set decision rules
Models and markets
Interpretability
Finding which internal components of a small language model carry a concept.
- layers 26–28causal band
- Concepts the model has to infer can be ablated, while concepts restated in the prompt repair themselves.
Techniques mechanistic interpretability, logit lens, lens vectors, ablation, coefficient-swap steering, Qwen3-4B
Paper
What doesn't match matters more: a paper on typed negative evidence for retrieval. All papers and articles, with DOIs, are on the publications page.