Models and tooling
- Question
- What do my own fine-tuned models add, and what does it cost to run and test open models locally?
- Status
- LoRA and full fine-tunes run with saved evaluations, from 0.6 to 9 billion parameters. Further fine-tunes are prepared but not run.
Training
- A sentence-support judge: given a sentence and its cited source, does the source support it? LoRA fine-tunes (rank 16 on all attention and MLP layers) of Qwen models from 0.6 to 9 billion parameters, and a full fine-tune of a multilingual NLI model, on a test set held out by entity.
- Token classifiers for multilingual NER, fine-tuned and compressed for in-browser inference.
- An XLM-RoBERTa classifier for detecting machine-written text in four languages, evaluated with story-level cross-validation on a small pilot corpus.
- A small adapter that feeds market-state embeddings into a frozen 4-billion-parameter language model, with a grounding test that breaks the pairing between example and embedding. The result was real but small, and much of the apparent effect was symbol recognition.
- Diffusion models: LoRA, DreamBooth and textual-inversion fine-tuning (see species identity).
- Prepared: a QLoRA fine-tune of a 27-billion-parameter open model on route-split data.
Evaluation and serving
- Local serving and benchmarking of open models on one GPU: multi-dimension bake-offs with the scorer itself audited and corrected, and speculative-decoding sweeps.
- Cost-aware model routing with budget caps, failure diagnosis, and a gate that blocks hosted requests matching configured sensitive-data patterns.
- An orchestrator that plans, decomposes, executes and reviews tasks with language-model agents, and a plugin harness for agent workflows with a delegation ledger.
- MCP servers for literature search (Crossref and Semantic Scholar) and prose analysis.
- Qwen3-Reranker 0.6B, untrained 0.657
- Qwen3-Reranker 0.6B, LoRA 0.781
- Qwen3-4B, untrained 0.682
- Qwen3-4B, LoRA 0.783
- Qwen3.5-4B, untrained 0.717
- Qwen3.5-4B, LoRA 0.824
- Qwen3.5-9B, untrained 0.723
- Qwen3.5-9B, LoRA 0.857
Area under the ROC curve on 945 to 1,403 held-out items per model (same task, same split within each model). Training uses 1,890 to 2,932 labelled items for one or two epochs. Labels come mostly from model readers and are not yet checked by a person, so read the gain as an improvement against that labelling.