Donald Murre

← Research

Models and tooling

Question
What do my own fine-tuned models add, and what does it cost to run and test open models locally?
Status
LoRA and full fine-tunes run with saved evaluations, from 0.6 to 9 billion parameters. Further fine-tunes are prepared but not run.

Training

Evaluation and serving

Fine-tuning a small support judge: AUC before and after
  • Qwen3-Reranker 0.6B, untrained 0.657
  • Qwen3-Reranker 0.6B, LoRA 0.781
  • Qwen3-4B, untrained 0.682
  • Qwen3-4B, LoRA 0.783
  • Qwen3.5-4B, untrained 0.717
  • Qwen3.5-4B, LoRA 0.824
  • Qwen3.5-9B, untrained 0.723
  • Qwen3.5-9B, LoRA 0.857

Area under the ROC curve on 945 to 1,403 held-out items per model (same task, same split within each model). Training uses 1,890 to 2,932 labelled items for one or two epochs. Labels come mostly from model readers and are not yet checked by a person, so read the gain as an improvement against that labelling.

Other research areas