Methods and evidence

Datasets, models, evaluation design, and published results. Metrics come from committed machine-readable experiment files.

Classical benchmark - v1.1.0

Size
Approximately 1,500 synthetic cases across five balanced classes
Splits
train 1000 / validation 200 / test 200 / challenge 100, family-isolated
Seed
42
Labels
Rule-derived with structural audit and stratified spot-check
Privacy
Synthetic organizations only; no customer or employer evidence

An early generator allowed TF-IDF to saturate at 1.0. Dataset v1.1 was hardened before the final classical run so simple lexical shortcuts no longer defined a trivial task.

Gemma experiment surface - v1.0.0

The completed Gemma zero-shot, few-shot, and QLoRA experiments use the earlier v1.0 synthetic packet surface. Free-tier compute constraints make a full v1.1 rerun outside the scope of this capstone.

Comparison ruleWithin-version comparisons are controlled. Cross-version classical-vs-Gemma scores are descriptive context only.

Published experiment paths

  1. C0Majority class / v1.1macro F1 0.067
  2. C1TF-IDF + logistic regression / v1.10.533 test / 0.524 challenge
  3. G0Gemma 3 1B zero-shot / v1.0macro F1 0.080
  4. G1Gemma 3 1B few-shot / v1.0macro F1 0.137
  5. G2Gemma 3 1B + QLoRA / v1.0macro F1 0.083

Controlled Gemma finding: few-shot is the strongest v1.0 Gemma run; QLoRA doesn't beat it and has poor label-parse reliability.

Evaluation controls