Methods and evidence
Datasets, models, evaluation design, and published results. Metrics come from committed machine-readable experiment files.
Classical benchmark - v1.1.0
An early generator allowed TF-IDF to saturate at 1.0. Dataset v1.1 was hardened before the final classical run so simple lexical shortcuts no longer defined a trivial task.
Gemma experiment surface - v1.0.0
The completed Gemma zero-shot, few-shot, and QLoRA experiments use the earlier v1.0 synthetic packet surface. Free-tier compute constraints make a full v1.1 rerun outside the scope of this capstone.
Comparison ruleWithin-version comparisons are controlled. Cross-version classical-vs-Gemma scores are descriptive context only.
Published experiment paths
- C0Majority class / v1.1macro F1 0.067
- C1TF-IDF + logistic regression / v1.10.533 test / 0.524 challenge
- G0Gemma 3 1B zero-shot / v1.0macro F1 0.080
- G1Gemma 3 1B few-shot / v1.0macro F1 0.137
- G2Gemma 3 1B + QLoRA / v1.0macro F1 0.083
Controlled Gemma finding: few-shot is the strongest v1.0 Gemma run; QLoRA doesn't beat it and has poor label-parse reliability.
Evaluation controls
- Family isolation: scenario families stay within one split to reduce leakage.
- Machine-readable metrics: public scores come from committed result JSON.
- Malformed outputs: invalid Gemma labels count as errors rather than being silently repaired.
- Failure analysis: aggregate scores are paired with concrete hard cases in the Failure Lab.
- Version disclosure: the public site never presents v1.1 classical and v1.0 Gemma as a controlled same-benchmark ranking.