Published experiment results
This page shows every published score and the benchmark version behind it. The main score is macro F1, a 0–1 measure that gives all five evidence labels equal weight. The challenge set contains harder cases from scenario families the model didn't see during training.
Important comparison limit
The traditional TF-IDF and majority-baseline results use benchmark v1.1.0. The completed Gemma runs use the earlier v1.0.0 benchmark. Both sets of results are valid, but they aren't a fair head-to-head ranking because the test sets differ. A clean comparison requires rerunning the approaches on the same benchmark version.
Published test macro F1
Full comparison table
| Experiment | Model | Test macro F1 | Accuracy | Challenge macro F1 |
|---|
Treat this table as a record of completed experiments, not a final overall ranking. Among the Gemma runs on v1.0.0, few-shot performed best. On v1.1.0, TF-IDF is the current traditional baseline.
Figures from the published results
These charts are generated from the committed result files. When a chart places v1.0 and v1.1 experiments side by side, read it as research history rather than a controlled same-test comparison.
How to read this
- A simple model can expose an easy benchmark If a basic word-pattern model scores extremely well on synthetic data, the dataset may contain shortcuts instead of requiring real evidence reasoning. The benchmark was hardened in v1.1.0 after an earlier generator made that problem visible.
- The three Gemma runs can be compared with each other Zero-shot, few-shot, and QLoRA were all tested on v1.0.0. Few-shot performed best of the three. QLoRA did worse than few-shot and often failed to return a label in the expected format.
- An overall winner still requires the same test To say whether the language model beats the traditional baseline, both approaches need to be run on the same benchmark version. ControlSift publishes the existing scores without pretending that comparison has already happened.