Published experiment results

This page shows every published score and the benchmark version behind it. The main score is macro F1, a 0–1 measure that gives all five evidence labels equal weight. The challenge set contains harder cases from scenario families the model didn't see during training.

Important comparison limit The traditional TF-IDF and majority-baseline results use benchmark v1.1.0. The completed Gemma runs use the earlier v1.0.0 benchmark. Both sets of results are valid, but they aren't a fair head-to-head ranking because the test sets differ. A clean comparison requires rerunning the approaches on the same benchmark version.
Status published Classical data v1.1.0 Gemma data v1.0.0 Seed 42 Source results/*/metrics_*.json

Published test macro F1

Full comparison table

Experiment Model Test macro F1 Accuracy Challenge macro F1

Treat this table as a record of completed experiments, not a final overall ranking. Among the Gemma runs on v1.0.0, few-shot performed best. On v1.1.0, TF-IDF is the current traditional baseline.

Figures from the published results

These charts are generated from the committed result files. When a chart places v1.0 and v1.1 experiments side by side, read it as research history rather than a controlled same-test comparison.

Model comparison chart
01 · Published experiment macro F1 (cross-version caveat applies)
Per-class F1 chart
02 · Per-class F1 (best available experiment)
Confusion matrix
03 · Confusion matrix
TF-IDF learning curve: test macro F1 vs training set size (seed 42)
05 · TF-IDF learning curve (v1.1.0)
Domain performance
07 · Performance by control domain
Challenge set comparison
10 · Published challenge results (cross-version caveat applies)

How to read this