BenchmarkThe fixed set of test cases used to measure how well each approach performs.
TF-IDF + logistic regressionA traditional machine-learning method that looks for useful word patterns. It's fast, inexpensive, and isn't a language model.
Gemma 3 1BThe small Google DeepMind language model tested in the project.
Zero-shot / few-shotZero-shot means asking the model with no worked examples. Few-shot means showing it a few examples first.
QLoRAA way to fine-tune a language model while using much less computing memory than retraining the whole model.
Macro F1A 0–1 summary score that gives each evidence label equal importance. Higher is better. It isn't the same thing as accuracy.