Model Card

Gemma 3 1B IT + QLoRA adapter for five-class evidence quality classification.

Base model
google/gemma-3-1b-it
Adaptation
QLoRA · 4-bit NF4 · double quant · r=16 · alpha=16 · dropout 0.05
Task
5-class control-evidence quality
Dataset
Gemma experiments: v1.0.0
Stack
PyTorch, Transformers, PEFT, TRL, bitsandbytes
Train entry
python -m controlsift.training.train

LoRA target modules are resolved from the loaded architecture, not copied from unrelated models.

Intended use

Research comparison of prompting versus PEFT on a synthetic benchmark; portfolio demonstration of evaluation rigor. Human judgment remains authoritative for any real control evidence.

Prohibited use

Evaluation results

Completed Gemma v1.0 experiments Zero-shot test macro F1: ~0.080. Few-shot: ~0.137. QLoRA: ~0.083. QLoRA test label-parse success: ~0.435. Within the controlled Gemma v1.0 comparison, few-shot is strongest and QLoRA doesn't beat it.

Result receipts are committed under results/gemma_zero_shot/, results/gemma_few_shot/, and results/gemma_qlora/. Classical experiments use dataset v1.1.0, so their scores aren't presented as a controlled same-benchmark comparison with these Gemma results.