Model Card
Gemma 3 1B IT + QLoRA adapter for five-class evidence quality classification.
google/gemma-3-1b-itpython -m controlsift.training.trainLoRA target modules are resolved from the loaded architecture, not copied from unrelated models.
Intended use
Research comparison of prompting versus PEFT on a synthetic benchmark; portfolio demonstration of evaluation rigor. Human judgment remains authoritative for any real control evidence.
Prohibited use
- Automated audit or compliance certification
- Regulatory decision-making
- Replacement for control owners or auditors
- Scoring real customer evidence in production without human review
- Claims of real-world compliance determination
Evaluation results
Completed Gemma v1.0 experiments Zero-shot test macro F1: ~0.080. Few-shot: ~0.137. QLoRA: ~0.083. QLoRA test label-parse success: ~0.435. Within the controlled Gemma v1.0 comparison, few-shot is strongest and QLoRA doesn't beat it.
Result receipts are committed under results/gemma_zero_shot/, results/gemma_few_shot/, and results/gemma_qlora/. Classical experiments use dataset v1.1.0, so their scores aren't presented as a controlled same-benchmark comparison with these Gemma results.