Research protocol
Frozen research contract and final execution boundary. Repo mirror: governance/PROTOCOL.md.
protocol-v1-lockedgoogle/gemma-3-1b-itFinal version boundary The locked v1.1 classical benchmark and the completed v1.0 Gemma runs are retained as separate experiment families. Cross-version scores are descriptive only and aren't a controlled same-benchmark leaderboard.
Hypotheses (frozen)
- H1 · Domain adaptation QLoRA Gemma beats the same model with a fixed zero-shot prompt on held-out macro F1.
- H2 · Prompting versus fine-tuning QLoRA beats a fixed five-class few-shot prompt.
- H3 · Generalization Measurable improvement remains on the challenge set.
- H4 · Data scaling Performance rises with train size, then diminishes.
- H5 · Remaining weaknesses Hard boundaries persist, especially PARTIAL vs SUFFICIENT and PARTIAL vs INSUFFICIENT.
Within the completed Gemma v1.0 experiments, few-shot is strongest and QLoRA doesn't beat it. Cross-family hypotheses requiring a common dataset version aren't treated as controlled tests.
Evaluation contract
- Deterministic inference
do_sample=false, temperature 0 where supported. - Strict parsing Ambiguous output becomes
UNPARSEABLE, never silently coerced to gold. - Uncertainty Bootstrap CIs; McNemar for paired comparisons where appropriate.
- Slices Domain, artifact type, difficulty, failure tags.
- Prompt freeze Canonical template in
src/controlsift/prompting/templates.py. Don't retune from test failures after lock.
Sealed v1.1 test hash
test.jsonl sha2569df94d33d042f3ad9d110f2115f53f470db0457930c21d07771ab59655f92e57
Full seal object: PROTOCOL_SEAL.json (mirrored from governance).