Can a Small Language Model Tell Proof from Paperwork?

Cybersecurity control-evidence classification with classical baselines, Gemma 3 experiments, failure analysis, and explicit research boundaries.

ControlSift Capstone Research Paper
MMC rubric option 4: Reduced Inequalities / official UN SDG 10
Primary metric: macro F1 · Seed 42 · Status: research and presentation complete

Abstract

Distinguishing operational proof from administrative paperwork is a recurring problem in cybersecurity assurance. ControlSift studies whether low-cost machine-learning methods can classify synthetic control evidence as SUFFICIENT, PARTIAL, INSUFFICIENT, IRRELEVANT, or CONTRADICTORY. The project combines a hardened classical benchmark, Gemma 3 prompting and QLoRA experiments, challenge evaluation, failure analysis, and responsible-use documentation. Classical results were produced on dataset v1.1.0, while Gemma zero-shot, few-shot, and QLoRA results were produced on dataset v1.0.0. Because those versions differ, cross-version scores are published for transparency but are descriptive rather than a controlled head-to-head comparison. Within Gemma v1.0, few-shot is strongest and QLoRA doesn't beat few-shot; QLoRA also shows poor label-parse reliability. The project is a bounded synthetic research artifact, not a production auditor.

1. Introduction

Security and governance, risk, and compliance programs depend on evidence that controls operate as intended. In practice, policies, screenshots, reports, exports, tickets, and attestations can all be filed under a control even when they don't prove operation. A policy requiring MFA is relevant to an access-control objective, for example, but it doesn't establish that privileged accounts were enrolled during the assessment period. An operational export can provide stronger proof, while an export showing a failed control condition can contradict the control itself.

ControlSift asks whether low-cost machine-learning methods can help classify those evidence-quality distinctions on controlled synthetic data. The project emphasizes research integrity over product claims: explicit dataset versions, held-out evaluation, negative-result reporting, machine-readable metrics, failure analysis, and a clear boundary between research triage and audit judgment.

2. Background and external grounding

The evidence-quality problem is grounded in three complementary assurance sources. NIST SP 800-53A Rev. 5 ties control assessment to implementation and intended outcomes. PCAOB AS 1105 distinguishes evidence quantity from quality and defines appropriateness through relevance and reliability. The IIA Global Internal Audit Standards require information supporting analysis to be relevant, reliable, and sufficient. ControlSift doesn't adopt any one source as its label taxonomy; together they support the broader premise that evidence quality is multidimensional and tied to an assessment objective.

The domain-use rationale extends beyond standards. Yang et al.'s peer-reviewed systematic review of LLMs in cybersecurity documents broad experimentation alongside unresolved evaluation and reliability challenges. Kokina et al.'s field study of AI in auditing reports practical use of AI/NLP for document-oriented and analytical support while identifying explainability, robustness, reliability, governance, and overreliance concerns. Those findings support studying assistive evidence triage while keeping human judgment authoritative.

The technical method is grounded in Hu et al.'s LoRA paper and Dettmers et al.'s QLoRA paper; model context comes from the Gemma 3 Model Card. Responsible-use boundaries draw on the NIST AI RMF and the EU AI Act's Article 14 human-oversight provisions. The access-to-capacity motivation remains grounded in NIST Small Business Cybersecurity, and the societal-goal mapping uses the United Nations' official SDG 10 page.

The Research Sources page provides a ten-source core research set with a separate "supports / doesn't prove" boundary for every source.

3. Problem formulation

Each case contains a control statement, environment, evidence type, evidence text, and one of five labels:

  • SUFFICIENT: complete, timely, on-objective operational proof.
  • PARTIAL: relevant operational evidence with material coverage or scope gaps.
  • INSUFFICIENT: on-topic material that doesn't establish execution.
  • IRRELEVANT: evidence that doesn't address the control objective.
  • CONTRADICTORY: evidence containing a fact that conflicts with the claimed control state.

The primary question is whether parameter-efficient adaptation of Gemma 3 1B can improve evidence-quality classification. Secondary questions examine prompting, output reliability, challenge generalization, lexical shortcuts, and difficult class boundaries. Macro F1 is the primary metric so all five classes contribute equally to the headline score.

4. Societal alignment

For the MMC rubric, ControlSift uses option 4: Reduced Inequalities. The official United Nations designation is SDG 10: Reduced Inequalities. The project connects that goal to unequal access to cybersecurity assurance capacity. The intended contribution is an open research and training artifact that can be inspected and reproduced at low cost, not an automated compliance substitute.

5. Benchmark design

The hardened classical benchmark contains approximately 1,500 synthetic evidence packets spanning multiple control domains. Dataset v1.1.0 uses family-isolated train, validation, test, and challenge splits of roughly 1000 / 200 / 200 / 100. Labels are rule-derived through controlled mutations rather than produced by a large human annotation program. Synthetic organizations and evidence avoid employer, customer, or patient data.

An early generator allowed TF-IDF to saturate at 1.0, revealing a lexical shortcut problem. The packet structure was hardened before sealing the v1.1 classical protocol, adding shared decoys, compositional scope cues, substance pointers, and row-level conflicts. The resulting TF-IDF score fell to approximately 0.53 macro F1, indicating a materially less trivial lexical task.

6. Methods

6.1 Classical path - dataset v1.1.0

The classical path includes a majority-class baseline and TF-IDF with logistic regression. These runs use the hardened v1.1 dataset and establish a low-cost lexical reference point.

6.2 Gemma path - dataset v1.0.0

The Gemma path uses google/gemma-3-1b-it with zero-shot prompting, fixed few-shot examples, and QLoRA adaptation. QLoRA uses 4-bit NF4 quantization and LoRA rank 16. Malformed generations are recorded as UNPARSEABLE and count against performance rather than being silently repaired.

6.3 Version boundary

The Gemma runs were completed on dataset v1.0.0 before the v1.1 classical hardening. Free-tier GPU constraints make a complete v1.1 rerun impractical for this capstone. Therefore, comparisons within the Gemma v1.0 ladder are controlled, and comparisons within the classical v1.1 path are controlled. Cross-version classical-vs-Gemma scores are shown only as descriptive context.

7. Results

7.1 Classical v1.1.0

ExperimentTest macro F1AccuracyChallenge macro F1
Majority0.06670.2000.0667
TF-IDF + LR0.53270.5450.5237

7.2 Gemma v1.0.0

ExperimentTest macro F1AccuracyChallenge macro F1Parse success
Gemma zero-shot0.07960.1400.10550.760
Gemma few-shot0.13650.2000.17410.985
Gemma QLoRA0.08270.0750.13090.435

The controlled Gemma conclusion is straightforward: few-shot is the strongest Gemma v1.0 run, and QLoRA doesn't beat few-shot. The low QLoRA parse-success rate shows that output-format reliability became a major failure mode after adaptation.

The v1.1 classical scores are numerically higher than the v1.0 Gemma scores, but the project does not interpret that difference as proof that TF-IDF beats Gemma on the same benchmark because the underlying dataset versions differ.

TF-IDF learning curve on the v1.1 classical benchmark
TF-IDF learning curve on the hardened v1.1 classical benchmark.

8. Failure analysis

The public Failure Lab treats errors as first-class research artifacts. Classical mistakes cluster around compositional coverage and evidence-substance distinctions, while the adapted Gemma run adds a different problem: many generations fail to produce a valid label at all. That parsing collapse isn't hidden or post-processed away.

9. Responsible innovation

ControlSift isn't an automated auditor, certification engine, or replacement for control owners or human assessors. Intended use is method comparison, education, and study of assistive evidence triage under synthetic conditions. The project publishes a Data Card, Model Card, AI risk register, intended-use policy, limitations, protocol information, and failure examples. Human judgment remains authoritative.

Practicality and sustainability come from low-cost components: synthetic data, an inexpensive classical baseline, a small language model, and free or commodity GPU infrastructure. Scaling into real operations would require governed real-world validation, stronger constrained output handling, security review, ongoing monitoring, and accountable human oversight. That boundary is consistent with the audit literature's warnings about reliability and overreliance, the NIST AI RMF's emphasis on defined human-AI roles, and EU Article 14's focus on understanding limitations, monitoring behavior, resisting automation bias, and retaining meaningful override authority.

10. Limitations

  • Synthetic packets aren't real multi-document audit binders.
  • Labels are rule-derived, not a large-scale human gold standard.
  • Compositional scaffolding may reward schema-reading behavior.
  • Gemma experiments use v1.0 while classical experiments use v1.1, preventing a controlled cross-family comparison.
  • QLoRA parse success is low, limiting interpretation of its score.
  • Results for Gemma 3 1B don't automatically generalize to other models or scales.
  • No production-auditor claim is supported by this work.

11. Reproducibility

pip install -e ".[dev]"
python scripts/validate_dataset.py
python scripts/run_tfidf_baseline.py
python scripts/run_evaluation.py
pytest

Committed experiment files under results/ feed the public results JSON and site. Secrets remain in platform secret stores rather than the repository.

12. Conclusion

ControlSift demonstrates an evidence-first applied AI research process rather than a product-success narrative. The project caught a lexical shortcut, hardened its classical benchmark, ran a small-model prompting and adaptation ladder, preserved an unsuccessful QLoRA result, published parsing failures, and disclosed a later dataset-version mismatch rather than overstating a cross-family comparison. The strongest controlled model conclusion is that few-shot prompting outperformed QLoRA within the Gemma v1.0 experiments. The larger contribution is the research process: explicit limits, reproducible receipts, failure visibility, and human-centered assurance boundaries.

References and accompanying artifacts