ControlSift: Reducing Inequality in Evidence Assurance with Small-Model Triage Research
1. Problem statement and objectives
Cybersecurity assurance depends on evidence that controls operate as intended. A policy that requires MFA is relevant to an access control but doesn't, by itself, prove that privileged accounts were enrolled during the assessment period. NIST SP 800-53A Rev. 5 frames control assessment around implementation and intended outcomes; PCAOB AS 1105 distinguishes evidence quantity from evidence quality and defines quality through relevance and reliability; and The IIA's Global Internal Audit Standards require information used for analysis to be relevant, reliable, and sufficient. Together, those sources support the project's core distinction between proof and merely related paperwork without turning ControlSift's five labels into a standards taxonomy.
Organizations with limited cybersecurity budgets also face uneven access to specialist capacity. NIST's Small Business Cybersecurity resources explicitly describe resource, budget, and staffing constraints. ControlSift studies a narrow response to that access problem: whether an openly reproducible small-model method can assist human evidence triage without lowering the assurance standard.
Objective: evaluate five-class evidence-quality classification - SUFFICIENT, PARTIAL, INSUFFICIENT, IRRELEVANT, CONTRADICTORY - using classical and small-language-model approaches on controlled synthetic data. Scope: research and education only; no production auditor, no private customer evidence, and human judgment remains authoritative.
2. Research and analysis
The external research set was deliberately broadened beyond one framework. Three assurance sources - NIST SP 800-53A, PCAOB AS 1105, and The IIA Global Internal Audit Standards - ground evidence quality and assessment. A peer-reviewed cybersecurity systematic review by Yang et al. maps current LLM use and unresolved evaluation challenges, while Kokina et al.'s field study of AI adoption in auditing documents practical NLP/document-analysis use alongside reliability, explainability, governance, and overreliance concerns. Original LoRA and QLoRA papers ground parameter-efficient adaptation; Google DeepMind's Gemma 3 documentation grounds model context; and the NIST AI RMF plus EU AI Act Article 14 ground explicit oversight, limitations, monitoring, and meaningful human override. These sources motivate the problem and method; they don't establish ControlSift's performance. See the ten-source annotated research set for claim-by-claim boundaries.
Core source families: NIST 800-53A · PCAOB AS 1105 · IIA Standards · LLMs & cybersecurity SLR · AI in auditing field study · LoRA · QLoRA · Gemma 3 · NIST AI RMF · EU AI Act Art. 14.
Modeling data is synthetic by design. The hardened classical benchmark uses dataset v1.1.0: approximately 1,500 compositional evidence packets with family-isolated train, validation, test, and challenge splits. An early generator allowed TF-IDF to saturate; v1.1 was hardened before the classical protocol seal.
The Gemma zero-shot, few-shot, and QLoRA runs were completed earlier on dataset v1.0.0. Free-tier compute constraints make a full rerun impractical for this capstone, so cross-version scores are published transparently but are descriptive rather than a controlled head-to-head comparison.
3. Solution development
The research compares a cheap classical baseline with prompted and adapted small-model methods while treating evidence quality as a compositional assurance problem rather than generic document relevance. The public artifact includes machine-readable results, a Failure Lab, Data Card, Model Card, AI risk register, intended-use boundaries, curriculum map, reproducibility instructions, and annotated research-source claim boundaries.
Practicality and sustainability: synthetic data avoids customer-data risk, the classical baseline is inexpensive, and Gemma experiments can run on free or commodity GPU infrastructure. The open method can scale as a research and training artifact. Production use would require real-world validation, stronger constrained output handling, security review, governance, and accountable human oversight.
4. Implementation plan and risk management
Completed work includes problem framing and SDG mapping; benchmark generation and hardening; classical baselines; Gemma zero-shot and few-shot runs; QLoRA training and evaluation; error analysis; public site; diversified research sources; report; slides; final narrated presentation; and submission packaging. The capstone artifact set is complete.
Resources included Python, scikit-learn, Gemma 3 1B, Hugging Face access, Kaggle/Colab free GPU paths, GitHub Pages, and the public repository. Risks addressed include dataset leakage, lexical shortcuts, test-set tuning, fabricated metrics, secret exposure, synthetic-to-real overgeneralization, output parsing failure, cross-version overclaiming, and users mistaking research triage for automated audit. Mitigations include family isolation, protocol seals, CI checks, machine-readable metrics, platform secrets, explicit limitations, Failure Lab publication, version-boundary disclosure, and mandatory human review.
5. Outcomes and recommendations
- Published reproducible classical and Gemma experiment records with explicit dataset-version boundaries.
- Hardened an initially trivial benchmark until the v1.1 lexical baseline fell from saturation to about 0.53 macro F1.
- Preserved a negative Gemma finding: on v1.0, QLoRA didn't beat few-shot and showed weak output-format reliability.
- Published failure analysis and assurance artifacts rather than reporting only aggregate scores.
- Grounded the work in assurance standards, peer-reviewed cyber/audit research, primary PEFT literature, model documentation, and independent human-oversight sources.
Recommendation: don't deploy ControlSift as an audit autopilot. Preserve human authority; improve constrained label generation before further model claims; validate on governed real-world evidence before discussing operational use; and never treat the v1.1 classical and v1.0 Gemma scores as a controlled same-dataset comparison.
6. Claim boundaries and authorship
External literature supports context and technical choices; ControlSift-specific result claims come only from committed experiment files. Labels are rule-derived on synthetic data, not a large human gold standard. Assistive tools were used for drafting and engineering support; problem selection, experimental design, benchmark decisions, execution, interpretation, and final submission responsibility remain with the author.