Problem and impact
Research problem, aim, stakeholders, evidence base, and intended outcomes.
Problem statement
Cybersecurity and GRC teams must distinguish artifacts that demonstrate control operation from material that's incomplete, irrelevant, stale, or contradictory. Treating a policy document as proof can create false assurance because relevance isn't the same thing as evidence of operation.
NIST SP 800-53A Rev. 5 grounds the assessment problem in verifying implementation and outcomes. NIST Small Business Cybersecurity guidance also recognizes real cybersecurity resource and budget constraints among smaller organizations. ControlSift studies whether low-cost evidence-triage methods can help without replacing human judgment.
Research aim
Evaluate a five-class evidence-quality task using a low-cost classical baseline and small-language-model methods, while publishing failures and research boundaries rather than forcing an AI-success narrative.
Stakeholders
- Intended beneficiaries: human evidence reviewers and learners who need low-cost triage research, not autopilots.
- Readers: ML and responsible-AI reviewers, hiring managers, and residency evaluators.
- Protected parties: organizations that must not treat synthetic model scores as audit conclusions.
Intended outcomes
The project delivers an open research artifact, a transparent experiment record, failure analysis, assurance documentation, and a measured answer to what the small-model experiments did and didn't show. For the MMC rubric the project uses option 4, Reduced Inequalities; the official United Nations designation is SDG 10.
Out of scope
- Automated pass/fail decisions on customer controls without human review
- Claims that synthetic-benchmark scores certify production control effectiveness
- Claims that v1.1 classical results definitively beat v1.0 Gemma results on the same benchmark
- Additional GPU reruns solely to manufacture a cleaner capstone leaderboard