Research sources & claim boundaries

ControlSift's model scores come from its own sealed synthetic benchmark. The external literature serves a different purpose: grounding what makes evidence persuasive, showing where AI and language models are already being studied in cybersecurity and audit-adjacent work, justifying the small-model and PEFT choices, and defining responsible-use and human-oversight boundaries.

Research-integrity rule No outside source is used to imply that ControlSift works in production. ControlSift-specific performance claims come only from results/* and docs/data/results.json. External literature supports problem framing, method selection, and governance boundaries - not the experiment's outcome.
3 assurance sourcesNIST, PCAOB, and The IIA on assessment and evidence quality.
2 domain-research sourcesPeer-reviewed cybersecurity and audit literature on AI/LLM use and limitations.
3 technical sourcesOriginal LoRA, QLoRA, and Gemma 3 method/model documentation.
2 oversight sourcesNIST AI RMF and the EU AI Act on governance and human oversight.

The ten sources below are the project's core research set. Small-business access and UN SDG materials are retained separately as capstone-program context rather than counted as technical evidence.

Evidence quality and assurance

01 · NIST · Control assessment

NIST SP 800-53A Rev. 5 - Assessing Security and Privacy Controls

Supports: Security and privacy control assessment is a structured process for determining whether controls are implemented, operating as intended, and producing desired outcomes. Assessment procedures are tied to objectives and can be tailored to context.

Doesn't prove: that a language model can perform control assessment, or that ControlSift's five labels are a NIST taxonomy.

NIST publication · DOI

02 · PCAOB · Audit evidence quality

PCAOB AS 1105 - Audit Evidence

Supports: Evidence quality isn't simply document presence. PCAOB distinguishes sufficiency (quantity) from appropriateness (quality), and defines appropriateness through relevance and reliability. The standard also notes that more low-quality evidence doesn't compensate for poor evidence quality and that contradictory or unreliable evidence requires further work.

Doesn't make: ControlSift a financial-statement audit tool. AS 1105 is used as an audit-adjacent evidence-quality analogue for why relevance, reliability, contradiction, and sufficiency are distinct concepts.

PCAOB AS 1105

03 · The IIA · Internal audit evidence

Global Internal Audit Standards - Standard 14.1

Supports: Internal auditors are expected to gather information that's relevant, reliable, and sufficient, apply professional skepticism to reliability, and obtain additional information when evidence can't support a reasonable basis for findings and conclusions.

Doesn't prescribe: ControlSift's label set or scoring scheme. It strengthens the broader assurance premise that evidence quality has multiple dimensions and still requires professional judgment.

The IIA · Complete Global Internal Audit Standards

AI and LLMs in cybersecurity and audit-adjacent work

04 · Peer-reviewed review · Cybersecurity

Yang et al. - When LLMs Meet Cybersecurity: A Systematic Literature Review

Supports: LLMs are being studied across a broad range of cybersecurity scenarios, while the literature also identifies continuing challenges around model construction, task fit, evaluation, reliability, security, and responsible deployment. That makes bounded empirical evaluation more defensible than assuming general language capability transfers cleanly to assurance tasks.

Doesn't establish: that LLMs are effective at cybersecurity control-evidence classification. ControlSift evaluates a much narrower task.

Yang et al. (2025) · Cybersecurity · DOI: 10.1186/s42400-025-00361-w

05 · Peer-reviewed field study · Auditing

Kokina et al. - Challenges and Opportunities for Artificial Intelligence in Auditing: Evidence from the Field

Supports: Interviews with experienced audit professionals found real use of simpler AI/NLP for document extraction and anomaly-oriented support while identifying transparency, explainability, bias, privacy, robustness, reliability, overreliance, and governance as persistent adoption challenges. The study describes AI as supporting audit work rather than eliminating the need for human auditors.

Doesn't prove: that ControlSift's approach is audit-ready. It provides empirical audit-domain context for assistive AI, document analysis, and the importance of human oversight.

Kokina, Blanchette, Davenport & Pachamanova (2025) · International Journal of Accounting Information Systems · DOI: 10.1016/j.accinf.2025.100734

Small-model and parameter-efficient adaptation context

06 · Primary technical paper · PEFT

Hu et al. - LoRA: Low-Rank Adaptation of Large Language Models

Supports: LoRA adapts pretrained models by freezing the original weights and training low-rank matrices, substantially reducing the number of trainable parameters and memory requirements compared with full fine-tuning.

Doesn't imply: that parameter efficiency guarantees better task performance. It explains the adapter mechanism that QLoRA builds on.

Hu et al. (2021/2022) · arXiv:2106.09685

07 · Primary technical paper · Quantized PEFT

Dettmers et al. - QLoRA: Efficient Finetuning of Quantized LLMs

Supports: The QLoRA method used in ControlSift: gradients are propagated through a frozen 4-bit quantized model into LoRA adapters, reducing memory requirements for fine-tuning. The paper introduces NF4 and related memory-saving techniques.

Doesn't prove: that QLoRA should improve this cybersecurity evidence task. ControlSift tests that hypothesis and preserves the negative result when it doesn't.

Dettmers, Pagnoni, Holtzman & Zettlemoyer (2023) · arXiv:2305.14314

08 · Google DeepMind · Model documentation

Gemma 3 Model Card and Technical Report

Supports: Gemma 3 includes a lightweight 1B text model intended for comparatively constrained deployment settings. The model documentation provides architecture, evaluation, intended-use, and limitations context for the base model selected in ControlSift.

Doesn't prove: domain suitability for cybersecurity evidence assessment. General benchmark performance isn't substituted for ControlSift's task-specific results.

Google model card · Gemma 3 Technical Report

Responsible AI and human oversight

09 · NIST · AI risk management

NIST AI Risk Management Framework 1.0

Supports: NIST calls for documented roles and responsibilities, explicit human-AI configurations, scoped use, test and evaluation, ongoing risk management, and defined human oversight. The framework also emphasizes documenting test sets, metrics, system limitations, and decision authority.

Doesn't prescribe: ControlSift's exact governance files. The Data Card, Model Card, risk register, intended-use boundary, Failure Lab, and human-review requirement are project-level implementations informed by those principles.

NIST AI RMF 1.0 · AI RMF Core

10 · European Union · Human oversight

EU AI Act - Article 14, Human Oversight

Supports: For systems within the Article's scope, effective human oversight includes understanding system capabilities and limitations, monitoring for unexpected behavior, remaining alert to automation bias, correctly interpreting output, and being able to disregard, override, reverse, or stop system output when appropriate.

Doesn't classify: ControlSift as a high-risk AI system under the EU AI Act. Article 14 is used as a strong external reference for what meaningful human oversight can require, not as a legal-status claim about this research project.

Regulation (EU) 2024/1689 · Article 14

Capstone-program context

These sources explain the access and SDG framing used by the capstone. They're intentionally separated from the ten-source technical and assurance research set above.

Context · NIST · Resource constraints

NIST Small Business Cybersecurity Corner

Supports: NIST explicitly recognizes that smaller businesses may operate with limited cybersecurity resources, budgets, and staffing and need practical, actionable, cost-conscious support.

Doesn't prove: a specific shortage of evidence reviewers or justify replacing specialists with AI. ControlSift uses the broader access constraint only to motivate low-cost, human-supervised research.

NIST Small Business Cybersecurity Corner

Context · United Nations · SDG

UN Sustainable Development Goal 10 - Reduced Inequalities

Supports: The official UN designation: Reduced Inequalities is SDG 10. ControlSift maps its low-cost, open research design to the capstone program's Reduced Inequalities option.

Doesn't claim: that the UN specifically identifies cybersecurity evidence review as an SDG 10 target. The assurance-access connection is the project's capstone interpretation.

United Nations Goal 10

How the literature connects to ControlSift

Evidence-quality problem
NIST, PCAOB, and The IIA converge on a useful idea: assurance depends on evidence that's fit for an objective, reliable enough to support a conclusion, and sufficient for the decision being made. That's the conceptual foundation for distinguishing proof from merely related paperwork.
AI-in-domain rationale
Cybersecurity and audit research show active use of AI and NLP for analysis and document-oriented support, while also documenting reliability, explainability, governance, and overreliance concerns. ControlSift therefore studies assistive classification rather than autonomous audit.
Technical design
LoRA and QLoRA ground parameter-efficient adaptation; Gemma documentation grounds the lightweight model choice. None of those sources are treated as evidence that fine-tuning must improve performance.
Responsible use
NIST AI RMF and EU human-oversight requirements support explicit scope, documented limitations, monitoring, meaningful override authority, and resistance to automation bias. ControlSift keeps human judgment authoritative.
Outcome evidence
Only ControlSift's sealed experiment files support statements such as “TF-IDF reaches 0.5327 macro F1 on v1.1” or “few-shot is strongest within the Gemma v1.0 ladder.” External literature never substitutes for experiment evidence.