How ControlSift was tested
ControlSift is a research experiment, not a chatbot product. It creates synthetic examples of cybersecurity evidence, keeps training and test cases separated, and records the exact dataset and experiment versions so the claims can be checked.
What the model has to decide
Each evidence example receives one of five labels. The label is the main thing being scored; any written explanation from the model is secondary.
The synthetic benchmark
- No employer or customer evidence is used. The examples are generated from controlled cybersecurity scenarios.
- Related scenarios stay together. Variations of the same scenario family are kept in only one split, so the model can't train on one version and then be tested on a close cousin.
- The expected labels come from rules. They were checked for structural consistency, and challenge cases received a 20-case narrative spot-check. This is not a large benchmark independently labeled by many human experts.
- The benchmark changed during the research. The hardened traditional-model benchmark is v1.1.0. The completed Gemma experiments use v1.0.0. Results are comparable within a version; cross-version scores are published only as research history.
Approaches tested
The experiment starts simple and adds complexity one step at a time.
- B0Majority class · v1.1.0complete
- B1TF-IDF + logistic regression · v1.1.0complete
- B2Gemma 3 1B zero-shot · v1.0.0complete
- B3Gemma 3 1B few-shot · v1.0.0complete
- M4Gemma 3 1B + QLoRA · v1.0.0complete
In plain English: the Gemma runs compare asking the model with no examples, giving it a few worked examples first, and fine-tuning it. Among those v1.0 runs, few-shot performed best.
Technical detail: QLoRA defaults were 4-bit NF4, double quant, LoRA r=16 / alpha=16, max length 512. Target modules were resolved from the loaded Gemma architecture rather than copied blindly. Few-shot test macro F1 was about 0.137; QLoRA was about 0.083, with test label-parse success around 0.435.
How performance is measured
- Main score: macro F1. This 0–1 score gives each of the five evidence labels equal weight. The project also reports accuracy, precision, recall, per-label scores, and confusion matrices.
- Malformed answers count as failures. If the model doesn't return a usable label, the result becomes
UNPARSEABLE. The evaluator doesn't quietly repair the answer. - Results are examined from several angles. The analysis includes uncertainty ranges and performance by control domain, artifact type, difficulty, and failure tag.
- The final test rules were frozen before final claims. The hypotheses, main metric, prompt, and test hash are recorded in the Protocol and Assurance pages.
What this project does not do
ControlSift isn't a live auditor, a production compliance engine, or a system that reads real evidence binders. It doesn't use OCR, retrieval over customer documents, or tuning against the sealed test set, and it doesn't publish invented metrics.