How ControlSift was tested

ControlSift is a research experiment, not a chatbot product. It creates synthetic examples of cybersecurity evidence, keeps training and test cases separated, and records the exact dataset and experiment versions so the claims can be checked.

Research question Can a small language model help classify the quality of cybersecurity evidence, and does giving it examples or fine-tuning it improve the result?

What the model has to decide

Each evidence example receives one of five labels. The label is the main thing being scored; any written explanation from the model is secondary.

SUFFICIENTPARTIALINSUFFICIENTIRRELEVANTCONTRADICTORY

The synthetic benchmark

Train 1000Validation 200Test 200Challenge 100Seed 42

Approaches tested

The experiment starts simple and adds complexity one step at a time.

  1. B0Majority class · v1.1.0complete
  2. B1TF-IDF + logistic regression · v1.1.0complete
  3. B2Gemma 3 1B zero-shot · v1.0.0complete
  4. B3Gemma 3 1B few-shot · v1.0.0complete
  5. M4Gemma 3 1B + QLoRA · v1.0.0complete

In plain English: the Gemma runs compare asking the model with no examples, giving it a few worked examples first, and fine-tuning it. Among those v1.0 runs, few-shot performed best.

Technical detail: QLoRA defaults were 4-bit NF4, double quant, LoRA r=16 / alpha=16, max length 512. Target modules were resolved from the loaded Gemma architecture rather than copied blindly. Few-shot test macro F1 was about 0.137; QLoRA was about 0.083, with test label-parse success around 0.435.

How performance is measured

What this project does not do

ControlSift isn't a live auditor, a production compliance engine, or a system that reads real evidence binders. It doesn't use OCR, retrieval over customer documents, or tuning against the sealed test set, and it doesn't publish invented metrics.