ControlSift

Can a small AI tell real security proof from paperwork?

ControlSift tests whether a traditional text classifier and a small AI model can distinguish strong cybersecurity evidence from weak or misleading paperwork. It publishes what worked, what failed, and what the results do and don't show.

Synthetic study · failures kept · comparison limits disclosed · human review required

ControlSift in 60 seconds

ControlSift asks whether AI can help a reviewer separate evidence that proves a security control worked from documents that merely sound relevant.

01

The security problem

A policy can describe what should happen without proving that it did happen. A screenshot may show one account but not the full population. A system export or log can provide much stronger operational proof. ControlSift studies those differences.

02

The experiment

The project built a synthetic set of security-evidence examples and tested several approaches: a simple baseline, a traditional text classifier, and a small language model used three different ways. Partway through the work, the benchmark was made harder, so some results come from version 1.0 and others from version 1.1.

03

What we can say now

On the original test set, giving the language model a few examples worked better than either giving it no examples or fine-tuning it. On the harder v1.1 test set, the traditional classifier scored about 0.53 on the project's main metric. Because those results came from different benchmark versions, ControlSift doesn't claim an overall winner.

Jargon decoder

Benchmark
The fixed set of test cases used to measure how well each approach performs.
TF-IDF + logistic regression
A traditional machine-learning method that looks for useful word patterns. It's fast, inexpensive, and isn't a language model.
Gemma 3 1B
The small Google DeepMind language model tested in the project.
Zero-shot / few-shot
Zero-shot means asking the model with no worked examples. Few-shot means showing it a few examples first.
QLoRA
A way to fine-tune a language model while using much less computing memory than retraining the whole model.
Macro F1
A 0–1 summary score that gives each evidence label equal importance. Higher is better. It isn't the same thing as accuracy.

Explore the whole artifact

Start with the question you care about. The technical detail is there when you want it.

R

Results

All published scores, which benchmark version each score came from, and which comparisons are fair.

M

Methods

How the synthetic test cases were built, how each approach was tested, and how performance was measured.

F

Failure Lab

Specific cases where the approaches disagree or fail. Start here if the mistakes matter more to you than the average score.

A

Assurance

What the project is for, where it can fail, how the research was governed, and why a human reviewer remains in charge.

X

Reproduce

Environment details and commands for validating the data and rerunning the public experiments.

C

Capstone hub

The complete MMC / Google DeepMind AI Research Foundations submission, including the report, sources, slides, video, and supporting work.

G

GitHub repository

Source code, notebooks, machine-readable results, governance files, scripts, tests, and project history.

The problem, with one example

Suppose the control says every privileged account must use multi-factor authentication.

ControlPrivileged accounts must use multi-factor authentication.

System export

An identity-system export showing every privileged account, whether MFA is enabled, when the data was captured, and the population covered.

SUFFICIENT

Policy document

A written standard that says privileged accounts must use MFA, but provides no account-level evidence that the requirement was actually followed.

INSUFFICIENT

Published result ledger

Important comparison limitThe traditional classifier was tested on benchmark v1.1.0. The completed Gemma runs used v1.0.0. Both sets of results are real, but because the test sets differ, these numbers aren't a fair head-to-head ranking.
Primary metric macro F1Classical v1.1.0Gemma v1.0.0Seed 42
TF-IDF v1.1 test macro F10.533
Gemma few-shot v1.0 test macro F10.137
TF-IDF v1.1 challenge macro F10.524

The honest headline

Among the Gemma tests run on v1.0.0, giving the model a few examples worked better than either giving it no examples or fine-tuning it. The fine-tuned run also had trouble returning usable labels consistently. On the harder v1.1.0 benchmark, the traditional TF-IDF model scored about 0.53 macro F1. Because those results come from different benchmark versions, ControlSift doesn't claim that the traditional model beats Gemma overall. That comparison still needs a same-version rerun.

Research boundaries

Receipts

Every important claim points back to a public artifact.