How the research keeps itself honest

This section documents what ControlSift is allowed to claim, how the data and tests are controlled, where the project can fail, and why a human reviewer remains in charge. The repository files under governance/ remain the source of truth; these pages make that material easier to read.

What ControlSift is not ControlSift is a research benchmark. It isn't an automated auditor, a compliance engine, or a replacement for professional judgment.

Integrity checklist

The main safeguards and research controls currently in place.

secured

Related scenarios can't leak across splits

Variations of one scenario family stay in only one of train, validation, test, or challenge. Automated checks fail if that boundary is broken.

secured

The benchmark was hardened against word shortcuts

An earlier synthetic generator made the task too easy for a simple text model. Version 1.1 was hardened, and TF-IDF now scores about 0.53 macro F1 instead of saturating the benchmark.

secured

Public scores must come from real result files

The site reads published experiment artifacts. Missing runs stay missing rather than being filled with placeholder numbers.

secured

Final test rules were frozen before final claims

The hypotheses, main metric, prompt, and test-set identity were recorded before the final interpretation was written.

secured

The expected labels were checked

The challenge set received a full structural audit plus a 20-case narrative spot-check. The project still does not claim a large, independently human-labeled gold standard.

secured

Gemma and QLoRA results are published

Zero-shot, few-shot, and QLoRA runs are complete for benchmark v1.0.0. Their scores and failures are published with the benchmark-version boundary clearly disclosed.

Assurance pages