Related scenarios can't leak across splits
Variations of one scenario family stay in only one of train, validation, test, or challenge. Automated checks fail if that boundary is broken.
This section documents what ControlSift is allowed to claim, how the data and tests are controlled, where the project can fail, and why a human reviewer remains in charge. The repository files under governance/ remain the source of truth; these pages make that material easier to read.
The main safeguards and research controls currently in place.
Variations of one scenario family stay in only one of train, validation, test, or challenge. Automated checks fail if that boundary is broken.
An earlier synthetic generator made the task too easy for a simple text model. Version 1.1 was hardened, and TF-IDF now scores about 0.53 macro F1 instead of saturating the benchmark.
The site reads published experiment artifacts. Missing runs stay missing rather than being filled with placeholder numbers.
The hypotheses, main metric, prompt, and test-set identity were recorded before the final interpretation was written.
The challenge set received a full structural audit plus a 20-case narrative spot-check. The project still does not claim a large, independently human-labeled gold standard.
Zero-shot, few-shot, and QLoRA runs are complete for benchmark v1.0.0. Their scores and failures are published with the benchmark-version boundary clearly disclosed.