Data Card

Synthetic cybersecurity control-evidence benchmark. Not production audit evidence.

Version
1.0.0
Total cases
1,500 (balanced five classes)
Splits
train 1000 / validation 200 / test 200 / challenge 100
Seed
42
Generation
deterministic rule-based (synthetic_rule_based)
Domains
13 control domains

Motivation

Support research on whether PEFT improves small-LM judgment of evidence sufficiency and relevance. The dataset isn't intended to train a production auditor.

Composition

Each case includes control statement, environment, evidence type, evidence text, label, failure tags, difficulty, scenario family id, and split. Hidden analytical metadata never enters model prompts.

SUFFICIENT PARTIAL INSUFFICIENT IRRELEVANT CONTRADICTORY

Ground truth

Important Primary labels are rule-derived inside the generator. They aren't fully human-labeled.

Split hashes

SplitnSHA-256
train100053b5c057f80f6ae5bd26c63dc2a03c7de96641b73094b4b464ff80ab39438d99
validation200d14f9bbe0c7d567c12f1a4f1d624bd0e47b1d12d0a6a3d916f9c6b95692ceef8
test2007476a67ec679b4e0198709a4fb812be43d43364b78d525dacee4ab6c1dab31d1
challenge100735cd57b874ae29e54f74e99c512df0cc0039bc5e0bd85dbf4f21f882b057cfe

Risks and limitations