Data Card
Synthetic cybersecurity control-evidence benchmark. Not production audit evidence.
synthetic_rule_based)Motivation
Support research on whether PEFT improves small-LM judgment of evidence sufficiency and relevance. The dataset isn't intended to train a production auditor.
Composition
Each case includes control statement, environment, evidence type, evidence text, label, failure tags, difficulty, scenario family id, and split. Hidden analytical metadata never enters model prompts.
SUFFICIENT
PARTIAL
INSUFFICIENT
IRRELEVANT
CONTRADICTORY
Ground truth
Important
Primary labels are rule-derived inside the generator. They aren't fully human-labeled.
- Structural audit 100% challenge + stratified ~10% development sample.
- Scripted spot-check 20 challenge cases (4 per label) - not a human GRC review.
- Tracking
data/review_log.csv- not fully human-labeled gold. - Research surface Compositional packets (v1.1); see
governance/RESEARCH_SURFACE.md.
Split hashes
| Split | n | SHA-256 |
|---|---|---|
| train | 1000 | 53b5c057f80f6ae5bd26c63dc2a03c7de96641b73094b4b464ff80ab39438d99 |
| validation | 200 | d14f9bbe0c7d567c12f1a4f1d624bd0e47b1d12d0a6a3d916f9c6b95692ceef8 |
| test | 200 | 7476a67ec679b4e0198709a4fb812be43d43364b78d525dacee4ab6c1dab31d1 |
| challenge | 100 | 735cd57b874ae29e54f74e99c512df0cc0039bc5e0bd85dbf4f21f882b057cfe |
Risks and limitations
- Synthetic-to-real gap Generator assumptions may not match real audit evidence.
- Schema-reading risk Strong scores may reflect packet grammar, not naturalistic judgment.
- Lexical baseline TF-IDF ~0.53 test macro F1 after harden; still a required comparison.
- Label assumptions Rule-derived labels encode generator mutation logic.