Limitations
Where ControlSift shouldn't be trusted, even after Gemma results land.
- Synthetic data Generator assumptions may not match real audit evidence distributions.
- Rule-derived labels Not fully human-labeled gold. Challenge uses structural audit + a 20-case spot-check.
- Compositional research surface (v1.1) Dual sections, decoys, and
SCOPEscaffolding mean strong scores may reflect schema reading, not naturalistic binder judgment. - Uneven class difficulty Classical CONTRADICTORY remains easier than SUFFICIENT / IRRELEVANT.
- Frozen prompt Zero-shot doesn't tutor the packet grammar; discovery failures are valid outcomes.
- Single small model Gemma 3 1B results don't automatically transfer to other sizes or families.
- Single-artifact cases v1 isn't a multi-document evidence binder task.
- No live inference product Static research artifact only.
- Parse failures Malformed outputs become
UNPARSEABLEand count against performance.