Reflection
What changed during the project, what the completed experiments taught, and which limits matter most.
What changed
- The benchmark had to earn its difficulty. An early generator let TF-IDF reach perfect macro F1. Hardening the classical benchmark to v1.1 reduced that shortcut and made the task materially less trivial.
- Fine-tuning wasn't automatically better. Within the Gemma v1.0 runs, few-shot prompting beat QLoRA.
- Output reliability became a model problem. QLoRA label parse success fell to about 0.435, making formatting failure a first-class result rather than a cleanup nuisance.
- Research boundaries matter as much as scores. The completed classical experiments use v1.1 while the Gemma experiments use v1.0, so cross-version numbers are descriptive rather than a controlled leaderboard.
What I learned
The most useful lesson wasn't that a particular model won. It was that an apparently successful experiment can be undermined by shortcut features, output-contract failure, or a dataset-version mismatch if the evidence trail isn't checked carefully. The project became stronger each time the process exposed an inconvenient result and kept it visible.
That's also the responsible-AI lesson: evaluation isn't a performance screenshot. It's a claim supported by a bounded dataset, protocol, metrics, failure analysis, and explicit limits on what the result can justify.
Remaining limitations
- Synthetic packets aren't real multi-document audit binders or OCR-heavy evidence collections.
- Labels are rule-derived rather than a large human gold standard.
- The Gemma v1.0 and classical v1.1 experiments can't support a controlled cross-family ranking.
- No live deployment or production evidence study was attempted.
Where I'd take it next
- Improve constrained decoding and label-output reliability before additional language-model claims.
- Use a single dataset version for any future classical-vs-LLM comparison.
- Validate carefully on governed real-world evidence only if a legitimate research context and privacy controls exist.
The MMC capstone itself is complete: research, documentation, presentation deck, final video, and supporting evidence are finished. The ideas above are intentionally future research, not unfinished capstone work.