MMC / Google DeepMind AI Research Foundations

ControlSift capstone hub

The complete project in one place: problem, external research, methods, experiment results, implementation, responsible-AI work, program alignment, report, slides, final presentation, reflection, weekly work, and submission package.

Plain-English project: ControlSift studies whether a small language model can distinguish strong cybersecurity control evidence from paperwork that's incomplete, irrelevant, or contradictory. The work includes a hardened synthetic benchmark, classical baselines, prompted Gemma runs, QLoRA, failure analysis, and research-governance artifacts.

Research-integrity noteClassical metrics use dataset v1.1.0; completed Gemma metrics use v1.0.0. The experiments are real and published, but cross-version scores are descriptive rather than a strict same-benchmark leaderboard.
Research artifactComplete
Published runs5 experiment rungs
Classical benchmarkv1.1.0
Gemma receiptsv1.0.0
Program alignmentMMC option 4 / UN SDG 10
Capstone packageComplete

Narrated capstone video

The completed eight-slide research presentation is hosted as the v1.0-capstone release asset.

What the grader needs

Every formal deliverable and administrative checkpoint is visible here.

RPT

8-page-or-less report

Rubric-facing report: problem, external research, analysis, solution, implementation, outcomes, recommendations, and claim boundaries.

Primary written submission
SLD

8-slide presentation

Exactly eight slides aligned to the final research story and version disclosure.

Presentation artifact
SUB

Submission package

Final bundle map for report, slides, video, links, and supporting evidence.

Why the experiment exists and what it found

Start with the problem, then inspect sources, methods, full paper, results, and failure cases.

SRC

Research sources

Annotated NIST, Google DeepMind, academic, and UN sources with claim boundaries.

MTH

Methods and evidence

Benchmark design, experiment paths, evidence handling, and integrity controls.

PPR

Research paper

Full technical narrative with final v1.0/v1.1 disclosure and references.

RES

Experiment ledger

Machine-readable metrics surfaced with the version boundary clearly labeled.

FAIL

Failure Lab

Hard boundaries and disagreement patterns hidden by aggregate scores.

How the work was executed and bounded

IMP

Implementation plan

Phases, compute, accounts, resources, risks, mitigations, and completed execution.

RAI

Responsible innovation

Privacy, human authority, synthetic-data boundaries, misuse risks, limitations, and non-claims.

AST

Assurance package

Protocol, Data Card, Model Card, AI risk register, intended use, limitations, and human-review requirements.

REP

Reproduce

Validation and rerun path for the public research artifact.

How ControlSift maps to the residency

MAP

Curriculum map

Maps AI Research Foundations content to concrete ControlSift work.

REF

Reflection

What changed, what the model results taught, and future work.

LAB

Weekly labs

Residency-facing weekly outputs and work trail.

Published experiment snapshot

Majority / v1.10.067
TF-IDF / v1.10.533
Gemma zero-shot / v1.00.080
Gemma few-shot / v1.00.137
Gemma QLoRA / v1.00.083
Defensible interpretationWithin Gemma v1.0, few-shot is strongest and QLoRA doesn't beat it. The hardened v1.1 classical baseline is about 0.53. A definitive classical-vs-Gemma ranking would require a matching-version experiment and is outside this capstone's scope.