Google DeepMind AI Research Foundations
Applied curriculum map showing how ControlSift demonstrates the eight-course learning path. Sources include the Google Skills path, the DeepMind collection, and the open labs.
How to read this page
This is an applied mapping, not a claim that ControlSift recreates every teaching lab exactly. "Meets" means the learning concept appears directly in the project. "Applied" means the project demonstrates the concept through a real research analogue rather than repeating the classroom exercise.
Final experiment boundaryClassical experiments use dataset v1.1.0. Gemma zero-shot, few-shot, and QLoRA use dataset v1.0.0. Cross-version scores are descriptive only.
At a glance
| Course | ControlSift demonstration | Status |
|---|---|---|
| 01 / Build Your Own Small Language Model | Problem framing, probability baseline, lexical-vs-language-model reasoning, research pipeline | meets |
| 02 / Represent Your Language Data | Dataset schema, synthetic text design, tokenization context, Data Card, privacy boundary | meets |
| 03 / Design and Train Neural Networks | Train/validation/test separation, learning curves, evaluation discipline, benchmark hardening | meets |
| 04 / Discover the Transformer Architecture | Gemma 3 model context, instruction-tuned inference, LoRA target reasoning | applied |
| 05 / Fine-Tune Your Model | Gemma 3 1B QLoRA with 4-bit NF4, fixed evaluation, negative-result reporting | meets |
| 06 / Align Your Model | Output contracts, parse reliability, intended use, human-review boundaries, risk register | applied |
| 07 / Accelerate Your Model | Free-tier GPU execution, quantization, resource constraints, practical reproducibility | meets |
| 08 / Capstone: Real-World Impact | Problem, research, solution, implementation, external sources, public artifact, report, slides, final narrated presentation | meets |
01 / Build Your Own Small Language Model
- Language-model problem: classify evidence quality rather than predict unrestricted prose.
- Probability baseline: majority-class macro F1 establishes a chance-like floor for the balanced five-class task.
- Lexical baseline: TF-IDF + logistic regression demonstrates what word-pattern methods can do before invoking a language model.
- Research pipeline: generate -> validate -> split -> evaluate -> analyze failures -> publish receipts.
02 / Represent Your Language Data
- Data design: synthetic compositional evidence packets rather than scraped customer binders.
- Privacy: no employer, patient, or customer evidence enters the repository.
- Representation: classical sparse features and Gemma tokenized instruction inputs provide two different text representations.
- Documentation: Data Card records provenance, structure, label generation, and limitations.
03 / Design and Train Neural Networks
- Split discipline: scenario families are isolated across train, validation, test, and challenge.
- Learning behavior: classical learning curves show how performance changes with training-set size.
- Benchmark hardening: an early lexical shortcut was treated as a dataset failure and corrected before the v1.1 classical seal.
04 / Discover the Transformer Architecture
- Applied transformer use: Gemma 3 1B instruction-tuned model is used rather than reimplementing attention from scratch.
- Model transparency: Model Card and project documentation record model family, size, intended use, and evaluation constraints.
- PEFT context: LoRA adapts selected model weights rather than fully retraining the model.
05 / Fine-Tune Your Model
- Method: QLoRA on Gemma 3 1B with 4-bit NF4 and LoRA rank 16.
- Controlled Gemma result: on dataset v1.0, few-shot macro F1 0.1365 is higher than QLoRA 0.0827.
- Research lesson: fine-tuning didn't automatically improve the task and introduced severe output-format fragility.
06 / Align Your Model
- Output contract: the model must return one of five valid labels.
- Alignment failure: QLoRA parse success of about 0.435 shows that task alignment includes reliable output behavior, not only classification scores.
- Human authority: intended-use documents prohibit autonomous audit conclusions.
07 / Accelerate Your Model
- Quantization: 4-bit NF4 reduces memory needs for QLoRA.
- Compute discipline: smoke tests precede full runs to conserve free GPU time.
- Practicality: the project deliberately uses free or commodity compute to keep the research reproducible for learners and smaller teams.
08 / Capstone: Real-World Impact
- Problem: distinguish operational proof from relevant-looking paperwork.
- Research: external grounding from NIST, Google DeepMind, academic QLoRA literature, NIST AI RMF, and the UN.
- Implementation: benchmark, classical path, Gemma path, failure analysis, assurance artifacts, public site.
- Responsible impact: MMC rubric option 4, Reduced Inequalities; official UN designation SDG 10.
- Submission: report, eight-slide deck, and narrated final presentation are complete.