Quality
Evaluation engineering with golden sets
Build golden datasets and metrics to measure AI system quality.
Target proficiency: Applied
Course outcomes
- AIN 401 CO4
Evaluate AI systems with golden sets, calibrated LLM-as-judge, cost/latency metrics and CI regression gates, and justify engineering decisions with evidence.
- AIN 402 CO4
Evaluate and prioritise AI use cases using value cases (baseline, benefits, TCO, ROI, sensitivity) and plan solution evaluation and benefits realisation.
- MST 501 CO1
Demonstrate Foundations & Agent Architecture competence: complete the Day 1–15 builds, explain the underlying theory, and defend the phase artefact under live questioning.
- MST 501 CO5
Demonstrate Evaluation, Observability & Security competence: complete the Day 61–75 builds, explain the underlying theory, and defend the phase artefact under live questioning.
- MST 502 CO4
Demonstrate Evaluation & Observability competence: complete the Day 46–60 builds, explain the underlying theory, and defend the phase artefact under live questioning.
- MST 505 CO2
Demonstrate Retrieval & Knowledge Context competence: complete the Day 16–30 builds, explain the underlying theory, and defend the phase artefact under live questioning.
- MST 505 CO5
Demonstrate Evaluation, Caching, Multimodal & Observability competence: complete the Day 61–75 builds, explain the underlying theory, and defend the phase artefact under live questioning.
- MST 506 CO1
Demonstrate Foundations — Data, Retrieval Theory & the Corpus competence: complete the Day 1–15 builds, explain the underlying theory, and defend the phase artefact under live questioning.
- MST 506 CO4
Demonstrate Evaluation, Observability & Retrieval Quality competence: complete the Day 46–60 builds, explain the underlying theory, and defend the phase artefact under live questioning.
- MST 507 CO4
Demonstrate Evaluation, Observability & the Verifier competence: complete the Day 46–60 builds, explain the underlying theory, and defend the phase artefact under live questioning.
- MST 508 CO4
Demonstrate Evaluation & Observability competence: complete the Day 46–60 builds, explain the underlying theory, and defend the phase artefact under live questioning.
- MST 508 CO6
Demonstrate Production LLMOps, Security & Mastery competence: complete the Day 76–90 builds, explain the underlying theory, and defend the phase artefact under live questioning.
Taught in weeks & phases
- AIN 401 · Week 9 — Evaluation Engineering: Golden Sets, LLM-as-Judge & Regression Evals
- AIN 402 · Week 13 — Solution Evaluation: KPIs, Experiments, Drift & Benefits Realisation
- MST 501 · Phase 1 — Foundations & Agent Architecture
- MST 501 · Phase 5 — Evaluation, Observability & Security
- MST 502 · Phase 4 — Evaluation & Observability
- MST 505 · Phase 2 — Retrieval & Knowledge Context
- MST 505 · Phase 5 — Evaluation, Caching, Multimodal & Observability
- MST 506 · Phase 1 — Foundations — Data, Retrieval Theory & the Corpus
- MST 506 · Phase 4 — Evaluation, Observability & Retrieval Quality
- MST 507 · Phase 4 — Evaluation, Observability & the Verifier
- MST 508 · Phase 4 — Evaluation & Observability
- MST 508 · Phase 6 — Production LLMOps, Security & Mastery