Quality
LLM-as-judge calibration
Use and calibrate model-graded evaluation against human labels.
Target proficiency: Applied
Course outcomes
- AIN 401 CO4
Evaluate AI systems with golden sets, calibrated LLM-as-judge, cost/latency metrics and CI regression gates, and justify engineering decisions with evidence.
- MST 502 CO4
Demonstrate Evaluation & Observability competence: complete the Day 46–60 builds, explain the underlying theory, and defend the phase artefact under live questioning.
- MST 504 CO3
Demonstrate Integration & graphs — agents, knowledge graphs, evals competence: complete the Day 16–23 builds, explain the underlying theory, and defend the phase artefact under live questioning.
- MST 506 CO4
Demonstrate Evaluation, Observability & Retrieval Quality competence: complete the Day 46–60 builds, explain the underlying theory, and defend the phase artefact under live questioning.
- MST 507 CO4
Demonstrate Evaluation, Observability & the Verifier competence: complete the Day 46–60 builds, explain the underlying theory, and defend the phase artefact under live questioning.
- MST 508 CO4
Demonstrate Evaluation & Observability competence: complete the Day 46–60 builds, explain the underlying theory, and defend the phase artefact under live questioning.
Taught in weeks & phases
- AIN 401 · Week 9 — Evaluation Engineering: Golden Sets, LLM-as-Judge & Regression Evals
- MST 502 · Phase 4 — Evaluation & Observability
- MST 504 · Phase 3 — Integration & graphs — agents, knowledge graphs, evals
- MST 506 · Phase 4 — Evaluation, Observability & Retrieval Quality
- MST 507 · Phase 4 — Evaluation, Observability & the Verifier
- MST 508 · Phase 4 — Evaluation & Observability