← All skills

Quality

LLM-as-judge calibration

Use and calibrate model-graded evaluation against human labels.

Target proficiency: Applied

Course outcomes

  • AIN 401 CO4

    Evaluate AI systems with golden sets, calibrated LLM-as-judge, cost/latency metrics and CI regression gates, and justify engineering decisions with evidence.

  • MST 502 CO4

    Demonstrate Evaluation & Observability competence: complete the Day 46–60 builds, explain the underlying theory, and defend the phase artefact under live questioning.

  • MST 504 CO3

    Demonstrate Integration & graphs — agents, knowledge graphs, evals competence: complete the Day 16–23 builds, explain the underlying theory, and defend the phase artefact under live questioning.

  • MST 506 CO4

    Demonstrate Evaluation, Observability & Retrieval Quality competence: complete the Day 46–60 builds, explain the underlying theory, and defend the phase artefact under live questioning.

  • MST 507 CO4

    Demonstrate Evaluation, Observability & the Verifier competence: complete the Day 46–60 builds, explain the underlying theory, and defend the phase artefact under live questioning.

  • MST 508 CO4

    Demonstrate Evaluation & Observability competence: complete the Day 46–60 builds, explain the underlying theory, and defend the phase artefact under live questioning.

Taught in weeks & phases

Textbook chapters