Advanced mini-project·3-4 hours

ML Serving Readiness Project

Turn quality evidence and traffic assumptions into a capacity model, SLO, rollout, and rollback plan.

Scenario

A model endpoint must serve 120 requests per second at peak while meeting a p95 latency target. One warm replica sustains 35 requests per second at the target latency.

You will demonstrate

  • Compute capacity with headroom.
  • Define complete serving SLOs.
  • Design canary and rollback evidence.

Project evidence

Show the work, not a checked box.

Each response is stored in this browser as you type. Include metrics, test output, or a decision rationale wherever the deliverable asks for it.

1

Build the capacity model

Compute replicas from peak traffic, measured capacity, and 30% headroom.

Required evidence: Passing capacity calculation and assumptions table.

0/80 minimum characters

2

Specify the SLO

Define latency percentile, error rate, availability, quality, and observation window.

Required evidence: A measurable SLO contract.

0/80 minimum characters

3

Design the load test

Include warm, cold, burst, overload, and dependency-failure cases.

Required evidence: A reproducible load profile and acceptance criteria.

0/80 minimum characters

4

Plan canary and rollback

Choose canary size, guardrails, rollback trigger, and fallback.

Required evidence: A staged release runbook.

0/80 minimum characters

Runnable Python lab

Calculate replica capacity with headroom

Compute the minimum whole number of replicas for peak traffic and a utilization headroom policy.

Project defense

Why is average latency insufficient for an online inference SLO?

Rubric self-review

Rate the evidence, not your effort: 0 missing, 1 weak, 2 adequate, 3 strong. All criteria must be reviewed, but a low honest score does not get hidden.

Capacity includes explicit headroom.

Tail latency and queueing are measured.

Quality remains a release dimension.

Rollback is automatic for defined critical failures.

Useful references

Project completion gate

Completion is controlled by stored evidence, deterministic tests, a decision defense, rubric review, and the artifact when required.

Evidence pendingCode pendingDefense pendingRubric pending

Optional cloud portfolio

Submit evidence across devices.

An account is required. Submit only when the local completion gate passes. AI review is advisory and separate from deterministic completion.

Account settings