Studio hour·60 minutes

Studio Hour: RAG Evaluation Table

Build a small evaluation table that separates retrieval failure, unsupported generation, and correct abstention.

Scenario

A grounded assistant looks fluent overall, but the team cannot tell whether bad answers begin in retrieval or generation.

You will demonstrate

  • Label evaluation cases.
  • Compute claim support.
  • Assign a failure owner.

Project evidence

Show the work, not a checked box.

Each response is stored in this browser as you type. Include metrics, test output, or a decision rationale wherever the deliverable asks for it.

1

Create the evaluation table

Record query type, relevant document, retrieved document, answer support, expected behavior, and failure owner for at least four cases.

Required evidence: A four-row RAG evaluation table plus one release threshold.

0/80 minimum characters

Runnable Python lab

Compute support rate

Measure the fraction of generated claims linked to allowed evidence.

Project defense

A relevant passage was retrieved but the answer adds an unsupported claim. Which component failed?

Rubric self-review

Rate the evidence, not your effort: 0 missing, 1 weak, 2 adequate, 3 strong. All criteria must be reviewed, but a low honest score does not get hidden.

The table separates retrieval evidence, generation evidence, expected behavior, and ownership.

Useful references

Project completion gate

Completion is controlled by stored evidence, deterministic tests, a decision defense, rubric review, and the artifact when required.

Evidence pendingCode pendingDefense pendingRubric pending

Optional cloud portfolio

Submit evidence across devices.

An account is required. Submit only when the local completion gate passes. AI review is advisory and separate from deterministic completion.

Account settings