Practical AI response evaluation portfolio with evidence-based scoring, worked reviews, and controlled response comparisons
-
Updated
Jul 28, 2026
Practical AI response evaluation portfolio with evidence-based scoring, worked reviews, and controlled response comparisons
Evidence-based evaluation of 27 real-world ChatGPT Go cases, covering instruction following, context handling, reasoning, tool use, and LLM response quality.
Model-agnostic AI response evaluation handbook with evidence-based rubrics, calibration guides, case studies, and reusable templates
Practical portfolio for AI response evaluation, data quality, annotation QA, and evidence-based evaluation.
To associate your repository with the response-evaluation topic, visit your repo's landing page and select "manage topics."