Builds evaluation harnesses for LLM products - golden datasets, deterministic checks, calibrated LLM-as-judge rubrics, and CI regression gates that turn "seems good" into a tracked number. Use when someone asks "how do I know this prompt change didn't break anything", "set up evals for my RAG pipeline", "is LLM-as-judge reliable", "why did quality drop after the model swap", or is shipping an LLM feature with no quality measurement. Do NOT use for classical ML model reporting (precision/recall, ROC curves, confusion matrices on trained classifiers) - use model-evaluation-report instead; do NOT use for analyzing online A/B experiments - use ab-test-analyzer instead.
Click to play with sound.
---
name: LLM Evaluation
description: Builds evaluation harnesses for LLM products - golden datasets, deterministic checks, calibrated LLM-as-judge rubrics, and CI regression gates that turn "seems good" into a tracked number. Use when someone asks "how do I know this prompt change didn't break anything", "set up evals for my RAG pipeline", "is LLM-as-judge reliable", "why did quality drop after the model swap", or is shipping an LLM feature with no quality measurement. Do NOT use for classical ML model reporting (precision/recall, ROC curves, confusion matrices on trained classifiers) - use model-evaluation-report instead; do NOT use for analyzing online A/B experiments - use ab-test-analyzer instead.
---
# LLM Evaluation
A prompt or model change that looks fine on the three examples you tried will break inputs you did not try, and without an eval you find out from users. The costly mistake this skill prevents is shipping on vibes: an uncalibrated judge or a ten-example "test" gives you a number that moves randomly, so real regressions hide inside noise. Build the dataset first, calibrate the judge before trusting it, and gate changes in CI.
## Operating procedure
Follow the order: dataset before metrics, calibration before judging, pipeline before iteration. A judge calibrated against nothing is a random-number generator with confidence.
### Step 1: Gather inputs
Collect these before building anything. Where the user cannot answer, use the default and label the value a guess.
1. Task definition - what the system produces, in one sentence.
2. What "good" means - 3-5 quality dimensions (correctness, groundedness, tone, format). Default: correctness plus format validity.
3. Real failure examples - at least 5 outputs the user considers bad, with why.
4. Traffic source - production logs, support tickets, or synthetic. Prefer production; label synthetic data as synthetic.
5. Human-label budget - how many examples a person will grade. Default: 50 for judge calibration.
### Step 2: Build the dataset
… load the full skill through Skill Me