📝AI Evals
| Dimension | Similarity/Generative Metrics | Exact Match Metrics |
|---|---|---|
| Evaluation Goal | Assess semantic/perceptual closeness | Assess precise correctness |
| Typical Output | Generative (text, image, audio) | Discriminative (classification) |
| Method | N-gram overlap, pixel difference | Binary match (right/wrong) |
| Example Metrics | BLEU, ROUGE, PSNR, SSIM | Accuracy, F1-score |
| Dimension | Mean Absolute Error (MAE) | Mean Squared Error (MSE) |
|---|---|---|
| Error Penalty | Linear penalty | Quadratic penalty (larger errors penalized more) |
| Outlier Sensitivity | Less sensitive; robust | Highly sensitive to outliers |
| Interpretability | Same units as target; easy to understand | Squared units; harder to directly interpret |
| Dimension | Automated Metrics | Human Evaluation |
|---|---|---|
| Speed/Scale | Very fast, high volume | Slow, limited scale |
| Objectivity | Highly objective, quantitative | Subjective, qualitative |
| Cost | Low after setup | High, ongoing |
| Nuance/Creativity | Struggles with nuance | Excels at nuance, creativity |
| Setup Effort | High initial effort | Lower initial effort |
| Dimension | Adversarial Evaluation | Standard Human Evaluation |
|---|---|---|
| Primary Goal | Discover unknown flaws, vulnerabilities | Validate expected behavior, performance |
| Evaluator Role | Active 'breaker' or attacker | Objective rater or assessor |
| Approach | Aggressive, creative, exploit-seeking | Passive, observational, task-oriented |
| Output Focus | Systemic vulnerabilities, misuse risks | Performance metrics, user experience |
| Dimension | Human Evaluation | Automated Evaluation |
|---|---|---|
| Subjectivity/Nuance | Excellent for complex, subjective quality | Poor; struggles with context and meaning |
| Cost | High due to labor and training | Low after initial setup |
| Speed | Slow; depends on human availability | Fast; instant results for large datasets |
| Reproducibility | Moderate; inter-rater agreement varies | High; consistent results given same inputs |
| Dimension | Automated Evals | Human Evals |
|---|---|---|
| Speed | Fast | Slow |
| Scale | High volume | Low volume |
| Subjectivity | Low | High |
| Nuance | Limited | High |
| Part of | MLOps | Ensures continuous quality in model deployment. |
| Depends on | Data Annotation | Requires labeled data for supervised metric calculation. |
| Made of | Evaluation Datasets | Comprises test sets, prompts, and ground truth labels. |
| Predecessor | Traditional Software QA | Focused on deterministic systems, less on emergent behavior. |
| Used in | Model Deployment | Gate for releasing models to production. |
| Confused with | Unit Testing | Unit testing checks code, evals check model behavior. |
| Limitation | Generalization Gap | Evals on one dataset may not extend to all real-world scenarios. |
| Dimension | Similarity/Generative Metrics | Exact Match Metrics |
|---|---|---|
| Evaluation Goal | Assess semantic/perceptual closeness | Assess precise correctness |
| Typical Output | Generative (text, image, audio) | Discriminative (classification) |
| Method | N-gram overlap, pixel difference | Binary match (right/wrong) |
| Example Metrics | BLEU, ROUGE, PSNR, SSIM | Accuracy, F1-score |
| Dimension | Mean Absolute Error (MAE) | Mean Squared Error (MSE) |
|---|---|---|
| Error Penalty | Linear penalty | Quadratic penalty (larger errors penalized more) |
| Outlier Sensitivity | Less sensitive; robust | Highly sensitive to outliers |
| Interpretability | Same units as target; easy to understand | Squared units; harder to directly interpret |
| Dimension | Automated Metrics | Human Evaluation |
|---|---|---|
| Speed/Scale | Very fast, high volume | Slow, limited scale |
| Objectivity | Highly objective, quantitative | Subjective, qualitative |
| Cost | Low after setup | High, ongoing |
| Nuance/Creativity | Struggles with nuance | Excels at nuance, creativity |
| Setup Effort | High initial effort | Lower initial effort |
| Dimension | Adversarial Evaluation | Standard Human Evaluation |
|---|---|---|
| Primary Goal | Discover unknown flaws, vulnerabilities | Validate expected behavior, performance |
| Evaluator Role | Active 'breaker' or attacker | Objective rater or assessor |
| Approach | Aggressive, creative, exploit-seeking | Passive, observational, task-oriented |
| Output Focus | Systemic vulnerabilities, misuse risks | Performance metrics, user experience |
| Dimension | Human Evaluation | Automated Evaluation |
|---|---|---|
| Subjectivity/Nuance | Excellent for complex, subjective quality | Poor; struggles with context and meaning |
| Cost | High due to labor and training | Low after initial setup |
| Speed | Slow; depends on human availability | Fast; instant results for large datasets |
| Reproducibility | Moderate; inter-rater agreement varies | High; consistent results given same inputs |
| Dimension | Automated Evals | Human Evals |
|---|---|---|
| Speed | Fast | Slow |
| Scale | High volume | Low volume |
| Subjectivity | Low | High |
| Nuance | Limited | High |
| Part of | MLOps | Ensures continuous quality in model deployment. |
| Depends on | Data Annotation | Requires labeled data for supervised metric calculation. |
| Made of | Evaluation Datasets | Comprises test sets, prompts, and ground truth labels. |
| Predecessor | Traditional Software QA | Focused on deterministic systems, less on emergent behavior. |
| Used in | Model Deployment | Gate for releasing models to production. |
| Confused with | Unit Testing | Unit testing checks code, evals check model behavior. |
| Limitation | Generalization Gap | Evals on one dataset may not extend to all real-world scenarios. |