📄 api-evaluation.md

← Vault

API Evaluation

Guide to evaluating OpenAI, Anthropic, and other API-based language models.

Overview

The lm-evaluation-harness supports evaluating API-based models through a unified TemplateAPI interface. This allows benchmarking of: