evaluatellm 0.1.0

Pre-submission audit fixes, ahead of the first CRAN release.

First release.

Statistical inference for language model evaluations, following Miller (2024) doi:10.48550/arXiv.2411.00640 for the core standard error and experiment design results, and the prediction-powered inference literature for the model judge functions.

Scoring

Comparison

Planning

Model judges

Leaderboards

Reporting