<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>Evals ML</title><description>Machine learning engineering reference for LLM evaluation benchmarks (MMLU, HumanEval, GSM8K), Pass@k statistical confidence, and AI token pricing.</description><link>https://evalsml.com/</link><language>en</language><item><title>Designing an LLM Evaluation That Actually Tells You Something</title><link>https://evalsml.com/posts/welcome/</link><guid isPermaLink="true">https://evalsml.com/posts/welcome/</guid><description>Task-grounded eval sets, grader choice, sampling variance and the contamination traps that make benchmark scores misleading.</description><pubDate>Sat, 01 Aug 2026 00:00:00 GMT</pubDate><category>llm-evaluation</category><category>benchmarks</category><category>pass-at-k</category><category>model-grading</category></item></channel></rss>