Task-grounded eval sets, grader choice, sampling variance and the contamination traps that make benchmark scores misleading.