Each block is one benchmark that names the skill. This counts sources, not tasks: a skill named by three benchmarks may still be thinly tested if each contributes a handful of examples.
This is a survey of the identified benchmarks. I haven't been able to find any that test the following. Where something is partly covered, I've named what covers it.
A skill covered by one benchmark is a single point of failure in an evaluation plan: you cannot tell a real capability gap from an artefact of that benchmark's construction. Counts are of benchmarks that list the skill, not of tasks, so a skill covered three times may still be thinly tested. Data as at September 2026.