Automated reproduction of comprehensive model evaluation benchmarks following the Benchmark Suite V3. Auto-activates for model benchmarking, comparison evaluation, or performance testing between AI mo
.claude/skills/model-evaluation-benchmark/SKILL.md(main)