This skill should be used when the user asks to "evaluate agent performance", "build test framework", "measure agent quality", "create evaluation rubrics", or mentions LLM-as-judge, multi-dimensional
skills/guanyang/antigravity-skills/evaluation/SKILL.md(main)