83,201 skills sorted by stars.
Formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles
克劳德代码会话的正式评估框架,实施评估驱动开发(EDD)原则
Formal evaluation framework for LLM features implementing Evaluation-Driven Development (EDD) principles.
A formal evaluation framework for Claude Code sessions, implementing eval-driven development (EDD) principles.
Formal evaluation framework for Antigravity Code sessions implementing eval-driven development (EDD) principles. Enforce...
Claude Code toolkit - agents, commands, skills, rules, and hooks for productive AI-assisted development
/============================================================================/
Build and run deterministic evaluation suites for agent workflows (single-turn or agentic). Use when you need reproducib...
Guide for diagnosing and improving MSBuild project evaluation performance. Only activate in MSBuild/.NET build context....
Evaluate implementation plan before execution - validates architecture, coverage, dependencies, and best practices. TRIG...
Run Microsoft's eval-recipes benchmarks to validate amplihack improvements against baseline agents. Activates when testi...
Generate evaluation report from session metrics, checkpoint evals, and pass@k data
Run eval scenarios to benchmark Mycelium effectiveness. Execute tasks using reflexion loop, validate against success cri...
Develop and run agent behavior evaluations. Use this skill when asked to "write evals", "test agent behavior", "create e...
Agent evaluation framework based on Anthropic's best practices. USE WHEN eval, evaluate, test agent, benchmark, verify b...
Write and analyze evaluations for AI agents and LLM applications. Use when building evals, testing agents, measuring AI...
Provides context about the Roo Code evals system structure in this monorepo. Use when tasks mention "evals", "evaluation...
Analyze a subject from five perspectives with assumption tracking and weighting. Use after research to examine all angle...
This SOP evaluates a completed brazil-bench attempt against the spec.md requirements, capturing metrics for comparison a...
Evaluate a generated diagram against a human reference using PaperBanana's VLM-as-Judge scoring.
計算資產價格相對長期指數成長趨勢線的偏離度,衡量當前是否處於歷史極端區間,並可選擇性地進行宏觀因子分析以判斷行情體質。
Critically assess external feedback (code reviews, AI reviewers, PR comments) and decide which suggestions to apply usin...
MMI評価エージェント - Modularity Maturity Indexによるモジュール成熟度評価(Cohesion/Coupling/Independence/Reusability)。/evaluate-mmi [対象パス] で呼...
Measure model performance on test datasets. Use when assessing accuracy, precision, recall, and other metrics.
评价 Skill 执行结果的 Skill。当需要评价产出质量、判断是否需要迭代时触发。触发词:评价、evaluate、打分、怎么样、效果如何。
Evaluate one or more traces against an existing Truesight live evaluation. Use when a deployed live evaluation already e...
Help users make better hiring decisions. Use when someone is evaluating job candidates, making hiring decisions, conduct...
Make an evidence-based hiring decision and produce a Candidate Evaluation Decision Pack (criteria + scorecard, signal lo...
Evaluates code generation models across HumanEval, MBPP, MultiPL-E, and 15+ benchmarks with pass@k metrics. Use when ben...
Evaluates Node.js packages before installation by checking bundle size impact, comparing alternatives, verifying latest...
Evaluate LLM systems using automated metrics, LLM-as-judge, and benchmarks. Use when testing prompt quality, validating...
Evaluates LLMs across 60+ academic benchmarks (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag). Use when benchmarking mod...
Evaluate trained machine learning models with the right metrics and comparison logic. Use for benchmark review, threshol...
Create a Technology Evaluation Pack (problem framing, options matrix, build vs buy, pilot plan, risk review, decision me...