Browse 138,453 skills across 70 categories. Mirrored from majiayu000/claude-skill-registry.
Claude Codeセッションの正式な評価フレームワークで、評価駆動開発(EDD)の原則を実装します
Formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles
克劳德代码会话的正式评估框架,实施评估驱动开发(EDD)原则
Formal evaluation framework for LLM features implementing Evaluation-Driven Development (EDD) principles.
A formal evaluation framework for Claude Code sessions, implementing eval-driven development (EDD) principles.
Formal evaluation framework for Antigravity Code sessions implementing eval-driven development (EDD) principles. Enforce...
Claude Code toolkit - agents, commands, skills, rules, and hooks for productive AI-assisted development
/============================================================================/
Build and run deterministic evaluation suites for agent workflows (single-turn or agentic). Use when you need reproducib...
Guide for diagnosing and improving MSBuild project evaluation performance. Only activate in MSBuild/.NET build context....
Evaluate implementation plan before execution - validates architecture, coverage, dependencies, and best practices. TRIG...
Run Microsoft's eval-recipes benchmarks to validate amplihack improvements against baseline agents. Activates when testi...
Run Microsoft's eval-recipes benchmarks to validate amplihack improvements against baseline agents. Auto-activates when...
Generate evaluation report from session metrics, checkpoint evals, and pass@k data
Run eval scenarios to benchmark Mycelium effectiveness. Execute tasks using reflexion loop, validate against success cri...
Develop and run agent behavior evaluations. Use this skill when asked to "write evals", "test agent behavior", "create e...
Agent evaluation framework based on Anthropic's best practices. USE WHEN eval, evaluate, test agent, benchmark, verify b...
Write and analyze evaluations for AI agents and LLM applications. Use when building evals, testing agents, measuring AI...
Provides context about the Roo Code evals system structure in this monorepo. Use when tasks mention "evals", "evaluation...
Analyze a subject from five perspectives with assumption tracking and weighting. Use after research to examine all angle...
This SOP evaluates a completed brazil-bench attempt against the spec.md requirements, capturing metrics for comparison a...
Evaluate a generated diagram against a human reference using PaperBanana's VLM-as-Judge scoring.
計算資產價格相對長期指數成長趨勢線的偏離度,衡量當前是否處於歷史極端區間,並可選擇性地進行宏觀因子分析以判斷行情體質。
Critically assess external feedback (code reviews, AI reviewers, PR comments) and decide which suggestions to apply usin...
MMI評価エージェント - Modularity Maturity Indexによるモジュール成熟度評価(Cohesion/Coupling/Independence/Reusability)。/evaluate-mmi [対象パス] で呼...
Measure model performance on test datasets. Use when assessing accuracy, precision, recall, and other metrics.
评价 Skill 执行结果的 Skill。当需要评价产出质量、判断是否需要迭代时触发。触发词:评价、evaluate、打分、怎么样、效果如何。
Evaluate one or more traces against an existing Truesight live evaluation. Use when a deployed live evaluation already e...
Help users make better hiring decisions. Use when someone is evaluating job candidates, making hiring decisions, conduct...
Make an evidence-based hiring decision and produce a Candidate Evaluation Decision Pack (criteria + scorecard, signal lo...
Evaluates code generation models across HumanEval, MBPP, MultiPL-E, and 15+ benchmarks with pass@k metrics. Use when ben...
Evaluates Node.js packages before installation by checking bundle size impact, comparing alternatives, verifying latest...
Evaluate LLM systems using automated metrics, LLM-as-judge, and benchmarks. Use when testing prompt quality, validating...
Evaluates LLMs across 60+ academic benchmarks (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag). Use when benchmarking mod...
Evaluate trained machine learning models with the right metrics and comparison logic. Use for benchmark review, threshol...
This skill allows AI assistant to evaluate machine learning models using a comprehensive suite of metrics. it should be...
Create a Technology Evaluation Pack (problem framing, options matrix, build vs buy, pilot plan, risk review, decision me...
Help users evaluate emerging technologies. Use when someone is assessing new tools, making build vs buy decisions, evalu...
Two-stage paper screening - abstract scoring then deep dive for specific data extraction
Evaluate skills by executing them across sonnet, opus, and haiku models using sub-agents. Use when testing if a skill wo...
開発ツール、フレームワーク、ライブラリの評価と比較を支援します。PoC計画、複数候補の比較分析、意思決定フレームワークを提供します。技術選定、ツール導入の判断が必要な場合に使用してください。
Help users make better decisions between competing options. Use when someone is weighing pros and cons, comparing altern...
Evaluate trade-offs and produce a Trade-off Evaluation Pack (trade-off brief, options+criteria matrix, all-in cost/oppor...
Build evaluation frameworks for agent systems
This skill should be used when the user asks to "evaluate agent performance", "build test framework", "measure agent qua...
Build evaluation frameworks for agent systems. Use when testing agent performance, validating context engineering choice...
Build evaluation frameworks for agent systems. Use when testing agent performance systematically, validating context eng...
Systematic evaluation and comparison of technologies, tools, frameworks, and approaches using weighted criteria and scor...
Evaluate agent systems with quality gates and LLM-as-judge. Use when you need to measure component quality or implement...
Guidelines for creating evaluation suites, including spec.md templates, rubric structures, and code-based validation pat...
- Overview
Builds repeatable evaluation systems with golden datasets, scoring rubrics, pass/fail thresholds, and regression reports...
PluginEval quality methodology — dimensions, rubrics, statistical methods, and scoring formulas. Use this skill when und...
name: evaluation-plan
Instrument evaluation metrics, quality scores, and feedback loops
Evaluation and reporting for code quality, performance, security, architecture, team processes, AI/LLM outputs, A/B test...
Systematic content evaluation framework progressing through Critique → Reinforcement → Risk Analysis → Growth. Use when...