~ / leaderboard / skill / evaluation
Skill · #1249
evaluation
This skill should be used when building agent evaluation systems: deterministic checks, regression suites, multi-dimensional rubrics, quality gates, production monitoring, baseline comparison, and out
6rigs use it
0%of all rigs
20.4kavg rig context tax
Get it
$ npx degit muratcankoylan/Agent-Skills-for-Context-Engineering/skills/evaluation .claude/skills/evaluation From the most-starred rig that ships it: muratcankoylan/Agent-Skills-for-Context-Engineering. Same-named skills in different rigs may differ — review before use.
Often used together with
tool-design · 6context-optimization · 5multi-agent-patterns · 5memory-systems · 5filesystem-context · 4advanced-evaluation · 4harness-engineering · 4context-compression · 4bdi-mental-states · 4context-degradation · 4
Rigs using evaluation (6)
muratcankoylan/Agent-Skills-for-Context-Engineering
A comprehensive collection of Agent Skills for context engineering, multi-agent architectures, and production agent systems. Use when building, optimizing, or debugging agent systems that require effective context management.
Skill Collector 3.8k tok ·
guanyang/open-agent-hub
A lightweight, zero-dependency CLI tool to manage and activate capabilities for AI coding assistants (such as Claude Code, Cursor, Trae, etc.).
MCP Hoarder 70.5k tok ·
shipshitdev/skills
Claude, Cursor, Codex skills and commands
Skill Collector 10.4k tok ·
greyhaven-ai/claude-code-config
Grey Haven's centralized repository for Claude Code agents, commands, and configs
Skill Collector 4.2k tok ·
docxology/template
GitHub template for reproducible research repos — two-layer Python 3.10+/uv monorepo splitting generic infrastructure from per-project code, with pytest coverage gates (60% infra, 90% projects), a 17-stage build pipeline, markdown-to-PDF re
Skill Collector 28.0k tok ·
jacobchenai/Claude-dotfiles
—
Skill Collector 5.6k tok ·