~ / leaderboard / skill / llm-evaluation
Skill · #1358
llm-evaluation
Implement comprehensive evaluation strategies for LLM applications using automated metrics, human feedback, and benchmarking. Use when testing LLM performance, measuring AI application quality, or est
6rigs use it
0%of all rigs
13.7kavg rig context tax
Get it
$ npx degit wshobson/agents/plugins/llm-application-dev/skills/llm-evaluation .claude/skills/llm-evaluation From the most-starred rig that ships it: wshobson/agents. Same-named skills in different rigs may differ — review before use.
Often used together with
rag-implementation · 5langchain-architecture · 5error-handling-patterns · 4async-python-patterns · 4microservices-patterns · 4ml-pipeline-workflow · 4prompt-engineering-patterns · 4code-reviewer · 3github-actions-templates · 3cost-optimization · 3
Rigs using llm-evaluation (6)
wshobson/agents
Multi-harness agentic plugin marketplace for Claude Code, Codex, Cursor, OpenCode, GitHub Copilot, Google Antigravity, and Pi
Skill Collector 8.4k tok ·
ckorhonen/claude-skills
A curated collection of skills for Claude Code and Codex — specialized capabilities across development, design, AI, security, and more. Part of the cdd.dev/skills family.
Orchestrator 6.1k tok ·
HermeticOrmus/LibreUIUX-Claude-Code
Complete UI/UX system for Claude Code — 67 specialized agents, design vocabulary, tested prompts
Skill Collector 7.6k tok ·
frank-luongt/faos-skills-marketplace
520+ AI-powered skills and 31 agent plugins for Claude Cowork, OpenAI Codex, Gemini CLI, GitHub Copilot, and Perplexity Computer. Built by the FAOS Framework.
Skill Collector 1.1k tok ·
carlopezzuto/agents
Claude Code configuration with 42 specialized agents, 33 skills, 28 commands, and a hook system for development best practices
Orchestrator 6.3k tok ·
spideynolove/claude-code-in-action
The process I experienced firsthand through the https://anthropic.skilljar.com/claude-code-in-action course using real-world daily projects
MCP Hoarder 52.5k tok ·