amitgambhir/agent-eval-kit
Behavioral evaluation for agentic AI systems. Test whether your agents honor their spec — role, scope, escalation, handoff. Claude-as-judge, markdown in, PR-gated CI.
ARCHETYPE
Pragmatist
A balanced, no-drama setup: some rules, some tools, nothing extreme.
Copy this rig
# review before running: this installs third-party code
$ npx degit amitgambhir/agent-eval-kit/.claude ./rig-agent-eval-kit # inspect, then merge into .claude/ MCP servers are added to Claude Code at local scope; env vars are shown as YOUR_… placeholders — we never store values. Files are fetched with degit into a separate folder so you can review before merging.
$ curl -fsSL --create-dirs -o .claude/agents/ada.md https://raw.githubusercontent.com/amitgambhir/agent-eval-kit/main/agents/ada.md $ curl -fsSL --create-dirs -o .claude/agents/curie.md https://raw.githubusercontent.com/amitgambhir/agent-eval-kit/main/agents/curie.md
This rig commits no guardrails. Here is the community baseline instead — the deny/ask rules most often found across all 6,974 rigs:
{
"permissions": {
"deny": [
"Read(./.env)",
"Read(**/.env)",
"Read(~/.ssh/**)",
"Bash(rm -rf *)",
"Read(**/*.pem)",
"Bash(rm -rf /)",
"Bash(git push --force:*)",
"Bash(sudo *)",
"Read(~/.aws/**)",
"Bash(rm -rf /*)",
"Read(./.env.*)",
"Read(.env)",
"Bash(git push --force*)",
"Bash(rm -rf:*)",
"Read(**/.env.*)",
"Read(**/*.key)",
"Bash(sudo:*)",
"Bash(git reset --hard*)",
"Bash(git reset --hard:*)",
"Read(.env.*)"
],
"ask": [
"Bash(git push:*)",
"Bash(git push *)",
"Bash(git commit:*)",
"Bash(rm *)",
"Bash(rm:*)",
"Bash(npm publish:*)",
"Bash(wget *)",
"Bash(git rebase *)",
"Bash(gh pr merge *)",
"Bash(git commit *)"
]
}
} Subagents (2)
| ada | — |
| curie | — |
Slash commands (4)
/eval-add/eval-case/eval-report/eval-run
Permissions
deny (0)
—
ask (0)
—
allow (17)
Bash(python scripts/generate_report.py:*)
Bash(python scripts/run_eval.py:*)
Bash(python3 scripts/generate_report.py:*)
Bash(python3 scripts/run_eval.py:*)
Bash(pip install -r scripts/requirements.txt)
Bash(pip install -r scripts/ci-requirements.txt)
Read(CLAUDE.md)
Read(README.md)
Read(.claude/commands/**)
Read(agents/**)
Read(eval-cases/**)
Read(schema/**)
Read(results/**)
Write(eval-cases/**)
Write(results/**)
Edit(eval-cases/**)
Edit(results/**)
Similar rigs
myICOR/myPKA
AI-powered Personal Knowledge Assistance in a folder. Built on the ICOR methodology. Plain markdown. Any LLM. Yours forever.
Orchestrator 8.3k tok ·
cdeust/zetetic-team-subagents
One epistemic standard on three hosts. Claude Code: 15 problem-shaped skills, 97 sourced reasoning patterns, 120 agents, commit-time gates. Codex and Gemini CLI: portable evidence-synthesis, design and goal-loop skills. Shared gates block u
Orchestrator 5.8k tok ·
anomalyco/opencode
The open source coding agent.
Pragmatist 4.2k tok ·
DietrichGebert/ponytail
Makes your AI agent think like the laziest senior dev in the room. The best code is the code you never wrote.
Pragmatist 1.6k tok ·