Data & ML starter
Notebooks, pipelines, evals, analysis
Install
Run in your project root. Existing files are never overwritten (unzip -n / [ -e … ] skip them). Review AGENTS.md and fill in the <placeholders>.
$ curl -fsSL https://agentrigs.dev/starters/agentrigs-data-ml.zip -o agentrigs-starter.zip && unzip -n agentrigs-starter.zip && chmod +x .claude/hooks/guard.sh && rm agentrigs-starter.zip $ npx degit affaan-m/ECC/.kiro/skills/python-patterns .claude/skills/python-patterns $ npx degit affaan-m/ECC/.agents/skills/eval-harness .claude/skills/eval-harness $ npx degit obra/superpowers/skills/systematic-debugging .claude/skills/systematic-debugging $ npx degit anthropics/skills/skills/xlsx .claude/skills/xlsx
Or download agentrigs-data-ml.zip (6 files). Skills are fetched from their source repos with npx degit.
mkdir -p .claude/hooks
[ -e AGENTS.md ] && echo "skip AGENTS.md (exists)" || cat > AGENTS.md <<'AGENTRIGS_EOF'
# Project
<one paragraph: the question or model, data sources, and what "good" looks like (metric + threshold)>
## Commands
- Env: `uv sync` · Tests: `uv run pytest -q` · Lint: `uv run ruff check .` · Pipeline: `uv run python -m pipeline`
## Conventions
- Raw data is read-only (`data/raw/`); write derived data to `data/processed/` with the script that made it.
- Fix random seeds; log dataset version, params and metrics for every run.
- Prefer scripts/modules over notebooks for anything reused; notebooks are for exploration only.
- Never upload data to third-party services without asking; treat all rows as potentially sensitive.
- Report results with the sample size and an uncertainty estimate, not just a point number.
## Checks before done
`uv run ruff check . && uv run pytest -q`; show the metric before and after your change.
## Workflow
1. Restate the task in one line and list the files you expect to touch.
2. Make the smallest change that works; keep diffs reviewable.
3. Run the checks below before saying you're done, and show the output.
4. If something is ambiguous, ask one precise question instead of guessing.
## Safety
- Never read or print secrets (.env, keys, ~/.ssh). Ask for values instead.
- No destructive commands (rm -rf, force-push, reset --hard) without explicit approval; the guard hook blocks them anyway.
- Don't add dependencies, services or paid APIs without asking.
AGENTRIGS_EOF
[ -e CLAUDE.md ] && echo "skip CLAUDE.md (exists)" || cat > CLAUDE.md <<'AGENTRIGS_EOF'
@AGENTS.md
# Claude Code notes
- Use plan mode for multi-file changes. Prefer the installed skills over ad-hoc procedures.
AGENTRIGS_EOF
[ -e .claude/settings.json ] && echo "skip .claude/settings.json (exists)" || cat > .claude/settings.json <<'AGENTRIGS_EOF'
{
"permissions": {
"deny": [
"Bash(rm -rf:*)",
"Bash(rm -fr:*)",
"Bash(sudo:*)",
"Bash(git push --force:*)",
"Bash(git push -f:*)",
"Bash(git reset --hard:*)",
"Bash(git clean -fd:*)",
"Read(./.env)",
"Read(./.env.*)",
"Read(./**/.env)",
"Read(./secrets/**)",
"Read(~/.ssh/**)",
"Read(~/.aws/**)",
"Read(./**/*.pem)"
],
"ask": [
"Bash(git push:*)",
"Bash(npm publish:*)",
"Bash(docker:*)",
"Bash(curl:*)"
],
"defaultMode": "default"
},
"hooks": {
"PreToolUse": [
{
"matcher": "Bash|Read|Edit|Write",
"hooks": [
{
"type": "command",
"command": "\"$CLAUDE_PROJECT_DIR\"/.claude/hooks/guard.sh"
}
]
}
]
}
}
AGENTRIGS_EOF
[ -e .claude/hooks/guard.sh ] && echo "skip .claude/hooks/guard.sh (exists)" || cat > .claude/hooks/guard.sh <<'AGENTRIGS_EOF'
#!/usr/bin/env bash
# .claude/hooks/guard.sh: PreToolUse guard. Exit code 2 blocks the tool call and shows the reason to the model.
# Requires jq. Make executable: chmod +x .claude/hooks/guard.sh
input=$(cat)
tool=$(printf '%s' "$input" | jq -r '.tool_name // empty')
cmd=$(printf '%s' "$input" | jq -r '.tool_input.command // empty')
path=$(printf '%s' "$input" | jq -r '.tool_input.file_path // .tool_input.path // empty')
if [ "$tool" = "Bash" ]; then
if printf '%s' "$cmd" | grep -Eq '(^|[;&| ])(sudo|mkfs|dd if=)|rm -[a-zA-Z]*r[a-zA-Z]*f|rm -[a-zA-Z]*f[a-zA-Z]*r|git push .*(--force|-f( |$))|git reset --hard|git clean -[a-z]*f|curl[^|]*\|[[:space:]]*(ba)?sh|chmod 777'; then
echo "Blocked by guard.sh: destructive command ($cmd). Ask the user to run it manually." >&2; exit 2
fi
fi
if printf '%s %s' "$path" "$cmd" | grep -Eq '(^|/|[[:space:]])\.env($|\.|[[:space:]])|id_rsa|\.pem($|[[:space:]])|\.ssh/|\.aws/credentials'; then
echo "Blocked by guard.sh: secrets file ($path$cmd)." >&2; exit 2
fi
exit 0
AGENTRIGS_EOF
[ -e .mcp.json ] && echo "skip .mcp.json (exists)" || cat > .mcp.json <<'AGENTRIGS_EOF'
{
"mcpServers": {
"context7": {
"type": "http",
"url": "https://mcp.context7.com/mcp"
}
}
}
AGENTRIGS_EOF
[ -e AGENTRIGS-STARTER.md ] && echo "skip AGENTRIGS-STARTER.md (exists)" || cat > AGENTRIGS-STARTER.md <<'AGENTRIGS_EOF'
# AgentRigs starter: Data & ML
Generated by https://agentrigs.dev/starters/data-ml from real, well-guarded public rigs:
- https://github.com/link7373/agentic-bi-team
- https://github.com/a-green-hand-jack/ml-project-repo-agent-native-template
- https://github.com/TakaGoto/rag-learning-academy
## Install
1. Unzip into your repo root (`unzip -n` won't overwrite existing files).
2. `chmod +x .claude/hooks/guard.sh` (needs `jq`).
3. Fill in the <placeholders> in AGENTS.md.
4. Install the recommended skills (third-party code, so review first):
```bash
npx degit affaan-m/ECC/.kiro/skills/python-patterns .claude/skills/python-patterns
npx degit affaan-m/ECC/.agents/skills/eval-harness .claude/skills/eval-harness
npx degit obra/superpowers/skills/systematic-debugging .claude/skills/systematic-debugging
npx degit anthropics/skills/skills/xlsx .claude/skills/xlsx
```
Optional MCP servers (each adds context tax every turn):
```bash
claude mcp add --transport http deepwiki https://mcp.deepwiki.com/mcp
```
Codex / Cursor / other harnesses read AGENTS.md directly; Claude Code reads CLAUDE.md, which imports AGENTS.md.
AGENTRIGS_EOF
chmod +x .claude/hooks/guard.sh
npx degit affaan-m/ECC/.kiro/skills/python-patterns .claude/skills/python-patterns # skill: python-patterns
npx degit affaan-m/ECC/.agents/skills/eval-harness .claude/skills/eval-harness # skill: eval-harness
npx degit obra/superpowers/skills/systematic-debugging .claude/skills/systematic-debugging # skill: systematic-debugging
npx degit anthropics/skills/skills/xlsx .claude/skills/xlsx # skill: xlsx Optional MCP servers (add only if you use them; each adds tool schemas to every turn):
$ claude mcp add --transport http deepwiki https://mcp.deepwiki.com/mcp Context tax
Claude Code pays for CLAUDE.md plus the imported AGENTS.md (~384 tok of instructions); Codex reads AGENTS.md alone (~353 tok). Index median: 2.2k tok. How we measure.
Guardrails
Files
# Project <one paragraph: the question or model, data sources, and what "good" looks like (metric + threshold)> ## Commands - Env: `uv sync` · Tests: `uv run pytest -q` · Lint: `uv run ruff check .` · Pipeline: `uv run python -m pipeline` ## Conventions - Raw data is read-only (`data/raw/`); write derived data to `data/processed/` with the script that made it. - Fix random seeds; log dataset version, params and metrics for every run. - Prefer scripts/modules over notebooks for anything reused; notebooks are for exploration only. - Never upload data to third-party services without asking; treat all rows as potentially sensitive. - Report results with the sample size and an uncertainty estimate, not just a point number. ## Checks before done `uv run ruff check . && uv run pytest -q`; show the metric before and after your change. ## Workflow 1. Restate the task in one line and list the files you expect to touch. 2. Make the smallest change that works; keep diffs reviewable. 3. Run the checks below before saying you're done, and show the output. 4. If something is ambiguous, ask one precise question instead of guessing. ## Safety - Never read or print secrets (.env, keys, ~/.ssh). Ask for values instead. - No destructive commands (rm -rf, force-push, reset --hard) without explicit approval; the guard hook blocks them anyway. - Don't add dependencies, services or paid APIs without asking.
@AGENTS.md # Claude Code notes - Use plan mode for multi-file changes. Prefer the installed skills over ad-hoc procedures.
{
"permissions": {
"deny": [
"Bash(rm -rf:*)",
"Bash(rm -fr:*)",
"Bash(sudo:*)",
"Bash(git push --force:*)",
"Bash(git push -f:*)",
"Bash(git reset --hard:*)",
"Bash(git clean -fd:*)",
"Read(./.env)",
"Read(./.env.*)",
"Read(./**/.env)",
"Read(./secrets/**)",
"Read(~/.ssh/**)",
"Read(~/.aws/**)",
"Read(./**/*.pem)"
],
"ask": [
"Bash(git push:*)",
"Bash(npm publish:*)",
"Bash(docker:*)",
"Bash(curl:*)"
],
"defaultMode": "default"
},
"hooks": {
"PreToolUse": [
{
"matcher": "Bash|Read|Edit|Write",
"hooks": [
{
"type": "command",
"command": "\"$CLAUDE_PROJECT_DIR\"/.claude/hooks/guard.sh"
}
]
}
]
}
}
#!/usr/bin/env bash
# .claude/hooks/guard.sh: PreToolUse guard. Exit code 2 blocks the tool call and shows the reason to the model.
# Requires jq. Make executable: chmod +x .claude/hooks/guard.sh
input=$(cat)
tool=$(printf '%s' "$input" | jq -r '.tool_name // empty')
cmd=$(printf '%s' "$input" | jq -r '.tool_input.command // empty')
path=$(printf '%s' "$input" | jq -r '.tool_input.file_path // .tool_input.path // empty')
if [ "$tool" = "Bash" ]; then
if printf '%s' "$cmd" | grep -Eq '(^|[;&| ])(sudo|mkfs|dd if=)|rm -[a-zA-Z]*r[a-zA-Z]*f|rm -[a-zA-Z]*f[a-zA-Z]*r|git push .*(--force|-f( |$))|git reset --hard|git clean -[a-z]*f|curl[^|]*\|[[:space:]]*(ba)?sh|chmod 777'; then
echo "Blocked by guard.sh: destructive command ($cmd). Ask the user to run it manually." >&2; exit 2
fi
fi
if printf '%s %s' "$path" "$cmd" | grep -Eq '(^|/|[[:space:]])\.env($|\.|[[:space:]])|id_rsa|\.pem($|[[:space:]])|\.ssh/|\.aws/credentials'; then
echo "Blocked by guard.sh: secrets file ($path$cmd)." >&2; exit 2
fi
exit 0
{
"mcpServers": {
"context7": {
"type": "http",
"url": "https://mcp.context7.com/mcp"
}
}
}
# AgentRigs starter: Data & ML Generated by https://agentrigs.dev/starters/data-ml from real, well-guarded public rigs: - https://github.com/link7373/agentic-bi-team - https://github.com/a-green-hand-jack/ml-project-repo-agent-native-template - https://github.com/TakaGoto/rag-learning-academy ## Install 1. Unzip into your repo root (`unzip -n` won't overwrite existing files). 2. `chmod +x .claude/hooks/guard.sh` (needs `jq`). 3. Fill in the <placeholders> in AGENTS.md. 4. Install the recommended skills (third-party code, so review first): ```bash npx degit affaan-m/ECC/.kiro/skills/python-patterns .claude/skills/python-patterns npx degit affaan-m/ECC/.agents/skills/eval-harness .claude/skills/eval-harness npx degit obra/superpowers/skills/systematic-debugging .claude/skills/systematic-debugging npx degit anthropics/skills/skills/xlsx .claude/skills/xlsx ``` Optional MCP servers (each adds context tax every turn): ```bash claude mcp add --transport http deepwiki https://mcp.deepwiki.com/mcp ``` Codex / Cursor / other harnesses read AGENTS.md directly; Claude Code reads CLAUDE.md, which imports AGENTS.md.
Recommended skills
Pythonic idioms, PEP 8 standards, type hints, and best practices for building robust, efficient, and maintainable Python applications.
Formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles
Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes
Use this skill any time a spreadsheet file is the primary input or output. This means any task where the user wants to: open, read, edit, or fix an existing .xlsx, .xlsm, .csv, or .tsv file (e.g., add
Built from these rigs
The rules, permissions and hook pattern were distilled from well-guarded public rigs that match this use case (guardrail score ≥ 3, lean context):