~ / starters / data-ml

Data & ML starter

Notebooks, pipelines, evals, analysis

1.8kcontext tax / turn · Featherweight
5/5guardrails
1 · 4MCP servers · skills
ARCHETYPEFort Knox

Install

Run in your project root. Existing files are never overwritten (unzip -n / [ -e … ] skip them). Review AGENTS.md and fill in the <placeholders>.

$ curl -fsSL https://agentrigs.dev/starters/agentrigs-data-ml.zip -o agentrigs-starter.zip && unzip -n agentrigs-starter.zip && chmod +x .claude/hooks/guard.sh && rm agentrigs-starter.zip
$ npx degit affaan-m/ECC/.kiro/skills/python-patterns .claude/skills/python-patterns
$ npx degit affaan-m/ECC/.agents/skills/eval-harness .claude/skills/eval-harness
$ npx degit obra/superpowers/skills/systematic-debugging .claude/skills/systematic-debugging
$ npx degit anthropics/skills/skills/xlsx .claude/skills/xlsx

Or download agentrigs-data-ml.zip (6 files). Skills are fetched from their source repos with npx degit.

mkdir -p .claude/hooks
[ -e AGENTS.md ] && echo "skip AGENTS.md (exists)" || cat > AGENTS.md <<'AGENTRIGS_EOF'
# Project
<one paragraph: the question or model, data sources, and what "good" looks like (metric + threshold)>

## Commands
- Env: `uv sync` · Tests: `uv run pytest -q` · Lint: `uv run ruff check .` · Pipeline: `uv run python -m pipeline`

## Conventions
- Raw data is read-only (`data/raw/`); write derived data to `data/processed/` with the script that made it.
- Fix random seeds; log dataset version, params and metrics for every run.
- Prefer scripts/modules over notebooks for anything reused; notebooks are for exploration only.
- Never upload data to third-party services without asking; treat all rows as potentially sensitive.
- Report results with the sample size and an uncertainty estimate, not just a point number.

## Checks before done
`uv run ruff check . && uv run pytest -q`; show the metric before and after your change.

## Workflow
1. Restate the task in one line and list the files you expect to touch.
2. Make the smallest change that works; keep diffs reviewable.
3. Run the checks below before saying you're done, and show the output.
4. If something is ambiguous, ask one precise question instead of guessing.

## Safety
- Never read or print secrets (.env, keys, ~/.ssh). Ask for values instead.
- No destructive commands (rm -rf, force-push, reset --hard) without explicit approval; the guard hook blocks them anyway.
- Don't add dependencies, services or paid APIs without asking.
AGENTRIGS_EOF
[ -e CLAUDE.md ] && echo "skip CLAUDE.md (exists)" || cat > CLAUDE.md <<'AGENTRIGS_EOF'
@AGENTS.md

# Claude Code notes
- Use plan mode for multi-file changes. Prefer the installed skills over ad-hoc procedures.
AGENTRIGS_EOF
[ -e .claude/settings.json ] && echo "skip .claude/settings.json (exists)" || cat > .claude/settings.json <<'AGENTRIGS_EOF'
{
  "permissions": {
    "deny": [
      "Bash(rm -rf:*)",
      "Bash(rm -fr:*)",
      "Bash(sudo:*)",
      "Bash(git push --force:*)",
      "Bash(git push -f:*)",
      "Bash(git reset --hard:*)",
      "Bash(git clean -fd:*)",
      "Read(./.env)",
      "Read(./.env.*)",
      "Read(./**/.env)",
      "Read(./secrets/**)",
      "Read(~/.ssh/**)",
      "Read(~/.aws/**)",
      "Read(./**/*.pem)"
    ],
    "ask": [
      "Bash(git push:*)",
      "Bash(npm publish:*)",
      "Bash(docker:*)",
      "Bash(curl:*)"
    ],
    "defaultMode": "default"
  },
  "hooks": {
    "PreToolUse": [
      {
        "matcher": "Bash|Read|Edit|Write",
        "hooks": [
          {
            "type": "command",
            "command": "\"$CLAUDE_PROJECT_DIR\"/.claude/hooks/guard.sh"
          }
        ]
      }
    ]
  }
}
AGENTRIGS_EOF
[ -e .claude/hooks/guard.sh ] && echo "skip .claude/hooks/guard.sh (exists)" || cat > .claude/hooks/guard.sh <<'AGENTRIGS_EOF'
#!/usr/bin/env bash
# .claude/hooks/guard.sh: PreToolUse guard. Exit code 2 blocks the tool call and shows the reason to the model.
# Requires jq. Make executable: chmod +x .claude/hooks/guard.sh
input=$(cat)
tool=$(printf '%s' "$input" | jq -r '.tool_name // empty')
cmd=$(printf '%s' "$input" | jq -r '.tool_input.command // empty')
path=$(printf '%s' "$input" | jq -r '.tool_input.file_path // .tool_input.path // empty')
if [ "$tool" = "Bash" ]; then
  if printf '%s' "$cmd" | grep -Eq '(^|[;&| ])(sudo|mkfs|dd if=)|rm -[a-zA-Z]*r[a-zA-Z]*f|rm -[a-zA-Z]*f[a-zA-Z]*r|git push .*(--force|-f( |$))|git reset --hard|git clean -[a-z]*f|curl[^|]*\|[[:space:]]*(ba)?sh|chmod 777'; then
    echo "Blocked by guard.sh: destructive command ($cmd). Ask the user to run it manually." >&2; exit 2
  fi
fi
if printf '%s %s' "$path" "$cmd" | grep -Eq '(^|/|[[:space:]])\.env($|\.|[[:space:]])|id_rsa|\.pem($|[[:space:]])|\.ssh/|\.aws/credentials'; then
  echo "Blocked by guard.sh: secrets file ($path$cmd)." >&2; exit 2
fi
exit 0
AGENTRIGS_EOF
[ -e .mcp.json ] && echo "skip .mcp.json (exists)" || cat > .mcp.json <<'AGENTRIGS_EOF'
{
  "mcpServers": {
    "context7": {
      "type": "http",
      "url": "https://mcp.context7.com/mcp"
    }
  }
}
AGENTRIGS_EOF
[ -e AGENTRIGS-STARTER.md ] && echo "skip AGENTRIGS-STARTER.md (exists)" || cat > AGENTRIGS-STARTER.md <<'AGENTRIGS_EOF'
# AgentRigs starter: Data & ML

Generated by https://agentrigs.dev/starters/data-ml from real, well-guarded public rigs:
- https://github.com/link7373/agentic-bi-team
- https://github.com/a-green-hand-jack/ml-project-repo-agent-native-template
- https://github.com/TakaGoto/rag-learning-academy

## Install
1. Unzip into your repo root (`unzip -n` won't overwrite existing files).
2. `chmod +x .claude/hooks/guard.sh` (needs `jq`).
3. Fill in the <placeholders> in AGENTS.md.
4. Install the recommended skills (third-party code, so review first):

```bash
npx degit affaan-m/ECC/.kiro/skills/python-patterns .claude/skills/python-patterns
npx degit affaan-m/ECC/.agents/skills/eval-harness .claude/skills/eval-harness
npx degit obra/superpowers/skills/systematic-debugging .claude/skills/systematic-debugging
npx degit anthropics/skills/skills/xlsx .claude/skills/xlsx
```

Optional MCP servers (each adds context tax every turn):

```bash
claude mcp add --transport http deepwiki https://mcp.deepwiki.com/mcp
```

Codex / Cursor / other harnesses read AGENTS.md directly; Claude Code reads CLAUDE.md, which imports AGENTS.md.
AGENTRIGS_EOF
chmod +x .claude/hooks/guard.sh
npx degit affaan-m/ECC/.kiro/skills/python-patterns .claude/skills/python-patterns  # skill: python-patterns
npx degit affaan-m/ECC/.agents/skills/eval-harness .claude/skills/eval-harness  # skill: eval-harness
npx degit obra/superpowers/skills/systematic-debugging .claude/skills/systematic-debugging  # skill: systematic-debugging
npx degit anthropics/skills/skills/xlsx .claude/skills/xlsx  # skill: xlsx

Optional MCP servers (add only if you use them; each adds tool schemas to every turn):

$ claude mcp add --transport http deepwiki https://mcp.deepwiki.com/mcp

Context tax

Instructions: 384MCP tool schemas (est.): 1.2kSkill metadata: 176Subagent metadata: 0

Claude Code pays for CLAUDE.md plus the imported AGENTS.md (~384 tok of instructions); Codex reads AGENTS.md alone (~353 tok). Index median: 2.2k tok. How we measure.

Guardrails

✓
Blocks destructive commands 9 deny/ask rule(s) e.g. Bash(rm -rf:*)
✓
Protects secrets Denies reads like Read(./.env)
✓
Pre-tool screening hook 1 PreToolUse hook(s)
✓
No YOLO mode Permission prompts stay on
✓
Sandbox or ask-first rules 4 ask rule(s), e.g. Bash(git push:*)

Files

# Project
<one paragraph: the question or model, data sources, and what "good" looks like (metric + threshold)>

## Commands
- Env: `uv sync` · Tests: `uv run pytest -q` · Lint: `uv run ruff check .` · Pipeline: `uv run python -m pipeline`

## Conventions
- Raw data is read-only (`data/raw/`); write derived data to `data/processed/` with the script that made it.
- Fix random seeds; log dataset version, params and metrics for every run.
- Prefer scripts/modules over notebooks for anything reused; notebooks are for exploration only.
- Never upload data to third-party services without asking; treat all rows as potentially sensitive.
- Report results with the sample size and an uncertainty estimate, not just a point number.

## Checks before done
`uv run ruff check . && uv run pytest -q`; show the metric before and after your change.

## Workflow
1. Restate the task in one line and list the files you expect to touch.
2. Make the smallest change that works; keep diffs reviewable.
3. Run the checks below before saying you're done, and show the output.
4. If something is ambiguous, ask one precise question instead of guessing.

## Safety
- Never read or print secrets (.env, keys, ~/.ssh). Ask for values instead.
- No destructive commands (rm -rf, force-push, reset --hard) without explicit approval; the guard hook blocks them anyway.
- Don't add dependencies, services or paid APIs without asking.
@AGENTS.md

# Claude Code notes
- Use plan mode for multi-file changes. Prefer the installed skills over ad-hoc procedures.
{
  "permissions": {
    "deny": [
      "Bash(rm -rf:*)",
      "Bash(rm -fr:*)",
      "Bash(sudo:*)",
      "Bash(git push --force:*)",
      "Bash(git push -f:*)",
      "Bash(git reset --hard:*)",
      "Bash(git clean -fd:*)",
      "Read(./.env)",
      "Read(./.env.*)",
      "Read(./**/.env)",
      "Read(./secrets/**)",
      "Read(~/.ssh/**)",
      "Read(~/.aws/**)",
      "Read(./**/*.pem)"
    ],
    "ask": [
      "Bash(git push:*)",
      "Bash(npm publish:*)",
      "Bash(docker:*)",
      "Bash(curl:*)"
    ],
    "defaultMode": "default"
  },
  "hooks": {
    "PreToolUse": [
      {
        "matcher": "Bash|Read|Edit|Write",
        "hooks": [
          {
            "type": "command",
            "command": "\"$CLAUDE_PROJECT_DIR\"/.claude/hooks/guard.sh"
          }
        ]
      }
    ]
  }
}
#!/usr/bin/env bash
# .claude/hooks/guard.sh: PreToolUse guard. Exit code 2 blocks the tool call and shows the reason to the model.
# Requires jq. Make executable: chmod +x .claude/hooks/guard.sh
input=$(cat)
tool=$(printf '%s' "$input" | jq -r '.tool_name // empty')
cmd=$(printf '%s' "$input" | jq -r '.tool_input.command // empty')
path=$(printf '%s' "$input" | jq -r '.tool_input.file_path // .tool_input.path // empty')
if [ "$tool" = "Bash" ]; then
  if printf '%s' "$cmd" | grep -Eq '(^|[;&| ])(sudo|mkfs|dd if=)|rm -[a-zA-Z]*r[a-zA-Z]*f|rm -[a-zA-Z]*f[a-zA-Z]*r|git push .*(--force|-f( |$))|git reset --hard|git clean -[a-z]*f|curl[^|]*\|[[:space:]]*(ba)?sh|chmod 777'; then
    echo "Blocked by guard.sh: destructive command ($cmd). Ask the user to run it manually." >&2; exit 2
  fi
fi
if printf '%s %s' "$path" "$cmd" | grep -Eq '(^|/|[[:space:]])\.env($|\.|[[:space:]])|id_rsa|\.pem($|[[:space:]])|\.ssh/|\.aws/credentials'; then
  echo "Blocked by guard.sh: secrets file ($path$cmd)." >&2; exit 2
fi
exit 0
{
  "mcpServers": {
    "context7": {
      "type": "http",
      "url": "https://mcp.context7.com/mcp"
    }
  }
}
# AgentRigs starter: Data & ML

Generated by https://agentrigs.dev/starters/data-ml from real, well-guarded public rigs:
- https://github.com/link7373/agentic-bi-team
- https://github.com/a-green-hand-jack/ml-project-repo-agent-native-template
- https://github.com/TakaGoto/rag-learning-academy

## Install
1. Unzip into your repo root (`unzip -n` won't overwrite existing files).
2. `chmod +x .claude/hooks/guard.sh` (needs `jq`).
3. Fill in the <placeholders> in AGENTS.md.
4. Install the recommended skills (third-party code, so review first):

```bash
npx degit affaan-m/ECC/.kiro/skills/python-patterns .claude/skills/python-patterns
npx degit affaan-m/ECC/.agents/skills/eval-harness .claude/skills/eval-harness
npx degit obra/superpowers/skills/systematic-debugging .claude/skills/systematic-debugging
npx degit anthropics/skills/skills/xlsx .claude/skills/xlsx
```

Optional MCP servers (each adds context tax every turn):

```bash
claude mcp add --transport http deepwiki https://mcp.deepwiki.com/mcp
```

Codex / Cursor / other harnesses read AGENTS.md directly; Claude Code reads CLAUDE.md, which imports AGENTS.md.

Recommended skills

python-patterns used in 36 rigs · from affaan-m/ECC

Pythonic idioms, PEP 8 standards, type hints, and best practices for building robust, efficient, and maintainable Python applications.

eval-harness used in 43 rigs · from affaan-m/ECC

Formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles

systematic-debugging used in 102 rigs · from obra/superpowers

Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes

xlsx used in 48 rigs · from anthropics/skills

Use this skill any time a spreadsheet file is the primary input or output. This means any task where the user wants to: open, read, edit, or fix an existing .xlsx, .xlsm, .csv, or .tsv file (e.g., add

Built from these rigs

The rules, permissions and hook pattern were distilled from well-guarded public rigs that match this use case (guardrail score ≥ 3, lean context):

Customised it? Paste your files into the Rig Doctor to re-check the tax and guardrails.

Other starters

copied ✓