Break tasks → limit permissions → sandbox dry-run — agents act, you accept
Break a repetitive job into agent steps → sandbox dry-run → accept output → name the reusable workflow
~25 min
Suggested tool: Bolt.new
Agent failures usually mean too much permission or fuzzy success criteria — “clean up the data” becomes prod edits. This scene trains a verifiable automation loop: break repetitive work into ≤6 steps (input/action/output/check each) → cap folders and network → sandbox dry-run → human confirm before writes → save as skill/workflow → document failure fallback. Good for: CSV summaries, batch renames, weekly Markdown reports, in-repo scripting. Bad for: unaudited prod DB access, auto-send email/payments, pasting API keys into prompts. Example: job=summarize examples/feedback.csv into output/report.md daily; sandbox=read examples/ only, write output/ only, no network. Acceptable: step table with success criteria; after dry-run report.md has three agreed headings; you know which steps need human OK. Tool pick: Claude Code in the IDE; Devin for longer standalone jobs; Bolt.new for quick prototypes. Start small — verify on official sites.
You are a senior automation engineer. Break this repetitive job into ≤6 agent steps: Job: summarize examples/feedback.csv into output/report.md weekly Sandbox: read examples/ only, write output/ only, no network Rules: project files/commands only; no prod; no network; no real business data edits; each step has input/action/output/check; mark external steps “needs human OK”. Markdown table: Step | Agent action | Permission | Success criteria | Failure fallback
Outcome checklist
Self rating (1–5)
You assist Claude Code / Devin. Run a **minimal sandbox dry-run** from this step list: Steps: [steps] Goal: [job] Sandbox: [sandbox] List paths before writes and wait; if files missing, list 3 samples to prepare; no out-of-scope paths. Output: 1) Steps run 2) Output paths 3) Verification (structure/headings/errors) 4) Reusable workflow name
Audit this agent plan for excessive permissions (review only — do not execute): Job: [job] Sandbox: [sandbox] Steps: [steps] List: high-risk actions, steps to delete or mark human-OK, least-privilege rewrites (folders/network/commands).
From the success criteria, make a human acceptance checklist for this agent output: Goal: [job] Expected structure: [success_criteria] Actual output summary (paste OK): [output_sample] 6–8 checks: file exists, headings/fields, no scope creep, no invented data, rerunnable, rollback on failure.
Dry-run passed. Turn these steps into a reusable workflow doc (Markdown): Job: [job] Sandbox: [sandbox] Steps: [steps] Include: workflow name, trigger (command/skill/manual), prerequisites, success criteria, failure fallback, variables to swap next time. Do not recommend new tools or invent prices.
Review what I’m about to paste to an agent for sensitive data (audit only): [steps] List secrets/tokens, internal hosts, prod paths, customer IDs. For each: delete/redact/OK. If none, say so — still remind me to human-check.
Anthropic's terminal-native AI coding assistant with deep codebase understanding, multi-file editing, test generation, and Git integration.
Cognition AI's fully autonomous AI software engineer that independently handles the full dev cycle from requirements to deployment.
StackBlitz's browser-based AI full-stack app generator that creates runnable web apps from prompts with one-click deploy.
After you leave with templates,
Continue your learning path