Orchestrate sandboxed coding agents in TypeScript with sandcastle.run()
by mattpocockTypeScript
Last 12 weeks · 0 commits
2 of 6 standards met
Docker continues to evolve their Docker Sandboxes feature, which runs AI agents in microVM-isolated environments via the sbx CLI. This proposal is to add it as a new provider — docker-sbx — following the same pattern as the existing docker and podman providers. Why microVMs over regular containers The docker and podman providers share the host kernel. For an agentic coding tool, this matters: agents regularly build images, run test containers, and execute arbitrary generated code. Docker Sandboxes addresses this with microVM-based isolation: Each sandbox gets its own kernel, filesystem, and Docker daemon — container escapes and kernel exploits can't reach the host Agents can run docker build/docker run inside the sandbox natively — no --privileged flag or DinD hacks needed Meaningfully stronger security posture for users running untrusted or AI-generated code Custom sandbox images Just like the existing docker/podman providers scaffold a Dockerfile and expose sandcastle docker build-image, docker-sbx would support custom sandbox templates. Users extend the official base image with their toolchain: The sandcastle docker-sbx build-image command would build and push to a registry (required since sbx pulls from OCI registries, not local storage), then sbx run -t claude would use it. This gives the same "optimize your starting environment for your codebase" workflow that Dockerfile-based providers already offer. ToS clarity (re: #191) The sbx CLI handles Claude Code authentication and subscription activation itself — Sandcastle wouldn't need to provide guidance on auth at all. Docker is an official partner running Claude Code inside their product with Anthropic's knowledge. This sidesteps the ambiguity in #191 entirely: the maintainer delegates the compliance surface to Docker rather than owning it. Implementation scope This would be a bind-mount provider modeled on src/sandboxes/docker.ts: New file: src/sandboxes/docker-sbx.ts (~200–300 lines) Entry in SANDBOX_PROVIDER_REGISTRY in src/InitService.ts CLI subcommand: sandcastle docker-sbx build-image (build + push to registry) Template: Dockerfile extending docker/sandbox-templates:claude-code-docker The provider abstraction already handles everything else (worktrees, env merging, lifecycle, streaming).
Summary Closes #863. The agent provider never populated : only extracted the final assistant text from the event and dropped the block that Pi includes on each assistant message. This change parses that usage and emits a stream event, mapping Pi's shape onto the existing Claude-shaped : Because carries the full conversation, usage is summed across every assistant message in the event so the iteration reflects all turns, not just the final one. The parser is also tolerant of older/malformed usage objects (returns rather than emitting a partial/NaN usage, matching ). Test plan — 219 passed Added three tests: extracts usage from an assistant message sums usage across multiple assistant messages omits the event when the usage object is missing/malformed — clean on changed files — clean Note on retries The issue's extra question about aggregating usage across retries is left out of scope: each attempt is a separate iteration, and is currently stored per iteration. Changing that is a design decision for the Orchestrator, not the Pi parser. Changeset Added ().
Problem Statement The current orchestration system for running AI coding agents (like Claude Code) inside Docker containers is implemented as a collection of bash scripts. These scripts are fragile, hard to test, and painful to debug — particularly the git sync logic (bundle in, format-patch out), which was the source of many bugs during development. There is no automated test coverage, so regressions are caught manually. The bash scripts are also not distributable — other developers can't a tool and start using it. Solution Build Sandcastle, a TypeScript CLI using Effect, that replaces the bash scripts with a well-tested, distributable tool. The core insight is that the sync logic (the hardest, most bug-prone part) can be tested against a filesystem-based sandbox without Docker, using real git repos in temp directories. Docker becomes an optional layer that implements the same interface. The CLI will be published to npm via changesets so anyone can install and use it. User Stories 1. As a developer, I want to install Sandcastle via npm, so that I can orchestrate AI agents without cloning a repo of bash scripts. 2. As a developer, I want to run to create and configure a Docker container for my project, so that I have an isolated environment ready for an agent. 3. As a developer, I want to run to execute multiple iterations of an AI agent against my repo's open GitHub issues, so that work gets done while I'm away. 4. As a developer, I want to run to open an interactive agent session inside the sandbox, so that I can collaborate with the agent in real time. 5. As a developer, I want to run to tear down the container and image, so that I can free up resources. 6. As a developer, I want to run to push my host repo state into the sandbox, so that I can debug sync issues in isolation. 7. As a developer, I want to run to pull sandbox changes back to my host, so that I can debug sync issues in isolation. 8. As a developer, I want sync-in to transfer my full repo state (including unpushed commits) into the sandbox via git bundle, so that the sandbox always matches my working state. 9. As a developer, I want sync-out to extract new commits from the sandbox via format-patch and apply them on my host with git am, so that I get clean commit history back. 10. As a developer, I want sync-out to also capture uncommitted changes (staged, unstaged, and untracked files) from the sandbox, so that I don't lose in-progress work. 11. As a developer, I want the iteration loop to stop early when the agent emits a completion signal, so that I don't waste iterations when all tasks are done. 12. As a developer, I want a config file in my repo with a command, so that I can run (or any setup command) automatically after code is synced into the sandbox. 13. As a developer, I want the agent to be swappable (Claude Code today, Codex tomorrow), so that I'm not locked into one AI tool. 14. As a contributor to Sandcastle, I want integration tests that verify the sync logic against real git repos on the filesystem, so that I can refactor the sync implementation with confidence. 15. As a contributor to Sandcastle, I want the filesystem sandbox to serve as a fast, Docker-free test environment, so that CI runs quickly and I can iterate locally without Docker. Implementation Decisions Module design Sandbox service: An Effect exposing three operations: , , . This is the adapter interface that decouples sync logic from the execution environment. Docker layer: Implementation of the Sandbox service using and . Handles container lifecycle (build, start, stop, destroy). Filesystem layer: Implementation of the Sandbox service using local shell commands and against a separate directory. Used for integration testing. Sync service: Built on top of the Sandbox service. Implements sync-in (git bundle → copyIn → fetch → reset → clean → postSyncIn) and sync-out (format-patch → copyOut → git am, plus uncommitted changes). This is the core domain logic. Orchestrator: Drives the iteration loop — sync-in, invoke agent, check for new commits, sync-out, check completion signal, repeat. Thin module. Config: Reads from the target repo. v1 schema: . CLI: Built with . Six commands: , , , , , . Thin wiring that composes the above modules. Technical decisions Effect as backbone: Every function returns . Services defined with , implementations provided as s. Full Effect program top to bottom. Git bundle for sync-in: Avoids GitHub round-trip, captures unpushed commits. Same strategy as the bash scripts. Format-patch for sync-out: Produces clean, replayable commits. Applied with for robust merging. CLI framework: Build tool: () Runtime: Node.js Publishing: npm via changesets Single package: Everything in the repo root, no monorepo. Testing Decisions What makes a good test Tests should verify external behavior of the sync service — given a host repo in state X, after sync-in and sync-out, the host repo should be in state Y. Tests should NOT assert on implementation details like which git commands were run or how many temp files were created. This allows swapping the sync mechanism later without rewriting tests. What gets tested Sync service only. The Sandbox service implementations (Docker layer, filesystem layer) are trivially thin. The Orchestrator is thin glue. The sync logic is where bugs hide. Integration test approach Each test creates real git repos in temp directories using the filesystem layer: A "host" repo (the developer's repo) A "sandbox" directory (simulating the container's filesystem) Tests exercise the sync service through the Sandbox service interface, using the filesystem layer. Test scenarios Sync-in (host → sandbox): 1. Clean repo syncs correctly — sandbox HEAD matches host HEAD 2. Repo with uncommitted changes — sandbox gets the committed state 3. Repo with unpushed commits — bundle captures them 4. Re-sync after sandbox has diverged — sandbox resets to host state Sync-out (sandbox → host): 1. Single new commit in sandbox — patch applies cleanly on host 2. Multiple new commits — all patches apply in order 3. Uncommitted staged changes come back 4. Uncommitted unstaged changes come back 5. Untracked files come back 6. No changes in sandbox — no-op, no error Round-trip: 1. Sync-in, make commit in sandbox, sync-out — host has the new commit 2. Sync-in, sync-out, sync-in again — stable, no drift Failure cases: 1. Patch conflicts on apply (host changed between sync-in and sync-out) 2. Empty repo / initial commit edge case Test framework Vitest with . Out of Scope Lifecycle hooks (onSetup, onSyncIn, beforeRun, etc.) — deferred to a future version. Documented in . Prompt system — the composable/overridable prompt architecture. v1 will use a static default prompt. Agent abstraction — v1 will invoke Claude Code directly. The agent service interface for swapping in Codex etc. is a future concern. Config beyond — prompt overrides, Docker settings, iteration settings will be added as needed. GitHub integration** — issue fetching is currently done inside the sandbox by the agent. No need to model it in Sandcastle v1. Further Notes The existing bash scripts in serve as the reference implementation. The TypeScript version should produce identical sync behavior. See for canonical terminology (Sandbox, Host, Agent, Sync-in, Sync-out, Bundle, Patch, Iteration, etc.). See for future config and hooks design notes.
Summary The template's Phase 3 (Merge) does locally and then . It never pushes. As a result, GitHub considers issues closed while is untouched — the "close-on-merge" invariant only holds for the local merge, not the merge that actually ships. Repro 1. Fresh repo. with the template. 2. Open a small -labelled issue with a trivial body (e.g. "add a file"). 3. . Let one full iteration complete. 4. Inspect: GitHub UI → issue is closed, comment "Completed by Sandcastle". → one or more local commits not on origin. Anyone running against sees no change. Expected When fires, the corresponding code should already be on (or at least somewhere the team can pull). "Closed" should imply "shipped." Why it matters State drift: GitHub Issues view says "done"; disagrees. Confusing for collaborators, and for any tooling that queries open issues from origin's perspective. No review surface: no PR, no CI run on the merged code, no bot reviews, no human merge button. Crash recovery: if the host dies after the local merge but before a manual push, the work lives on a local branch only — but issues are already marked done. Suggested fix In increasing scope: 1. Minimal: have the merge prompt also after the merge commit and before . Aligns the two states for GitHub-backed repos. 2. Better**: replace the local-merge phase with a per-issue host-publish step: Optional follow-up: drive-to-green agent that waits for status checks and addresses bot reviews before exiting. GitHub then closes the issue automatically when the PR merges. The manual goes away and close-on-merge becomes mechanical. We've been running variant (2) across two repos and it removes the state-drift class of bugs entirely. Happy to send a PR against the template if you'd like it upstreamed. Related #585 (practical-guide thread — opener wrestling with similar mental-model issues) #480 (closed feedback issue; one note mentions a successful GitLab+MR flow, suggesting (2)-style integration is feasible)
Repository: mattpocock/sandcastle. Description: Orchestrate sandboxed coding agents in TypeScript with sandcastle.run() Stars: 8205, Forks: 881. Primary language: TypeScript. Languages: TypeScript (99.7%), JavaScript (0.2%), Dockerfile (0.1%), Shell (0%). License: MIT. Latest release: v0.12.0 (3mo ago). Open PRs: 62, open issues: 113. Last activity: 3mo ago. Community health: 42%. Top contributors: mattpocock, github-actions[bot], dichioniccolo, bwcampbell9, jerome-benoit, mellson, lautarosegura, NitayRabi, screenfluent, aravindcm49.