Technology & society

Harden your AI agent

There is rarely a good reason to give your AI agent access to files outside the folder you're operating in - it's a huge security risk. Not just because your AI agent can misuse them "locally", but also because you risk the AI agent sending sensitive data to others, e.g. by including them in a prompt to an AI provider or saving them on an obscure German Wikipedia page - AI works in mysterious ways. You can easily prevent your agent from doing that by using Agent Safehouse, which is an AI sandbox for macOS that blocks the AI agent's access to files and folders at the kernel level. You can use Agent Safehouse by starting your preferred AI harness with e.g. safehouse pi. I have it as an alias in my .zshrc file: alias safepi='safehouse pi'. But that doesn't prevent the AI agent from reading sensitive files in the folder it operates in, e.g. .env files.

To prevent that, I have, in my scripts folder, the configuration file sensitive-deny.sb, which prevents the AI agent from reading various sensitive files. Here is the contents of sensitive-deny.sb:

;; ---------------------------------------------------------------------------
;; sensitive-deny.sb - terminal sensitive-file deny layer
;;
;; WHAT THIS IS
;;   A profile fragment that is loaded AFTER all path grants so its denies win.
;;   SBPL rule ordering is "later rules win", so appending this last is the
;;   supported way to make must-not-read paths stick even though the workdir
;;   (or tool config dirs) are readable.
;;
;; USAGE - upstream Safehouse (append-profile is its documented final layer):
;;   safehouse --append-profile="$HOME/.config/sandbox-exec/sensitive-deny.sb" -- pi
;;
;; USAGE - the standalone wrapper from this setup:
;;   run-sandboxed.sh appends this file after the launch-time workdir grants.
;;
;; This file uses HOME_DIR/home-literal from Safehouse's 00-base profile
;; (agent.sb defines the same helpers), which is in scope in both usages
;; because the file is always loaded after the base profile.
;; ---------------------------------------------------------------------------
;;
;; SCOPE
;;   Severity 1: `.env` family, `.dev.vars`, key material, credential stores,
;;               service-account JSONs, terraform state, direnv. Denied below.
;;   Severity 2: global tool credentials (~/.npmrc, gh hosts.yml, ...).
;;               Commented out at the bottom - enabling them breaks those
;;               tools inside the sandbox.
;;   Everything else (SSH keys, cloud dirs, keychain, ...) is already covered
;;   by `(deny default)`; add explicit denies only as defense-in-depth.
;;
;; CAVEATS (name-based deny limits)
;;   - APFS is case-insensitive: `.ENV` can still resolve to `.env`.
;;   - hardlinked/renamed copies of a secret are readable under the new name.
;;   - a committed .env stays readable from git history/objects (`.git/`).
;;   - writes are denied too, so `cp .env.example .env` will fail; drop
;;     `file-write*` from a deny rule if you need template generation.

;; ===========================================================================
;; 1. Environment files and direnv
;; ===========================================================================
(deny file-read* file-read-metadata file-write*
    ;; .env and *.env (empty prefix before the basename is intentional)
    (regex #"^.*/[^/]*\.env$")
    ;; .env.local, .env.production, foo.env.staging, ...
    (regex #"^.*/[^/]*\.env\..*$")
    ;; Cloudflare Workers / Wrangler local secrets (incl. built dist copies)
    (regex #"^.*/\.dev\.vars$")
    (regex #"^.*/\.dev\.vars\..*$")
    ;; direnv config, frequently carries secrets
    (regex #"^.*/\.envrc$")
)

;; ===========================================================================
;; 2. Key material and credential stores
;; ===========================================================================
(deny file-read* file-read-metadata file-write*
    ;; keys and certs (remove *.key if test fixtures need it)
    (regex #"^.*/[^/]*\.pem$")
    (regex #"^.*/[^/]*\.key$")
    (regex #"^.*/[^/]*\.p12$")
    (regex #"^.*/[^/]*\.pfx$")
    (regex #"^.*/[^/]*\.jks$")
    (regex #"^.*/[^/]*\.keystore$")
    (regex #"^.*/id_rsa$")
    (regex #"^.*/id_ed25519$")
    (regex #"^.*/id_ecdsa$")
    ;; credential stores
    (regex #"^.*/\.netrc$")
    (regex #"^.*/_netrc$")
    (regex #"^.*/\.git-credentials$")
    (regex #"^.*/\.htpasswd$")
    (regex #"^.*/\.sentryclirc$")
)

;; ===========================================================================
;; 3. Cloud / infrastructure secrets
;; ===========================================================================
(deny file-read* file-read-metadata file-write*
    (regex #"^.*/service[-_]?account[^/]*\.json$")
    (regex #"^.*/firebase-adminsdk[^/]*\.json$")
    (regex #"^.*/secrets\.ya?ml$")
    (regex #"^.*/[^/]*\.tfvars$")
    (regex #"^.*/[^/]*\.tfvars\.json$")
    (regex #"^.*/[^/]*\.tfstate$")
    (regex #"^.*/[^/]*\.tfstate\.backup$")
)

;; ===========================================================================
;; 4. Non-secret templates re-opened (must come after the denies above)
;;    Remove if your repo keeps real values in these files.
;; ===========================================================================
(allow file-read*
    (regex #"^.*/[^/]*\.env\.example$")
    (regex #"^.*/[^/]*\.env\.sample$")
    (regex #"^.*/[^/]*\.env\.template$")
    (regex #"^.*/[^/]*\.env\.dist$")
)

;; ===========================================================================
;; 5. OPT-IN: global tool credentials (commented by default)
;;    Each one breaks the corresponding tool inside the sandbox.
;;    home-literal is available because this file always loads after 00-base.
;; ===========================================================================
;;
(deny file-read* file-read-metadata file-write*
    (home-literal "/.npmrc")                    ;; npm auth token -> breaks npm publish/private installs
    (home-literal "/.pypirc")                   ;; PyPI token -> breaks twine upload
    (home-literal "/.config/gh/hosts.yml")      ;; gh OAuth token -> breaks gh auth/PR automation
    (home-literal "/.config/git/credentials")   ;; git credential store
    (home-literal "/.git-credentials")
    (home-literal "/.netrc")                    ;; breaks curl/heroku-style tooling
)

I include the configuration file in my alias in ~/.zshrc like this:

alias safepi='safehouse --append-profile="$HOME/Scripts/sensitive-deny.sb" -- pi'

When you now ask your AI agent to print DATABASE_URL from the .env file, Agent Safehouse will block it. If I, for example, run the prompt Can you read the DATABASE_URL from the .env file?, the AI agent tries in several different ways to read the .env file, which it can see exists, but cannot read. In the end the AI agent concludes (in English):

No — I can't. The file /Users/.../.env exists, but every access attempt is denied at the OS level:

- `cat`, `grep`, `ls`, `stat`, `xattr` → Operation not permitted (EPERM)
- Node's fs.lstatSync('.env') → EPERM
- Pi's own read tool → EPERM
- Other files in the repo (e.g. .gitignore) are readable, so it's specific to this file