Run a maintenance task with the harness
The harness (CADS-agent-marketplace Phase 2, ct-agent harness run) lets a signed,
publisher-authored task drive a bounded local-LLM agent against a service you’ve already
installed from a manifest — think
“apply this fix,” not “run whatever code I feel like.” It is a separate mechanism from
manifest install — that page gets a
service running; this one lets a trusted publisher’s signed instructions maintain it afterward,
without ever handing that publisher shell access to your machine. Every command and every output
below was actually run, against a real locally-built ct-agent and manifest-core — including
every rejection path.
Why this is safe to run at all
Three things bound what a task can actually do, source-grounded in
CADS-agent-marketplace/crates/harness-core:
- No shell, three tools total.
read_file,write_file, andrebuild(docker compose buildonly — neverup/down) are the entire attack surface (tools.rs). There is no bash tool, no arbitrary command execution, at all. - Real containment, not a lexical check. Every file path is resolved and symlink-canonicalized
against the bundle directory before use (
containment.rs) — a..or absolute path is refused before it’s even joined, and a symlink a malicious bundle planted to point outside the bundle is caught too, because containment is checked against the real, resolved filesystem path, not the string..env— the installer’s own secrets file — is refused by name at any depth, regardless of what a task’s prompt asks for. - A hard local ceiling on top of the signed one. A task’s own
max_turnsis part of what it signs (tampering with it after signing invalidates the signature), but nothing upstream bounds how high a compromised or buggy publisher key could set it — so the harness itself refuses anything over 200 turns, regardless of what’s signed. This matters specifically because therebuildtool never touches your LiteLLM spend budget at all, so a task that callsrebuildevery turn burns realdocker compose buildtime completely unbounded by any token-spend cap.
Two independent allowlists gate a run before any of this even starts: the manifest’s own publisher
trust allowlist (same one manifest activate uses) and a separate harness-side model allowlist
— even a trusted publisher’s task naming a model you haven’t allowed is refused, so a compromised
trust-allowlisted key can’t be used to drive spend against an arbitrary, expensive model.
1. Have an activated manifest to run against
The harness needs a real, already-installed bundle — reusing exactly the manifest install recipe:
mkdir bundle && cd bundle
printf '#!/bin/sh\necho "docs-example manifest installed successfully"\n' > hello.sh
printf '#!/bin/sh\necho "verify.sh: ok"\nexit 0\n' > verify.sh
chmod +x hello.sh verify.sh
cd .. && tar -czf bundle.tar.gz -C bundle .
./ct-agent channel init # if you don't already have a holder identity
CT_MANIFEST_NAME=docs-harness-example CT_MANIFEST_VERSION=0.1.0 \
CT_MANIFEST_BUNDLE_URL="$PWD/bundle.tar.gz" \
CT_MANIFEST_BUNDLE_SHA256=$(sha256sum bundle.tar.gz | cut -d' ' -f1) \
CT_MANIFEST_KIND=binary CT_MANIFEST_COMPOSE_FILE=hello.sh CT_MANIFEST_VERIFY_SCRIPT=verify.sh \
CT_MANIFEST_VERIFY_TIMEOUT_SECS=30 ./ct-agent manifest create > unsigned.json
CT_MANIFEST_HOLDER_KEY=<from channel init> CT_MANIFEST_IN=unsigned.json \
./ct-agent manifest sign > signed.json
mkdir work
CT_MANIFEST_URL="$PWD/signed.json" \
CT_MANIFEST_ALLOW_LOCAL_PATH=1 \
CT_MANIFEST_TRUST_ALLOWLIST=<your holder pubkey> \
CT_MANIFEST_PROJECT_NAME=docs-harness-proof \
CT_MANIFEST_WORK_DIR="$PWD/work" \
./ct-agent manifest activate
CT_MANIFEST_ALLOW_LOCAL_PATH=1 set on the agent, as above -- without it, a bare path
is refused the same way http:// is. See
[Install an agent manifest](/how-to/install-an-agent-manifest/) for the
full explanation and the related sandbox-activation requirement (#183) for Binary-kind manifests.
work/ is the parent activate unpacks into — since ct-agent v0.7.27 (#165) the manifest’s
files actually land in work/docs-harness-proof/ (<CT_MANIFEST_WORK_DIR>/<CT_MANIFEST_PROJECT_NAME>,
created fresh; a non-empty target is refused, naming what’s already there). That per-project
subdirectory — not CT_MANIFEST_WORK_DIR itself — is what the harness is later containment-scoped
to (CT_HARNESS_BUNDLE_DIR below), and it’s where you’ll find the .ct-agent-activation.json
marker the harness checks before it will run against it.
2. Sign a task
manifest-core's own examples/
dev_sign_task.rs — its own doc comment literally calls itself a "local dev tool." A real
ct-agent task sign-style subcommand (mirroring manifest create/manifest
sign's shape) doesn't exist as of this writing. Until it does, a publisher wanting to sign a
task for real needs this example binary (or their own small program calling
SignedTask::sign_new directly) — not a gap in this page, a gap in the tooling.
git clone https://github.com/scimbe/CADS-agent-marketplace.git
cd CADS-agent-marketplace
CT_TASK_HOLDER_KEY=<your holder key, same one that signed the manifest, or any trusted publisher key> \
CT_TASK_MANIFEST_ID=<manifest_id from signed.json> \
CT_TASK_PROMPT="Say hello" \
CT_TASK_MODEL="gpt-4o-mini" \
CT_TASK_MAX_TURNS=6 \
cargo run --example dev_sign_task -p manifest-core > task.json
CT_TASK_NOW/CT_TASK_EXPIRES_IN_SECS/CT_TASK_MAX_OUTPUT_TOKENS/CT_TASK_ID all have defaults —
see the example’s own source for the exact fallback values. Real output:
{
"publisher_pubkey": "f8c7fafde5c2521fa30ecfd92af6478fd0d275ad4091c1f8317f927819b61c7b",
"task_id": "0808080808080808080808080808080808080808080808080808080808080808",
"manifest_id": "7349ea7913a1f2eaefc3b828a95f72fe6617e83c636bb00b510fec402f4e6a12",
"prompt": "Say hello",
"model": "gpt-4o-mini",
"max_turns": 6,
"max_output_tokens": 2048,
"issued_at": 1788352163,
"expires_at": 1788355763,
"signature": "f4c4ee7f…"
}
3. Run it
CT_HARNESS_TASK_URL_OR_PATH="$PWD/task.json" \
CT_HARNESS_MANIFEST_URL_OR_PATH="$PWD/signed.json" \
CT_HARNESS_TRUST_ALLOWLIST=<your holder pubkey> \
CT_HARNESS_ALLOWED_MODELS=gpt-4o-mini \
CT_HARNESS_BUNDLE_DIR="$PWD/work" \
CT_HARNESS_LITELLM_URL=<your own LiteLLM proxy's base URL> \
CT_HARNESS_LITELLM_KEY_FILE=<path to a file holding a budget-capped LiteLLM virtual key> \
./ct-agent harness run
CT_HARNESS_LITELLM_KEY_FILE is a file, never an inline env var — the same file-based-secret
discipline ct-agent’s own CT_AGENT_CAPABILITY_OUT uses, so the key never lands in a ps/
process-env dump. CT_HARNESS_MANIFEST_URL_OR_PATH is re-fetched and re-verified here (signature,
expiry, trust allowlist) even though the manifest was already activated in step 1 — the harness
re-confirms CT_HARNESS_BUNDLE_DIR really looks installed from the manifest the task claims to be
scoped to before trusting anything in it, a defense-in-depth check independent of what
manifest activate already did.
read_file/write_file/rebuild and finishing with
"status": "ok" — needs a real LiteLLM proxy to click-test against, which this pass
didn't have. What's verified live below is everything up to and including the first real HTTP call
to it (every rejection path, plus the exact request shape once nothing else is left to reject) —
not the full successful round-trip. If you run this against a real LiteLLM deployment and hit
something this page gets wrong, the source above is the fastest way to check what actually happens
next.
Confirmed live, in order — nothing in the bundle is ever touched once any check fails:
Manifest trust allowlist doesn’t match (checked before the task’s own publisher check even runs):
"the manifest fetched from CT_HARNESS_MANIFEST_URL_OR_PATH is signed by a publisher not on CT_HARNESS_TRUST_ALLOWLIST -- refusing to trust its bundle.compose_file"
Model not on the harness’s own allowlist (CT_HARNESS_ALLOWED_MODELS doesn’t include the
task’s model, even with the trust allowlist satisfied):
{
"status": "rejected",
"reason": "model 'gpt-4o-mini' is not on this host's harness model allowlist",
"task_id": "0808080808080808080808080808080808080808080808080808080808080808"
}
max_turns over the local ceiling (a task signed with max_turns=500):
{
"status": "rejected",
"reason": "task.max_turns (500) exceeds this harness's local ceiling of 200 turns -- refusing regardless of the LiteLLM budget cap, since the rebuild tool never touches that budget",
"task_id": "0808080808080808080808080808080808080808080808080808080808080808"
}
Everything above passes, then the actual model call (against a deliberately unreachable
LiteLLM URL, http://127.0.0.1:9) — this is the exact HTTP shape a real run makes, an OpenAI-
compatible /chat/completions POST:
{
"status": "failed",
"task_id": "0808080808080808080808080808080808080808080808080808080808080808",
"manifest_id": "7349ea7913a1f2eaefc3b828a95f72fe6617e83c636bb00b510fec402f4e6a12",
"turns_used": 0,
"reason": "model call failed: POST http://127.0.0.1:9/chat/completions: error sending request for url (http://127.0.0.1:9/chat/completions)"
}
harness run exits 0 exactly when status is "ok", same convention as manifest activate —
ct-agent harness run && … scripts correctly. A "failed" report (task passed every check, then
something went wrong mid-run) still carries turns_used and, on later turns, whichever files the
loop had already changed before the failure — check <bundle_dir>/.harness-transcript.jsonl for the
full turn-by-turn record either way (report.rs’s TranscriptEntry log, append-only, written
regardless of the run’s final outcome).
Reference
CT_HARNESS_TASK_URL_OR_PATH—https://URL or local path to the signed task JSON.CT_HARNESS_MANIFEST_URL_OR_PATH— the same manifest reference used atmanifest activatetime.CT_HARNESS_TRUST_ALLOWLIST(comma-separated 64-hex publisher pubkeys) orCT_HARNESS_TRUST_ALLOWLIST_FILE(one per line) — exactly one of the two, required; an empty allowlist is refused outright rather than silently allowing everything.CT_HARNESS_BUNDLE_DIR— the manifest’s own already-activated project directory, i.e.<CT_MANIFEST_WORK_DIR>/<CT_MANIFEST_PROJECT_NAME>from step 1, notCT_MANIFEST_WORK_DIRitself (since ct-agent v0.7.27). The harness refuses to run if this directory’s.ct-agent-activation.jsonmarker names a different manifest, is corrupt, or is missing entirely — the last case only bypassable withCT_HARNESS_ALLOW_UNMARKED_BUNDLE=1, for a directory activated by an older ct-agent.CT_HARNESS_LITELLM_URL/CT_HARNESS_LITELLM_KEY_FILE— your own LiteLLM proxy and a budget-capped virtual key file for it.CT_HARNESS_ALLOWED_MODELS— comma-separated model names the harness may call; no default, ever.- The three tools a task can call:
read_file,write_file(both containment-checked, 4 MiB cap each),rebuild(docker compose -f <compose_file> build, 300s timeout, process-group-killed on timeout — neverup/down).
Full reference for manifest’s own subcommands and env vars:
Install an agent manifest.