Platform services
exo-verify
Code verification as a platform service — a verify contract, sandboxed runs, live staging checks, and an AI review pass.
Source: ADR 0005, accepted (owner-approved) on 2026-09-27. It is on main and not yet on
staging: read it on main.
Why
Each App carries ad-hoc unit tests, and staging failures slip through. The motivating example: the AI router passed all unit tests and review but fell back 100% of the time on staging, because a fix never merged. Only a live behavioral check against the deployed environment catches that class of failure.
Two constraints shape the design:
- The owner is moving from GitHub to Entire, so GitHub Actions and GitHub-app bots are out.
- Most commits are made by agents (~300/month across the portfolio), so per-seat pricing is the wrong unit.
Four layers
Standard verify contract
Every repository declares verify commands (typecheck, lint, test, smoke) in
project.yaml. Exo runs what is declared, and a missing command is a visible gap, not a silent
pass.
Sandboxed runs on Railway workers
The preferred trigger is an Entire webhook, if Entire exposes one. The fallbacks are an
exo verify CLI that agents run before pushing, then Exo polling. Inside the sandbox, an agent may
also write and run targeted tests for the change.
Post-deploy live checks against staging
Health, auth (Exo Key JWKS and OAuth callbacks), and two or three declared user journeys per App, driven by Playwright or a browser agent.
AI review pass
An open-source engine (in the style of PR-Agent or Kodus) or Exo's own prompts. Every model call goes through exo-router, so cost is logged per App and the router keeps easy diffs on cheap models.
Results are stored in Convex as release gates, attached to the Entire trail, and gate the staging → production promotion.
Unit economics
| Option | Pricing | Est. monthly at ~300 agent commits |
|---|---|---|
| exo-verify (own) | ~$0.007 compute per 5-min 2 vCPU / 2 GB Railway run + $0.01–0.10 model tokens per run | ~$6–30 |
| Greptile Pro | $30/seat/mo, 50 credits/seat, $1 per extra credit; non-GitHub/GitLab forges Enterprise-only | ~$300–900 |
| CodeRabbit | $24–30/user; CLI free at 3 reviews/hr on local git | Per seat; CLI usable as an interim |
| Graphite | $40/user | Per seat, GitHub-bound |
Cost per verified commit is roughly $0.02–0.10, recorded per App so that it shows up in each App's cost to serve.
Consequences
- Every App gains a
verifyblock inproject.yaml, and onboarding includes declaring its two or three golden journeys. - Third-party review engines start out better tuned. Exo tracks catch rate against false positives per layer and will revisit buying the review layer if the gap persists.
- Live checks need a staging-only sign-in path for bots (seeded Exo Key test identities), because browser agents cannot pass Google's bot detection.
- Open question: does Entire expose webhooks or a checks API?
Status
Built in part on staging (2026-10-01). What runs today:
- Contract + CLI (
packages/exo-verify, zero dependencies): averify:block inproject.yaml(typecheck, lint, a list of test commands, build, smoke checks per env,mode: enforce|report).exo-verify run [--gate] [--json <path>]andexo-verify smoke --env staging|production. A missing command counts as a gap, never a pass. Only enforce mode exits non-zero. - Results:
POST /v1/verify/runs(service token) stores rows in theverifyRunsConvex table.GET /v1/verify/statusreturns the latest status for each app. - Build gates: Exo's and Observatory's Railway builds run
exo-verify run --gatein enforce mode. Construct (12 existing lint errors) and Rampart (tests assume a root base path) run it in report mode. Apps use a vendored copy (scripts/exo-verify.mjs) fromexo-verify vendor. - Live checks: a Convex cron runs every 15 minutes. It makes HTTP GETs to each app's declared
smoke routes on staging and production. For drift, it records whether production's
/api/healthcommit differs from staging's and since when. It also checks production domains over DNS-over-HTTPS. - Pre-push hook:
exo-verify installturns on a versioned.githooks/for one worktree. The hook runs the gate and then hands off to the existing.git/hooks(Entire). To skip once, setEXO_VERIFY_SKIP=1. The skip is always written to a local log. It is also posted as askiprun (best effort) whenEXO_VERIFY_URLandEXO_VERIFY_TOKENare set. - AI review:
exo-verify review --base <ref>runs PR-Agent (MIT, pinnedpr-agent==0.46.0, plain-diff mode) through exo-router's OpenAI-compatiblePOST /v1/chat/completions, which logs per-app tokens and cost inexoAiBrokerJobs. Off by default: the Convex AI gateway is not enabled for the team, so the endpoint returns 503exo_router_unavailable, records the failed job, and the CLI printsreview skipped: <reason>.
Measured cost per run:
| Check | Time measured | Cost |
|---|---|---|
| Exo Railway build gate (typecheck 15.4s, lint 12.7s, Vitest 14.0s) | 42 s added to each deploy | No new service. If billed at Railway's 2 vCPU / 2 GB list rate: about $0.001 per deploy |
| App gates on Railway (Observatory: typecheck 8.7s, lint 9.7s, tests 1.1s) and locally (Construct, Rampart) | 20 s on Railway for Observatory; 7–9 s locally | About $0.0005 per deploy at the same rates |
| Live sweep (4 apps, both envs) | One Convex action plus about 40 GETs (smoke, health, and DNS-over-HTTPS; Exo alone declares 14) | About 96 sweeps a day, well inside the Convex plan's function-call quota |
| AI review | Not measured: the gateway is not enabled | Recorded per app once it is |
Not built yet: browser journeys (Playwright, Apache-2.0), Entire triggers, and sandboxed Railway workers. Vitest (MIT) is the test runner for Exo itself.