exoDocs

Platform services

exo-verify

Code verification as a platform service — a verify contract, sandboxed runs, live staging checks, and an AI review pass.

Source: ADR 0005, accepted (owner-approved) on 2026-09-27. It is on main and not yet on staging: read it on main.

Why

Each App carries ad-hoc unit tests, and staging failures slip through. The motivating example: the AI router passed all unit tests and review but fell back 100% of the time on staging, because a fix never merged. Only a live behavioral check against the deployed environment catches that class of failure.

Two constraints shape the design:

  • The owner is moving from GitHub to Entire, so GitHub Actions and GitHub-app bots are out.
  • Most commits are made by agents (~300/month across the portfolio), so per-seat pricing is the wrong unit.

Four layers

Standard verify contract

Every repository declares verify commands (typecheck, lint, test, smoke) in project.yaml. Exo runs what is declared, and a missing command is a visible gap, not a silent pass.

Sandboxed runs on Railway workers

The preferred trigger is an Entire webhook, if Entire exposes one. The fallbacks are an exo verify CLI that agents run before pushing, then Exo polling. Inside the sandbox, an agent may also write and run targeted tests for the change.

Post-deploy live checks against staging

Health, auth (Exo Key JWKS and OAuth callbacks), and two or three declared user journeys per App, driven by Playwright or a browser agent.

AI review pass

An open-source engine (in the style of PR-Agent or Kodus) or Exo's own prompts. Every model call goes through exo-router, so cost is logged per App and the router keeps easy diffs on cheap models.

Results are stored in Convex as release gates, attached to the Entire trail, and gate the staging → production promotion.

Unit economics

OptionPricingEst. monthly at ~300 agent commits
exo-verify (own)~$0.007 compute per 5-min 2 vCPU / 2 GB Railway run + $0.01–0.10 model tokens per run~$6–30
Greptile Pro$30/seat/mo, 50 credits/seat, $1 per extra credit; non-GitHub/GitLab forges Enterprise-only~$300–900
CodeRabbit$24–30/user; CLI free at 3 reviews/hr on local gitPer seat; CLI usable as an interim
Graphite$40/userPer seat, GitHub-bound

Cost per verified commit is roughly $0.02–0.10, recorded per App so that it shows up in each App's cost to serve.

Consequences

  • Every App gains a verify block in project.yaml, and onboarding includes declaring its two or three golden journeys.
  • Third-party review engines start out better tuned. Exo tracks catch rate against false positives per layer and will revisit buying the review layer if the gap persists.
  • Live checks need a staging-only sign-in path for bots (seeded Exo Key test identities), because browser agents cannot pass Google's bot detection.
  • Open question: does Entire expose webhooks or a checks API?

Status

Built in part on staging (2026-10-01). What runs today:

  • Contract + CLI (packages/exo-verify, zero dependencies): a verify: block in project.yaml (typecheck, lint, a list of test commands, build, smoke checks per env, mode: enforce|report). exo-verify run [--gate] [--json <path>] and exo-verify smoke --env staging|production. A missing command counts as a gap, never a pass. Only enforce mode exits non-zero.
  • Results: POST /v1/verify/runs (service token) stores rows in the verifyRuns Convex table. GET /v1/verify/status returns the latest status for each app.
  • Build gates: Exo's and Observatory's Railway builds run exo-verify run --gate in enforce mode. Construct (12 existing lint errors) and Rampart (tests assume a root base path) run it in report mode. Apps use a vendored copy (scripts/exo-verify.mjs) from exo-verify vendor.
  • Live checks: a Convex cron runs every 15 minutes. It makes HTTP GETs to each app's declared smoke routes on staging and production. For drift, it records whether production's /api/health commit differs from staging's and since when. It also checks production domains over DNS-over-HTTPS.
  • Pre-push hook: exo-verify install turns on a versioned .githooks/ for one worktree. The hook runs the gate and then hands off to the existing .git/hooks (Entire). To skip once, set EXO_VERIFY_SKIP=1. The skip is always written to a local log. It is also posted as a skip run (best effort) when EXO_VERIFY_URL and EXO_VERIFY_TOKEN are set.
  • AI review: exo-verify review --base <ref> runs PR-Agent (MIT, pinned pr-agent==0.46.0, plain-diff mode) through exo-router's OpenAI-compatible POST /v1/chat/completions, which logs per-app tokens and cost in exoAiBrokerJobs. Off by default: the Convex AI gateway is not enabled for the team, so the endpoint returns 503 exo_router_unavailable, records the failed job, and the CLI prints review skipped: <reason>.

Measured cost per run:

CheckTime measuredCost
Exo Railway build gate (typecheck 15.4s, lint 12.7s, Vitest 14.0s)42 s added to each deployNo new service. If billed at Railway's 2 vCPU / 2 GB list rate: about $0.001 per deploy
App gates on Railway (Observatory: typecheck 8.7s, lint 9.7s, tests 1.1s) and locally (Construct, Rampart)20 s on Railway for Observatory; 7–9 s locallyAbout $0.0005 per deploy at the same rates
Live sweep (4 apps, both envs)One Convex action plus about 40 GETs (smoke, health, and DNS-over-HTTPS; Exo alone declares 14)About 96 sweeps a day, well inside the Convex plan's function-call quota
AI reviewNot measured: the gateway is not enabledRecorded per app once it is

Not built yet: browser journeys (Playwright, Apache-2.0), Entire triggers, and sandboxed Railway workers. Vitest (MIT) is the test runner for Exo itself.

Source: content/docs/platform/exo-verify.mdx

On this page