Decisions and roadmap
ADR 0005: exo-verify — code verification as a platform service
- Status: accepted (owner-approved 2026-09-27)
- Date: 2026-09-27
- Roadmap:
exo-verify
Context
Exo provisions shared platform services to every app so apps do not spend time on basic plumbing. Verification is not one of them yet: each app carries ad-hoc unit tests, and staging failures slip through. Example: the AI router passed all unit tests and review but fell back 100% of the time on staging, because a fix never merged. Only a live behavioral check against the deployed environment catches that class of failure.
Two constraints shape the design. The owner is moving off GitHub to Entire as the forge, so GitHub Actions and GitHub-app PR bots are out. And most commits are made by agents (~300/month across the portfolio), so per-seat vendor pricing is the wrong unit.
Decision
Exo provides verification as a shared service, exo-verify, in four layers:
- Standard verify contract. Every repo declares
verifycommands (typecheck,lint,test,smoke) inproject.yaml. Exo runs what is declared; a missing command is a visible gap, not a silent pass. - Sandboxed runs on Railway workers (not GitHub Actions). Trigger order of preference:
an Entire webhook if Entire exposes one; otherwise an
exo verifyCLI that agents run before pushing; otherwise Exo polling. In the sandbox an agent may also write targeted tests for the change and run them (TREX-style runtime validation). - Post-deploy live checks against staging: health, auth (Exo ID JWKS, OAuth callbacks), and 2–3 declared user journeys per app driven by Playwright or a browser agent.
- AI review pass using an open-source engine (PR-Agent / Kodus-style) or our own
prompts. All model calls go through the Exo AI broker, so cost is logged per app in
exoAiBrokerJobsand the router keeps easy diffs on cheap models.
Results are stored in Convex as release gates, attached to the Entire trail, and gate staging → production promotion.
Unit economics
| Option | Pricing | Est. monthly at ~300 agent commits |
|---|---|---|
| exo-verify (own) | ~$0.007 compute per 5-min 2 vCPU / 2 GB Railway run + $0.01–0.10 model tokens per run | ~$6–30 |
| Greptile Pro | $30/seat/mo, 50 credits/seat, $1/extra credit; TREX runtime validation (beta) costs more credits; non-GitHub/GitLab forges are Enterprise-only | ~$300–900 |
| CodeRabbit | $24–30/user; CLI free at 3 reviews/hr on local git | per seat; CLI usable as interim |
| Graphite | $40/user | per seat, GitHub-bound |
Cost per verified commit for the own service is roughly $0.02–0.10, recorded per app so it shows up in each app's cost-to-serve.
Alternatives considered
- Greptile — strongest vendor review and has runtime validation, but priced per seat plus credits and requires Enterprise for a non-GitHub/GitLab forge. Kept as a later option for the review layer (Greptile Enterprise) if our review quality falls short.
- CodeRabbit — the free CLI reviews a local diff, so agents can use it pre-push as an interim second opinion while layer 4 is built. Not the system of record.
- Graphite — GitHub-centric stacking/review; does not fit the Entire forge.
- GitHub Actions — rejected; the owner is leaving GitHub.
Consequences
- Every app gains a
verifyblock inproject.yaml; onboarding an app includes declaring its 2–3 golden journeys. - Vendor review engines are more tuned than ours at first. We track catch rate against false positives per layer and revisit buying the review layer if the gap persists.
- Live checks need a staging-only sign-in path for bots (seeded Exo ID test identities), because browser agents cannot pass Google's bot detection.
- Open question: does Entire expose webhooks or a checks API? Verify before building the trigger; the CLI and polling fallbacks keep the design unblocked either way.