Platform services
exo-router
Exo's smart model routing — one metered AI egress where an assessor scores each task, the utility function weighs capability against cost, and the Convex AI Gateway executes. Apps never hold provider keys.
Sources: AI routing, convex/aiBroker.ts, convex/ai/*,
services/arch-router, and the roadmap item
ai-gateway-jev-routing.
exo-router is the name for Exo's own smart model routing: the whole path from an App's prompt
to a metered model call. Older code and docs call the HTTP half the "broker" (convex/aiBroker.ts,
EXO_AI_BROKER_TOKEN, exoAiBrokerJobs); those technical names are unchanged.
Two pieces
The endpoint (POST /v1/ai/generate)
An HTTP action on the Exo Convex deployment. Construct, and later every Exo App, posts generation jobs here instead of holding provider credentials. exo-router owns model routing, budgets, and provider keys, and records every job for unit economics.
- Auth: an Exo Key token (
exo_tk_…) issued per tenant, optionally bound to one app, with scopes (ai,usage:read) and a 90-day default expiry. Only the token's SHA-256 hash is stored. The legacy shared token (EXO_AI_BROKER_TOKEN) still works and bills to thelegacy-brokertenant. Phase 1 of ADR 0008 (proposal); see Metering and usage. - Job types:
page.draftonly, today. - Provider: with
EXO_AI_MODELset, jobs go to the Convex AI Gateway, which holds the provider credentials. Without it, a deterministic local composer (exo-local-composer) serves the request so the loop works end to end. - Failure policy: gateway failures fail the job loudly. A silent fallback would hide a spend or configuration problem.
- Metering: every job writes
exoAiBrokerJobs(job type, status, provider, model, source count, duration, input and output tokens), plus tenant, vendor cost, markup, and price.
The router (convex/ai/router.ts)
surface (help desk / chat)
│
▼
convex/ai/router.ts → assessor (Arch-Router or Jev)
│ utility = intelligence^w / blended cost
▼
convex/ai/chat.ts → Convex AI Gateway (OpenAI-compatible, service-token auth)
│
▼
chosen model from aiModelPool- Fast path: jobs that already know their category and difficulty (every broker job type:
page.draftis writing,classify.choiceextraction,design.reviewvisual-review) skip the assessor; the value score still picks from the live pool. Decisions are memoised for 10 minutes per pool version, and every call still writes its ownaiRoutingLogrow. Routing adds a few milliseconds before the model call starts. - Assessors (
EXO_AI_ASSESSOR, for prompts without a caller category):jev= TypeSafe's Jev through the Convex AI Gateway Decisions API (same gateway token, 600 ms hard timeout),arch= self-hosted Arch-Router (ARCH_ROUTER_URL),none= pure utility at difficulty weight 0.5 with every decision flaggedfallback: true. Unset defaults toarchwhenARCH_ROUTER_URLis set. - Guardrail:
utility = intelligence^w / blendedCost^(1-w)per pool model, wherewis the assessed difficulty (0–1) andblendedCostblends input and output prices 3:1. The assessor's choice gets a confidence-scaled bonus (1 + 0.25 × confidence), utility keeps it honest against cost, and the highest score wins. - Hot-swappable pool:
aiModelPoolrows carry intelligence (0–100) and per-million-token costs. Edits take effect on the next request, with no redeploy. - Audit: every decision lands in
aiRoutingLog: surface, task, chosen model, utility, assessor provider, raw assessor response, the fallback flag, decision latency, and memo hits. - Surfaces today: apps reach the router only through the broker endpoint
POST /v1/ai/generate(decisions log assurface: "broker:<jobType>").ai/chat:chat(routed) andai/chat:chatWithModel(pinned) are internal actions for trusted server-side Exo code; they are not callable from clients. - Endpoint jobs: with
EXO_AI_MODEL=auto(or a per-requestrouting: "auto"),/v1/ai/generatejobs go through the router, and every job records its tier, routing-log id, and estimated cost inexoAiBrokerJobs. UnsetEXO_AI_MODELkeeps the $0 local composer.
What is Arch-Router?
Arch-Router is a small open-weights model from Katanemo (Arch-Router-1.5B, 1.5 billion parameters) built for one job: reading a prompt and choosing which route, and so which model, should handle it. It does not answer the prompt itself.
In exo-router it is the difficulty assessor. Exo runs it on its own Railway service from
services/arch-router/, served by llama.cpp as an OpenAI-compatible endpoint
(POST /v1/chat/completions). It is CPU-only, runs comfortably in ~1 GB RAM, and needs no
external account, so there is no per-call cost. The router sends it the task plus one route per
pool model and it answers {"route": "<model>"}. Arch-Router does not score difficulty directly,
so on staging the router derives w from the model it reaches for (that model's intelligence ÷
100, in convex/ai/assessors.ts), then applies the utility guardrail before picking a model. It replaced the
dormant Jev integration when TypeSafe closed signups.
Model benchmarks
The catalog names Artificial Analysis as the source of independent intelligence, speed, and
price benchmarks for scoring models. The code does not use it yet: the five seeded pool models
(GPT-5 mini, Claude Haiku 4.5, Gemini 2.5 Flash, Claude Sonnet 4.5, GPT-5) carry intelligence
estimates of 72–94 that seedPool describes as "ours and meant to be calibrated from aiRoutingLog
over time". How and when Artificial Analysis scores feed aiModelPool: TBD.
Privacy boundary
Exo may receive aggregate quality, latency, token, cost, reliability, and rollback findings. It must not receive prompts, responses, conversation content, user identifiers, provider keys, or application source (Convex framework).
Metering and usage
Every AI call is attributed to the tenant that owns the calling token. If the token is bound to an app, that app is recorded and the app the request claims is ignored.
| Field | Meaning |
|---|---|
vendorCostMicros | What Exo pays upstream: the gateway-reported cost, or pool price × tokens |
markupBps | The tenant's markup in basis points (default 2000 = 20%) |
priceMicros | What the tenant is charged: vendor cost × (1 + markup), rounded up |
- Ledger: each metered call appends one
usageLedgerrow (tenant, dimensionai, app, model, tokens, cost, price). A running monthly total per tenant makes spend checks a single read. - Usage API:
GET /v1/usage?month=YYYY-MMwith ausage:readtoken returns the tenant's events, tokens, vendor cost and price for the month, broken down by app and model. - Spend cap: a tenant can have a hard monthly AI cap. Once spend reaches it, calls return
HTTP 402 (
spend_cap_exceeded) until the cap is raised or the month rolls over. - Issuing tokens (phase 1): an Exo operator mints and revokes tokens with the deployment key
(
npx convex run exoKeyTokens:mint,exoKeyTokens:revoke,exoKeyTokens:setBilling). The plaintext token is shown once. Self-serve issuing comes with the Exo Key account UI.
Cost notes
- The assessor has no per-call cost.
- Vendor cost and price are recorded per call and per tenant (see above). Fixed costs (Convex plan, Arch-Router hosting) are not yet allocated per tenant: TBD.
Status
- Live: the
/v1/ai/generateendpoint, used by Construct. - Built: router, assessors, routing log, and model pool.
- Live on staging: Exo Key tokens, per-tenant metering, the usage ledger,
/v1/usage, and monthly spend caps (ADR 0008 phase 1). - TBD: Artificial Analysis benchmarks in the model pool.