The stack
12 · Artificial intelligence
Apps send prompts to exo-router, Exo's smart model routing. It scores how hard each task is and picks the least costly capable model through the Convex AI Gateway.
- Dimension
- 12 · Artificial intelligence
- Platform
- Convex AI Gateway + Artificial Analysis + Arch-Router
- Convex AI Gateway
- Provider-neutral model access; Exo never hands provider keys to apps. Core standard · authority: Exo
- Artificial Analysis
- Independent intelligence, speed, and price benchmarks used to score models. Optional adapter · authority: Exo
- Arch-Router
- Self-hosted open model that rates how hard each prompt is, so exo-router can pick a model. Core standard · authority: EXO
- Status
- exo-router endpoint live on Convex; difficulty assessor and model pool built; Artificial Analysis benchmarks and per-App cost attribution TBD.
What it is
exo-router is Exo's own smart model routing. Apps send prompts to exo-router and never hold provider keys. It scores how hard each task is, weighs model capability against cost, and sends the call to the chosen model through the Convex AI Gateway. Every job is metered. In short: an assessor decides, and the gateway executes (AI routing). The design is on the exo-router page.
Platform
- Convex AI Gateway is the OpenAI-compatible egress with service-token auth. Convex holds the provider credentials.
- Artificial Analysis is named in the catalog as the source of independent intelligence,
speed, and price benchmarks used to score models. It is not wired into the code yet: the
aiModelPoolintelligence scores (0–100) are Exo's own seed estimates inconvex/ai/router.ts, "meant to be calibrated from aiRoutingLog over time". The benchmark import is TBD. - Arch-Router (Katanemo Arch-Router-1.5B, open weights) is the difficulty assessor,
self-hosted on Railway via llama.cpp from
services/arch-router/. See What is Arch-Router?. - TypeSafe (Jev, System One) is a dormant alternate assessor: TypeSafe closed new signups (verified 2026-09-23).
Boundary and replaceability
- Hot-swappable model pool.
aiModelPoolis a table, so pool changes are data edits, not deploys. - Pluggable assessors. The first configured assessor wins (
ARCH_ROUTER_URL, thenTYPESAFE_API_KEY). With neither configured, exo-router uses pure utility and flags every decisionfallback: true. - No silent fallback. Gateway failures fail the job and are recorded.
- Aggregate findings only. Exo may receive aggregate quality, latency, token, cost, and reliability findings from Apps. It never receives prompts, responses, user identifiers, or provider keys (Convex framework).
Cost notes
- Assessor: no per-call cost; CPU-only, ~1 GB RAM on Railway.
- Routing policy: on
staging,convex/ai/router.tscomputesutility = intelligence^w / blendedCost, wherewis the assessed difficulty (0–1) andblendedCostis the 3:1 input:output blend of the pool row's per-million-token prices. Open PR #115 (againstmain) changes it tointelligence^w / blendedCost^(1-w). Easy tasks let cost dominate, and hard tasks weight intelligence. - Metering: each pool row carries input and output cost per million tokens (USD).
exoAiBrokerJobsrecords tokens and duration per job. Onstagingthe endpoint does not yet computecostMicrosUsdand rows carry no App identifier, so per-App AI spend is TBD.
Status
- Live: the exo-router HTTP endpoint on the Exo Convex deployment (
POST /v1/ai/generate), used by Construct. WithEXO_AI_MODELunset, a deterministic local composer serves requests so the loop works without keys. - Built:
convex/ai/router.ts,assessors.ts, andchat.tswith the routing log. - In progress: sending
/v1/ai/generatejobs through the router (PR #115), and per-App cost attribution (roadmap: ai-gateway-jev-routing). - TBD: Artificial Analysis benchmark import into
aiModelPool.
Source: content/docs/stack/ai.mdx