exoDocs

The stack

12 · Artificial intelligence

Apps send prompts to exo-router, Exo's smart model routing. It scores how hard each task is and picks the least costly capable model through the Convex AI Gateway.

Dimension
12 · Artificial intelligence
Platform
Convex AI Gateway + Artificial Analysis + Arch-Router
Convex AI Gateway
Provider-neutral model access; Exo never hands provider keys to apps. Core standard · authority: Exo
Artificial Analysis
Independent intelligence, speed, and price benchmarks used to score models. Optional adapter · authority: Exo
Arch-Router
Self-hosted open model that rates how hard each prompt is, so exo-router can pick a model. Core standard · authority: EXO
Status
exo-router endpoint live on Convex; difficulty assessor and model pool built; Artificial Analysis benchmarks and per-App cost attribution TBD.

What it is

exo-router is Exo's own smart model routing. Apps send prompts to exo-router and never hold provider keys. It scores how hard each task is, weighs model capability against cost, and sends the call to the chosen model through the Convex AI Gateway. Every job is metered. In short: an assessor decides, and the gateway executes (AI routing). The design is on the exo-router page.

Platform

  • Convex AI Gateway is the OpenAI-compatible egress with service-token auth. Convex holds the provider credentials.
  • Artificial Analysis is named in the catalog as the source of independent intelligence, speed, and price benchmarks used to score models. It is not wired into the code yet: the aiModelPool intelligence scores (0–100) are Exo's own seed estimates in convex/ai/router.ts, "meant to be calibrated from aiRoutingLog over time". The benchmark import is TBD.
  • Arch-Router (Katanemo Arch-Router-1.5B, open weights) is the difficulty assessor, self-hosted on Railway via llama.cpp from services/arch-router/. See What is Arch-Router?.
  • TypeSafe (Jev, System One) is a dormant alternate assessor: TypeSafe closed new signups (verified 2026-09-23).

Boundary and replaceability

  • Hot-swappable model pool. aiModelPool is a table, so pool changes are data edits, not deploys.
  • Pluggable assessors. The first configured assessor wins (ARCH_ROUTER_URL, then TYPESAFE_API_KEY). With neither configured, exo-router uses pure utility and flags every decision fallback: true.
  • No silent fallback. Gateway failures fail the job and are recorded.
  • Aggregate findings only. Exo may receive aggregate quality, latency, token, cost, and reliability findings from Apps. It never receives prompts, responses, user identifiers, or provider keys (Convex framework).

Cost notes

  • Assessor: no per-call cost; CPU-only, ~1 GB RAM on Railway.
  • Routing policy: on staging, convex/ai/router.ts computes utility = intelligence^w / blendedCost, where w is the assessed difficulty (0–1) and blendedCost is the 3:1 input:output blend of the pool row's per-million-token prices. Open PR #115 (against main) changes it to intelligence^w / blendedCost^(1-w). Easy tasks let cost dominate, and hard tasks weight intelligence.
  • Metering: each pool row carries input and output cost per million tokens (USD). exoAiBrokerJobs records tokens and duration per job. On staging the endpoint does not yet compute costMicrosUsd and rows carry no App identifier, so per-App AI spend is TBD.

Status

  • Live: the exo-router HTTP endpoint on the Exo Convex deployment (POST /v1/ai/generate), used by Construct. With EXO_AI_MODEL unset, a deterministic local composer serves requests so the loop works without keys.
  • Built: convex/ai/router.ts, assessors.ts, and chat.ts with the routing log.
  • In progress: sending /v1/ai/generate jobs through the router (PR #115), and per-App cost attribution (roadmap: ai-gateway-jev-routing).
  • TBD: Artificial Analysis benchmark import into aiModelPool.

Source: content/docs/stack/ai.mdx

On this page