Skip to main content
Relay decides which model should handle each request. It routes per turn rather than locking an entire session to one model, using company-defined choices around cost, quality, latency, policy, and fallback behavior. A gateway passes traffic. Relay provides the routing intelligence. It can also operate as the gateway itself.

Choose how Relay fits

Start with the gateway

Create a project and send the first OpenAI-compatible request.

Compare integration modes

Choose an execution or observation path without changing more than necessary.

Routing control

Teams can route automatically or pin a model, choose candidate models, bias cost, quality, or latency, set fallbacks, and inspect why each request was routed. Custom routers can be built for chat, code, or company-specific workloads. Visibility gives Relay real company traffic and use-case context. That lets routers reflect the work the company actually performs rather than a generic benchmark alone.

Providers stay yours

Relay supports customer-controlled provider accounts and keys. A company can change providers, models, or gateway architecture without rewriting the applications that call its OpenAI-compatible endpoint.

Measure before and after

Traffic logs, decision replay, model comparisons, quality evaluations, cost projections, latency reports, and agent-run grouping show the effect of routing on real workloads. Activity and results return to Visibility as part of the shared operating map.

Configure routing

Set a pinned or automatic policy, candidate set, tradeoff, and fallback.

Measure outcomes

Replay decisions, compare live and shadow results, and run quality evaluations.

Inspect traffic

Trace requests, understand decisions, group agent runs, and classify signals.

Review the architecture

See how Relay fits with Visibility, Sidekick, customer gateways, and deployment models.