Where the project is headed, and what it already does. For per-release detail, see CHANGELOG.md.
What's next, roughly in priority order:
- Plugin SDK & vendor observability bridges — external guardrail and transform plugins, plus bridges for LangSmith, Langfuse, Datadog, New Relic, Honeycomb, Grafana, and more, shipped from a companion
ai-gateway-pluginsrepo so the core binary stays slim. - Webhook notifications — configurable alerts for budget limits, error spikes, and circuit-breaker events.
- Semantic & Redis-backed caching — beyond the built-in in-memory cache.
- Broader provider coverage and an official Go client, driven by community demand.
- Deeper deployment guidance — Kubernetes operators and Terraform modules.
Have a plugin idea, or a recipe worth sharing? Plugin proposals and cookbook-style usage recipes are discussed in the open — ideas there help shape what lands next:
Everything below is available today.
- 8 strategies: single, fallback, load balance, least latency, cost-optimized, content-based, A/B test, conditional
- Per-target retry (jittered backoff, status-code filters) honoured under every routing mode; provider failover in the pool modes, which move to a sibling only after a failover-safe failure (v1.5.1)
- Per-request model aliases, operator-declared
targets[].models, and per-targetmodel_map— one visible model key, a different upstream id on each provider (v1.5.1) - One
gateway.routing.attemptobservation per physical provider call and A/B variant attribution on every attempt, opt-in for exporters (v1.5.1) - One ranker for every surface, so a routing mode orders its targets the same way for chat, streams, embeddings, images, rerank, moderation, transcription and speech (v1.5.2)
- Least-latency samples keyed by target and upstream model that expire, measure time to first chunk, and keep exploring; cost-optimized ranking on input plus output price with weighted tie-break (v1.5.2)
targets[].timeoutper attempt, a429cooldown honouringRetry-After, a typed context-length failover class, andstrategy.failover_on_status_codes(v1.5.2)- Sticky hashing on the request
userfor load balancing and A/B tests, ruletarget_keyschains with a hard boundary, and bounded conditional predicates (user,stream,has_tools, one metadata header) (v1.5.2) - Routing configs that loaded and silently misbehaved are refused at load: out-of-range
retry.on_status_codes, empty or paddedmodel/model_prefixand emptyusercondition values, duplicate A/B variant labels under Unicode case folding (v1.5.5) strategy.modespelledload-balancelike the other seven modes, with the legacyloadbalancestill accepted and normalised on load; circuit-breaker fields distinguish omitted from written zero, so a written threshold or timeout must be positive and an omitted one keeps the default (v1.5.6)X-Gateway-Provider/-Target/-Model/-Attemptsattribution headers on every routed surface,ferro.routing.attempton the request span, andaigateway.WithCatalogfor a host-owned price catalog (v1.5.2)- Per-target concurrency limits with a bounded queue and 429 shedding
- Per-provider circuit breaker shared across every surface
- 30 providers behind one OpenAI-compatible API — OpenAI-compatible and native-wire alike (Anthropic, Gemini, Bedrock, Vertex AI, Cohere, …)
- Capability matrix and
GET /v1/capabilities— machine-readable, per-provider parameter support - Live model discovery plus a shared model catalog powering
/v1/models
- Chat (with streaming), embeddings, image generation, and legacy completions
- Audio — speech-to-text and text-to-speech — plus rerank and moderations
- Batch and files, and a priced, governed Responses API
- Transparent pass-through proxy for any other
/v1/*route
- Built-in: word-filter, max-token, response-cache, request-logger, rate-limit, budget
- Staged middleware (before / after / on-error) with an explicit deny-versus-fail policy
- stdio and Streamable-HTTP transports with an agentic tool-call loop
- Tool allowlists, bounded call depth, cross-server deduplication, and subprocess environment isolation
- OpenTelemetry tracing (OTLP gRPC/HTTP, W3C propagation, GenAI +
ferro.*attributes, privacy levels) — zero-allocation when off - Prometheus metrics;
/health,/livez, and/readyz; a single trace ID unified across logs, spans, and theX-Request-IDheader - Structured request logging with SQLite or PostgreSQL persistence
- A React/TypeScript console compiled into the binary and served at
/from the same port — one artifact, no second origin - Overview, analytics, providers, routing, plugins, playground, tracing, request logs, audit trail, configuration, and API keys
- Scoped API keys, dashboard sessions, an audit trail, and config history with rollback
- Security headers, trusted-proxy client-IP resolution, secret redaction, and
${VAR}references resolved at construction (never stored), plus production-mode startup guards
- Single static binary, ~32 MB base memory, 13,925 RPS at 1,000 concurrent users
- memory / SQLite / PostgreSQL backends
- Multi-arch container images, a Helm chart, GoReleaser packaging, and Railway & Render deploy templates
- Official Python and TypeScript SDKs
- Importable from Go (v1.5.0):
run.Main()composes the wholeferrogwprogram into your own binary with extra plugins compiled in,run.Run(ctx, …)runs it under a context you own, andhttpgatewaymounts the gateway's HTTP surfaces behind your own middleware