<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>gateway on tomrochette.com</title>
    <link>https://tomrochette.com/tags/gateway/</link>
    <description>Recent content in gateway on tomrochette.com</description>
    <generator>Hugo -- gohugo.io</generator>
    <language>en</language>
    <managingEditor>tom@tomrochette.com (Tom Rochette)</managingEditor>
    <webMaster>tom@tomrochette.com (Tom Rochette)</webMaster>
    <copyright>© 2026 Tom Rochette</copyright>
    <lastBuildDate>Wed, 07 Oct 2026 06:38:49 -0400</lastBuildDate><atom:link href="https://tomrochette.com/tags/gateway/index.xml" rel="self" type="application/rss+xml" />
    
    <item>
      <title>Bifrost</title>
      <link>https://tomrochette.com/agents/model-access/bifrost/</link>
      <pubDate>Wed, 07 Oct 2026 00:00:00 +0000</pubDate>
      <author>tom@tomrochette.com (Tom Rochette)</author>
      <guid>https://tomrochette.com/agents/model-access/bifrost/</guid>
      <category>research-note</category><category>agent-curated</category><category>fully-ai-generated</category><category>llm=glm-5.3-flash</category><category>model-access</category><category>gateway</category><category>self-hosted</category>
      <description>&lt;p&gt;Bifrost is Maxim&amp;rsquo;s high-performance open-source AI gateway: one Go service unifying 20-plus providers behind an OpenAI-compatible API with automatic failover, load balancing, semantic caching, budgets, and an MCP gateway, self-hosted or clustered.&lt;/p&gt;&#xA;&lt;p&gt;&lt;strong&gt;Bifrost is the self-hosted-gateway family&amp;rsquo;s performance challenger, the LiteLLM sibling built in Go instead of Python, and its governing claim is overhead: the vendor&amp;rsquo;s sustained benchmark puts added latency at 11 microseconds per request at 5,000 requests per second.&lt;/strong&gt;&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;What it is&#xA;    &lt;div id=&#34;what-it-is&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#what-it-is&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;An Apache-2.0 Go gateway by Maxim (the LLM-evaluation company), deployable with npx, Docker, or Kubernetes Helm charts, with a built-in web UI for configuration and monitoring (8,598 stars, pushed 2026-10-07, as of 2026-10-07).&#xA;The core routes across 20-plus providers (OpenAI, Anthropic, Bedrock, Vertex, Azure, and more) with weighted key distribution, automatic fallbacks, and semantic caching; governance layers virtual keys, budgets, and rate limits per consumer, team, or customer.&#xA;Around the core sit an MCP gateway that is both client and server (with OAuth, tool filtering, and an agent mode), a plugin system in Go or WASM, Prometheus and OpenTelemetry instrumentation, and SDK drop-in replacements for the OpenAI, Anthropic, Bedrock, and Google GenAI SDKs.&#xA;The enterprise tier adds clustering with gossip-based sync, RBAC with Okta and Entra identity providers, guardrails through Bedrock and Model Armor, in-VPC deployment, and audit logs aimed at SOC 2, GDPR, and HIPAA.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Status&#xA;    &lt;div id=&#34;status&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#status&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;Active and fast-releasing: created 2025-03-19, 8,598 stars, component-tagged releases shipping through 2026-10-06 (transports v2.2.6, telemetry v1.8.5, semanticcache v1.6.9), as of 2026-10-07.&lt;/p&gt;&#xA;&lt;picture&gt;&#xA;  &lt;source media=&#34;(prefers-color-scheme: dark)&#34; srcset=&#34;https://api.star-history.com/chart?repos=maximhq/bifrost&amp;type=date&amp;theme=dark&amp;legend=top-left&#34; /&gt;&#xA;  &lt;source media=&#34;(prefers-color-scheme: light)&#34; srcset=&#34;https://api.star-history.com/chart?repos=maximhq/bifrost&amp;type=date&amp;legend=top-left&#34; /&gt;&#xA;  &lt;img alt=&#34;Star History Chart&#34; src=&#34;https://api.star-history.com/chart?repos=maximhq/bifrost&amp;type=date&amp;legend=top-left&#34; /&gt;&#xA;&lt;/picture&gt;&#xA;&lt;p&gt;The repo description positions it as &amp;ldquo;50x faster than LiteLLM&amp;rdquo;, a vendor claim whose published support is the docs&amp;rsquo; own 11-microsecond overhead benchmark; its Show HN threads stayed small (3 points in December 2025), so the adoption case is stars, Docker pulls, and the Trendshift badge rather than independent load-testing.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Strengths&#xA;    &lt;div id=&#34;strengths&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#strengths&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;The Go core buys a latency and concurrency profile the Python gateways in this family cannot match on the same hardware, at least per the vendor&amp;rsquo;s own numbers.&lt;/li&gt;&#xA;&lt;li&gt;Governance is first-class, with hierarchical budgets, per-consumer virtual keys, and MCP tool allow-lists, not bolted on.&lt;/li&gt;&#xA;&lt;li&gt;Drop-in SDK replacement means adoption is a base-URL change, and the LiteLLM compatibility layer offers a migration path from the incumbent.&lt;/li&gt;&#xA;&lt;li&gt;Plugin extensibility in Go or WASM covers the custom middleware cases that send teams toward building their own gateway.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Cautions&#xA;    &lt;div id=&#34;cautions&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#cautions&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;The performance and 50x claims are vendor-published; no independent replication exists, as of 2026-10-07.&lt;/li&gt;&#xA;&lt;li&gt;Smaller ecosystem than LiteLLM: fewer providers, fewer integrations, and a shorter community track record.&lt;/li&gt;&#xA;&lt;li&gt;Enterprise governance (clustering, RBAC, SSO) sits behind contact-sales, so the open core alone may not carry a large production deployment.&lt;/li&gt;&#xA;&lt;li&gt;A young project from a company whose main product is an evaluation platform, so the gateway&amp;rsquo;s long-term priority is an inference about Maxim&amp;rsquo;s roadmap, not a fact.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Pricing&#xA;    &lt;div id=&#34;pricing&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#pricing&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;Free and open source under Apache-2.0; the gateway itself has no paid tier, so pricing does not apply.&#xA;Enterprise governance features are priced through contact-sales conversations with unpublished numbers.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Compared to&#xA;    &lt;div id=&#34;compared-to&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#compared-to&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/model-access/litellm/&#34; &gt;LiteLLM&lt;/a&gt;: the family incumbent at roughly 60k stars, Python, broadest provider and integration coverage; choose Bifrost for throughput-critical self-hosting, LiteLLM for ecosystem breadth.&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/model-access/experiential/&#34; &gt;Experiential&lt;/a&gt;: the YC-backed zero-markup gateway that mines agent traces; Bifrost keeps your traces local instead of trading them for routers.&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/model-access/plano/&#34; &gt;Plano&lt;/a&gt;: the Envoy-based gateway and agent data plane; Bifrost is the application-level gateway, Plano the infrastructure-level one.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Bottom line&#xA;    &lt;div id=&#34;bottom-line&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#bottom-line&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;&lt;strong&gt;Recommended for teams that need a self-hosted gateway fast enough to sit in the hot path and want budgets, failover, and MCP governance in one deployable.&lt;/strong&gt;&#xA;Not for the broadest provider matrix or the largest community, where LiteLLM still leads, and not for buyers who need independently benchmarked performance claims before committing.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Changes&#xA;    &lt;div id=&#34;changes&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#changes&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;2026-10-07 - Created.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;See also&#xA;    &lt;div id=&#34;see-also&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#see-also&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/model-access/litellm/&#34; &gt;LiteLLM&lt;/a&gt; - the Python incumbent Bifrost positions itself against&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/model-access/experiential/&#34; &gt;Experiential&lt;/a&gt; - the zero-markup hosted-gateway sibling in this family&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/model-access/plano/&#34; &gt;Plano&lt;/a&gt; - the Envoy-based self-hosted gateway with the router lineage&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/model-access/model-access-feature-matrix/&#34; &gt;Model Access Feature Matrix&lt;/a&gt; - the category comparison this note joins&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;References&#xA;    &lt;div id=&#34;references&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#references&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://github.com/maximhq/bifrost&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=github.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://github.com/maximhq/bifrost&lt;/a&gt; - repository, Apache-2.0 license, the 23-plus-provider README, and the 50x-LiteLLM positioning (fetched 200, 2026-10-07)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://api.github.com/repos/maximhq/bifrost&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=api.github.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://api.github.com/repos/maximhq/bifrost&lt;/a&gt; - stars, created date, push date, and license for the as-of status (fetched 200, 2026-10-07)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://docs.getbifrost.ai&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=docs.getbifrost.ai&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://docs.getbifrost.ai&lt;/a&gt; - the feature surface: failover, virtual keys, budgets, semantic caching, the MCP gateway, plugins, the 11-microsecond benchmark claim, and the enterprise tier (fetched 200, 2026-10-07)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://api.github.com/repos/maximhq/bifrost/releases&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=api.github.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://api.github.com/repos/maximhq/bifrost/releases&lt;/a&gt; - the component-tagged release line through 2026-10-06 (fetched 200, 2026-10-07)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://hn.algolia.com/api/v1/items/46203228&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=hn.algolia.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://hn.algolia.com/api/v1/items/46203228&lt;/a&gt; - the December 2025 Show HN thread, 3 points, grounding the small-HN-footprint observation (fetched 200, 2026-10-07)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://raw.githubusercontent.com/maximhq/bifrost/HEAD/README.md&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=raw.githubusercontent.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://raw.githubusercontent.com/maximhq/bifrost/HEAD/README.md&lt;/a&gt; - the quickstart surfaces (npx, Docker) and the provider list behind the gateway cells (fetched 200, 2026-10-07)&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;</description>
      
    </item>
    
  </channel>
</rss>
