Architecture
What actually happens between an HTTP request arriving and a tool result going back out.
One process, four parts
The Express app sets trust proxy from TRUSTED_PROXIES, mounts /livez unauthenticated, mounts the auth router, and then registers two routes per configured server plus /hub.
The OAuth authorization server is oidc-provider, mounted on the hub's own paths and backed by the same JSON file everything else uses. What is not the library's: password login and per-client approval are the hub's pages, metadata documents are resolved by the hub's address-pinned client rather than the library's fetcher, and RFC 7592 registration management stays on the hub's stricter handlers. Access tokens are opaque, which is what makes a revocation take effect immediately.
The supervisor owns one long-lived MCP client per configured server and keeps it alive. What sits under that client is the only thing that differs between the four kinds: a child process (stdio), an HTTP/SSE client (remote), a container's attached stdin/stdout (docker) or a Unix/TCP socket (socket). The last two carry the stdio framing over a byte stream, which is what lets an untrusted server live in its own container without an HTTP listener. Everything above the transport — ping, backoff, hot reload, the proxy layer, /hub, /health — treats them identically.
The proxy layer builds a throwaway MCP Server per HTTP request that forwards requests verbatim to the supervisor's client.
Request pipeline
The order of the middleware is deliberate:
- Bearer verification — two shapes: an opaque OAuth access token, looked up in the store so a withdrawn one stops working at once, or an admin-minted API token, which is an EdDSA JWT with a pinned algorithm. There is no per-IP rate limit on the MCP routes; that budget guards the OAuth endpoints, where a caller has no token yet.
- Resource check — the token's audience must match this endpoint;
/healthshares the/hubresource. - Body parsing — capped at
MCP_BODY_LIMIT, and only now, so an unauthenticated request never allocates a megabyte. - Per-client gate — requests per minute, in-flight concurrency and open streams, keyed by OAuth client rather than IP. After the parse, because a
subscriptions/listenPOST belongs to the stream budget and only the parsed body says which kind of POST this is. - Routing — to
/hub, to one server's proxy, or 404.
An unauthenticated request costs one token lookup and nothing more: no bcrypt, no allocation proportional to the body.
Stateless transport
Each MCP request gets a fresh Server and a StreamableHTTPServerTransport with sessionIdGenerator: undefined — no session ID, no server-side session table. When the HTTP response closes, both are closed and forgotten.
One object outlives the request: the handler serving 2026-07-28, because it owns any open subscriptions/listen stream. That is not a session table — it holds no record of who you are, only the sockets currently open — and when a socket closes its subscription goes with it.
The reason is concrete: claude.ai reconnects roughly every five minutes and does not send a session DELETE first. Any per-session state would accumulate one entry per reconnect, forever, and take processes or memory with it. Statelessness makes that impossible by construction.
The cost lands on one era only, and it is worth being precise about which.
On 2025-11-25 server-initiated messages have nowhere to go. listChanged notifications, resource subscriptions and sampling are not delivered — and, since they cannot be, they are not advertised either. Request/response traffic — tools, resources, prompts, completions — is forwarded in full, and the proxy advertises only the capabilities its child actually declared.
On 2026-07-28 the same statelessness is the reason a thing works rather than the reason it does not. That revision removed the server→client request channel outright: a server that needs input answers input_required, the call ends, and the client retries carrying the answer. There is nothing to hold between the two legs, which is precisely what this transport is good at — so elicitation travels end to end, through /hub and /<name>/mcp alike. What is still missing on both eras is the push traffic: listChanged and subscriptions/listen to a child. The full split is the capability matrix.
Why a 2025 client is not bridged to a modern child
A child on 2026-07-28 can raise a question that a 2025-11-25 client, over HTTP, has no way to receive: this transport builds one Server per request, so it never saw an initialize and holds no client capabilities to route an answer back to. Bridging it anyway would need a pending registry, a JSON-RPC id translation, and — because a 2025 elicitation carries no field naming the call that triggered it — a lock serialising every tool call per child across all clients. A gateway that makes itself a bottleneck, with a five-minute hang as its failure mode.
So the hub does not announce the capability to the child, and the child takes its own fallback. Same rule as listChanged above: say only what can be delivered. Over stdio the question does not arise — mcp-hub-stdio is spawned per client session, so both eras reach a person.
Supervisor lifecycle
The numbers, and the environment variable that moves each one (they exist for tests that cannot wait a minute; a deployment has no reason to touch them):
| Ping interval | 60 s | MCP_PING_INTERVAL_MS |
| Ping timeout | 30 s | MCP_PING_TIMEOUT_MS |
| Initial backoff | 1 s | MCP_BACKOFF_INITIAL_MS |
| Maximum backoff | 5 min | MCP_BACKOFF_MAX_MS |
| Backoff reset | after 5 min of uptime | MCP_BACKOFF_RESET_AFTER_MS |
A ping failure is treated as death: the client is closed, which triggers the same restart path an exit would. There is no separate "unhealthy but running" state to reason about.
While a server is not up, its path answers a JSON-RPC error with HTTP 503 naming the state. A client gets a clear failure instead of a hanging request.
Configuration hot reload
The config file is watched two ways: fs.watch on the parent directory, and fs.watchFile polling the file itself every 3 seconds. Both funnel into a 300 ms debounce.
The directory watch is what makes the recommended directory mount (./config:/config:ro) catch every kind of edit, including editors that save via rename. The poller is not belt-and-braces either: with a single-file bind mount — -v ./mcp.json:/config/mcp.json — an edit on the host produces no inotify event inside the container, so only polling sees in-place changes. And a rename-style save under a single-file mount is invisible to both — the mount binds one inode, and the new file is a new inode. That setup logs a startup warning.
On a change the new file is parsed and diffed against the running configuration. Added servers start, removed servers stop, changed servers restart, untouched servers keep their connections. A file that fails to parse is logged and ignored — the previous configuration stays live.
The /hub aggregate
Registering nine connectors puts nine servers' worth of tool schemas into the model's context before a question is asked. /hub inverts that: one connector, six meta-tools, and schemas fetched only when needed.
The hub keeps a per-server tool cache, refreshed when a child sends tools/list_changed, so list_tools answers without a round trip to the child. call_tool forwards with a five-minute deadline that is absolute by default — a progress notification does not extend it unless MCP_RESET_TIMEOUT_ON_PROGRESS says so.
Servers marked "hub": false are invisible here: list_servers omits them and call_tool refuses them. Their own paths are unaffected.
See the meta-tool reference for the exact schemas.
State on disk
/data holds everything that must survive a restart:
| File | Contents |
|---|---|
jwt-key.pem | the Ed25519 signing key, generated on first boot |
state.json | registered OAuth clients, approvals, refresh-token families, revocation markers |
mcp-hub.log | only if LOG_FILE points there |
There is no database and no migration step — deliberately: staying light enough for a single-board computer like a Raspberry Pi is a project goal, and a database would be the first thing to outgrow one. A corrupt state.json is moved aside as state.json.corrupt-<timestamp> and the hub boots with empty state rather than crash-looping — connectors then have to authorize again, which is recoverable, unlike a hub that will not start.
Losing /data invalidates every connector authorization. Treat both files as secrets: anyone holding jwt-key.pem can mint access tokens.