High availability
High availability requires the Obot application, database, and storage to remain available together. Configure and test each layer before adding replicas.
Application replicas and storage
The Kubernetes deployment guide describes replicaCount and external PostgreSQL. If replicas need the same local files, use a ReadWriteMany Obot data volume. A shared ReadWriteOnce claim is not a multi-replica storage solution.
Failure behavior
Tunnel peers route requests to connected tunnel clients across replicas. Failed in-flight requests are not replayed automatically. Peer-token rotation can temporarily interrupt tunneled traffic during a rolling update.
Validate application restart, workload rescheduling, database availability, and storage recovery in a staging environment. Keep backup and recovery procedures separate from availability configuration.
Plan each layer
| Layer | Required planning |
|---|---|
| Obot replicas | Configure the Helm replica count, capacity, and ingress routing. Review pod placement so one node failure does not remove every replica. |
| PostgreSQL | Use an external production database with its own availability and recovery plan. Additional Obot replicas do not replicate the database. |
| MCP workloads | Assess their own restart behavior and persistent storage. More Obot replicas do not make each workload highly available. |
| Tunnel clients | Operate clients where they can reach the private service; test client and Obot replica loss. |
Verify failover
In a staging environment, establish a client connection and perform a read-only tool call. Replace an Obot pod, reconnect, and verify a new call and its audit record. Separately test database failover and access to any shared files from each replica. Inspect failed in-flight requests; do not assume the gateway retries them.
For workload sizing, see Capacity. For recovery after data loss, see Backup and recovery.