Deploy replicas with DNS routing

This page deploys a primary and two read-only replicas on three hosts, with one DNS name per node. Each node serves HTTPS itself, with its own certificate. There is no load balancer and no reverse proxy: the browser picks the node, as described in How the browser routes requests, so DNS only has to map each name to its host.

Read Read-only replicas first; the requirements listed there are what this setup is built to satisfy.

Topology

flowchart TB subgraph internet[Public DNS] d1["workspace.example.com → 203.0.113.10"] d2["replica-1.workspace.example.com → 203.0.113.11"] d3["replica-2.workspace.example.com → 203.0.113.12"] end subgraph h0["Host 1 · 10.0.1.10"] p["Primary<br/>HTTPS :443<br/>gRPC :42999"] end subgraph h1["Host 2 · 10.0.1.11"] r1["Replica<br/>HTTPS :443"] end subgraph h2["Host 3 · 10.0.1.12"] r2["Replica<br/>HTTPS :443"] end d1 --> p d2 --> r1 d3 --> r2 r1 -- "private network<br/>10.0.1.10:42999, :443" --> p r2 -- "private network" --> p
Name Resolves to Used by
workspace.example.com Host 1 public address; on the replica hosts, host 1's private address Users, data connectors, API clients. Always the primary. The replicas use it as MSK_PRIMARY_ADDRESS.
replica-1.workspace.example.com Host 2 public address Browsers, for read-only-compatible requests.
replica-2.workspace.example.com Host 3 public address Browsers, for read-only-compatible requests.

The replicas reach the primary by its public name, so the primary's certificate is valid for the connection, but that name must resolve to the primary's private address on the replica hosts. The Compose files below do this with extra_hosts; a split-horizon DNS zone does the same for a whole network.

The three hosts need a private network between them. If they only share the public internet, run the replication traffic over a VPN or WireGuard tunnel and use the tunnel addresses as the private addresses below: port 42999 carries the whole database without TLS.

What not to do

These setups look reasonable for DNS and break the front-end's routing:

  • One name for all nodes (round-robin A records, or a load balancer pool with the primary and the replicas). The front-end treats the server that served the page as the primary. Writes that land on a replica are rejected with 303, and the front-end retries them with no limit.
  • Health-checked DNS failover of the workspace name to a replica. The same problem, triggered exactly when the primary is down. It also leaves users on a read-only node that looks like a primary, with no banner.
  • One name balanced across several replicas. Long-running endpoint calls poll the URL they started on, and the poll can land on a node that never saw the call.
  • MSK_PRIMARY_ADDRESS resolving to the primary's public address. The replica opens its gRPC channel to port 42999 of that host, which must stay closed on the public side.

Steps

1

Create the DNS records

workspace.example.com.            300  IN  A  203.0.113.10
replica-1.workspace.example.com.  300  IN  A  203.0.113.11
replica-2.workspace.example.com.  300  IN  A  203.0.113.12

A short TTL does not give you failover (see above). It only shortens the wait when you move a name to a new host.

2

Get TLS certificates

Each node presents the certificate for its own name, through MSK_CERT_FILE and MSK_CERT_FILE_PRIVATE_KEY (PEM). Either one certificate per host, or a certificate for workspace.example.com and *.workspace.example.com copied to all three. A wildcard does not cover the bare workspace.example.com, so list both names.

With a certificate set, Curiosity serves HTTPS on MSK_PORT, 443 by default. The replicas must be served over HTTPS when the primary is: a page loaded over https:// cannot call an http:// replica.

3

Open the replication port to the replicas only

On the primary host, allow port 42999 from the replicas' private addresses, and nothing else. Port 443 is open to users. Do it in the network firewall or security group in front of the host where you can.

Docker publishes container ports through its own iptables rules, which ufw does not filter. Bind the replication port to the private address, as the Compose file in the next step does, and add a rule in the network firewall:

# Network firewall / cloud rule, expressed as iptables for a host without Docker
iptables -A INPUT -p tcp -s 10.0.1.11 --dport 42999 -j ACCEPT
iptables -A INPUT -p tcp -s 10.0.1.12 --dport 42999 -j ACCEPT
iptables -A INPUT -p tcp --dport 42999 -j DROP

The replicas need no inbound rule from the primary. The primary never connects to a replica.

4

Start the primary

Generate the shared key once and keep it in your secret store. Every node gets the same value.

openssl rand -base64 48

/etc/curiosity/secrets.env, on all three hosts:

MSK_JWT_KEY=<shared key>
MSK_GRAPH_MASTER_KEY=<if you encrypt the graph>

/srv/curiosity/compose.yml on host 1:

services:
curiosity:
image: curiosityai/curiosity:<version>
restart: unless-stopped
ports:
- "443:443"                  # users, and the replicas' forwarded endpoint calls
- "10.0.1.10:42999:42999"    # replicas: replication (gRPC, no TLS)
volumes:
- /srv/curiosity/data:/data
- /etc/ssl/workspace:/certs:ro
env_file: /etc/curiosity/secrets.env
environment:
MSK_GRAPH_STORAGE: /data/curiosity
MSK_PUBLIC_ADDRESS: https://workspace.example.com
MSK_CERT_FILE: /certs/fullchain.pem
MSK_CERT_FILE_PRIVATE_KEY: /certs/privkey.pem
docker compose -f /srv/curiosity/compose.yml up -d
curl -s https://workspace.example.com/api/graph/ready    # true once the graph is loaded

If this is an existing workspace that ran without MSK_JWT_KEY, adding it invalidates every token issued so far: users sign in again, and API tokens used by connectors have to be re-created. Schedule it.

5

Start each replica

/srv/curiosity/compose.yml on host 2:

services:
curiosity:
image: curiosityai/curiosity:<same version as the primary>
restart: unless-stopped
ports:
- "443:443"
extra_hosts:
- "workspace.example.com:10.0.1.10"   # reach the primary over the private network
volumes:
- /srv/curiosity/data:/data
- /etc/ssl/workspace:/certs:ro
env_file: /etc/curiosity/secrets.env
environment:
MSK_REPLICA: "true"
MSK_PRIMARY_ADDRESS: https://workspace.example.com                  # resolves to 10.0.1.10 here
MSK_PUBLIC_ADDRESS: https://replica-1.workspace.example.com         # what browsers call
MSK_SERVER_ADDRESS: https://10.0.1.11                               # required, must not be empty
MSK_CORS: https://workspace.example.com                             # the primary's public origin
MSK_CERT_FILE: /certs/fullchain.pem
MSK_CERT_FILE_PRIVATE_KEY: /certs/privkey.pem
MSK_GRAPH_STORAGE: /data/curiosity

Host 3 is the same with replica-2.workspace.example.com and 10.0.1.12.

docker compose -f /srv/curiosity/compose.yml up -d
docker compose -f /srv/curiosity/compose.yml logs -f

The first start copies the whole database, so it takes as long as transferring the primary's storage folder over the private network. The log shows each file as it is fetched, then Starting to replicate WAL... and Registered with primary. A replica that fails to register or to copy the database exits; see Troubleshooting.

6

Verify

# From a replica host: the primary is reachable over the private network
curl -s --resolve workspace.example.com:443:10.0.1.10 https://workspace.example.com/api/graph/ready
nc -zv 10.0.1.10 42999

# Each replica is up and reachable on its public name
curl -s https://replica-1.workspace.example.com/api/graph/ready    # true

# The primary advertises the replicas to browsers
curl -s https://workspace.example.com/ | grep -o "MSKREPLICAURLS = '[^']*'"
# MSKREPLICAURLS = 'https://replica-1.workspace.example.com/;https://replica-2.workspace.example.com/'

# The primary tracks both replicas and their lag
curl -s https://workspace.example.com/health/replica

# The replica accepts cross-origin calls from the primary's origin
curl -si -X OPTIONS https://replica-1.workspace.example.com/api/graph/ready \
-H "Origin: https://workspace.example.com" \
-H "Access-Control-Request-Method: POST" \
-H "Access-Control-Request-Headers: authorization,content-type" | grep -i '^access-control'

The last command must print an Access-Control-Allow-Origin: https://workspace.example.com line. If it prints nothing, MSK_CORS is missing on that replica, and browsers will fail every request they send to it.

Finally, sign in at https://workspace.example.com, run a search, and open Manage → Developer tools: both replicas are listed as available, and the search request went to one of the three servers.

When a node is down

Situation What users see
Primary restarting or upgrading (process up, graph loading or unloading) The primary answers 503. Open browsers switch reads to the replicas, and writes fail. Once the primary is back, a page reload restores writes.
Primary process stopped, or host 1 unreachable Requests to workspace.example.com fail with a network error. Open browsers show errors instead of switching to read-only mode, and new visitors cannot load the app from that name.
A replica stopped Browsers that load the app afterwards find it unavailable and do not use it. Browsers already using it get errors on read-only-compatible requests until they reload.

When the primary is down, a replica URL still serves the whole application read-only, with a READ-ONLY REPLICA banner. Tell users to open it. Replicas keep serving the data they had when the primary went down. To turn a replica into the new primary, follow Failover.

Adding and removing replicas

  • Add: create the DNS record and certificate, allow its private address to reach the primary on 42999 and 443, and start it with its own MSK_PUBLIC_ADDRESS. The primary advertises it once it has caught up. Browsers with the app open receive the URL over the WebSocket, but only use it after they next check their servers, which in practice means after a reload.
  • Remove: stop it. The primary keeps it registered, and keeps advertising its URL, for 10 minutes after its last report, then drops it. A browser that loads the app in that window probes it, finds it unavailable and does not use it. Remove the DNS record afterwards.
  • Replace a replica's host: storage is disposable. A new host with empty storage copies the database from the primary. To shorten the copy of a large database, seed the new storage folder from a snapshot of the primary's storage: the replica keeps every .sst file whose name and size match the primary's. Do not seed it from another replica's folder, whose .sst files were written by that replica and can share names with different files on the primary.
© 2026 Curiosity. All rights reserved.
Powered by Neko