Operate replicas
This page covers checking that replicas are healthy and current, upgrading a deployment with replicas, promoting a replica when the primary is lost, and diagnosing a replica that does not join. For how replication works, see Read-only replicas.
Health endpoints
| Endpoint | Auth | Answers |
|---|---|---|
GET /api/graph/ready |
none | true once the graph is loaded. What the front-end uses to decide whether a server is available. |
GET /api/available-status |
none | 200 available once the graph is loaded, 503 not available before. Use it for load balancer health checks. |
GET /health |
none | Aggregate report: memory, disk, readiness. |
GET /health/replica |
none | The node's role, and on the primary, the lag of each replica. |
GET /api/graph/replica |
system admin | The node's role and replica id. |
/health/replica on the primary:
{
"isReplica": false,
"hasReplicas": true,
"currentLag": 42,
"averageLag": 35.2,
"emaLag": 33.1,
"recommendedDelayMs": 200,
"replicas": [
{ "replicaUID": "…", "replicaName": "replica-cheerful-orion", "lagVersions": 42 },
{ "replicaUID": "…", "replicaName": "replica-quiet-fern", "lagVersions": 17 }
]
}
On a replica:
{
"isReplica": true,
"hasReplicas": false,
"replicaUID": "…",
"replicaName": "replica-cheerful-orion"
}
| Field | Meaning |
|---|---|
hasReplicas |
At least one replica is registered. |
replicas[].lagVersions |
How many RocksDB sequence numbers the replica is behind the primary, as of its last report. One write batch advances the sequence by the number of operations in it, so this counts operations, not seconds. |
currentLag |
The largest lagVersions across replicas, at the last report. |
averageLag, emaLag |
Mean and exponential moving average of currentLag over recent reports. Alert on emaLag: it ignores single spikes. |
recommendedDelayMs |
The commit delay, from 0 to 2000 ms, that the primary's lag controller recommends for the current lag. The primary reports it but does not currently slow its writes by it. |
The lag fields and replicas are only present on a primary with MSK_JWT_KEY set. A replica reports every 250 ms while it is receiving changes. A replica that stops reporting keeps its last lagVersions until the primary drops it, 10 minutes after its last report, so alert on a replica missing from replicas[] as well as on lag.
replicaName is derived from the replica id, which is new every time the replica process starts. A restarted replica therefore appears under a new name, and its old entry stays listed until it is dropped.
Monitoring
In Manage → Monitoring, each replica has its own series for its replication version and its replication error count, labelled with its replica name. On the primary, also watch:
RocksDB_ObsoleteSstFilesSize. While any replica is registered, the primary does not let RocksDB delete obsolete files. Compaction output then accumulates on the primary's disk until no replica is registered. Leave disk headroom, and alarm on free space.- The WAL archive. The primary keeps write-ahead log files until every registered replica has reported a sequence number past them, and deletes them after that. A replica that is registered but not reporting holds them for up to 10 minutes.
Manage → Developer tools shows the server list from the point of view of the browser it runs in: which replicas it knows, whether it considers each available, and which one it is currently using.
Logs
A replica logs each step of joining. Replica log lines are prefixed [SEC] on the console.
Registering with primary 'http://primary.workspace.internal:8080' with Replica ID …
Registered with primary 'http://primary.workspace.internal:8080' with Replica ID …
Starting initial catching-up of replica with with primary … on folder: '…/rocksdb'
Fetching file: 000123.sst with 67108864 bytes (Reason: File not yet sync'd)
Done catching up up with primary in 42 seconds
Loading read-only graph from storage
Starting front-end sync with primary into …
Done front-end sync with primary in 812ms: fetched 214 of 214 files
Starting to replicate WAL...
Destination DB sequence number: 1234567. Starting WAL sync...
Registering with primary 'http://primary.workspace.internal:8080': public address: 'https://replica-1.workspace.example.com', internal address: 'http://10.0.1.11:8080'
Registered with primary 'http://primary.workspace.internal:8080': public address: '…', internal address: '…'
On the primary, the same join shows as Registering new secondary read-only replica: … and Registering replica with public address '…'. The primary's replication channel logs to its own folder, <MSK_LOG_PATH>/replication, with [RPC] on the console.
Troubleshooting
| Symptom | Cause and fix |
|---|---|
Replica fails at startup with Setting a common JWT key using MSK_JWT_KEY between primary and replicas is mandatory for replication |
MSK_JWT_KEY is not set in the replica's environment. |
Replica fails to start right after Registering with primary, with a gRPC Unavailable error |
Port 42999 on the host in MSK_PRIMARY_ADDRESS is not reachable. Check the firewall or security group, that the address names the primary and not a load balancer, and that MSK_JWT_KEY is set on the primary: without it, the primary does not listen on 42999. |
Replica fails with Replication is not authenticated; the primary logs Replica failed to connect with invalid secret |
The two MSK_JWT_KEY values differ. |
Replica exits with Failed to catch-up with primary database, exiting... (exit code 173 on Linux) |
The database copy failed part-way: network interruption, disk full on the replica, or the primary restarted during the copy. Fix the cause and restart; files already copied are kept. |
Replica exits with Failed to register with replica (exit code 173); the primary logs UriFormatException |
MSK_PUBLIC_ADDRESS or MSK_SERVER_ADDRESS is empty on the replica. Both are required. |
Replica exits after loading, with License does not cover replica usage |
The workspace license does not include the Replicas feature. |
| Replica joins, but browsers never use it | Check the replica's CORS: an OPTIONS request with Origin: <primary's public origin> must return Access-Control-Allow-Origin. Set MSK_CORS on the replica. Then check that its MSK_PUBLIC_ADDRESS is reachable from a browser, over HTTPS if the primary is served over HTTPS. |
MSKREPLICAURLS is missing from the primary's page |
The replica has not completed its join, or was dropped. Check /health/replica and the replica's log. |
| Browsers keep calling a replica that is down | The front-end only re-checks its servers on page load and on a 503; a stopped replica fails with a network error instead. Reloading the page restores normal routing. See Known limitations. |
Writes hang in the browser, repeating a request that answers 303 |
The request is reaching a replica: either the workspace hostname routes to a replica, or the tab was opened on a replica URL. Route the workspace name to the primary only. |
| Reads are stale | Expected for a few milliseconds after a write, longer under heavy write load or on an under-sized replica. Compare lagVersions over time; a value that only grows means the replica cannot keep up. |
| Primary disk fills up while replicas are connected | Obsolete RocksDB files are kept while any replica is registered. Check RocksDB_ObsoleteSstFilesSize and add disk. |
Upgrades
The replica does not check the primary's version. Keep every node on the same version, and keep the window in which they differ short:
- Pull the new image on every node.
- Restart the primary on the new version. While it is down, the replicas keep serving reads from their current state, and open browsers switch to read-only mode.
- When the primary answers
trueon/api/graph/ready, restart the replicas on the new version, one at a time, so at least one keeps serving reads. Each registers again and copies only the files that changed. - Reload the app in the browser to leave read-only mode, if it did not reload by itself.
Failover
There is no automatic promotion. If the primary is lost, you can turn a replica into the new primary:
- Pick a replica. If the primary still answers, pick the one with the lowest
lagVersionsin/health/replica. Otherwise any replica will do, and writes it had not received when the primary was lost are gone. Replication is asynchronous: under normal load, that is the last few milliseconds of writes. - Make sure the old primary cannot come back as a primary. Stop it, and keep it stopped. Two primaries accepting writes cannot be merged.
- Stop the chosen replica and start it again as a primary: same storage folder, same
MSK_JWT_KEYandMSK_GRAPH_MASTER_KEY, withoutMSK_REPLICAandMSK_PRIMARY_ADDRESS, with the primary'sMSK_PUBLIC_ADDRESS, and with port42999open to the other replicas. - Route the workspace name to it: the DNS record, or the load balancer target.
- Point the other replicas at it, with empty storage. Change their
MSK_PRIMARY_ADDRESS(or the private DNS name it uses), clear their storage folders, and start them. Their existing files were written by themselves and can share names with different files on the new primary, so do not let them reuse those files. - Bring the old primary back as a replica, with empty storage, once it is repaired, or retire it.
Rehearse this before you need it, and keep backups: a replica is not one. See Backup and restore.
Rotating the shared key
MSK_JWT_KEY signs every user session and API token, and the replication secret is derived from it. Rotating it:
- invalidates every session and API token, including tokens used by data connectors;
- needs every node restarted with the new value, primary first. A replica still on the old key cannot register until it is restarted.
There is no overlap period with two valid keys. Schedule the rotation in a maintenance window and re-issue API tokens afterwards.
Known limitations
- One primary, manual promotion. See Failover.
- Routing depends on hostnames. The browser treats the server that served the page as the primary, and a request that a replica rejects with
303is retried without a limit. See Assumptions and requirements. - The browser re-checks servers only on page load and on a
503. A replica that fails with a network error, and a primary that comes back after an outage, are not noticed until then. A reload restores normal routing. - New replicas reach open browsers late. Open browsers receive a new replica's URL over the WebSocket but only use it after their next server check, which in practice is the next page load.
- Replicas do not deregister. A stopped replica stays registered, and its URL stays advertised, for 10 minutes. During that time it also holds WAL on the primary.
- Only the workspace front-end routes to replicas. Connectors, the CLI, the libraries and REST clients use the URL they are given.