Read-only replicas

A read-only replica is a second Curiosity Studio process that keeps a full copy of a primary workspace, follows the primary's writes as they happen, and serves read traffic: search, node and graph reads, and code endpoints marked read-only. Replicas spread read load off the primary, and they keep search and browsing available while the primary restarts.

This page explains how replication works, how the browser decides which server a request goes to, and what a deployment has to provide for both to work. The pages after it apply this to concrete setups:

Licensed feature

Replicas require the Replicas feature on the workspace license. A replica that loads a license without it logs License does not cover replica usage and exits.

Architecture

flowchart LR subgraph users[Browsers] b[Workspace front-end] end subgraph clients[Other clients] c[Data connectors, CLI, API tokens] end b -- "all requests, writes, WebSocket<br/>https://workspace.example.com" --> P b -- "read-only-compatible requests<br/>https://replica-1.workspace.example.com" --> R1 b -- "read-only-compatible requests<br/>https://replica-2.workspace.example.com" --> R2 c -- "reads and writes" --> P P[("Primary<br/>read-write")] R1[("Replica 1<br/>read-only")] R2[("Replica 2<br/>read-only")] R1 -- "gRPC :42999 (private)<br/>file copy, WAL stream, reports" --> P R2 -- "gRPC :42999 (private)" --> P R1 -. "HTTP (private)<br/>endpoints forwarded to the primary" .-> P

A deployment has exactly one primary, the only process that writes, and any number of replicas. The primary does not need to be configured with a list of replicas: each replica connects to the primary, registers itself, and tells the primary the public URL it can be reached at. The primary then advertises that URL to browsers.

Every replica keeps its own complete copy of the data on its own disk. Nothing is shared between the nodes: no shared filesystem, no shared block device, no database server.

What runs where

Work Primary Replica
Writes to the graph: nodes, edges, schema, configuration Yes No. A write throws on the replica and the request answers 303 See Other.
Data connectors, scheduled tasks, file uploads and extraction Yes No
Administration: settings, users, licenses, SSO configuration Yes No
The WebSocket the front-end keeps open (notifications, live updates) Yes No
Search, node and folder reads, graph queries, file text and blobs Yes Yes, for requests marked read-only-compatible
Code endpoints Yes Only endpoints marked Read Only, called with readOnlyCompatible
Sign-in Yes Yes. The replica issues tokens with the shared key, but does not update the user node (last-active time, password-hash upgrade).
Audit, search and click-through events caused by requests on the replica Recorded Forwarded to the primary over gRPC. The replica also writes its own audit files (…_audit_logs_replica-<uid>_…).

A replica has everything it needs to serve reads, because all of it lives in the replicated RocksDB database: nodes, edges, schemas, users and access groups, configuration, and the data of every index (text, vector and numeric). A replica does not rebuild indexes; it reads the index data the primary wrote. Files whose content is stored in the graph are replicated with it. When the workspace keeps file blobs in S3 (external blob storage), the bucket is shared, and each replica needs the same MSK_BLOB_STORAGE_AWS_* settings and access to it.

How a replica joins and stays current

sequenceDiagram autonumber participant R as Replica participant P as Primary (gRPC :42999) participant B as Browsers R->>P: RegisterReplica(new replica id) Note over P: starts tracking the replica,<br/>pins WAL and obsolete files R->>P: SyncInitialState P-->>R: file list of a RocksDB checkpoint loop every file that is missing, resized, or not an .sst R->>P: GetFile P-->>R: file content (1 MB chunks) end R->>P: DoneWithInitialState Note over R: opens the copy read-only R->>P: SyncFrontEndFiles / GetFrontEndFile R->>P: SyncUpdates(last sequence + 1) P-->>R: WAL chunks, streamed as they are written R->>P: RegisterReplicaIsOnline(public URL, internal URL) P->>B: index.html carries the replica URL (MSKREPLICAURLS) P->>B: REPLICAS_CHANGED over the WebSocket loop every 250 ms while chunks arrive R->>P: ReportLastSyncSequenceNumber end loop every second R->>P: ReplicaMetrics end
  1. Register. Each time the replica process starts it picks a new replica id and registers with the primary over gRPC, on port 42999 of the host named in MSK_PRIMARY_ADDRESS. The call is authenticated with a secret derived from MSK_JWT_KEY, so both sides need the same key.
  2. Copy the database. The primary takes a RocksDB checkpoint (hard links in <storage>/rep, so no data is copied on the primary) and streams the file list. The replica downloads what it lacks. SST files are immutable, so an .sst already on the replica's disk with the same name and size is kept; everything else is always downloaded. A replica restarted on its existing storage therefore downloads only what changed, and a new replica seeded from a snapshot of the primary's storage downloads only what changed since the snapshot.
  3. Open and sync the front-end. The replica opens its copy read-only and mirrors the primary's front-end folder, so it serves the same application, including a custom front-end. It re-syncs whenever the primary's front-end changes.
  4. Tail the write-ahead log. The replica asks for every change after its last sequence number. The primary streams write batches as it commits them, and the replica applies each chunk with a single write. The caches a read is served from are refreshed 5 ms after a chunk lands, and search readers reopen 250 ms after a replicated index commit.
  5. Go online. The replica sends its public and internal addresses. The primary adds the public address to the MSKREPLICAURLS value it injects into index.html, and tells connected browsers over the WebSocket.
  6. Report. The replica reports the last sequence number it applied every 250 ms while changes arrive, and a heartbeat with its replication counters every second. The primary uses these reports to measure lag and to decide which WAL files it can delete.

If the primary restarts, the next heartbeat tells the replica it is unknown. The replica registers again and resumes the WAL stream from its last sequence number, without copying the database again. It keeps serving reads, from its last state, while the primary is down.

If a replica stops reporting for 10 minutes, the primary drops it: it stops keeping WAL for it and removes its URL from the list sent to browsers. A replica that shuts down cleanly is dropped the same way, 10 minutes later, because replicas do not deregister.

Consistency

Replication is asynchronous. A write returns once the primary has committed it, before any replica has it. On a local primary and replica pair, a write became readable on the replica after a median of 4.3 ms (maximum 10 ms, 20 writes); network latency, sustained write load and an under-sized replica all add to that.

Two things follow:

  • No read-your-writes guarantee. A user who saves something and immediately reads it back through a replica can briefly see the previous version. In a custom front-end, leave a read that must see a write the user just made unmarked, so it goes to the primary.
  • Search results lag slightly more than node reads, because the search readers reopen 250 ms after an index commit replicates.

How the browser routes requests

The routing is done by the front-end in the browser, not by the server or a load balancer. The front-end keeps a list of servers and marks one of them as the primary:

Server Where the URL comes from
Primary The page's own origin plus /api. The front-end assumes the server that served the page is the primary.
Replicas MSKREPLICAURLS, injected into index.html by the primary: the public address each online replica registered with, ;-separated. Updated live by a REPLICAS_CHANGED WebSocket message when a replica comes online or is dropped.

A custom front-end can set AppSettings.ReplicaURLs in its settings callback instead. That list replaces MSKREPLICAURLS; the URLs must include the /api suffix.

The front-end checks each server with GET <server>/api/graph/ready on its first request, and again when a request receives 503 Service Unavailable (at most once every 5 seconds). A server counts as available when that returns true.

Each request is then routed like this:

flowchart TD A[Request] --> B{".ForcePrimary()?"} B -- yes --> P[Primary] B -- no --> C{"Marked .ReadOnlyCompatible()<br/>or primary unavailable?"} C -- no --> P C -- yes --> D{Any server available?} D -- no --> P D -- yes --> E["Random available server,<br/>primary included,<br/>kept for 5 to 30 minutes"] E --> F{Response 303?} F -- no --> G[Done] F -- yes --> H[Retry without the read-only flag] H --> C
  • Only marked requests use replicas. The built-in front-end marks its read calls with ReadOnlyCompatible(): search, node and folder reads, graph queries, file text and blobs, and exports. Everything else goes to the primary.
  • The primary is part of the pool. A replica is picked at random from all available servers, including the primary, and the choice is kept for 5 to 30 minutes (random per browser) while the server stays available. Load spreads across browsers, not across the requests of one browser.
  • A 303 sends the request back to the primary. A replica answers 303 See Other to a request that would write, and to a code endpoint that is not marked read-only. The front-end then repeats the request without the read-only flag.
  • When the primary is not ready, reads fail over. When the app starts and the primary is unreachable or still loading, the front-end checks the replicas. If one is available the app opens in read-only mode, with a Read-only mode banner, and reloads when the primary is ready again. While the app is already open, a 503 from the primary (which is what it answers while loading or shutting down) triggers a re-check of the servers, after which reads go to replicas. Writes fail, and keep failing after the primary is back until the page is reloaded.
  • Network errors do not trigger failover. A refused connection, a 502 or a 504 fails the request without re-checking the servers. Open browsers therefore switch to replicas while the primary is starting or shutting down, but not when its process has stopped or its host is unreachable; they recover on the next page load.
  • Opening a replica's URL directly loads the same application, with a READ-ONLY REPLICA banner. In that tab every request goes to replicas, and writes cannot reach the primary. Use it for diagnostics, or as the fallback address during a primary outage; do not send users there otherwise.

Manage → Developer tools shows the server list as the current browser sees it, with each server's availability.

Code endpoints

A code endpoint runs on a replica only when both sides allow it:

  • The endpoint has Read Only enabled in the endpoint editor. It then compiles against the read-only scope, which has no write methods.
  • The caller marks the call: API.Endpoints.CallAsync<T>(name, body, readOnlyCompatible: true) in the front-end SDK, or .ReadOnlyCompatible() on a raw REQ.

An endpoint that is not marked Read Only answers 303 on a replica, and the front-end repeats the call on the primary. A read-only endpoint running on a replica can call RunEndpointOnPrimaryAsync to run another endpoint on the primary; the replica forwards that call over HTTP to MSK_PRIMARY_ADDRESS, authenticated with a token the primary issued at registration. See Global scope.

AI tools that wrap an endpoint follow the endpoint's flag.

Clients other than the browser

The replica routing lives only in the workspace front-end. Data connectors, curiosity-cli, the C# and Python libraries and anything calling the REST API with a token talk to the URL they are given, and most of them write. Point them at the primary.

Assumptions and requirements

The routing and replication above only work when the deployment provides the following. Each item names what breaks without it.

Requirement Why
One fixed primary. There is no election and no automatic promotion. Promoting a replica is a manual restart; see Failover.
The workspace hostname always reaches the primary, and only the primary. The front-end treats the server that served the page as the primary. If that name can land on a replica (DNS round-robin, a shared load-balancer pool, health-checked DNS failover), writes that reach the replica are rejected with 303, and the front-end retries the same URL with no limit.
Every replica has its own stable public URL, and that URL reaches only that replica. Browsers call replicas directly by the URL each replica registered with. Long-running endpoint calls poll the same URL for their result, so one name must not be balanced across several replicas.
Replica URLs are reachable from users' browsers, over HTTPS when the primary is served over HTTPS. Browsers block http:// requests from an https:// page (mixed content).
Each replica allows the primary's public origin in CORS. The page is served from the primary's origin, so calls to a replica are cross-origin. A replica automatically allows the origin of MSK_PRIMARY_ADDRESS, which is usually a private address, not the one users browse to. Set MSK_CORS on each replica to the primary's public origin.
A private network path from each replica to the primary, on port 42999 and on the primary's HTTP port. The replication channel is gRPC over HTTP/2 without TLS (h2c), authenticated by a shared secret sent in a header. MSK_PRIMARY_ADDRESS must name the primary itself, not a load balancer: the replica connects to port 42999 of that host and sends forwarded endpoint calls to the URL. Never expose 42999 outside the private network; anyone who can read that traffic can copy the database.
The same MSK_JWT_KEY on every node, set in the environment. It signs user and API tokens, so a token issued by one node is accepted by all of them, and the replication secret is derived from it. The primary only opens port 42999 when the key is set. Adding a key to an existing workspace invalidates every token issued before, including API tokens.
The same Curiosity Studio version on every node. The replica does not check the primary's version. Upgrade them together; see Upgrades.
The same MSK_GRAPH_MASTER_KEY, and the same external blob storage settings, when the primary uses them. The replica opens a byte-for-byte copy of the primary's data.
Fast local disk on each replica, sized like the primary's. Each replica holds the whole database and does the same random reads as the primary. Use local SSD or NVMe.
Disk headroom on the primary. While any replica is registered, the primary does not let RocksDB delete obsolete files, so compaction output accumulates until no replica is registered. Watch the RocksDB_ObsoleteSstFilesSize metric.

Configuration

Replica

Variable Required Value
MSK_REPLICA yes true
MSK_PRIMARY_ADDRESS yes A URL of the primary whose host name resolves to the primary's private address on the replica, for example http://primary.workspace.internal:8080, or https://workspace.example.com when the primary serves HTTPS itself and that name points at its private address on the replica hosts. The replica connects to port 42999 of this host for replication, and posts forwarded endpoint calls to this URL.
MSK_JWT_KEY yes Same value as the primary. The replica refuses to start without it.
MSK_PUBLIC_ADDRESS yes The URL users' browsers reach this replica at, for example https://replica-1.workspace.example.com. This is what the primary advertises.
MSK_SERVER_ADDRESS yes The replica's private URL. Must be set: an empty value makes registration fail and the replica exits.
MSK_CORS yes, in practice The primary's public origin, for example https://workspace.example.com. ;-separated when there are several.
MSK_GRAPH_STORAGE yes Local storage for the replica's copy. Never the primary's folder.
MSK_GRAPH_MASTER_KEY when the primary sets it Same value as the primary.
MSK_BLOB_STORAGE_AWS_* when the primary sets them Same values as the primary. At startup the replica writes, reads and deletes a test object in the bucket.
MSK_PORT no HTTP port, 8080 by default.

Settings stored in the workspace (Manage pages) replicate with the data, so a replica needs no copy of them. The license is part of that stored configuration unless it is passed with MSK_LICENSE, in which case pass it to the replicas too.

Primary

Variable Required Value
MSK_JWT_KEY yes The shared key. Replication is enabled when it is set: the primary then listens on 42999 on all interfaces.
MSK_PUBLIC_ADDRESS recommended The workspace URL, for example https://workspace.example.com. Not used by replication, but used in links the workspace generates.

The primary needs no replica-specific setting. It accepts any replica that presents the secret derived from its key.

Ports

From To Port Protocol Purpose
Browsers Primary (directly or through a load balancer) 443 HTTPS, WebSocket The application and every non-replica request
Browsers Each replica (directly or through a load balancer) 443 HTTPS Read-only-compatible requests
Replica Primary 42999 gRPC, HTTP/2 cleartext Registration, database copy, WAL stream, reports, forwarded events
Replica Primary MSK_PORT (8080, or 443 when the node serves HTTPS itself) HTTP or HTTPS Endpoints forwarded with RunEndpointOnPrimaryAsync
Primary Replica none The primary never connects to a replica

Where to go next

© 2026 Curiosity. All rights reserved.
Powered by Neko