Deploy replicas on AWS

This page deploys a primary and two read-only replicas on EC2 in one VPC, behind one Application Load Balancer. It follows the same rules as Deploy with DNS routing: one public name per node, and the browser picks the node. The load balancer only terminates TLS and maps each name to one instance.

Read Read-only replicas for how replication and routing work, and Deploying on AWS for the single-node setup this extends.

Architecture

flowchart TB users[Browsers] -->|"HTTPS 443<br/>workspace / replica-1 / replica-2<br/>.workspace.example.com"| alb subgraph vpc["VPC 10.0.0.0/16"] subgraph pub["Public subnets (3 AZs)"] alb["Application Load Balancer<br/>ACM certificate<br/>host-header rules"] nat[NAT gateway] end subgraph a["Private subnet · AZ a"] p["EC2 primary<br/>curiosity :8080<br/>gRPC :42999<br/>EBS gp3"] end subgraph b["Private subnet · AZ b"] r1["EC2 replica-1<br/>curiosity :8080"] end subgraph c["Private subnet · AZ c"] r2["EC2 replica-2<br/>curiosity :8080"] end end alb -->|tg-primary| p alb -->|tg-replica-1| r1 alb -->|tg-replica-2| r2 r1 -->|"42999, 8080<br/>primary.workspace.internal"| p r2 -->|"42999, 8080"| p p & r1 & r2 -.->|egress| nat
Component Configuration
VPC Public subnets for the ALB and NAT gateway, private subnets for the instances, in at least two Availability Zones. Put the replicas in different AZs from the primary, so an AZ outage leaves a replica serving reads.
ALB Internet-facing (or internal, for a private workspace), HTTPS listener with an ACM certificate for workspace.example.com and *.workspace.example.com, HTTP listener that redirects to HTTPS.
Target groups One per node, each with exactly one instance: tg-primary, tg-replica-1, tg-replica-2.
Listener rules Host: replica-1.workspace.example.com → tg-replica-1, Host: replica-2.workspace.example.com → tg-replica-2, default → tg-primary.
EC2 One instance per node, running the curiosityai/curiosity image with Docker. The ALB sends traffic straight to the container on port 8080; TLS ends at the ALB.
Route 53, public zone Alias records for all three names, pointing at the ALB.
Route 53, private zone primary.workspace.internal → the primary's private IP. Replicas use it as MSK_PRIMARY_ADDRESS.
Secrets Manager MSK_JWT_KEY and MSK_GRAPH_MASTER_KEY, read by every instance through its instance role.

What the load balancer must not do

The ALB is tempting to use for things the front-end already does, and each of these breaks it:

  • Do not put the primary and the replicas in one target group. The front-end treats whichever server answers workspace.example.com as the primary. A write that lands on a replica is rejected with 303, and the front-end repeats it with no limit.
  • Do not put two replicas in one target group behind a shared replica name. Long-running endpoint calls poll the URL they started on, and ALB stickiness does not help: the front-end's cross-origin requests do not send cookies.
  • Do not use a Route 53 failover record that moves workspace.example.com to a replica when the primary is unhealthy. It causes the first problem at the moment the primary is down.
  • Do not route replication through the ALB. Port 42999 is gRPC over cleartext HTTP/2, between instances, inside the VPC. MSK_PRIMARY_ADDRESS names the primary instance, not the load balancer.

Security groups

flowchart LR internet((Internet)) -->|443, 80| sgalb[sg-alb] sgalb -->|8080| sgp[sg-primary] sgalb -->|8080| sgr[sg-replica] sgr -->|"42999, 8080"| sgp
Group Inbound Attached to
sg-alb 443 and 80 from 0.0.0.0/0, or from your corporate ranges ALB
sg-primary 8080 from sg-alb; 42999 and 8080 from sg-replica Primary instance
sg-replica 8080 from sg-alb Replica instances

Nothing is inbound from the primary to the replicas: the primary never opens a connection to a replica. Outbound traffic from the instances (LLM providers, the Docker registry) leaves through the NAT gateway. Use Systems Manager Session Manager for shell access instead of opening port 22.

Port 42999 carries the whole database without TLS, authenticated by a secret derived from MSK_JWT_KEY. Keep it limited to sg-replica, and turn on VPC Flow Logs if you need a record of who connected to it.

Terraform

The load balancer, target groups, listener rules and security groups. VPC, subnets, NAT, instances and the ACM certificate are assumed to exist.

alb.tf
locals {
  domain   = "workspace.example.com"
  replicas = { "replica-1" = aws_instance.replica_1.id, "replica-2" = aws_instance.replica_2.id }
}

resource "aws_security_group" "alb" {
  name   = "curiosity-alb"
  vpc_id = var.vpc_id

  ingress {
    from_port       = 443
    to_port         = 443
    protocol        = "tcp"
    cidr_blocks     = ["0.0.0.0/0"]
  }

  ingress {
    from_port       = 80
    to_port         = 80
    protocol        = "tcp"
    cidr_blocks     = ["0.0.0.0/0"]
  }

  egress {
    from_port       = 0
    to_port         = 0
    protocol        = "-1"
    cidr_blocks     = ["0.0.0.0/0"]
  }
}

resource "aws_security_group" "replica" {
  name   = "curiosity-replica"
  vpc_id = var.vpc_id

  # HTTP, from the load balancer
  ingress {
    from_port       = 8080
    to_port         = 8080
    protocol        = "tcp"
    security_groups = [aws_security_group.alb.id]
  }

  egress {
    from_port       = 0
    to_port         = 0
    protocol        = "-1"
    cidr_blocks     = ["0.0.0.0/0"]
  }
}

resource "aws_security_group" "primary" {
  name   = "curiosity-primary"
  vpc_id = var.vpc_id

  # HTTP, from the load balancer
  ingress {
    from_port       = 8080
    to_port         = 8080
    protocol        = "tcp"
    security_groups = [aws_security_group.alb.id]
  }

  # replication, from replicas only
  ingress {
    from_port       = 42999
    to_port         = 42999
    protocol        = "tcp"
    security_groups = [aws_security_group.replica.id]
  }

  # endpoint calls forwarded by replicas
  ingress {
    from_port       = 8080
    to_port         = 8080
    protocol        = "tcp"
    security_groups = [aws_security_group.replica.id]
  }

  egress {
    from_port       = 0
    to_port         = 0
    protocol        = "-1"
    cidr_blocks     = ["0.0.0.0/0"]
  }
}

resource "aws_lb" "curiosity" {
  name               = "curiosity"
  load_balancer_type = "application"
  subnets            = var.public_subnet_ids
  security_groups    = [aws_security_group.alb.id]
}

# One target group per node, one target per group.
resource "aws_lb_target_group" "node" {
  for_each = merge({ "primary" = aws_instance.primary.id }, local.replicas)

  name     = "curiosity-${each.key}"
  port     = 8080
  protocol = "HTTP"
  vpc_id   = var.vpc_id

  health_check {
    path                = "/api/available-status"   # 200 once the graph is loaded, 503 otherwise
    matcher             = "200"
    interval            = 15
    healthy_threshold   = 2
    unhealthy_threshold = 2
  }
}

resource "aws_lb_target_group_attachment" "node" {
  for_each         = aws_lb_target_group.node
  target_group_arn = each.value.arn
  target_id        = each.key == "primary" ? aws_instance.primary.id : local.replicas[each.key]
  port             = 8080
}

resource "aws_lb_listener" "https" {
  load_balancer_arn = aws_lb.curiosity.arn
  port              = 443
  protocol          = "HTTPS"
  ssl_policy        = "ELBSecurityPolicy-TLS13-1-2-2021-06"
  certificate_arn   = var.certificate_arn   # covers workspace.example.com and *.workspace.example.com

  default_action {
    type             = "forward"
    target_group_arn = aws_lb_target_group.node["primary"].arn
  }
}

resource "aws_lb_listener_rule" "replica" {
  for_each     = local.replicas
  listener_arn = aws_lb_listener.https.arn
  priority     = 10 + index(keys(local.replicas), each.key)

  condition {
    host_header { values = ["${each.key}.${local.domain}"] }
  }

  action {
    type             = "forward"
    target_group_arn = aws_lb_target_group.node[each.key].arn
  }
}

resource "aws_lb_listener" "http" {
  load_balancer_arn = aws_lb.curiosity.arn
  port              = 80
  protocol          = "HTTP"

  default_action {
    type = "redirect"
    redirect {
      port        = "443"
      protocol    = "HTTPS"
      status_code = "HTTP_301"
    }
  }
}

resource "aws_route53_record" "public" {
  for_each = toset(concat([local.domain], [for k in keys(local.replicas) : "${k}.${local.domain}"]))

  zone_id = var.public_zone_id
  name    = each.value
  type    = "A"

  alias {
    name                   = aws_lb.curiosity.dns_name
    zone_id                = aws_lb.curiosity.zone_id
    evaluate_target_health = false
  }
}

resource "aws_route53_record" "primary_internal" {
  zone_id = var.private_zone_id              # private hosted zone associated with the VPC
  name    = "primary.workspace.internal"
  type    = "A"
  ttl     = 60
  records = [aws_instance.primary.private_ip]
}

The instances

Sizing and storage

  • Same instance type for every node. A replica serves the same queries as the primary, from its own copy of the data and its own caches. See Deploying on AWS and Scaling for sizing the primary.
  • Primary: EBS gp3, with snapshots through Data Lifecycle Manager, as in the single-node setup. Leave headroom: while any replica is registered, the primary does not delete obsolete RocksDB files. Watch the RocksDB_ObsoleteSstFilesSize metric.
  • Replicas: EBS gp3 of the same size, or instance-store NVMe. A replica's storage is disposable, since it can always be copied again from the primary. Instance store (r6id, m6id, i4i) is faster and costs nothing extra, but it is erased when the instance stops, and the next start copies the whole database from the primary again. Replicas need no snapshots.

Instance role

Every instance reads its secrets at boot:

{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Effect": "Allow",
      "Action": "secretsmanager:GetSecretValue",
      "Resource": "arn:aws:secretsmanager:eu-central-1:123456789012:secret:curiosity/prod-*"
    }
  ]
}

If the workspace keeps file blobs in S3 (MSK_BLOB_STORAGE_AWS_S3_BUCKET), the replicas need the same bucket settings and the same object permissions as the primary: at startup every node writes, reads and deletes a test object to check the bucket works.

{
  "Effect": "Allow",
  "Action": ["s3:GetObject", "s3:PutObject", "s3:DeleteObject"],
  "Resource": "arn:aws:s3:::curiosity-prod-blobs/*"
}

Environment

Store the shared values in one Secrets Manager secret, so every node gets the same MSK_JWT_KEY:

aws secretsmanager create-secret --name curiosity/prod \
  --secret-string "$(jq -n --arg jwt "$(openssl rand -base64 48)" \
                           --arg mk  "$(openssl rand -base64 48)" \
                           '{MSK_JWT_KEY:$jwt, MSK_GRAPH_MASTER_KEY:$mk}')"

At boot, each instance writes it to an env file:

aws secretsmanager get-secret-value --secret-id curiosity/prod \
  --query SecretString --output text \
  | jq -r 'to_entries[] | "\(.key)=\(.value)"' > /etc/curiosity/secrets.env
chmod 600 /etc/curiosity/secrets.env

Primary, /srv/curiosity/compose.yml:

services:
  curiosity:
    image: curiosityai/curiosity:<version>
    restart: unless-stopped
    ports:
      - "8080:8080"      # the ALB, and forwarded endpoint calls from replicas (sg-primary)
      - "42999:42999"    # replication from replicas (sg-primary)
    volumes:
      - /srv/curiosity/data:/data
    env_file: /etc/curiosity/secrets.env
    environment:
      MSK_GRAPH_STORAGE: /data/curiosity
      MSK_PUBLIC_ADDRESS: https://workspace.example.com

Replica replica-1, /srv/curiosity/compose.yml:

services:
  curiosity:
    image: curiosityai/curiosity:<same version as the primary>
    restart: unless-stopped
    ports:
      - "8080:8080"      # the ALB (sg-replica)
    volumes:
      - /srv/curiosity/data:/data
    env_file: /etc/curiosity/secrets.env
    environment:
      MSK_REPLICA: "true"
      MSK_PRIMARY_ADDRESS: http://primary.workspace.internal:8080
      MSK_PUBLIC_ADDRESS: https://replica-1.workspace.example.com
      MSK_SERVER_ADDRESS: http://${PRIVATE_IP}:8080           # required, must not be empty
      MSK_CORS: https://workspace.example.com
      MSK_GRAPH_STORAGE: /data/curiosity

Set PRIVATE_IP from the instance metadata before starting Compose:

TOKEN=$(curl -s -X PUT http://169.254.169.254/latest/api/token -H "X-aws-ec2-metadata-token-ttl-seconds: 60")
export PRIVATE_IP=$(curl -s -H "X-aws-ec2-metadata-token: $TOKEN" http://169.254.169.254/latest/meta-data/local-ipv4)
docker compose -f /srv/curiosity/compose.yml up -d

The security groups are what restrict the ports. Docker publishes them on all interfaces, the instances have no public address, and only the ALB (and, on the primary, the replicas) can reach them.

When a node's process is down

While a Curiosity process is starting or shutting down, it answers 503 itself: /api/available-status fails the health check, and open browsers switch reads to the replicas when the primary does it. When the process is not running at all, the ALB answers 502 or 504 (when every target in a group is unhealthy, the ALB still sends requests to them). Browsers that already have the app open then show errors instead of switching to read-only mode, and one that was using a stopped replica keeps getting errors on read-only-compatible requests until it reloads. A replica URL still serves the application read-only while the primary is down.

Verify

From a replica, through Session Manager:

curl -s http://primary.workspace.internal:8080/api/graph/ready     # true
nc -zv primary.workspace.internal 42999

From anywhere:

curl -s https://replica-1.workspace.example.com/api/graph/ready    # true
curl -s https://workspace.example.com/ | grep -o "MSKREPLICAURLS = '[^']*'"
curl -s https://workspace.example.com/health/replica
curl -si -X OPTIONS https://replica-1.workspace.example.com/api/graph/ready \
  -H "Origin: https://workspace.example.com" \
  -H "Access-Control-Request-Method: POST" | grep -i '^access-control-allow-origin'

Then sign in at https://workspace.example.com and open Manage → Developer tools to see the replicas as the browser sees them.

Monitoring

Signal Source Alarm when
Node health UnHealthyHostCount per target group (CloudWatch, AWS/ApplicationELB) ≥ 1 for any group. Each group holds one node, so this is per-node health.
Replication lag GET /health/replica on the primary, emaLag and replicas[] emaLag above your tolerance, or a replica missing from replicas[]. Scrape it with a CloudWatch Synthetics canary or the CloudWatch agent and publish a custom metric.
Primary disk CloudWatch agent disk metrics on the data volume; RocksDB_ObsoleteSstFilesSize in the workspace's monitoring Free space under 25 %.
Replica restarts Container restart count, or the replica's log line Registering with primary More than one per day.

See Operate replicas for the health endpoints and what the lag numbers mean.

Cost notes

  • Inter-AZ data transfer. A replica in another AZ receives every write the primary makes, plus the whole database on its first start and on every start with empty storage. That traffic is billed as inter-AZ transfer in both directions. It is the price of surviving an AZ outage; a replica in the primary's AZ avoids it and does not survive one.
  • Instance-store replicas avoid EBS cost, but every stop and start copies the database again.
  • One ALB serves all nodes. Each extra replica adds one target group and one listener rule, not another load balancer.

Failover on AWS

The general procedure is in Failover. On AWS, promoting replica-1 to primary also means:

  1. Stop the old primary instance, if it is still running, so it cannot accept writes.
  2. Restart replica-1's container with the primary's Compose file: without MSK_REPLICA and MSK_PRIMARY_ADDRESS, with MSK_PUBLIC_ADDRESS=https://workspace.example.com, and publishing 42999.
  3. Attach sg-primary to the instance.
  4. Register the instance in tg-primary and deregister it from tg-replica-1.
  5. Point primary.workspace.internal at its private IP.
  6. Restart the remaining replicas with empty storage, so they copy the database from the new primary. Their existing files were written by themselves, and can share names with different files on the new primary.
  7. Re-create a replica for the replica-1 name, or remove its listener rule and DNS record.

Rehearse this before you need it. Snapshots of the old primary's EBS volume are your fallback if the promoted replica turns out to be missing writes it had not received yet.

On EKS

The same rules apply to a Kubernetes deployment with the AWS Load Balancer Controller:

  • A separate StatefulSet (or Deployment with one replica and its own volume) per node, so each node has its own pod and PVC.
  • A ClusterIP Service for the primary exposing 8080 and 42999. Replicas use MSK_PRIMARY_ADDRESS=http://curiosity-primary.<namespace>.svc.cluster.local:8080.
  • One Service and one Ingress host rule per node, grouped onto one ALB with alb.ingress.kubernetes.io/group.name. Never a Service that selects more than one node.
  • A NetworkPolicy that admits 42999 on the primary only from the replica pods.

See Kubernetes for the base manifests.

© 2026 Curiosity. All rights reserved.
Powered by Neko