Install Iskue as a distributed cluster

This page covers Iskue's distributed-cluster deployment: a self-contained Docker Compose stack that runs the application on a Citus distributed Postgres cluster behind two load-balanced application replicas. Everything lives under deploy/ in the repository and does not touch the repo's single-node dev docker-compose.yml. Two compose files exist: deploy/docker-compose.yml (the dev/e2e cluster on port 8090) and deploy/docker-compose.prod.yml (the production variant). Run all commands from the repository root on a Linux host with Docker Engine and Compose v2 installed — the compose files resolve build contexts (../backend, ../frontend) and migration mounts relative to their own location.

Topology

ComponentService(s)Notes
Databasecitus-coordinator, citus-worker1, citus-worker2Citus 14 / PG 17 (citusdata/citus:14.0.0-pg17). Dev file publishes the coordinator on host :5432.
Worker registrationcitus-register (one-shot)Runs deploy/citus/register.sql: citus_set_coordinator_host + citus_add_node.
Migrationsflyway (one-shot)flyway/flyway:11 runs the full chain from backend/src/main/resources/db/migration once, on the coordinator. The app boots with Flyway disabled.
Shardingcitus-distribute (one-shot)Applies deploy/citus/distribute.sql (Scheme B) after migrations.
Applicationbackend1, backend2Two Spring Boot replicas (multi-stage Maven build, JDK 21), health-checked on /actuator/health.
Load balancerlbnginx serves the built SPA and round-robins /api/* across both backends. Dev host port :8090; prod binds 127.0.0.1 only.

Development cluster (:8090)

Bring up the dev cluster from the repository root
docker compose -f deploy/docker-compose.yml up -d --build

Then open http://localhost:8090. The first build is long (Maven build plus image pulls). The dev cluster backends run with APP_DEV_LOGIN_ENABLED: "true", so the login page shows a "Dev login (no password)" button that issues a JWT for a seeded dev-admin@issuehub.local admin. That flag defaults to false everywhere else, where the endpoint returns 404.

The dev cluster is for local development and e2e testing only. It has no database volumes (any container recreate destroys all data), hard-codes the database password and JWT_SECRET, publishes the coordinator on host :5432, and hands any visitor an admin JWT via dev login. Never expose it on a public interface — use deploy/docker-compose.prod.yml for that.

Reset everything, including database state
docker compose -f deploy/docker-compose.yml down -v

Bring-up order

You do not orchestrate the sequence yourself — it is enforced by depends_on conditions inside the compose file. Each one-shot has restart: "no" and gates the next stage with condition: service_completed_successfully; long-running services gate on condition: service_healthy.

Enforced startup sequence
coordinator + 2 workers (healthy)
  -> citus-register   (register workers)        [one-shot]
  -> flyway           (migrate, exactly once)   [one-shot]
  -> citus-distribute (apply sharding scheme B) [one-shot]
  -> backend1 + backend2 (healthy, Flyway disabled)
  -> lb               (serves :8090)
  • citus-register waits for all three database nodes to pass pg_isready, then runs deploy/citus/register.sql against the coordinator: citus_set_coordinator_host('citus-coordinator', 5432) followed by citus_add_node for citus-worker1 and citus-worker2. citus_add_node is idempotent — re-running returns the existing node id.
  • flyway mounts backend/src/main/resources/db/migration read-only and runs migrate against jdbc:postgresql://citus-coordinator:5432/issuehub with -connectRetries=30. Migrations run exactly once, here; the application replicas start with SPRING_FLYWAY_ENABLED: "false" so they never race each other on schema changes.
  • citus-distribute applies deploy/citus/distribute.sql with ON_ERROR_STOP — this converts the freshly migrated schema into the Scheme B layout described below.
  • backend1 and backend2 start only after distribution completes. Each carries a per-replica identity (APP_INSTANCE_ID, APP_NODE_NAME) on top of a shared environment block, shares one attachments volume mounted at /data/attachments (STORAGE_DIR), and is health-checked with wget -qO- http://localhost:8080/actuator/health | grep -q UP.
  • lb starts last, once both backends are healthy. Its image builds the SPA with VITE_API_URL="" (same-origin API) and serves it via nginx, round-robining /api/* across backend1:8080 and backend2:8080 with SSE-friendly settings (proxy_buffering off, proxy_read_timeout 1h).

Production cluster (docker-compose.prod.yml)

deploy/docker-compose.prod.yml is derived from the dev file with the same topology and bring-up order. The differences all matter on a public host:

  • Postgres data lives in named volumes (coordinator-data, worker1-data, worker2-data). The dev file has none, so any container recreate would silently destroy the whole database.
  • Nothing is published on the public interface. The load balancer binds 127.0.0.1:${LB_PORT:-8092}:80 and is reached through a host reverse proxy; the Citus coordinator is not published at all.
  • Dev login is off (APP_DEV_LOGIN_ENABLED: "false").
  • Secrets come from deploy/.env (git-ignored, mode 0600) instead of being hard-coded. The compose file fails fast if a required variable is missing (${DB_PASSWORD:?set DB_PASSWORD in deploy/.env}).
  • Rate limits are the production defaults, and APP_RATELIMIT_TRUST_FORWARDED_FOR: "true" is set so the limiter keys buckets on real client IPs rather than the proxy's address (the client entry is taken APP_RATELIMIT_TRUSTED_PROXY_HOPS positions — default 1 — from the right of X-Forwarded-For, so fabricated entries a client prepends cannot displace it).
  • Self-registration ships closed (APP_REGISTRATION_ENABLED: "false"). To bootstrap a fresh install, temporarily set it to "true", register the first account (it becomes the global ADMIN), then set it back to "false" and re-run the compose.
Create deploy/.env before first bring-up
install -m 600 /dev/null deploy/.env
cat > deploy/.env <<'EOF'
DB_PASSWORD=<strong-random-password>
JWT_SECRET=<long-random-secret>
PUBLIC_ORIGIN=https://issues.example.com
# LB_PORT=8092   # optional; host loopback port for the LB
EOF
VariableRequiredUsed for
DB_PASSWORDYesPOSTGRES_PASSWORD on all Citus nodes, the Flyway one-shot, and the backends' DB_PASSWORD.
JWT_SECRETYesSigning key for the backends' stateless JWTs.
PUBLIC_ORIGINYesCORS_ALLOWED_ORIGINS, PORTAL_BASE_URL, OAUTH2_SUCCESS_REDIRECT (${PUBLIC_ORIGIN}/oauth2/callback) and OAUTH2_FAILURE_REDIRECT (${PUBLIC_ORIGIN}/login?error=sso).
LB_PORTNo (default 8092)Host loopback port the nginx LB binds to: 127.0.0.1:${LB_PORT:-8092}:80.
Bring up the production cluster
docker compose -f deploy/docker-compose.prod.yml up -d --build

Front the loopback-bound LB with a TLS reverse proxy on the host. deploy/apache-iskue.conf.example is a working Apache vhost to adapt: it proxies / to http://127.0.0.1:8092/, sets ProxyPreserveHost On (the upstream name issuehub_backend contains an underscore that Tomcat would reject as a Host header), sets X-Forwarded-Proto "https" explicitly (the backend relies on it to build correct https URLs for OAuth redirects and portal links), and uses ProxyTimeout 3600 because Server-Sent Events connections stay open for a long time.

The first account created becomes a global ADMIN. The dev cluster leaves self-registration open; the production file ships it closed (APP_REGISTRATION_ENABLED: "false"), so bootstrap by enabling it temporarily, registering the admin, and closing it again. Either way, do not make the cluster world-reachable before the admin exists — the example vhost restricts access with Require ip lines for exactly this reason.

Intra-cluster Postgres auth is POSTGRES_HOST_AUTH_METHOD: trust in both files: Citus coordinator-to-worker connections are made without secrets. This is acceptable only because the database ports are never published outside the compose network in production.

What Scheme B sharding means operationally

deploy/citus/distribute.sql converts the plain migrated schema into the cluster layout without changing any application code or columns. The scheme: ticket is distributed by id; ticket-child tables (ticket_comment, ticket_label, ticket_change, ticket_watcher, attachment, git_link, csat_rating, ticket_sla and the rest) are distributed by ticket_id and colocated with ticket (issue links by source_ticket_id); every other config/global table is a reference table, replicated in full to every node.

  • Child primary keys that were (id) become composite (ticket_id, id) — Citus requires the distribution column in every PK and UNIQUE constraint. The columns are already populated by the app.
  • The ticket self-FK (parent_id → id) cannot hold across shards and is dropped permanently; parent_id remains as a plain column.
  • The per-project UNIQUE (project_id, number) constraint cannot include the distribution column and is dropped — ticket-number uniqueness is now enforced by the application only (race-safe sequential keys via pessimistic lock).
  • Child ticket_id → ticket(id) FKs are dropped before distribution and re-created afterwards as valid colocated distributed FKs with ON DELETE CASCADE.
  • The script sets citus.multi_shard_modify_mode TO 'sequential' at both database and role level (ALTER DATABASE issuehub / ALTER ROLE issuehub). This is required: ticket creation combines a parallel multi-shard operation on ticket with a write to the project reference table in one transaction, which the default parallel mode aborts.
  • The full-text helper function ticket_fts is created on the workers via run_command_on_workers, so FTS predicates can be pushed down to the ticket shards.

Rule for developers: every new table needs a distribute.sql entry

Migrations are append-only in backend/src/main/resources/db/migration, but on the cluster a migration alone leaves a new table as a plain local table on the coordinator. Every new table needs a deliberate entry in deploy/citus/distribute.sql — decide whether it is a reference table or distributed-by-ticket, and add the statement so a fresh cluster bring-up produces the correct layout.

  • Config/global table (the common case): add SELECT create_reference_table('your_table'); in FK-dependency order — a table must be converted after every table it references.
  • Ticket-child table: give it a ticket_id column, ensure the PK includes it, add SELECT create_distributed_table('your_table', 'ticket_id', colocate_with => 'ticket');, and re-add its ticket_id → ticket(id) FK in the post-distribution section of the script.
  • FKs between two distributed tables are only valid on the distribution column with colocation; FKs from a distributed table to a reference table are fine; an FK from a reference table to a distributed table is not — model such links as app-validated columns (as goal_ticket.ticket_id does).
  • FKs between two reference tables must be added after both sides are reference tables; the script's post-distribution section (e.g. project_permission_scheme_fk) shows the pattern.

On ticket-child JPA entities the ticket_id column must be @Column(updatable = false) — otherwise the first UPDATE fails on the distribution column. The bug is latent in insert-only entities and only the cluster e2e suite catches it.

Rate-limit headroom

Iskue runs a bucket4j per-IP/per-user rate limiter (per-IP on the auth, portal and webhook surfaces; per authenticated user — or per-IP when anonymous — on the rest of the API), configurable via app.ratelimit.* properties (bound to APP_RATELIMIT_* environment variables). The production defaults are tuned per real user, but the cluster Playwright suite funnels ~50 specs through one dev-admin user and one client IP, which trips 429s mid-suite. The dev compose therefore keeps abuse protection on but raises the limits; the production file leaves the defaults in place.

Environment variablePropertyDefaultDev clusterProd cluster
APP_RATELIMIT_AUTH_PER_MINapp.ratelimit.auth-per-min10200default
APP_RATELIMIT_API_PER_MINapp.ratelimit.api-per-min2402000default
APP_RATELIMIT_PORTAL_PER_MINapp.ratelimit.portal-per-min60defaultdefault
APP_RATELIMIT_WEBHOOK_PER_MINapp.ratelimit.webhook-per-min120defaultdefault
APP_RATELIMIT_TRUST_FORWARDED_FORapp.ratelimit.trust-forwarded-forfalsedefault"true"

Enable APP_RATELIMIT_TRUST_FORWARDED_FOR only when the backends are unreachable except through a trusted proxy chain that sets X-Forwarded-For (in this deployment: host proxy → nginx lb → backend). Trusting the header on a directly reachable backend lets clients spoof their IP and bypass the limiter.

Verify the cluster is up

All services: coordinator, workers, backend1, backend2 and lb healthy; citus-register, flyway and citus-distribute exited with code 0
docker compose -f deploy/docker-compose.yml ps -a
Application health and load balancing (use http://127.0.0.1:8092 on the prod cluster)
# Health through the load balancer (expect {"status":"UP"}):
curl -s http://localhost:8090/actuator/health

# Round-robin proof: X-Served-By names the replica that answered.
# Repeated calls should alternate between the two backend addresses.
curl -sI http://localhost:8090/actuator/health | grep -i x-served-by
curl -sI http://localhost:8090/actuator/health | grep -i x-served-by
Citus-level checks (container names are identical in both compose files)
# Two active worker nodes:
docker exec issuehub-cl-coordinator psql -U issuehub -d issuehub \
  -c "SELECT * FROM citus_get_active_worker_nodes() ORDER BY 1, 2;"

# Every table classified as distributed or reference, with its distribution column:
docker exec issuehub-cl-coordinator psql -U issuehub -d issuehub \
  -c "SELECT table_name, citus_table_type, distribution_column, colocation_id FROM citus_tables ORDER BY citus_table_type, table_name;"

# Sequential multi-shard modify mode picked up by new connections:
docker exec issuehub-cl-coordinator psql -U issuehub -d issuehub \
  -c "SHOW citus.multi_shard_modify_mode;"

Expect citus_get_active_worker_nodes() to return citus-worker1 and citus-worker2; citus_tables to list ticket and its child tables as distributed (children sharing ticket's colocation_id) and everything else as reference; and the modify mode to be sequential. If a one-shot failed, inspect it directly with docker compose -f deploy/docker-compose.yml logs citus-register flyway citus-distribute. Finally, open the site (http://localhost:8090 on dev, your public origin on prod) and sign in — on the dev cluster, via the Dev login button.

docker compose -f deploy/docker-compose.prod.yml down -v destroys the named volumes coordinator-data, worker1-data, worker2-data and attachments — the entire database and all uploaded files. On production, use down without -v unless you intend a full reset.