Upgrading Iskue

An Iskue upgrade has two parts: application images rebuilt from the updated source, and database migrations that bring the schema up to the level the new code expects. The schema is managed by Flyway as an append-only chain of migrations, so upgrading is always a forward move — take a backup before you start. The mechanics differ between the two supported deployment shapes:

DeploymentCompose fileEntry pointDatabase containerMigrations applied by
Single hostdocker-compose.yml (profile full)http://localhost:8088issuehub-postgresthe backend container, on startup
Cluster (dev/e2e)deploy/docker-compose.ymlhttp://localhost:8090issuehub-cl-coordinatorone-shot flyway container at bring-up
Cluster (production)deploy/docker-compose.prod.yml127.0.0.1:8092 (default LB_PORT), behind a host TLS proxyissuehub-cl-coordinatorone-shot flyway container at bring-up

How schema migrations are applied

Migrations live in backend/src/main/resources/db/migration — currently 115 files, V1__identity.sql through V115__git_smart_commits_opt_in.sql, named V<n>__<description>.sql. The chain is append-only: a release only ever adds new files on top, never edits existing ones. Flyway records each applied file in the flyway_schema_history table and, on any subsequent run, applies exactly the files your database has not seen yet.

  • Single host — the backend service has Flyway enabled (the application config sets spring.flyway.enabled: true and the root docker-compose.yml does not override it). The container applies any pending migrations automatically while starting, before it serves requests.
  • Cluster — both replicas run with SPRING_FLYWAY_ENABLED: "false". Migrations are applied once per bring-up by the one-shot flyway service (image flyway/flyway:11), which mounts backend/src/main/resources/db/migration read-only and runs migrate against the Citus coordinator. A second one-shot, citus-distribute, applies the sharding layout afterwards, and backend1/backend2 only start once both one-shots have completed successfully.

Never edit a migration that has already been applied. Flyway stores a checksum per applied file and validates the whole chain on every run; a modified historical file fails validation and blocks the upgrade. Schema changes always arrive as a new V<n> file.

Check the current schema version

The newest successful row is the current schema version — on the current release it reports 115. Compare it with the highest V<n> file shipped in backend/src/main/resources/db/migration.
# Single host
docker exec issuehub-postgres psql -U issuehub -d issuehub \
  -c "SELECT version, description, installed_on FROM flyway_schema_history WHERE success ORDER BY installed_rank DESC LIMIT 1;"

# Cluster (either compose file)
docker exec issuehub-cl-coordinator psql -U issuehub -d issuehub \
  -c "SELECT version, description, installed_on FROM flyway_schema_history WHERE success ORDER BY installed_rank DESC LIMIT 1;"

Back up first

Two places hold state: the database, and uploaded attachments (STORAGE_DIR: /data/attachments, backed by the named volume attachments in every compose file). The migration chain is forward-only and the repository ships no undo migrations, so the pre-upgrade backup is the rollback path. Take one every time, and test the restore before you rely on it.

Single host.
# Database (logical dump)
docker exec issuehub-postgres pg_dump -U issuehub issuehub > iskue-db-$(date +%F).sql

# Attachments volume
docker run --rm --volumes-from issuehub-backend -v "$PWD":/backup alpine \
  tar czf /backup/iskue-attachments-$(date +%F).tar.gz -C /data/attachments .
Cluster, online: a logical dump taken on the coordinator (rows of the distributed tables are fetched through it from the workers).
docker exec issuehub-cl-coordinator pg_dump -U issuehub issuehub > iskue-db-$(date +%F).sql
Cluster, cold: archives the production data volumes (coordinator-data, worker1-data, worker2-data) plus attachments. Resume with start, not up — up would re-run the one-shot containers.
docker compose -f deploy/docker-compose.prod.yml stop lb backend1 backend2 \
  citus-coordinator citus-worker1 citus-worker2

for c in coordinator worker1 worker2; do
  docker run --rm --volumes-from issuehub-cl-$c -v "$PWD":/backup alpine \
    tar czf /backup/iskue-$c-data-$(date +%F).tar.gz -C /var/lib/postgresql/data .
done
docker run --rm --volumes-from issuehub-cl-backend1 -v "$PWD":/backup alpine \
  tar czf /backup/iskue-attachments-$(date +%F).tar.gz -C /data/attachments .

docker compose -f deploy/docker-compose.prod.yml start citus-coordinator citus-worker1 citus-worker2
docker compose -f deploy/docker-compose.prod.yml start backend1 backend2 lb

Upgrade a single host

Run from the repository root, with the same .env you installed with — Compose reads it for JWT_SECRET and the other required variables. Fetch the release, rebuild, then restart:

git pull

# Build first, so the restart window is only the container swap + boot,
# not the whole image build.
docker compose --profile full build
docker compose --profile full up -d

A single host runs one backend replica, so this is not a zero-downtime upgrade: the API is unavailable from the moment the old container stops until the new one has booted and applied its pending migrations — typically well under a minute with pre-built images, longer when a release ships heavy migrations. The postgres container and its pgdata volume are untouched, so data survives the restart. The one-liner docker compose --profile full up -d --build works too, but adds the build time to the outage window.

Verify: health reports UP and the log shows the newly applied versions.
# Health (the backend port is not published on the host, so exec into the container)
docker exec issuehub-backend wget -qO- http://localhost:8080/actuator/health

# What Flyway did during boot
docker logs issuehub-backend 2>&1 | grep -i -E "flyway|migrat" | tail -n 10

Upgrade a production cluster

First bring-up runs the compose file's one-shot chain in order: coordinator and workers healthy, then citus-register, then flyway, then citus-distribute, then backend1/backend2, then lb. On a cluster that already holds data, upgrade with the --no-deps flow below, which touches only the services you name. All commands assume the repository root; secrets are read from deploy/.env, exactly as at install time.

Do not re-run the whole file (docker compose -f deploy/docker-compose.prod.yml up -d with no service list) against a database that holds data. citus-register is idempotent and flyway only applies what is pending, but deploy/citus/distribute.sql is not re-runnable — create_reference_table fails on tables that are already distributed, and services that depend on citus-distribute completing successfully will then refuse to start.

1. Fetch and build

backend1 and backend2 share the image issuehub-backend:prod, so one build covers both replicas.
git pull
docker compose -f deploy/docker-compose.prod.yml build backend1 lb

2. Apply pending migrations

Re-runs the one-shot Flyway migrate against the coordinator; only migrations above the flyway_schema_history high-water mark are applied. The output prints directly to your terminal.
docker compose -f deploy/docker-compose.prod.yml run --rm --no-deps flyway

The old replicas keep serving while migrations run. Purely additive migrations coexist with the previous release, but a migration that restructures existing tables can break the running replicas until step 3 replaces them — keep the gap between steps 2 and 3 short. For a strict-safety window, stop the backends first (docker compose -f deploy/docker-compose.prod.yml stop backend1 backend2) and accept the brief outage.

3. Roll the backends, then the load balancer

--no-deps recreates only the named service, so the one-shots stay untouched; --wait returns once the container's healthcheck (GET /actuator/health) reports healthy, so backend2 only goes down after backend1 is back.
docker compose -f deploy/docker-compose.prod.yml up -d --no-deps --wait backend1
docker compose -f deploy/docker-compose.prod.yml up -d --no-deps --wait backend2
docker compose -f deploy/docker-compose.prod.yml up -d --no-deps lb

Rolled this way the API stays available: the nginx upstream in deploy/lb/lb.conf (max_fails=3 fail_timeout=10s) directs traffic to the surviving replica while the other restarts. Treat it as near-zero downtime, not zero — in-flight requests and SSE streams on the restarting replica are dropped and clients reconnect, and recreating lb itself interrupts service for the second or two nginx takes to come back.

4. Apply Citus entries for new tables

citus-distribute runs only at first bring-up, so any tables created by this release's migrations exist after step 2 as ordinary local tables on the coordinator. Every new table has a deliberate entry in deploy/citus/distribute.sql — reference table, or distributed by ticket_id and colocated with ticket — added in the same release. Check what the release changed in that file and apply just the new statements on the coordinator. Some releases also carry post-distribution statements (FKs between converted tables) that the file itself marks as applied manually on the coordinator after their Flyway one-shots; apply those the same way.

Example: the release that shipped V110__wiki_page_comments.sql also added wiki_page_comment to distribute.sql.
# See which distribute.sql statements this release added
git log -p --oneline -- deploy/citus/distribute.sql

# Apply each new statement on the coordinator, e.g. for a new reference table:
docker exec issuehub-cl-coordinator psql -U issuehub -d issuehub \
  -c "SELECT create_reference_table('wiki_page_comment');"

Dev/e2e cluster: reset instead

The dev cluster (deploy/docker-compose.yml, load balancer on :8090, dev login enabled) exists for development and the e2e suite. Its database containers have no named volumes — the data lives inside the containers, and any container recreate destroys it — so treat that data as disposable and upgrade by resetting. The one-shots then run in dependency order against a fresh database and no manual Citus step is needed.

Full reset. The first build is long (Maven + image pulls).
docker compose -f deploy/docker-compose.yml down -v
docker compose -f deploy/docker-compose.yml up -d --build

Verify the upgrade

127.0.0.1:8092 is the production default (LB_PORT). docker logs issuehub-cl-flyway shows the original bring-up migration run; step 2's output appeared directly in your terminal.
# Through the load balancer (production; dev cluster: http://localhost:8090/actuator/health)
curl -fsS http://127.0.0.1:8092/actuator/health

# Confirm both replicas answer — the LB adds X-Served-By; repeat to see both addresses
curl -si http://127.0.0.1:8092/actuator/health | grep -i x-served-by

# Schema version matches the highest V<n> file in the release
docker exec issuehub-cl-coordinator psql -U issuehub -d issuehub \
  -c "SELECT version, description, installed_on FROM flyway_schema_history WHERE success ORDER BY installed_rank DESC LIMIT 1;"

Rolling back

There are no undo migrations: rolling back means restoring state, not migrating down.

  • Keep the previous images before building: docker tag issuehub-backend:prod issuehub-backend:prod-previous and docker tag issuehub-lb:prod issuehub-lb:prod-previous.
  • To roll back a cluster, stop the stack, restore the database — replay the logical dump into a fresh cluster, or restore the cold volume archives — and restore the attachments archive if uploads must be rewound with it.
  • Start the previous code against the restored database: check out the previously deployed revision and rebuild, or retag the *-previous images back and up -d --no-deps the backends and lb.
  • On a single host, note that docker compose down -v removes both pgdata and attachments; restore the SQL dump and the attachments archive before starting the previous build.