Troubleshooting

This page is a symptom → cause → fix guide for operators. It covers all compose shapes: the dev stack in docker-compose.yml (frontend on host port 8088), the cluster stack in deploy/docker-compose.yml (nginx load balancer on 8090), the production variant deploy/docker-compose.prod.yml (LB on loopback 8092, behind Apache TLS) and the sandbox variant (127.0.0.1:8093).

Where to look first

  • Container logs — startup failures, stack traces and the fail-fast guard messages all land here.
  • /actuator/health — public (no token required) and proxied by the shipped nginx configs, so it works through the front door. The backend containers' Docker healthchecks poll it too.
  • flyway_schema_history — the migration ledger in the database. Any row with success = false explains a boot or upgrade failure.
The three evidence sources
# Logs
docker compose logs -f backend                                        # dev stack
docker compose -f deploy/docker-compose.yml logs -f backend1 backend2 # cluster

# Healthcheck state as Docker sees it (healthy / unhealthy / starting)
docker compose -f deploy/docker-compose.yml ps

# Application health (public endpoint)
curl -s http://localhost:8088/actuator/health   # dev full stack
curl -s http://localhost:8090/actuator/health   # cluster LB

# Migration history (dev stack: container issuehub-postgres instead)
docker exec -it issuehub-cl-coordinator psql -U issuehub -d issuehub -c \
  "SELECT installed_rank, version, description, success, installed_on
     FROM flyway_schema_history ORDER BY installed_rank DESC LIMIT 10;"

On the cluster, responses to /api/* and /actuator/health carry an X-Served-By header naming the replica that answered — useful for telling whether one backend or both are misbehaving. The lb service only starts once backend1 and backend2 are healthy, so "the LB never came up" almost always means a backend is stuck.

Compose refuses to start: the JWT_SECRET guard

The dev docker-compose.yml declares JWT_SECRET: ${JWT_SECRET:?Set JWT_SECRET (32+ chars) in .env}. The :? interpolation guard makes Compose abort while parsing the file whenever the variable is unset or empty — Compose interpolates before it looks at profiles or the requested services, so the error appears even for docker compose up -d postgres.

# The error
required variable JWT_SECRET is missing a value: Set JWT_SECRET (32+ chars) in .env

# The fix: create .env next to docker-compose.yml
cat > .env <<EOF
JWT_SECRET=$(openssl rand -base64 48)
EOF

The other files guard their own variables the same way: deploy/docker-compose.prod.yml requires DB_PASSWORD, JWT_SECRET and PUBLIC_ORIGIN in deploy/.env, and the sandbox file requires them in deploy/.env.sandbox. Compose reads the .env file from the directory containing the compose file. The cluster dev file (deploy/docker-compose.yml) hard-codes a throwaway secret and needs no .env at all.

Setting the variable is not enough if you set it to the committed placeholder. The backend fails fast at startup with an IllegalStateException — "JWT_SECRET is unset — the app is using the committed dev-only secret" — whenever the secret equals the shipped default dev-only-secret-change-me-0123456789-0123456789 and dev login is off. A backend container in a restart loop with that message means the real secret never reached the container.

401 and 403 responses

The API is stateless: every protected request needs Authorization: Bearer <JWT>. Requests that fail authentication receive a body-less 401 from the security entry point; 403 means the request was understood but refused.

StatusCauseFix
401Missing, malformed or expired token. Session JWTs are valid for JWT_EXPIRATION_MINUTES (default 480, i.e. 8 hours); anything that fails verification is discarded and the request is treated as anonymous.Sign in again. Raise JWT_EXPIRATION_MINUTES if 8 hours is too short for your users.
401 for a user who just signed inThe account was deactivated — tokens are re-validated against the live user row on every request, so deactivation and role changes take effect immediately.Reactivate the account via admin user management.
401 for everyone after a redeployJWT_SECRET changed, so previously issued tokens no longer verify.Expected — users sign in again. Keep the secret stable across restarts.
401 between the password and 2FA stepsThe intermediate MFA challenge token (valid 5 minutes) carries no role claim and cannot authenticate API calls.Complete POST /api/v1/auth/login/mfa to obtain the real session token.
403 from POST /api/v1/auth/registerAPP_REGISTRATION_ENABLED=false. The body says "Self-service registration is disabled on this instance". This is deliberate and unconditional — there is no bootstrap escape hatch.To onboard someone, set the flag to true briefly, then back to false. There is currently no invite or password-reset flow.
403 with an empty body on login, while the page loads fineCORS rejection — the serving origin is missing from CORS_ALLOWED_ORIGINS. Browsers send an Origin header on same-origin POSTs too.See the CORS section below.
403 on /api/v1/admin/**The account lacks the ADMIN role.Use an admin account, or have one grant the role.

Dev login returns 404

POST /api/v1/auth/dev-login returning 404 is by design: the whole controller is conditional on app.dev-login.enabled (@ConditionalOnProperty), so with the flag off the route simply does not exist — it is not a routing problem. Enable it with APP_DEV_LOGIN_ENABLED=true only in throwaway environments: it issues an ADMIN JWT for the seeded dev-admin@issuehub.local with no credentials. The cluster dev file ships it on; the default everywhere else is off. When enabled, GET on the same path returns {"enabled": true} so the SPA can show the button.

HTTP 429: rate limits

RateLimitFilter enforces in-memory per-minute buckets as defence-in-depth. Over-limit requests receive status 429 with body {"status":429,"message":"Too many requests — slow down"}. The buckets, their keys and their defaults:

Path prefixBucket keyDefault / minuteEnvironment override
/api/v1/auth/**client IP10APP_RATELIMIT_AUTH_PER_MIN
/api/v1/portal/**client IP60APP_RATELIMIT_PORTAL_PER_MIN
/api/v1/email/webhook/**, /api/v1/git/webhook/**client IP120APP_RATELIMIT_WEBHOOK_PER_MIN
everything elseuser id (client IP when anonymous)240APP_RATELIMIT_API_PER_MIN

Behind a reverse proxy, the classic failure is spurious 429s for everyone. The backend's socket peer is always the proxy, so all clients share one IP bucket and the tight per-IP caps on /api/v1/auth/** trip almost immediately. The fix is two settings: APP_RATELIMIT_TRUST_FORWARDED_FOR=true (default false) makes the filter read X-Forwarded-For, and APP_RATELIMIT_TRUSTED_PROXY_HOPS (default 1) says how many entries from the right of that header to skip — the leftmost entries are client-forgeable and are never trusted. The default of 1 is correct both for the single-nginx shapes and for the Apache → nginx production chain (Apache records the client, nginx appends Apache's address; skipping one hop from the right lands on the real client). The production compose file sets APP_RATELIMIT_TRUST_FORWARDED_FOR: "true" already.

Only enable APP_RATELIMIT_TRUST_FORWARDED_FOR when the backends are unreachable except through your proxy chain. If a client can reach a backend directly, it can forge X-Forwarded-For, mint a fresh bucket per request and defeat every cap.

Buckets are in-memory and per backend instance: on the cluster each replica counts separately (so the effective limit behind round-robin is up to twice the configured value), and a restart resets all counters. SSE requests (paths ending /live) are exempt — a stream is one long-lived request, not many. The cluster dev file ships APP_RATELIMIT_AUTH_PER_MIN: "200" and APP_RATELIMIT_API_PER_MIN: "2000" to absorb e2e traffic.

Live (SSE) updates not arriving behind a reverse proxy

Live board and ticket updates stream over Server-Sent Events from GET /api/v1/projects/{projectKey}/live; the browser's EventSource cannot set headers, so it passes the JWT as an access_token query parameter. An SSE stream holds one HTTP response open for a long time, so every proxy on the path must not buffer the response and must not time out the idle read. The shipped nginx images already do both — frontend/nginx.conf and deploy/lb/lb.conf carry identical directives:

deploy/lb/lb.conf (excerpt) — frontend/nginx.conf sets the same directives
location /api/ {
    proxy_pass http://issuehub_backend;
    ...
    # SSE: disable buffering, long timeout
    proxy_buffering off;
    proxy_read_timeout 1h;
    proxy_http_version 1.1;
    proxy_set_header Connection "";
}

So when live updates fail, the culprit is almost always a proxy you added in front. Symptoms: the stream connects but silently drops after the proxy's idle timeout (updates stop until a page refresh), or events arrive late in bursts (response buffering). For Apache, the repo's example vhost sets ProxyTimeout 3600 to match nginx's one-hour read timeout, and forwards the headers the backend needs (forward-headers-strategy: framework relies on X-Forwarded-Proto to build correct https URLs):

deploy/apache-iskue.conf.example (excerpt)
ProxyPreserveHost On
RequestHeader set X-Forwarded-Proto "https"

# Server-Sent Events (boards, tickets) hold connections open for a long time.
ProxyTimeout 3600

ProxyPass        / http://127.0.0.1:8092/
ProxyPassReverse / http://127.0.0.1:8092/

Swagger UI returns 404

/swagger-ui.html, /swagger-ui/** and /v3/api-docs/** are gated by a single flag, APP_API_DOCS_ENABLED (default true). When it is false the routes 404 outright rather than merely requiring auth — deliberately, because together they publish a complete machine-readable map of every endpoint. The production compose file ships APP_API_DOCS_ENABLED: "false"; the dev and cluster files leave it on. To restore Swagger on a non-public instance, set the variable to true and recreate the backend containers — do not re-enable it on an internet-facing host.

CORS errors

Allowed origins come from CORS_ALLOWED_ORIGINS (property app.cors.allowed-origins) as a comma-separated list of exact origins — scheme, host and port must all match, no paths, no trailing slash. Defaults per shape: http://localhost:5173 (bare backend), http://localhost:8088 (dev full stack), http://localhost:8090 (cluster). Two symptom shapes: the obvious one is a CORS error in the browser console when the SPA is served from an origin the backend does not list; the subtle one is sign-in failing with a bare 403 while the page loads fine — browsers send an Origin header on same-origin POSTs too, so the list must contain *every* origin your front door answers to, including bare-domain aliases and the raw IP, not just the canonical hostname.

The prod compose maps CORS_ALLOWED_ORIGINS from CORS_ORIGINS, falling back to PUBLIC_ORIGIN
# deploy/.env — list every public origin, comma-separated
PUBLIC_ORIGIN=https://www.example.com
CORS_ORIGINS=https://www.example.com,https://example.com,https://203.0.113.10

Port conflicts

Compose fileHost ports published
docker-compose.yml (dev)5432 → Postgres; 8088 → frontend (profile full)
deploy/docker-compose.yml (cluster)5432 → Citus coordinator; 8090 → nginx LB
deploy/docker-compose.prod.yml127.0.0.1:8092 → LB (override with LB_PORT); database not published
deploy/docker-compose.sandbox.yml127.0.0.1:8093 → LB (override with LB_PORT); database not published

Both the dev stack and the cluster stack claim host port 5432, so they cannot run simultaneously on one machine — and either collides with a Postgres already running on the host. Docker reports Bind for 0.0.0.0:5432 failed: port is already allocated. Fix: stop the other stack (docker compose down), stop the host service, or change the *host* side of the mapping (for example "15432:5432"); container-side ports must stay as shipped. The production and sandbox files avoid the problem by not publishing the database at all and binding the LB to loopback.

Slow first build

Expected. The backend image is a multi-stage Maven build (maven:3.9-eclipse-temurin-21) that runs mvn -q dependency:go-offline and then mvn -q package -DskipTests — the first build downloads the entire dependency tree; the frontend runs npm ci plus a Vite production build; and the base images are pulled once. The deploy README says it plainly: the first build is long (Maven + image pulls). Subsequent builds are fast: pom.xml and package*.json are copied before the source, so Docker's layer cache re-downloads dependencies only when those files change. On the cluster, backend1 and backend2 share one image (issuehub-backend:cluster), so the backend is compiled once, not twice.

Flyway failures

Migrations run in one of two modes. Dev shapes: Spring runs Flyway inside the backend at startup (spring.flyway.enabled: true is the default), so a failed migration means the backend exits and restart-loops — the SQL error is in docker compose logs backend. Cluster shapes: a one-shot flyway container (flyway/flyway:11) migrates the coordinator exactly once; the backends boot with SPRING_FLYWAY_ENABLED: "false" and Hibernate merely validates the schema (ddl-auto: validate). The bring-up order is enforced by depends_on conditions — citus-register → flyway → citus-distribute → backends → lb — so when the flyway service fails, everything downstream never starts. Check docker compose -f deploy/docker-compose.yml logs flyway first.

A migration that fails mid-run leaves a row with success = false in flyway_schema_history, and every later attempt refuses with a "detected failed migration" error. Inspect the table with the query under *Where to look first*, fix the underlying cause, then either delete the failed row or run Flyway's repair command with the same container image, and bring the stack up again to re-run the one-shots (migrate is a no-op for versions already applied). A backend that boots but fails Hibernate validation ("Schema-validation: missing table …") means migrations did not run at all — verify SPRING_FLYWAY_ENABLED and the history table before touching anything else.

In the cluster dev file the Citus nodes have no data volumes — database state lives in the container filesystem and is destroyed by down -v or any container recreate. Production uses named volumes (coordinator-data, worker1-data, worker2-data) precisely so recreates are safe. Never point a production host at the dev cluster file: the dev and prod files share a Compose project name and every container_name, so running the wrong one hijacks the live containers.