The only copy of every customer, order and one-of-a-kind item was the live Postgres data directory, plus whatever the deploy checklist's manual pg_dump happened to have caught. That dump is good and stays, but it only runs when someone deploys: a quiet week meant the newest copy of real customer data was a week old, and nothing bounded the gap. Two services rather than one, because they are different jobs. The database is small, changes constantly and wants a logical dump — daily, gzipped, 7/4/6 daily-weekly-monthly retention. Uploads are large and append-mostly and want an archive — weekly, 56 days, mounted read-only so a backup process cannot damage the thing it is backing up. Forcing both through one tool would serve one of them badly. The dumper is pinned to postgres-backup-local:16 to match the server. pg_dump refuses to dump a server newer than itself, so a floating tag is a backup that stops working the day Postgres is upgraded — silently, because nothing reads a dump until it is needed. It depends_on the database's existing pg_isready healthcheck, which is the constraint that shaped this: a dumper is a client, and without that the first run after a NAS reboot races Postgres coming up. Both carry a staleness healthcheck rather than trusting the schedule. A regime that stopped a month ago is indistinguishable from a working one until a restore is attempted, and `find -mmin` is the cheapest thing that tells them apart. It surfaces in Portainer beside the app rather than somewhere separate to remember to look. Windows are the interval plus grace — 26 hours daily, 9 days weekly — so a late run is not a failure, and start_period covers the first cycle when nothing has been written yet. Verified rather than assumed. Both images were pulled and checked to have a shell and `find`, since a CMD-SHELL healthcheck against an image without one reports unhealthy forever. postgres-backup-local:16 ships pg_dump 16.10 against the postgres:16 server. The healthcheck expression was exercised three ways in the image itself — empty directory, fresh artifact, and one aged three days — and returns unhealthy, healthy, unhealthy. The compose file parses and the #118 drift guard still passes over it. Three things these deliberately do not cover, written into the compose file and the doc rather than left to be discovered: They run while the stack runs, so they cannot protect the stack's own teardown. Deleting the Portainer stack deletes them too. That is why the deploy checklist's manual dump stays, and README now says so where the checklist is. They write to the same volume as the data they protect. That survives a bad migration, a dropped table, a bad deploy and a stack deletion, and not the disk. Getting a copy off /volume1 is a Synology-side job and is what turns this from a convenience into a guarantee. Daily database against weekly uploads leaves a window where a restore pairs the two from different moments. An orphaned image is harmless; a row without its image is a broken thumbnail on one recent item, usually still on the admin's machine. Neither is data loss, and a synchronised snapshot is not worth the complexity to avoid it. The restore procedure leads with practising on a throwaway database, because an untested backup is a file of unknown validity and a truncated dump looks exactly like a good one until it matters. It checks row counts and the pgmigrations head — the migration check being the one most easily skipped and most likely to bite, since a dump older than the code restores a schema the app will fail against. Both open items are listed as unticked in the doc: no off-volume copy exists yet, and no restore has been performed. Until the second is done this documents an untested procedure, and it says so. Refs #147
332 lines
15 KiB
YAML
332 lines
15 KiB
YAML
# Production stack — the live storefront. See issue #118 for why this file is
|
|
# in the repository at all.
|
|
#
|
|
# It used to live only in Portainer's web editor, which meant no test could read
|
|
# it. `backend/tests/unit/composeEnvironment.test.ts` asserts that the deploying
|
|
# environment sets every name in ALWAYS_REQUIRED, and it could only ever check
|
|
# QA. Production was unguarded, and on 2026-08-23 it refused to boot because
|
|
# UPLOADS_DIR had no line here — while being set in Portainer's stack variables,
|
|
# where it does nothing. Committing the file is what lets the guard cover it.
|
|
#
|
|
# Name the Portainer stack `redefined-designs`, NOT `redefined-designs-qa`.
|
|
# The stack name becomes the compose project name, and reusing QA's would make
|
|
# compose reconcile the two against each other.
|
|
#
|
|
# DEPLOY THIS AS A GIT REPOSITORY STACK, not from the web editor — otherwise the
|
|
# file here and the file that actually runs drift apart again, which is the
|
|
# whole problem this is solving.
|
|
#
|
|
# Repository: https://gitea.bermudalamb.synology.me/bermudalamb/redefined-designs
|
|
# Reference: refs/heads/main
|
|
# Compose path: docker-compose.prod.yml
|
|
#
|
|
# THIS STACK DOES NOT BUILD. It runs the image tagged redefined-designs:latest,
|
|
# which has to already exist on the NAS before the stack starts. A first deploy,
|
|
# or a NAS that has pruned images, fails with "image not found" rather than
|
|
# quietly building one — see #146.
|
|
#
|
|
# That is deliberate. QA builds from this repository and is reviewed; production
|
|
# then runs the image that was reviewed, promoted by hand:
|
|
#
|
|
# docker tag redefined-designs:qa redefined-designs:latest
|
|
#
|
|
# Building here instead would look simpler and would ship something else. The
|
|
# Dockerfile copies package.json without package-lock.json and installs with
|
|
# `npm install`, so two builds of the same commit can resolve different
|
|
# transitive dependencies. "Same git ref" is therefore not "same image", and the
|
|
# reviewed bytes are the only thing that is.
|
|
#
|
|
# The promotion is load-bearing, not a convenience. Forgetting it means a
|
|
# redeploy that reuses the previous redefined-designs:latest and appears to
|
|
# succeed while running old code — the same failure `pull_policy: build` guards
|
|
# against in QA, arriving by a different route. State which image is being
|
|
# promoted as part of the deploy.
|
|
#
|
|
# Leave any Portainer option that re-pulls images turned OFF — there is no
|
|
# registry to pull this image from.
|
|
#
|
|
# WHY MOST VALUES ARE HARDCODED HERE RATHER THAN INTERPOLATED
|
|
#
|
|
# Portainer's stack variables are substituted into this file; they are not
|
|
# handed to the container. A variable set in Portainer with no line here never
|
|
# reaches the app, and the failure reads as "I set it and it says it is not
|
|
# set". Only secrets are interpolated below, because only secrets have a reason
|
|
# not to be in the repository. Everything else is written out, so there is one
|
|
# place to look and one thing that can be wrong.
|
|
#
|
|
# Required stack environment variables — all secrets, all must be set in
|
|
# Portainer for this stack:
|
|
#
|
|
# DB_PASSWORD Postgres password for the `redefined` database.
|
|
# SMTP_USER Brevo SMTP login.
|
|
# SMTP_PASSWORD
|
|
# SMTP_FROM The From address customers see.
|
|
# ADMIN_GATE_SECRET The shared secret Nginx Proxy Manager injects as the
|
|
# X-Admin-Gate header on the gated location. Both sides
|
|
# must hold the same value or the admin API returns 403.
|
|
# See #63. Without it, /api/admin is protected only by
|
|
# the proxy — anything reaching the container directly
|
|
# can administer the store.
|
|
# PAYPAL_CLIENT_ID Live PayPal credentials. Required because DEMO_MODE is
|
|
# PAYPAL_CLIENT_SECRET false below; the app refuses to start without them.
|
|
# PAYPAL_WEBHOOK_ID
|
|
# BACKUP_PASSPHRASE Optional. Set it and the uploads archives are
|
|
# encrypted at rest; leave it empty and they are not.
|
|
# See docs/ops/backup-and-restore.md before setting it —
|
|
# an archive nobody can decrypt is not a backup.
|
|
# USPS_CLIENT_ID Optional. Leave unset to run without address
|
|
# USPS_CLIENT_SECRET validation; the app degrades gracefully rather than
|
|
# failing, so an empty value is a working configuration.
|
|
#
|
|
# If the existing stack in Portainer uses different names for any of these,
|
|
# rename them there to match — the names above are what this file reads.
|
|
|
|
services:
|
|
redefined-designs:
|
|
# No `build:` — see the note at the top of this file. This tag is produced
|
|
# by promoting the image QA was reviewed against, not by rebuilding here.
|
|
image: redefined-designs:latest
|
|
container_name: redefined-designs-syn
|
|
environment:
|
|
- TZ=America/Chicago
|
|
- PORT=3000
|
|
|
|
# NODE_ENV is deliberately absent. The Dockerfile sets it to `production`,
|
|
# and that value gates the `secure` flag on the session cookie. Setting it
|
|
# to anything else here would silently serve session cookies over plain
|
|
# HTTP. Do not add a line for it.
|
|
|
|
- PGHOST=redefined-designs-db-syn
|
|
- PGPORT=5432
|
|
- PGUSER=redefined
|
|
- PGPASSWORD=${DB_PASSWORD}
|
|
- PGDATABASE=redefined
|
|
|
|
# Real payments. This is the difference between production and QA, and it
|
|
# is why the three PayPal secrets are required rather than optional — the
|
|
# app refuses to start without them when this is false.
|
|
#
|
|
# To bring the stack up before PayPal is configured, set this to `true`
|
|
# and the three PAYPAL_ lines can be removed. The full cart and checkout
|
|
# flow then works end to end and NOBODY IS EVER CHARGED. That is a
|
|
# deliberate interim state and a quiet disaster if it is left on.
|
|
- DEMO_MODE=false
|
|
- PAYPAL_ENV=live
|
|
- PAYPAL_CLIENT_ID=${PAYPAL_CLIENT_ID}
|
|
- PAYPAL_CLIENT_SECRET=${PAYPAL_CLIENT_SECRET}
|
|
- PAYPAL_WEBHOOK_ID=${PAYPAL_WEBHOOK_ID}
|
|
|
|
# Host, port and secure are not secrets and are pinned rather than
|
|
# inherited: the mailer's fallbacks are Gmail's (smtp.gmail.com, 465, TLS)
|
|
# and Brevo needs 587 with STARTTLS, which is why SMTP_SECURE is false.
|
|
# Getting these wrong fails at send time, not at boot.
|
|
- SMTP_HOST=smtp-relay.brevo.com
|
|
- SMTP_PORT=587
|
|
- SMTP_SECURE=false
|
|
- SMTP_USER=${SMTP_USER}
|
|
- SMTP_PASSWORD=${SMTP_PASSWORD}
|
|
- SMTP_FROM=${SMTP_FROM}
|
|
|
|
# MAIL_ALLOWLIST is deliberately absent, and this is the one environment
|
|
# where that is correct. It restricts delivery to named recipients, which
|
|
# is what keeps QA from emailing real customers. Production has to be able
|
|
# to reach real customers, so it is unrestricted here on purpose. The
|
|
# boot-time warning about it is expected and should not be silenced.
|
|
|
|
# Not a secret, and hardcoded rather than interpolated so it cannot go
|
|
# missing: every link in a verification, password-reset, favorite-alert
|
|
# and cart-reminder email is built from it, and an unset value renders
|
|
# them all as "undefined".
|
|
- PUBLIC_URL=https://redefined-designs.bermudalamb.synology.me
|
|
|
|
- SITE_CURRENCY=USD
|
|
|
|
# Must match the right-hand side of the volume mapping below. Hardcoded
|
|
# for that reason — splitting it across two places is how they drift, and
|
|
# its absence is what stopped this stack booting on 2026-08-23.
|
|
- UPLOADS_DIR=/app/uploads
|
|
|
|
# Optional. Address validation is skipped when these are empty, rather
|
|
# than failing, so an unset pair is a working configuration.
|
|
- USPS_ENV=production
|
|
- USPS_CLIENT_ID=${USPS_CLIENT_ID}
|
|
- USPS_CLIENT_SECRET=${USPS_CLIENT_SECRET}
|
|
|
|
- ADMIN_GATE_SECRET=${ADMIN_GATE_SECRET}
|
|
volumes:
|
|
# Production's own uploads directory. QA writes to
|
|
# /volume1/configs/redefined-designs-qa/uploads; sharing this one would
|
|
# let a QA teardown delete real product images.
|
|
- /volume1/configs/redefined-designs/uploads:/app/uploads
|
|
ports:
|
|
# 32750, not QA's 32751.
|
|
#
|
|
# An unqualified host port binds to 0.0.0.0, so this answers directly on
|
|
# http://<nas-ip>:32750 from anywhere on the LAN, bypassing Nginx Proxy
|
|
# Manager and its TLS. The admin gate still fails closed — a direct
|
|
# request arrives without the X-Admin-Gate header — but the storefront,
|
|
# the customer API, login and registration are all reachable in the clear.
|
|
#
|
|
# That is #117, and the fix is not a loopback binding: NPM runs as its own
|
|
# container, so 127.0.0.1 would stop the proxy reaching this at all. The
|
|
# fix is a shared external network with this block removed entirely, which
|
|
# also requires repointing the proxy host entry at the container name and
|
|
# port 3000. Left as-is here so this file matches what is deployed today;
|
|
# changing it is #117's job, not this file's.
|
|
- 32750:3000
|
|
depends_on:
|
|
redefined-designs-db-syn:
|
|
condition: service_healthy
|
|
# Unlike QA's `no`: production is meant to come back after a NAS reboot.
|
|
restart: unless-stopped
|
|
# Docker's default json-file driver has no size cap. POST /api/client-errors
|
|
# is unauthenticated, so an unrotated log is a disk-filling vector on its
|
|
# own. QA has carried these options for a while; production could not,
|
|
# because this file did not exist.
|
|
logging:
|
|
driver: json-file
|
|
options:
|
|
max-size: 10m
|
|
max-file: "3"
|
|
|
|
redefined-designs-db-syn:
|
|
image: postgres:16
|
|
container_name: redefined-designs-db-syn
|
|
environment:
|
|
- POSTGRES_USER=redefined
|
|
- POSTGRES_PASSWORD=${DB_PASSWORD}
|
|
- POSTGRES_DB=redefined
|
|
- PGDATA=/var/lib/postgresql/data/pgdata
|
|
volumes:
|
|
# THE REAL DATA. Distinct from QA's
|
|
# /volume1/configs/redefined-designs-qa/postgres. Never point a QA stack
|
|
# at this path.
|
|
- /volume1/configs/redefined-designs/postgres:/var/lib/postgresql/data
|
|
healthcheck:
|
|
test: ["CMD-SHELL", "pg_isready -U redefined -d redefined"]
|
|
interval: 10s
|
|
timeout: 5s
|
|
retries: 10
|
|
restart: unless-stopped
|
|
logging:
|
|
driver: json-file
|
|
options:
|
|
max-size: 10m
|
|
max-file: "3"
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# Backups (#147)
|
|
#
|
|
# Two services rather than one, because they are different jobs on different
|
|
# cadences. The database is small, changes constantly, and wants a logical
|
|
# dump. Uploads are large, append-mostly, and want an archive. Forcing both
|
|
# through one tool serves one of them badly.
|
|
#
|
|
# WHAT THESE DO NOT COVER, and it matters:
|
|
#
|
|
# They run while the stack runs, so they cannot protect the stack's own
|
|
# teardown. Deleting the Portainer stack deletes these containers along with
|
|
# everything else. The manual pg_dump in README's deploy steps therefore
|
|
# stays exactly where it is — a routine regime and a snapshot taken before a
|
|
# risky operation are different jobs, and neither replaces the other.
|
|
#
|
|
# And they write to the same volume as the data they protect. That survives a
|
|
# bad migration, a dropped table, a bad deploy and a stack deletion. It does
|
|
# not survive the disk. Getting a copy off /volume1 is a Synology-side job —
|
|
# Hyper Backup to another volume, an external disk, or offsite — and it is
|
|
# what turns this from a convenience into a guarantee. See
|
|
# docs/ops/backup-and-restore.md.
|
|
# ---------------------------------------------------------------------------
|
|
|
|
redefined-designs-db-backup-syn:
|
|
# Pinned to 16 to match the server. pg_dump refuses to dump a server newer
|
|
# than itself, so a floating tag here is a backup that stops working on the
|
|
# day Postgres is upgraded — silently, since nothing reads the dumps until
|
|
# they are needed.
|
|
image: prodrigestivill/postgres-backup-local:16
|
|
container_name: redefined-designs-db-backup-syn
|
|
environment:
|
|
- TZ=America/Chicago
|
|
- POSTGRES_HOST=redefined-designs-db-syn
|
|
- POSTGRES_PORT=5432
|
|
- POSTGRES_DB=redefined
|
|
- POSTGRES_USER=redefined
|
|
- POSTGRES_PASSWORD=${DB_PASSWORD}
|
|
# Daily at 03:00. Late enough that a deploy is unlikely to be in flight,
|
|
# and pg_dump takes a consistent snapshot anyway, so a dump running while
|
|
# customers are shopping is fine.
|
|
- SCHEDULE=@daily
|
|
- BACKUP_KEEP_DAYS=7
|
|
- BACKUP_KEEP_WEEKS=4
|
|
- BACKUP_KEEP_MONTHS=6
|
|
# --clean --if-exists so the dump can be restored over an existing
|
|
# database without hand-dropping it first, which is the state a real
|
|
# restore happens in.
|
|
- POSTGRES_EXTRA_OPTS=--clean --if-exists
|
|
volumes:
|
|
- /volume1/configs/redefined-designs/backups/postgres:/backups
|
|
depends_on:
|
|
redefined-designs-db-syn:
|
|
# The dumper is a client and needs a server accepting connections. This
|
|
# is the constraint that shaped the design: without it the first run
|
|
# after a NAS reboot races Postgres coming up.
|
|
condition: service_healthy
|
|
healthcheck:
|
|
# Unhealthy when nothing has been written inside the window. A backup
|
|
# regime that stopped a month ago is indistinguishable from a working one
|
|
# until a restore is attempted, and this is the cheapest thing that tells
|
|
# them apart. It shows in Portainer beside the app rather than somewhere
|
|
# separate to remember to look.
|
|
#
|
|
# 1560 minutes is 26 hours: the daily interval plus two hours of grace, so
|
|
# a dump that runs a little late is not reported as a failure.
|
|
test: ["CMD-SHELL", "find /backups -name '*.sql.gz' -mmin -1560 | grep -q ."]
|
|
interval: 1h
|
|
timeout: 30s
|
|
retries: 3
|
|
# Nothing exists until the first scheduled run, so without this the
|
|
# container reports unhealthy for its first day on every fresh deploy.
|
|
start_period: 25h
|
|
restart: unless-stopped
|
|
logging:
|
|
driver: json-file
|
|
options:
|
|
max-size: 10m
|
|
max-file: "3"
|
|
|
|
redefined-designs-uploads-backup-syn:
|
|
image: offen/docker-volume-backup:v2
|
|
container_name: redefined-designs-uploads-backup-syn
|
|
environment:
|
|
- TZ=America/Chicago
|
|
# Weekly, not daily. Uploads are append-mostly and much larger than the
|
|
# database, so a daily full archive would mostly be copies of itself.
|
|
- BACKUP_CRON_EXPRESSION=0 4 * * 0
|
|
- BACKUP_FILENAME=uploads-%Y-%m-%dT%H-%M-%S.tar.gz
|
|
- BACKUP_ARCHIVE=/archive
|
|
# Eight weeks. Shorter than the database's tail because each archive is
|
|
# far bigger, and an image that was deleted two months ago is not
|
|
# something anyone is restoring.
|
|
- BACKUP_RETENTION_DAYS=56
|
|
- BACKUP_PRUNING_PREFIX=uploads-
|
|
# Optional. An empty value means no encryption, which is the default.
|
|
- GPG_PASSPHRASE=${BACKUP_PASSPHRASE}
|
|
volumes:
|
|
# Read-only. A backup process with write access to the thing it is backing
|
|
# up is a way to lose both at once.
|
|
- /volume1/configs/redefined-designs/uploads:/backup/uploads:ro
|
|
- /volume1/configs/redefined-designs/backups/uploads:/archive
|
|
healthcheck:
|
|
# Nine days: the weekly interval plus two days of grace.
|
|
test: ["CMD-SHELL", "find /archive -name 'uploads-*' -mmin -12960 | grep -q ."]
|
|
interval: 6h
|
|
timeout: 30s
|
|
retries: 3
|
|
start_period: 8d
|
|
restart: unless-stopped
|
|
logging:
|
|
driver: json-file
|
|
options:
|
|
max-size: 10m
|
|
max-file: "3"
|