The constraints
The production instance runs ERPNext v13 — deployed 2021, a Docker Compose stack from the official frappe_docker repo, unmodified since. The database holds years of sales invoices, contractor payments, and chart-of-accounts entries. Small by ERPNext standards, large enough that losing it is not an option.
Three constraints shaped everything:
- Frappe refuses to skip major versions.
bench migrateon a v13 database with v15 code exits early with a version check. The only supported path is v13 → v14 → v15, running each version’s patches in order. That means we need a working v14 environment and a working v15 environment, each with the right code. - No maintenance window on prod. We don’t get a weekend to take the existing instance down and experiment. Everything has to happen on a separate stack, with production continuing to run (frozen, but serving) until we’re confident.
- HRMS and Payments have to be installed. In v13, HR and payment-gateway code lived inside erpnext. In v15, they’re separate apps. The v15 upgrade expects to find them — and our custom v15 image (part 1) ships with both apps’ code pre-installed in the bench.
Approach: the workbench
The obvious approach is “spin up a v15 stack, restore backup, done.” That fails at the version-gap problem. The equally obvious approach is “spin up a v13 stack, upgrade in place.” That fails because you don’t have a v13 image handy and building one is another afternoon of Docker archaeology.
The approach that worked is a hybrid we called the workbench: a minimal, single-purpose Compose stack whose volumes persist across image swaps. Only the image tag changes between phases.
v13 backup ──restore──▶ [v14 image] ──bench migrate──▶ [v14 backup]
│
[v15 image] ◀──swap image only──────┘
│
bench migrate──▶ install HRMS+Payments──▶ [v15 backup]
The Compose file (overrides/compose.migration.yaml) keeps only what a migration needs:
backend— the app server, running whichever image we’re onconfigurator— a one-shot container that writescommon_site_config.json(DB host, Redis endpoints)db(MariaDB 10.11),redis-cache,redis-queue
Everything else — nginx frontend, websocket server, scheduler, queue workers — is stripped. A migration doesn’t serve web pages; it churns through Python patch scripts as fast as the database will let it.
Because the sites named volume survives docker compose down, moving between v14 and v15 is:
MIGRATION_IMAGE=frappe-erpnext MIGRATION_TAG=v14-hrms docker compose up -d
# ... restore + migrate ...
docker compose down
MIGRATION_IMAGE=frappe-erpnext MIGRATION_TAG=v15-hrms docker compose up -d
# ... migrate ...
Same volumes, same site, new code. Downtime between phases: the container restart, seconds.
Restore: the easy part, with two footguns
The v13 backup from the production instance contains four files: the gzipped SQL dump, public files tarball, private files tarball, and a site_config_backup.json. Restoring needs the file paths, not flags:
bench --site <your-site> restore \
--db-root-password "$DB_PASSWORD" \
--admin-password admin \
--with-public-files backups/...-files.tar \
--with-private-files backups/...-private-files.tar \
backups/...-database.sql.gz
Two footguns worth noting:
Footgun 1: the site gets created before the restore, on the v14 image, with --no-mariadb-socket. The restore then overwrites the admin password with the one from the backup. The --admin-password admin flag you just passed? It applies to the fresh site, not the restored one. Budget one “incorrect password” login round before remembering to run bench set-admin-password after restore. (And budget a five-minute timeout after a few failed attempts — Frappe lockouts accounts on repeated failures, which is the correct behavior at the worst possible time.)
Footgun 2: bench new-site flag names moved between versions. v14 uses --no-mariadb-socket; v15 wants --mariadb-user-host-login-scope=%. Copy-pasting the wrong one into the other version’s container produces a connection error that looks like a DB permission problem.
v14 migrate: the is_virtual failure
The first bench migrate on the restored v13 data dies partway:
pymysql.err.OperationalError: (1054, "Unknown column 'is_virtual' in 'field list'")
The failing patch is one of the chat/communication removals in v14 — it tries to update a column that, on our database, was already gone. This is the classic Frappe migration scenario: the patch log says one thing, the actual schema says another, and neither is lying (a v13 variant or a prior manual cleanup explains the drift).
Frappe provides an escape hatch for exactly this:
bench --site <your-site> migrate --skip-failing
--skip-failing logs failing patches and continues, letting the run finish. It’s a blunt instrument — you’re choosing to skip a patch rather than understand it — so the discipline is to note which patches got skipped, grep the post-migration data for anything they’d have touched, and move on only if the impact is nil. In our case: chat-related tables nobody used, skipped cleanly.
After v14 migrate completes and clears cache, we take a full v14 backup before touching v15. That backup is the rollback point; if v15 migrate explodes, we restore the v14 state and we’re out maybe an hour, having lost nothing from production.
v15 migrate and the orphaned DocTypes
Swap the image tag, up -d, and run bench migrate — no --skip-failing this time. It runs clean, then prints something that looks alarming:
Deleting orphaned DocTypes...
…followed by roughly fifteen DocType names. Read that as good news, not bad. Those DocTypes were registered by v13-era code (old HR module, old payment gateway integrations) whose code no longer exists in the v15 bench. The tables held data our new stack can’t even represent, and leaving them registered would leave the Desk showing doctypes that 500 on open. The migration deletes the registrations; the tables themselves remain in MariaDB (check SHOW TABLES two versions later — they’re still there, just unreferenced).
Then install the split apps:
bench --site <your-site> install-app hrms
bench --site <your-site> install-app payments
Because the apps’ code was baked into the image (part 1), the install is pure database wiring — modules, DocTypes, patches — and takes minutes, not a git-clone-and-hope session. bench list-apps afterwards:
frappe 15.x.x
erpnext 15.x.x
hrms 15.x.x
payments 0.0.x
Final migrate, final backup, done. Four years of data, three versions, forty-ish minutes of compute and the better part of a day of babysitting.
What the workbench buys you
The alternative to a persistent-volume workbench is the “full stack swap” — for each phase, tear down a complete stack (frontend, workers, scheduler, the works), restore, migrate, tear down again. It works, but every cycle multiplies the places things can break: Redis state mismatch, configurator ordering, site naming. The workbench collapses all of that into one stack where the only variable across phases is the image tag. When something breaks mid-migration, you’re debugging one thing, not the whole compose matrix.
The script (scripts/migrate-workbench.sh) encodes the whole pipeline as phases — setup, restore, migrate-v14, backup-v14, upgrade-v15, migrate-v15, install-apps, backup-v15 — so a re-run after a failure resumes at the broken phase instead of starting over. That design earned its keep twice: once when a bad apps.json broke the v15 image (rebuild, re-up, resume), and once when the DB container died mid-migrate (restart, resume).
Next post: validation. The REST API test suite we built to check every split-app boundary — and the CSS disaster that taught us not to run bench build on one container of a multi-container stack.