Skip to main content

Migrating to the PostgreSQL Queue

Task processing moves from RabbitMQ to PostgreSQL in v0.180.0. This page covers that one upgrade. If you are already on v0.180.0 or later, use the normal Upgrading steps instead.

Do not upgrade past v0.180.0 with work still queued

Going from v0.179.0 or earlier straight to a release after v0.180.0 removes the RabbitMQ workers in the same step that switches processing over. Anything still queued at that moment is lost — silently, with no error.

Make sure the old queues are empty before you cross that line. Pick a path below.

Pick a path​

Want to avoid downtime?PathWhat it costs
Yes (recommended)B — upgrade to v0.180.0 first, then to the latest versionTwo upgrades, no downtime
NoA — drain the queues, then upgrade straight to the latest versionOne upgrade, with downtime and possible loss of work submitted during it

Followed as written, neither path loses work that is already queued.

The first on-prem release after v0.180.0 is v0.181.1001

v0.181.1001 has no Celery workers, so it is the release both paths end on. Path A goes straight to it. Path B goes through v0.180.0 first.

What downtime means here​

A maintenance window you announce to your users. It runs from the start of the drain until you re-enable your producers — the drain, the upgrade itself, and the checks after it — so size it for all three, not just the drain.

The platform does not stop accepting work on its own. There is no server-side switch for this: stopping your producers at step 1 is your job, and it is why the window matters twice over.

  • Work submitted after the check is lost. Anything that reaches the queue after the check prints READY is stranded by the upgrade — silently, with no error, exactly as described in the scheduled-pipeline warning below.
  • Work running during the deployment is unpredictable. RabbitMQ and the workers restart, and the backend APIs differ between versions, so executions in flight can fail in ways that are hard to attribute afterwards.

Announce the window to your users and to anyone running API or ETL integrations against the platform. Path B avoids the window entirely.

Path B — no downtime​

  1. Upgrade to v0.180.0. No configuration changes needed.
  2. Check both worker sets are running — the new PostgreSQL workers and your existing Celery workers.
  3. Let it run. New work goes to PostgreSQL immediately; the Celery workers finish what was already queued.
  4. Check the queues have drained. Run the queue check until it prints READY.
  5. Upgrade to the latest supported version when convenient.

Path A — with downtime​

  1. Stop every producer — the UI, the API, your scheduled pipelines (see below), and any ETL job or integration that submits work.
  2. Wait for running executions to finish.
  3. Check the queues are empty. Run the queue check until it prints READY.
  4. Upgrade to the latest supported version.
  5. Reopen the UI and the API, re-enable your scheduled pipelines, and check the pipelines are listed and firing.

Step 3 is the one that matters: anything still queued when you upgrade is lost.

Disable your scheduled pipelines first

A scheduled pipeline is a producer, so leaving it enabled undoes the drain. If one fires between steps 3 and 4, its work goes to RabbitMQ and the upgrade leaves it with no consumer. That execution is never recovered — it stays unfinished in your execution history.

Disable the pipelines before step 2 and re-enable them after step 4. Keep that order — re-enabling on the new version is what makes it safe, because a pipeline resumed there does not fire a catch-up run against the schedule it held while disabled. Schedules themselves survive the upgrade and resume from their next due time.

Checking the queues​

Run this against your RabbitMQ pod. rabbitmqctl ships with the broker and is not installed in the worker containers — running it there gives you command not found.

# find the pod first:  kubectl get pods -n unstract | grep rabbit
kubectl -n unstract exec <rabbitmq-pod> -- \
rabbitmqctl -q --no-table-headers list_queues name messages \
| awk '
{ lines++ }
$2 !~ /^[0-9]+$/ { printf " unexpected output: %s\n", $0; bad=1; next }
$1 ~ /\.pidbox$|^celery_periodic_logs$|^celery_log_task_queue$|^dashboard_metric_events$/ { next }
$2 > 0 { printf " still draining: %s (%s messages)\n", $1, $2; n++ }
END {
if (bad) { print "CHECK FAILED - unexpected output. Do not upgrade."; exit 1 }
if (lines==0) { print "CHECK FAILED - no queues listed. Do not upgrade."; exit 1 }
if (n) { print "NOT READY - wait a few minutes and run again"; exit 1 }
print "READY - the old queues are empty"
}'

Only READY means it is safe to upgrade. The command also exits non-zero unless it prints READY, so it can gate a script.

OutputWhat it means
READY - the old queues are emptySafe to upgrade.
NOT READY + the queues still holding workWait a few minutes and run it again.
CHECK FAILED - no queues listedWrong pod name, a pod that is not RabbitMQ, or a broker that is down. Not an all-clear.
CHECK FAILED - unexpected outputSomething other than queue<TAB>count rows arrived. Read the lines it echoes.

The CHECK FAILED cases matter because kubectl and rabbitmqctl write errors to stderr, which never reaches the pipe. Without them, a failed command would look exactly like a drained broker.

The check ignores three queues, because none of them hold execution work:

  • celery_periodic_logs — no consumer, never drains. See below.
  • dashboard_metric_events — metric aggregation. Loses its consumer at v0.180.0, so from then on the stuck-drain command below flags it as BACKLOG, NO CONSUMER. That is expected: anything left in it is aggregation that will not run, and it affects neither execution nor your data.
  • celery_log_task_queue — log traffic. Its consumer stays up through v0.180.0 to drain what is already there, so messages here are expected on both paths. Log traffic moves to Redis after the upgrade, so it should then fall to zero and stay there — if it keeps growing on v0.180.0, contact support.

If you cannot use kubectl exec​

Some hardened clusters deny pods/exec. The broker's management API gives the same numbers over pods/portforward, which is a separate permission:

kubectl -n unstract port-forward svc/unstract-rabbitmq 15672:15672

In another terminal, load the broker credentials — the same ones you set as CELERY_BROKER_USER / CELERY_BROKER_PASS, also available from the Secret:

RABBIT_USER=$(kubectl -n unstract get secret unstract-rabbitmq-default-user \
-o jsonpath='{.data.username}' | base64 -d)
RABBIT_PASS=$(kubectl -n unstract get secret unstract-rabbitmq-default-user \
-o jsonpath='{.data.password}' | base64 -d)

Then query the queues:

curl -sf -u "$RABBIT_USER:$RABBIT_PASS" \
'http://localhost:15672/api/queues/%2F?columns=name,messages,consumers'

curl -sf returns nothing and a non-zero exit on an authentication failure, so an empty result means the credentials are wrong — not that the queues are empty. The management UI in a browser works over the same port-forward.

If the drain looks stuck​

To see every queue and whether a consumer is attached:

kubectl -n unstract exec <rabbitmq-pod> -- \
rabbitmqctl -q --no-table-headers list_queues name messages consumers \
| sort -k2 -nr \
| awk 'BEGIN{printf "%-38s %9s %9s\n","QUEUE","MESSAGES","CONSUMERS"
print "------------------------------------------------------------"}
{printf "%-38s %9s %9s%s\n",$1,$2,$3,
($2>0 && $3==0 ? " <-- BACKLOG, NO CONSUMER" : ($3==0 ? " <-- no consumer" : ""))}'

A queue holding messages with no consumer is an orphan, not a slow drain — waiting will not change it. On an execution queue that means work is queued with nothing to run it. no consumer on an empty queue is harmless: that is a worker you have not deployed.

Also check the broker for a disk alarm, which blocks every publisher:

kubectl -n unstract exec <rabbitmq-pod> -- rabbitmqctl status | grep -A3 -i alarm

If an alarm is set, contact support before continuing.

The celery_periodic_logs backlog​

You do not need to drain or clear it before upgrading, whatever number it shows. It has no consumer, it never drains, and it holds no execution work — the check excludes it deliberately.

It is worth knowing about for one reason: on a version that still publishes to it, the backlog grows by roughly 86,000 messages a day and can fill enough broker disk to trip RabbitMQ's disk alarm — which blocks every publisher, and on v0.179.0 that means every execution dispatch. Freeing disk or purging that one queue restores normal draining — contact support before you purge anything.

What not to do​

Each of these fails without an error message.

  1. Do not disable the Celery workers in the same upgrade as Path B step 1. That strands whatever is still queued — no consumer, no error, no failure on the sending side.
  2. Do not disable the schedule mirror. The chart refuses to deploy, which is the safe outcome. Working around it by re-enabling the old scheduler causes schedules to fire twice.
  3. Do not disable the PostgreSQL worker fleet as a way to revert. It looks like the undo lever and is not — you get a platform that accepts uploads, reports success, and runs nothing. To revert, redeploy the previous image and chart together.
  4. Do not roll back and assume schedules follow. They must be handed back explicitly — contact support for the procedure.
  5. Do not re-enable the old scheduler afterwards. Your schedules already belong to the new one, so anything the old scheduler still holds will fire twice.

Notes​

  • RabbitMQ is still required for v0.180.0. It is removed in a later release.
  • Queue names are unchanged in v0.180.0, so execution work already queued is picked up by the same consumers as before. That is why Path B needs no downtime. The exception is the three queues the check excludes — dashboard_metric_events loses its consumer at v0.180.0 and celery_periodic_logs never had one, so anything left in those never runs.
  • Schedules move automatically. They are copied into the new scheduler and ownership is handed over during the upgrade. They start from their next due time — there is no catch-up burst.
  • Upgrade from v0.179.0 or earlier, or from a fully current version. A database left partway through an earlier PostgreSQL-queue rollout has migration records with no matching files, and the upgrade will not resolve them correctly. If you are unsure which state you are in, contact support first.

If the upgrade reports a schedule-migration failure​

The schedule migration runs after the database migration. On a jump across several versions, database index rebuilds can still be running when it starts, and the schedule step fails — loudly, so you will see it.

Re-run the same helm upgrade command once the database migration has finished. The schedule migration is a post-upgrade hook, so it runs again with the settings your deployment needs. Both the upgrade and the migration are idempotent, so repeating them is safe.

Do not run the migration command by hand

It takes flags the chart derives from your values. The wrong ones can hand your metrics schedules to a scheduler with no worker to run them — which stops those jobs with no error, the exact failure the migration exists to avoid.

If it fails again, capture the Job's output and contact support:

kubectl -n unstract logs job/pg-schedule-mirror

Schedules that were not transferred fire on neither scheduler, so escalate this rather than working around it.