Skip to main content

Using an External or Managed Redis

By default the Unstract Helm chart deploys a Redis server for you. You can point it at your own Redis instead — either a managed service such as Google Memorystore, AWS ElastiCache or Azure Cache for Redis, or a server you run yourself.

info

Available from v0.182.0. Two behaviours described here — Sentinel turning off automatically under multi-AZ HA, and the TLS keys clearing themselves when you revert — need v0.183.0 or newer. On v0.182.0 an HA installation renders cleanly and then fails at runtime, and a revert leaves REDIS_SSL in force.

Quick start​

Two changes to your values file.

1. Stop the chart deploying its own Redis:

redis:
enabled: false

This replaces the bundled Redis only. If you pass values-multi-az.yaml, keep passing it — dropping it would also turn off the RabbitMQ and MinIO HA topologies. Use a copy taken from the chart version you are installing: older saved copies set Sentinel and port keys on each service that the chart now derives, and the upgrade fails if a hand-set value disagrees with it. See Redis Sentinel.

2. Point the platform at your server:

global:
sharedConfigs:
redis:
REDIS_HOST: "redis.example.internal"
REDIS_PORT: "6379"
REDIS_PASSWORD: "<your auth string>"
REDIS_USER: ""

Deploy, and every component follows automatically — the backend, the workers, the platform service, the runner, and the startup checks that wait for Redis before a pod begins.

note

A managed endpoint's AUTH string is a password, not a username, so REDIS_USER normally stays unset. default — which sample.on-prem.values.yaml ships — is equally fine: it is Redis's built-in user, and AUTH default <password> is identical to AUTH <password>.

What is not supported is a named ACL user. The cache client drops the username, so the server sees the password alone and rejects it. If your provider issues per-application ACL users, use the default user's password instead.

That is the whole configuration for a plaintext connection. Everything below is for specific situations.

Requirements​

Your Redis server must be:

RequirementWhy
Not in cluster modeUnstract connects with a standard Redis client. A cluster endpoint returns MOVED errors on most keys. This is about the protocol, not the number of databases.
Redis 6.2 or newerLog streaming uses the BLMOVE command.
noeviction or volatile-lruAn allkeys-* policy can discard queued data that has no expiry.
Highly available, if you need the platform to beUnstract has no independent Redis failover.

Most managed tiers meet all four without changes.

Connecting with a URL​

Instead of separate host and port values you can give a single URL, which takes precedence over REDIS_HOST and REDIS_PORT:

global:
sharedConfigs:
redis:
REDIS_URL: "redis://redis.example.internal:6379/0"
REDIS_PASSWORD: "<password>"

The chart reads the host and port out of the URL for the components that cannot accept one, so you do not need to repeat them.

Keep the password out of the URL and set REDIS_PASSWORD instead. A URL with no credentials is authenticated from that value. Both forms work — a URL that already carries a password is used exactly as written, and REDIS_PASSWORD fills only a gap it leaves — but a password in the URL travels wherever the URL goes, including logs and diagnostic output.

note

An IPv6 address cannot be used, in a URL or in REDIS_HOST. The chart refuses to render a REDIS_URL containing one. REDIS_HOST is not a workaround: the startup check splits host:port on the first colon, so 2001:db8::1 is checked as host 2001, port db8, and the pods sit in Init indefinitely with nothing explaining why. Use a hostname that resolves to the IPv6 address.

TLS​

TLS is off unless you turn it on. Use either form:

global:
sharedConfigs:
redis:
# with host and port
REDIS_SSL: "true"

# or in the URL, using the rediss:// scheme
REDIS_URL: "rediss://<host>:<port>/0?ssl_cert_reqs=required"

Every snippet below goes under the same global.sharedConfigs.redis block.

The TLS port differs by provider​

There is no standard TLS port for Redis. Check yours:

ProviderTLS portNotes
Google Memorystore6378Non-TLS instances use 6379
AWS ElastiCache6379Same port; encryption is in transit
Azure Cache for Redis6380

For Memorystore you can read it directly from the instance:

gcloud redis instances describe <name> --region <region> \
--format="value(host,port,transitEncryptionMode)"
warning

A wrong port does not produce an error message. Pods wait for Redis before starting, so they stay in Init indefinitely and it looks like a firewall or networking problem. If pods hang at startup after enabling TLS, check the port first.

Certificate verification​

By default the platform verifies the server's certificate against the system trust store and checks that it matches the hostname you configured.

Providers using public certificate authorities work with no extra configuration — this includes AWS ElastiCache and Azure Cache for Redis.

Providers using their own certificate authority are not supported yet. This includes Google Memorystore, whose certificate is signed by a Google-managed CA. The chart cannot currently mount a CA certificate into the pods, so such a server cannot be verified. You can still connect with verification disabled:

REDIS_SSL_CERT_REQS: "none"

The connection remains encrypted, but the server is not authenticated. Treat this as a temporary measure on a trusted private network.

Connecting by IP address​

Hostname verification compares the certificate against the address you configure. Managed providers issue certificates for their DNS names, not their IP addresses, so an IP-addressed endpoint fails verification.

Prefer the DNS name. Where that is not possible:

REDIS_SSL_CHECK_HOSTNAME: "false"

This still verifies the certificate chain and skips only the name check. It is a smaller step than REDIS_SSL_CERT_REQS: "none", which disables verification entirely.

All TLS settings​

KeyDefaultPurpose
REDIS_SSL"false"Enable TLS with host and port
REDIS_SSL_CERT_REQS"required""none" disables certificate verification
REDIS_SSL_CHECK_HOSTNAME"true""false" skips the hostname check only

A privately signed certificate cannot be verified in this version: there is no supported way to mount your own CA into the deployments that connect to Redis. Use REDIS_SSL_CHECK_HOSTNAME: "false" if only the name fails to match, or REDIS_SSL_CERT_REQS: "none" to disable verification entirely — the traffic is still encrypted, but the server is not authenticated.

Servers that provide only one database​

Unstract uses two Redis databases on-premise. A few non-clustered services expose database 0 only, such as Azure Managed Redis or a single-shard Redis Cloud database. Move everything onto database 0:

backend:
configMap:
FILE_ACTIVE_CACHE_REDIS_DB: "0"
METRICS_REDIS_DB: "0"

workerV2ConfigMap:
shared:
CACHE_REDIS_DB: "0"
METRICS_REDIS_DB: "0"

Set all four or none. The installation refuses to proceed on a half-edit, because every one of these failures is silent at runtime:

  • CACHE_REDIS_DB and FILE_ACTIVE_CACHE_REDIS_DB must match, or duplicate files stop being recognised and get processed again.
  • The two METRICS_REDIS_DB values must match each other, or one process's timings are lost.
  • On a single-database endpoint the LLMWhisperer bridge's REDIS_DB must be 0 too.

Nothing needs migrating. The affected keys all expire on their own, so the old database empties itself.

note

If you pin application images to a different version than the chart, older images may not read METRICS_REDIS_DB and will keep using database 1. On a server that provides database 0 only, that costs the per-run timing metric and nothing else — extraction, deduplication and results are unaffected.

Leave all four unset for the normal layout. It needs no configuration.

Keeping the password in a Secret​

The Redis password can come from a Kubernetes Secret instead of your values file, the same as any other shared-config group.

One rule is specific to Redis, and it is easy to get half-right: the endpoint has to appear in both places.

Setting existingSecret stops the chart creating its own Redis Secret, so your Secret is the only thing the pods receive. It must carry every key they need, not just the password:

REDIS_HOST, REDIS_PORT, REDIS_PASSWORD, REDIS_USER, and any REDIS_SSL* keys you use

REDIS_HOST and REDIS_PORT must also stay in your values file, because several things resolve them at render time rather than reading them from the Secret — the startup checks, the workers' CACHE_REDIS_* and the LLMWhisperer bridge among them. The installation refuses to proceed without them:

global:
sharedConfigs:
redis:
REDIS_HOST: "redis.example.internal"
REDIS_PORT: "6379"
existingSecret: "my-redis-secret"

Neither is secret, so there is nothing lost by repeating them. Keep the two copies in agreement. If they disagree, the startup check reaches the inline endpoint and reports success while the application connects to the one in the Secret.

Pods are not restarted automatically on this path — see the note under Going back to the in-cluster Redis.

Moving an existing installation​

The new server starts empty and nothing is copied across.

  • Let the platform go idle first. Queued execution logs are written out every few seconds; anything still queued when you switch is not recovered.
  • Everything else is cache and rebuilds itself — deduplication state, worker caches and execution trackers all carry expiry times.
  • Running both servers during the move is fine. Change the values only once the old server is quiet.

A fresh installation has none of this to consider.

Going back to the in-cluster Redis​

Remove the keys you added and set redis.enabled: true (or drop the override). The TLS settings clear themselves — REDIS_SSL, REDIS_SSL_CERT_REQS, REDIS_SSL_CHECK_HOSTNAME and REDIS_URL all have declared defaults, so removing them from your values file resets them rather than leaving them in force.

The bundled Redis takes a minute or two to become ready — longer under multi-AZ HA, which waits for its replicas to sync. The platform returns errors until then; that is the server starting, not a failed rollback.

note

It does not come back empty. The bundled Redis keeps its data on PersistentVolumeClaims, and disabling the subchart removes its pods, not those claims. It also runs with AOF persistence in both the standalone and multi-AZ topologies, so re-enabling reattaches the same volumes and Redis comes up holding the data from before you switched away. That is harmless for a cache, but it can read as live state. To come back clean, list the leftover claims while the subchart is still disabled and nothing is using them, then delete them before re-enabling:

kubectl get pvc -n <namespace> \
-l app.kubernetes.io/name=redis,app.kubernetes.io/instance=<your-helm-release-name>

Match on the instance label, not on the name. Anything else running a Redis in that namespace — LLM Whisperer deploys its own — labels its claims name=redis too, and only the release name tells them apart. Check each claim belongs to this release before deleting it; deleting another release's claim destroys the storage of a Redis that is still in use.

note

The chart restarts the affected pods for you — unless the endpoint comes from a Secret. Each service carries a checksum of the Redis config, so changing it in your values file changes the pod template and Kubernetes rolls those pods. The checksum is computed from the values, so if the endpoint is delivered by existingSecret it does not change and the pods are not rolled. Restart them yourself there. A pod reads its configuration once at startup, so if part of the platform still reaches the old address, check those pods actually restarted.

Troubleshooting​

SymptomLikely cause
Pods stay in Init after enabling TLSWrong TLS port — see the table above
Pods stay in Init after switching serversEndpoint unreachable: firewall, VPC peering, or a wrong host
MOVED errors in logsThe server is in cluster mode, which is not supported
Authentication failures on a managed endpointREDIS_USER names a named ACL user, which is not supported. Use the default user's password
Certificate verification failuresA private CA (see above), or an IP-addressed endpoint
Part of the platform still uses the old serverThose pods have not restarted
Install fails naming a hand-set *_SENTINEL_MODE or REDIS_PORTAn old local values-multi-az.yaml — see HA Deployment
Duplicate files reprocessedCACHE_REDIS_DB and FILE_ACTIVE_CACHE_REDIS_DB disagree

Checking what the pods actually got​

The install refuses to proceed when CACHE_REDIS_DB and FILE_ACTIVE_CACHE_REDIS_DB differ, so that pair cannot drift on a current chart. It can still differ on an installation upgraded from a chart older than this feature, and the values reaching a pod are worth reading directly in any case — a value set in more than one place resolves to whichever source the pod sees last.

Replace <namespace> with the namespace the chart was installed into (unstract in the deployment guide's examples).

kubectl exec -n <namespace> deploy/unstract-backend -- \
printenv REDIS_HOST REDIS_PORT FILE_ACTIVE_CACHE_REDIS_DB METRICS_REDIS_DB

# Any worker pod will do -- they share one ConfigMap for these keys.
kubectl get pods -n <namespace> -o name | grep unstract-worker | head -1 | \
xargs -I{} kubectl exec -n <namespace> {} -- \
printenv CACHE_REDIS_HOST CACHE_REDIS_PORT CACHE_REDIS_DB METRICS_REDIS_DB

CACHE_REDIS_DB from the second command must equal FILE_ACTIVE_CACHE_REDIS_DB from the first. The two *_HOST values must name the same server — the workers reach the endpoint through their own CACHE_REDIS_* keys, which the chart derives from the same connection values as the backend's.