Using an External or Managed Redis
By default the Unstract Helm chart deploys a Redis server for you. You can point it at your own Redis instead — either a managed service such as Google Memorystore, AWS ElastiCache or Azure Cache for Redis, or a server you run yourself.
Available from v0.182.0. Two behaviours described here — Sentinel turning off
automatically under multi-AZ HA, and the
TLS keys clearing themselves when you revert — need v0.183.0 or newer. On
v0.182.0 an HA installation renders cleanly and then fails at runtime, and a revert
leaves REDIS_SSL in force.
Quick start
Two changes to your values file.
1. Stop the chart deploying its own Redis:
redis:
enabled: false
This replaces the bundled Redis only. If you pass values-multi-az.yaml, keep passing it —
dropping it would also turn off the RabbitMQ and MinIO HA topologies. Use a copy taken from
the chart version you are installing: older saved copies set Sentinel and port keys on each
service that the chart now derives, and the upgrade fails if a hand-set value disagrees with
it. See Redis Sentinel.
2. Point the platform at your server:
global:
sharedConfigs:
redis:
REDIS_HOST: "redis.example.internal"
REDIS_PORT: "6379"
REDIS_PASSWORD: "<your auth string>"
REDIS_USER: ""
Deploy, and every component follows automatically — the backend, the workers, the platform service, the runner, and the startup checks that wait for Redis before a pod begins.
A managed endpoint's AUTH string is a password, not a username, so REDIS_USER
normally stays unset. default — which sample.on-prem.values.yaml ships — is
equally fine: it is Redis's built-in user, and AUTH default <password> is identical
to AUTH <password>.
What is not supported is a named ACL user. The cache client drops the username, so the server sees the password alone and rejects it. If your provider issues per-application ACL users, use the default user's password instead.
That is the whole configuration for a plaintext connection. Everything below is for specific situations.
Requirements
Your Redis server must be:
| Requirement | Why |
|---|---|
| Not in cluster mode | Unstract connects with a standard Redis client. A cluster endpoint returns MOVED errors on most keys. This is about the protocol, not the number of databases. |
| Redis 6.2 or newer | Log streaming uses the BLMOVE command. |
noeviction or volatile-lru | An allkeys-* policy can discard queued data that has no expiry. |
| Highly available, if you need the platform to be | Unstract has no independent Redis failover. |
Most managed tiers meet all four without changes.
Connecting with a URL
Instead of separate host and port values you can give a single URL, which takes
precedence over REDIS_HOST and REDIS_PORT:
global:
sharedConfigs:
redis:
REDIS_URL: "redis://redis.example.internal:6379/0"
REDIS_PASSWORD: "<password>"
The chart reads the host and port out of the URL for the components that cannot accept one, so you do not need to repeat them.
Keep the password out of the URL and set REDIS_PASSWORD instead. A URL with
no credentials is authenticated from that value. Both forms work — a URL that
already carries a password is used exactly as written, and REDIS_PASSWORD fills
only a gap it leaves — but a password in the URL travels wherever the URL goes,
including logs and diagnostic output.
An IPv6 address cannot be used, in a URL or in REDIS_HOST. The chart refuses to
render a REDIS_URL containing one. REDIS_HOST is not a workaround: the startup check
splits host:port on the first colon, so 2001:db8::1 is checked as host 2001, port
db8, and the pods sit in Init indefinitely with nothing explaining why. Use a
hostname that resolves to the IPv6 address.
TLS
TLS is off unless you turn it on. Use either form:
global:
sharedConfigs:
redis:
# with host and port
REDIS_SSL: "true"
# or in the URL, using the rediss:// scheme
REDIS_URL: "rediss://<host>:<port>/0?ssl_cert_reqs=required"
Every snippet below goes under the same global.sharedConfigs.redis block.
The TLS port differs by provider
There is no standard TLS port for Redis. Check yours:
| Provider | TLS port | Notes |
|---|---|---|
| Google Memorystore | 6378 | Non-TLS instances use 6379 |
| AWS ElastiCache | 6379 | Same port; encryption is in transit |
| Azure Cache for Redis | 6380 |
For Memorystore you can read it directly from the instance:
gcloud redis instances describe <name> --region <region> \
--format="value(host,port,transitEncryptionMode)"
A wrong port does not produce an error message. Pods wait for Redis before
starting, so they stay in Init indefinitely and it looks like a firewall or
networking problem. If pods hang at startup after enabling TLS, check the port
first.
Certificate verification
By default the platform verifies the server's certificate against the system trust store and checks that it matches the hostname you configured.
Providers using public certificate authorities work with no extra configuration — this includes AWS ElastiCache and Azure Cache for Redis.
Providers using their own certificate authority are not supported yet. This includes Google Memorystore, whose certificate is signed by a Google-managed CA. The chart cannot currently mount a CA certificate into the pods, so such a server cannot be verified. You can still connect with verification disabled:
REDIS_SSL_CERT_REQS: "none"
The connection remains encrypted, but the server is not authenticated. Treat this as a temporary measure on a trusted private network.
Connecting by IP address
Hostname verification compares the certificate against the address you configure. Managed providers issue certificates for their DNS names, not their IP addresses, so an IP-addressed endpoint fails verification.
Prefer the DNS name. Where that is not possible:
REDIS_SSL_CHECK_HOSTNAME: "false"
This still verifies the certificate chain and skips only the name check. It is a
smaller step than REDIS_SSL_CERT_REQS: "none", which disables verification
entirely.
All TLS settings
| Key | Default | Purpose |
|---|---|---|
REDIS_SSL | "false" | Enable TLS with host and port |
REDIS_SSL_CERT_REQS | "required" | "none" disables certificate verification |
REDIS_SSL_CHECK_HOSTNAME | "true" | "false" skips the hostname check only |
A privately signed certificate cannot be verified in this version: there is no
supported way to mount your own CA into the deployments that connect to Redis. Use
REDIS_SSL_CHECK_HOSTNAME: "false" if only the name fails to match, or
REDIS_SSL_CERT_REQS: "none" to disable verification entirely — the traffic is still
encrypted, but the server is not authenticated.
Servers that provide only one database
Unstract uses two Redis databases on-premise. A few non-clustered services expose database 0 only, such as Azure Managed Redis or a single-shard Redis Cloud database. Move everything onto database 0:
backend:
configMap:
FILE_ACTIVE_CACHE_REDIS_DB: "0"
METRICS_REDIS_DB: "0"
workerV2ConfigMap:
shared:
CACHE_REDIS_DB: "0"
METRICS_REDIS_DB: "0"
Set all four or none. The installation refuses to proceed on a half-edit, because every one of these failures is silent at runtime:
CACHE_REDIS_DBandFILE_ACTIVE_CACHE_REDIS_DBmust match, or duplicate files stop being recognised and get processed again.- The two
METRICS_REDIS_DBvalues must match each other, or one process's timings are lost. - On a single-database endpoint the LLMWhisperer bridge's
REDIS_DBmust be0too.
Nothing needs migrating. The affected keys all expire on their own, so the old database empties itself.
If you pin application images to a different version than the chart, older
images may not read METRICS_REDIS_DB and will keep using database 1. On a
server that provides database 0 only, that costs the per-run timing metric and
nothing else — extraction, deduplication and results are unaffected.
Leave all four unset for the normal layout. It needs no configuration.
Keeping the password in a Secret
The Redis password can come from a Kubernetes Secret instead of your values file, the same as any other shared-config group.
One rule is specific to Redis, and it is easy to get half-right: the endpoint has to appear in both places.
Setting existingSecret stops the chart creating its own Redis Secret, so your Secret is
the only thing the pods receive. It must carry every key they need, not just the password:
REDIS_HOST, REDIS_PORT, REDIS_PASSWORD, REDIS_USER, and any REDIS_SSL* keys you use
REDIS_HOST and REDIS_PORT must also stay in your values file, because several
things resolve them at render time rather than reading them from the Secret — the
startup checks, the workers' CACHE_REDIS_* and the LLMWhisperer bridge among them.
The installation refuses to proceed without them:
global:
sharedConfigs:
redis:
REDIS_HOST: "redis.example.internal"
REDIS_PORT: "6379"
existingSecret: "my-redis-secret"
Neither is secret, so there is nothing lost by repeating them. Keep the two copies in agreement. If they disagree, the startup check reaches the inline endpoint and reports success while the application connects to the one in the Secret.
Pods are not restarted automatically on this path — see the note under Going back to the in-cluster Redis.
Moving an existing installation
The new server starts empty and nothing is copied across.
- Let the platform go idle first. Queued execution logs are written out every few seconds; anything still queued when you switch is not recovered.
- Everything else is cache and rebuilds itself — deduplication state, worker caches and execution trackers all carry expiry times.
- Running both servers during the move is fine. Change the values only once the old server is quiet.
A fresh installation has none of this to consider.
Going back to the in-cluster Redis
Remove the keys you added and set redis.enabled: true (or drop the override). The TLS
settings clear themselves — REDIS_SSL, REDIS_SSL_CERT_REQS, REDIS_SSL_CHECK_HOSTNAME
and REDIS_URL all have declared defaults, so removing them from your values file resets
them rather than leaving them in force.
The bundled Redis takes a minute or two to become ready — longer under multi-AZ HA, which waits for its replicas to sync. The platform returns errors until then; that is the server starting, not a failed rollback.
It does not come back empty. The bundled Redis keeps its data on PersistentVolumeClaims, and disabling the subchart removes its pods, not those claims. It also runs with AOF persistence in both the standalone and multi-AZ topologies, so re-enabling reattaches the same volumes and Redis comes up holding the data from before you switched away. That is harmless for a cache, but it can read as live state. To come back clean, list the leftover claims while the subchart is still disabled and nothing is using them, then delete them before re-enabling:
kubectl get pvc -n <namespace> \
-l app.kubernetes.io/name=redis,app.kubernetes.io/instance=<your-helm-release-name>
Match on the instance label, not on the name. Anything else running a Redis in that
namespace — LLM Whisperer deploys its own — labels its claims name=redis too, and only the
release name tells them apart. Check each claim belongs to this release before deleting it;
deleting another release's claim destroys the storage of a Redis that is still in use.
The chart restarts the affected pods for you — unless the endpoint comes from a
Secret. Each service carries a checksum of the Redis config, so changing it in your
values file changes the pod template and Kubernetes rolls those pods. The checksum is
computed from the values, so if the endpoint is delivered by existingSecret it
does not change and the pods are not rolled. Restart them yourself
there. A pod reads its configuration once at startup, so if part of the platform still
reaches the old address, check those pods actually restarted.
Troubleshooting
| Symptom | Likely cause |
|---|---|
Pods stay in Init after enabling TLS | Wrong TLS port — see the table above |
Pods stay in Init after switching servers | Endpoint unreachable: firewall, VPC peering, or a wrong host |
MOVED errors in logs | The server is in cluster mode, which is not supported |
| Authentication failures on a managed endpoint | REDIS_USER names a named ACL user, which is not supported. Use the default user's password |
| Certificate verification failures | A private CA (see above), or an IP-addressed endpoint |
| Part of the platform still uses the old server | Those pods have not restarted |
Install fails naming a hand-set *_SENTINEL_MODE or REDIS_PORT | An old local values-multi-az.yaml — see HA Deployment |
| Duplicate files reprocessed | CACHE_REDIS_DB and FILE_ACTIVE_CACHE_REDIS_DB disagree |
Checking what the pods actually got
The install refuses to proceed when CACHE_REDIS_DB and
FILE_ACTIVE_CACHE_REDIS_DB differ, so that pair cannot drift on a current
chart. It can still differ on an installation upgraded from a chart older than
this feature, and the values reaching a pod are worth reading directly in any
case — a value set in more than one place resolves to whichever source the pod
sees last.
Replace <namespace> with the namespace the chart was installed into
(unstract in the deployment guide's examples).
kubectl exec -n <namespace> deploy/unstract-backend -- \
printenv REDIS_HOST REDIS_PORT FILE_ACTIVE_CACHE_REDIS_DB METRICS_REDIS_DB
# Any worker pod will do -- they share one ConfigMap for these keys.
kubectl get pods -n <namespace> -o name | grep unstract-worker | head -1 | \
xargs -I{} kubectl exec -n <namespace> {} -- \
printenv CACHE_REDIS_HOST CACHE_REDIS_PORT CACHE_REDIS_DB METRICS_REDIS_DB
CACHE_REDIS_DB from the second command must equal
FILE_ACTIVE_CACHE_REDIS_DB from the first. The two *_HOST values must name
the same server — the workers reach the endpoint through their own
CACHE_REDIS_* keys, which the chart derives from the same connection values as
the backend's.