Deployment Guide
This guide takes you from a prepared Kubernetes cluster to a running RootCause Platform. It assumes your infrastructure team has completed the checklist in the Requirements doc.
There are four CLI steps to bootstrap the operator, then everything else happens in the Admin UI.
Before you start
Confirm you have:
If anything is missing, go back to the Requirements doc and hand the checklist to your infrastructure team.
Step 1: Create the namespace
All RootCause components will be deployed into this namespace.
Verify:
Step 2: Create registry credentials
The operator and platform images are hosted on GitLab Container Registry. Create an image pull secret:
Use the GitLab deploy token provided by RootCause. It needs read_registry scope. The operator reuses this secret for both image pulls and OCI chart pulls — no separate chart registry secret is needed.
Verify:
Step 3: Install the unified MongoDB Kubernetes Operator (MCK)
The platform uses the unified MongoDB Kubernetes Operator (MCK — MongoDB Controllers for Kubernetes) to manage its MongoDB replica set and MongoDB Search (mongot). MongoDB Search is part of every deployment, and only MCK can deploy it — the legacy MongoDB Community Operator cannot, and the RootCause Operator's preflight checks reject it.
The RootCause Operator pins MCK chart version 1.9.1 for compatibility with the dependency charts it deploys.
Skip this step if a compatible MCK is already installed cluster-wide. It should be configured to watch all namespaces. If an MCK is present but too old, the Admin UI's Rollout panel offers an Upgrade MongoDB operator button that remediates it in place.
Verify:
You should see one mongodb-kubernetes-operator-... pod in Running state.
Step 4: Install the RootCause Operator
The chart defaults imagePullSecrets to regcred, so no extra flags are needed.
Verify:
You should see:
Step 5: Access the Admin UI
Preflight checks: on first startup the operator seeds a
RootCauseInstallationinmode: preflight. Nothing is deployed yet, but the reconcile loop immediately runs live cluster checks — required CRDs, registry pull secret, MongoDB operator (MCK) version and schema compatibility — and populates the release catalog, so the Bootstrap page shows live verdicts and version dropdowns before you configure anything. Applying the wizard configuration in Step 6 switches the installation tomode: activeand starts the actual deployment.
Port-forward the Admin UI to your local machine:
Open http://localhost:3000 in your browser.
Log in
The operator generates a master password during installation. Retrieve it:
Enter this password on the login page.
Step 6: Configure via the Bootstrap Wizard
The Admin UI presents a wizard with several sections. Walk through each one.
Release Versions
Dependencies chart version and Platform chart version: Use the latest versions unless RootCause support has told you otherwise.
Installation
Namespace: Shown read-only — the namespace the operator was installed into.
Client ID: Your organization identifier (provided by RootCause, e.g.,
acme-corp).
Storage
Choose a storage backend and enter its connection details:
S3 / S3-compatible
Region, access key ID, secret access key; optional endpoint URL for S3-compatible stores
Azure Blob Storage
Storage account, endpoint suffix (core.windows.net), and one auth method: connection string, account key, or service principal (tenant ID, client ID, client secret)
Google Cloud Storage
Project ID and service account credentials
Garage (self-hosted, in-cluster)
Nothing — fully automated
For cloud backends, enter a single bucket/container name. Datasets, digital twins, and ML models are stored as sub-directories of this one bucket (ROOTCAUSE_BUCKET).
For Garage, the operator deploys an S3-compatible object store inside the cluster, creates the bucket, generates credentials, and wires everything to the platform. Pick a mode:
Single node — 1 replica, RF=1. For development and single-node clusters.
Production — 3 replicas, RF=3, pod anti-affinity. Requires 3+ schedulable nodes.
Set the storage capacity (PVC size per replica, default 50Gi).
Also set the Kubernetes storage class for persistent volumes (e.g., managed-csi, gp3, standard-rwo).
Networking
Base domain
Your base domain (e.g., rootcause.example.com)
Ingress class
nginx for most deployments. azure/application-gateway for Azure App Gateway.
Subdomains
Platform, Auth, and LiteLLM subdomains (e.g., platform, auth, litellm)
Ingress annotations
Add annotations under All ingresses (base) based on your ingress controller:
nginx:
nginx.ingress.kubernetes.io/ssl-redirect
true
For TLS with a pre-existing wildcard certificate, select Manual (bring your own secret) under TLS mode and provide the secret name.
Azure Application Gateway:
appgw.ingress.kubernetes.io/ssl-redirect
false
appgw.ingress.kubernetes.io/use-private-ip
true
For each ingress (Platform, Auth, LiteLLM), also add:
appgw.ingress.kubernetes.io/appgw-ssl-certificate
(your certificate name)
Azure Application Gateway note: App Gateway does not resolve loopback external URLs from within the cluster. You will need a
hostAliasesresource patch in the Advanced section — see Azure Application Gateway patches below.
Identity & Access
Choose based on the decision you made in the Requirements doc:
External OIDC (recommended)
Authentication mode
OIDC
Deploy FusionAuth
No
Issuer
Your IdP's issuer URL
Client ID
Your application's client ID
Client secret
Your application's client secret
Well-known URL
Your IdP's OpenID configuration URL
Logout URL
Your IdP's logout endpoint (include a post_logout_redirect_uri back to your platform URL)
Azure EntraID specifics: The issuer URL is
https://login.microsoftonline.com/<tenant-id>/v2.0. In the Azure Portal, ensure your app registration has:
Redirect URI:
https://<platform-subdomain>.<base-domain>/api/auth/callback/login(type: Web)Token configuration: Include
profile, andopenidscopesAPI permissions:
openid,profile,
External SAML
Authentication mode
SAML
Deploy FusionAuth
No
(remaining fields)
Metadata URL or XML, entity ID, and certificate from your IdP
Managed FusionAuth (POC / no existing IdP)
Authentication mode
Built-in (no SSO)
Deploy FusionAuth
Yes
FusionAuth API key
(leave blank)
FusionAuth Tenant ID
(leave blank)
Leave credentials blank. The operator auto-provisions everything: API keys, tenant, application, and OIDC configuration.
Email (optional)
SMTP credentials for outbound platform email: user, passkey, host, port, and an optional CA certificate (only needed when the SMTP server uses a private CA). Host defaults to smtp.rootcau.se and port to 30025. Skip this section if you don't need outbound email.
Telemetry (optional)
An opt-in, in-cluster observability stack — off by default. When enabled, logs and metrics flow through an OpenTelemetry Collector gateway into local Loki + Prometheus, with a passwordless Grafana over both (reach it with kubectl port-forward). Configure retention days (default 7) and volume sizes, and optionally forward everything to an external OTLP collector (gRPC or HTTP).
Dependencies
Each dependency is deployed by the operator by default, or connected to an external instance:
PostgreSQL
Deployed by the operator
Host, port, username, password, SSL mode
MongoDB
Set replicas: 1 for testing, 3 for production. MongoDB Search (mongot) is deployed alongside by default.
Connection URI
Redis
Set replicas: 1 for testing, 3 for production
Connection URL or JSON sentinel config
Temporal
Deployed by the operator
Frontend URL (host:7233)
LiteLLM
Deployed by the operator
URL of your existing LiteLLM instance
Node Selector (optional)
Constrain all workloads — every pod across both the dependencies and platform charts — to nodes matching label key/value pairs. For example, usage: rootcause to match a dedicated AKS node pool label.
Compute Node Pool (optional)
Route heavy ML jobs and digital twins (causal discovery, digital twin training, simulation) onto a dedicated node pool that scales to zero when idle. This applies only to the ephemeral Job pods the Data Service spawns — not the long-running services. Leave empty to run all ML jobs on the default nodes.
The contract is the same on every cloud:
You provision a node group that is tainted
rootcause.ai/workload=compute:NoSchedule, labeledrootcause.ai/nodepool=compute, and autoscaling with min = 0.The wizard fields stamp a matching node selector + toleration onto the spawned ML Job pods. Those pods stay
Pendinguntil a compute node exists.Your cluster autoscaler sees the
Pendingpod and scales the pool0 → 1, then back to0when the job finishes. The operator does not create node groups.
AKS example (built-in autoscaler, nothing extra to deploy):
On GKE, use the built-in autoscaler the same way. On EKS, self-deploy Cluster Autoscaler (or Karpenter) with node-template tags for the taint/label so the group can scale from zero.
Optional fields:
Pool name — logical name used for pool-scoped capacity checks (e.g.,
compute)Max node memory GiB — lets the platform reject physically impossible workloads while the pool is at zero
Job kinds — which ML job kinds are routed to the pool. Default: the heavy kinds (
causal_discovery,causal_discovery_aggregate,digital_twin,simulation); small, frequent ontology jobs stay on the default pool to avoid scale-from-zero latency.
Components
Configure replicas, resources, and extra environment variables per long-running component. Defaults work for testing. For production, use these as a starting point:
Platform
2-3
500m
1 Gi
Data Service
2-3
1
2 Gi
ML jobs are not configured here — they run as ephemeral Kubernetes Job pods with per-job resources set automatically by the Data Service. See the Upgrades & Operations doc for detailed scaling guidance.
Service Accounts (optional)
Create Kubernetes service accounts with annotations and assign them to the platform and/or data-service deployments. Useful for workload identity (e.g., AWS IAM Roles for Service Accounts).
Advanced
For most deployments, you can skip this section. It's here for edge cases.
Resource patches: Deep-merge patches into rendered Kubernetes manifests (annotations, tolerations, node selectors,
hostAliases, etc.)Raw Helm overrides: Free-form YAML merged over computed values for any chart setting not covered by the wizard
Azure Application Gateway patches
Azure Application Gateway does not resolve external URLs from within the cluster. The platform pod needs a hostAliases entry to route auth traffic to the Application Gateway's IP directly.
Add a Deployment resource patch:
Chart
platform
Kind
deployment
Object name
rootcause-platform
Patch content:
App Gateway also requires explicit path definitions. Add these in Raw Helm Overrides:
Platform chart overrides:
Dependencies chart overrides:
Without both path types, App Gateway may return 502 errors on some requests.
Effective Values Preview
Before applying, this section shows the final Helm values the operator computes from your configuration (wizard fields, patches, and raw overrides merged) — use it to sanity-check what will actually be deployed.
Deploy
Review the Deployment Summary at the bottom of the wizard, then click Apply configuration.
The operator will:
Create required secrets (storage credentials, platform config)
Create the
RootCauseInstallationcustom resourceDeploy infrastructure dependencies (PostgreSQL, Redis, MongoDB, Temporal, LiteLLM, and optionally FusionAuth)
Deploy platform services (Platform, Data Service)
Watch the Deployment status panel on the right. The phase progresses through Reconciling to Ready, typically in 2-5 minutes.
Step 7: Configure LLM models
After the platform is deployed, configure LLM access on the LLM page in the Admin UI. Configure at least two providers (e.g., one OpenAI, one Anthropic) so fallbacks keep the platform working through provider outages.
If you configured an external LiteLLM instance in the wizard, manage models in your own instance instead, then continue to Step 8.
The LLM page has four sections. Work top to bottom:
Provider Keys
Add an API key per LLM provider. The keys are stored as a Kubernetes secret in the installation namespace — you never need to log in to the LiteLLM UI or create virtual keys; the operator reads the LiteLLM master key itself.
Models
Register the LiteLLM models backed by those provider keys: pick the provider, the upstream model, and a model name. Repeat for each model you want available.
Tiers
Map the platform's model tiers (small / medium / large) to default models. The platform picks a tier per task; these defaults decide which registered model serves each tier.
Fallbacks
Define per-model fallback chains so the platform always has a working LLM, even during provider outages: select a primary model and add one or more fallbacks in order of preference.
Tip: Match fallbacks by weight class — if your primary is GPT-4, fall back to Claude Sonnet, not to a smaller model. The exact model doesn't matter as much as matching capability.
Where this configuration lives
The model registry and fallback configuration are persisted declaratively in the RootCauseInstallation CR (spec.config.dependenciesConfig.litellm.config); the LiteLLM ConfigMap is rendered from that spec on every reconcile. Tier routing is stored in MongoDB. This configuration survives upgrades and reinstalls — there is nothing to export or screenshot. Exporting the CR (see Upgrades & Operations) backs it up along with everything else.
Verify: the Models section lists your models without runtime sync errors, and each tier has a default model assigned.
Step 8: Create platform users
Navigate to the Users page in the Admin UI.
With managed FusionAuth:
Click + Add user
Enter email, password, first name, last name
Click Create user
The operator creates a FusionAuth account, registers it to the platform application, and adds the email to the admin list automatically.
With external OIDC or SAML:
Create users in your external identity provider (EntraID, Okta, etc.), then add their email addresses in the Platform admin emails section on the Users page. These emails are stored in the CR and mounted as PLATFORM_ADMIN_EMAILS on the platform deployment.
Step 9: Log in and verify
Navigate to https://<platform-subdomain>.<base-domain>.
With FusionAuth: click "Continue with OIDC SSO" and log in with the credentials from Step 8
With external OIDC/SAML: you'll be redirected to your identity provider
Post-install smoke test
Run through these checks to confirm everything is working:
1
All pods running
kubectl get pods -n rootcause — all should be Running or Completed
3
Platform login works
Navigate to platform URL, complete login flow, reach the home page
4
LLM responds
Admin UI LLM page shows your models without sync errors; ask the platform's AI a question and get a response
5
Data works end-to-end
Create a workspace, open a Data View, confirm it loads
If any check fails, see the Troubleshooting section in the Upgrades & Operations doc.
Appendix: Migrating from legacy Helm deployment
If you are migrating from an older Helm-based deployment (pre-operator), follow these steps before starting at Step 1 above.
Before you begin
Note down your LiteLLM model configuration — in legacy deployments this lives in PostgreSQL (via the LiteLLM UI) and will not survive the migration. Record your providers, models, and fallback rules; after migrating you will re-enter them once on the Admin UI LLM page, where they are stored in the installation CR from then on.
Export current Helm values —
helm get values rootcause-platform -n <namespace>andhelm get values rootcause-dependencies -n <namespace>. You'll use these as reference when filling in the bootstrap wizard.Backup databases (recommended):
mongodump --host <host> --out ./mongo-backuppg_dump -h <host> -U postgres -d appdb > postgres-backup.sql
Clean up existing installation
Choose one method:
Method A: Clean uninstall (preserves namespace, use if namespace is shared)
Method B: Delete and recreate namespace (simpler, but destroys everything in the namespace)
After cleanup
Start the deployment guide from Step 1 (Step 3 replaces the legacy MongoDB Community Operator with the unified MCK)
When filling in the bootstrap wizard, use your exported Helm values as reference
After deployment, re-enter your LiteLLM model configuration on the Admin UI LLM page (Step 7) — from now on it is stored in the CR and survives future reinstalls
Verify with the smoke test in Step 9
This is a one-time procedure. Once migrated to the operator, all future updates happen through the Admin UI.
Next step
Proceed to the Upgrades & Operations doc for day-2 operations: applying updates, rolling back, scaling, and troubleshooting.
Last updated

