OPC Foundation Cloud Initiative Open-Source Reference Solution
- Why This Solution
- Reference Edge Hardware
- Deploying the Software Stack
- Uninstalling
- Simulated Production Line
- Automatic Certificate Provisioning (GDS Server Push)
- Enabling TLS
- Accessing the Web UIs
- Managing the Cluster with Portainer
- Inspecting the Broker with MQTT Explorer
- UA Cloud Library for Digital Product Passports
- UA Data Processor (PCF and Battery Passport)
- Browsing the Data as a Graph (i3X)
- Asking Questions with AI (MCP)
- Pre-Provisioned Grafana Dashboards
- Tutorials
- Security Analysis (STRIDE)
Industrial data is trapped. It sits in machines that speak dozens of incompatible protocols, behind gateways you don't control, in platforms that charge for every tag and make leaving expensive. Connecting a factory for data analytics and AI today usually means picking a vendor — and then living with their protocols, their data model, their pricing, and their roadmap for decades.
This reference solution shows there is another way. It is a complete, higly scalable, end-to-end industrial IoT stack — from the sensor on the shop floor to a queryable time-series dashboard and back down to a command that actuates a machine — built entirely from open standards and open-source components. You can deploy it today, on hardware you own, with no subscription fee.
| Problem | How this solution addresses it |
|---|---|
| Protocol fragmentation — every machine speaks something different (Modbus, BACnet, OPC DA/AE, Siemens S7, Rockwell, Beckhoff, Mitsubishi, IEC 61850, OCPP, LoRaWAN, Matter, Redfish, HTTP/REST…) | The UA Edge Translator normalises all of them into a single OPC UA information model, using W3C Web of Things (WoT) Thing Descriptions as the declarative, vendor-neutral mapping format — no custom code per device. |
| Data without meaning — most IoT pipelines ship anonymous numbers that need out-of-band documentation to interpret | Data flows as OPC UA PubSub with accompanying metadata messages, so every value arrives with its type, semantics, and source. Full OPC UA Information Models can be imported from the UA Cloud Library so you know what a machine could report, not just what it happens to send. |
| Vendor lock-in — proprietary apps, proprietary payload formats, and egress/ingest pricing that grows with your data | Every component is open source and speaks standard MQTT with standard OPC UA PubSub JSON payloads. Point it at any broker, any database, any cloud — or keep it entirely on-premises. |
| Read-only pipelines — telemetry goes up, but nothing can come back down safely | UA Cloud Commander implements the spec-compliant OPC UA PubSub Actions request/response pattern, so cloud or local applications can securely read, write, call methods, and read history on shop-floor servers. |
| No closed loop — insights stay in dashboards instead of driving action | UA Cloud Action watches the time-series data and automatically triggers OPC UA method calls when thresholds are crossed — a genuine digital feedback loop, running at the edge with no cloud dependency. |
| Certificate management pain — OPC UA security is often disabled because provisioning trust is tedious | GDS Server Push provisions certificates and trust lists automatically, so the stack runs secure by default instead of secure-in-theory. |
| Hard to evaluate — pilots stall because getting any data flowing takes days if not weeks | A simulated production line ships with the stack. Apply two manifests and real OPC UA telemetry is flowing into dashboards within minutes — no hardware required to start. |
- Connect a brownfield machine without learning a custom UI and a proprietary asset description language from a connectivity provider. Describe it once in a WoT Thing Description and it appears as a fully-modelled OPC UA server automatically.
- Move your entire data pipeline between clouds — or off the cloud — in an afternoon. Because the wire formats are open standards, the broker, database, and dashboards are all replaceable parts, not a platform you're bound to.
- Close the loop from analytics back to the machine, using a standardised, auditable command pattern rather than a bespoke integration.
- Build your own applications against a standard REST API. The OPC UA Web API (OpenAPI-based) lets any language talk to your OPC UA estate over plain HTTP/JSON.
- Run the whole thing in under a minute on a less than $300 industrial PC — or scale the same manifests across a fleet, or split them between edge and cloud clusters. Same code, same standards.
- 100% open source. Every component — UA Edge Translator, UA Cloud Publisher,
UA Cloud Commander, UA Cloud Action, Mosquitto, Telegraf, InfluxDB, Grafana,
Portainer — is open source. There are no subscription fees, no per-message
costs, no seat counts, and no expiring trials. Fork it, audit it, extend it.
You still own the total cost of ownership: the hardware and the effort to operate, patch, and support the stack yourself. What you avoid is recurring licence and consumption billing — and the dependency that comes with it.
- Runs anywhere. It is plain Kubernetes (K3s). Deploy it on a Raspberry Pi in a control cabinet, a rack server in your own data center, or a managed Kubernetes service in any public cloud. The edge/cloud split lets you draw the boundary wherever your architecture and data-sovereignty rules require — including fully air-gapped.
- Vendor independence. Nothing here depends on a specific PLC vendor, cloud provider, historian, or dashboard tool. Each box in the pipeline is swappable because the interfaces between them are public specifications, not private APIs.
- No proprietary payloads. What goes over the wire is OPC UA PubSub JSON over MQTT — documented, inspectable, and consumable by anything.
Open standards are used throughout the stack, not just at the edges:
| Standard | Where it is used |
|---|---|
| OPC UA (IEC 62541 series) | The information model and the security model for all shop-floor connectivity. |
| OPC UA PubSub (IEC 62541-14) | The telemetry wire format (JSON over MQTT), including metadata messages. |
| OPC UA Actions (IEC 62541-14) | The command & control request/response pattern used by Cloud Commander and Cloud Action. |
| OPC UA GDS Server Push (IEC 62541-12) | Automated certificate and trust-list provisioning. |
| OPC UA Web API (IEC 62541-4, OpenAPI) | The RESTful interface for building custom applications — an OpenAPI representation of the OPC UA Services. |
| W3C Web of Things (WoT) (W3C Recommendation) | Thing Descriptions that declaratively map non-OPC UA assets into OPC UA. |
| MQTT 5.0 (OASIS) | The messaging transport, with TLS and authentication. MQTT v5 features (Correlation Data, Response Topic, Message Expiry) carry the request/response correlation for OPC UA Actions. |
| EN 18222 (CEN/CENELEC) | Digital Product Passport data model and unique identifiers — the structure of the DPPs that UA Data Processor generates and stores in the UA Cloud Library. |
| EN 18223 (CEN/CENELEC) | Digital Product Passport system architecture and data exchange — how DPPs are stored and retrieved by downstream consumers (customer, recycler, regulator) over the Cloud Library's REST API. |
| i3X | The vendor-neutral REST API for browsing industrial data as a connected ISA-95 graph — enterprise, site, area, line, station — instead of as flat time series, with typed relationships and current/historical values behind one interface. |
| Kubernetes (CNCF) | The deployment and operations model. |
| MCP (Model Context Protocol) | The agentic AI interface, exposing the plant as a set of tools to AI agents. |
Because these are published specifications rather than product features, any conforming tool — from any vendor — can participate in this architecture. That is the difference between an integration and an ecosystem.
Evaluating this for your organisation? Start with Deploying the Software Stack: two
kubectl applycommands bring up the full pipeline plus a simulated production line, so you can see live data in Grafana before committing any hardware. A STRIDE security analysis and production hardening guidance are included to support an enterprise architectural review.
The reference solution runs on any 64-bit Linux host capable of running K3s. For a validated, industrial-grade edge gateway we recommend a fanless Raspberry Pi Compute Module 5 (CM5) industrial PC — see hardware.md for the recommended bill of materials, SSD imaging, assembly, and first-boot instructions.
The reference workload is split into two manifests that run on a lightweight Kubernetes cluster (K3s):
| Manifest | Namespace | Components |
|---|---|---|
edge.yaml |
edge |
UA Edge Translator, UA Cloud Publisher, UA Cloud Commander |
edge.yaml |
munich |
Simulated production line (MES, assembly, test, packaging stations) |
cloud.yaml |
cloud |
Mosquitto, Telegraf, InfluxDB, Grafana, Portainer, UA Cloud Action |
The edge part contains the components that sit next to the machines and speak OPC UA / industrial protocols, plus a simulated production line so the stack produces real OPC UA telemetry out of the box (see Simulated Production Line). The cloud part contains the broker, storage, visualization, and management components that would typically run in a data center or public cloud.
For convenience, everything can be installed on a single K3s instance Simply apply both manifests to the same cluster; the two namespaces keep the edge and cloud workloads logically separated while they share one node. In a distributed deployment, apply
edge.yamlto the edge cluster andcloud.yamlto the cloud cluster, and update the cross-namespace service DNS names (see Apply the Stack Manifests) to point at the remote endpoints.
Together, edge.yaml and cloud.yaml deploy the following components, forming an
end-to-end pipeline from industrial protocols to a time-series database.
Workloads
| Component | Namespace | Image | Ports |
|---|---|---|---|
| ua-edgetranslator | edge |
ghcr.io/opcfoundation/ua-edgetranslator:main |
4840, 5000/5001, 19520/19521 (TCP on node); 8080 UI (ClusterIP; HTTPS via ingress) |
| ua-cloudpublisher | edge |
ghcr.io/barnstee/ua-cloudpublisher:main |
8081 (ClusterIP; HTTPS via ingress) |
| ua-cloudcommander | edge |
ghcr.io/opcfoundation/ua-cloudcommander:main |
— |
| mes, assembly, test, packaging | munich |
ghcr.io/digitaltwinconsortium/manufacturingontologies:main |
4840 (each) |
| modbus-simulator | munich |
python:3.12-slim |
502 (Modbus TCP) |
| mosquitto | cloud |
eclipse-mosquitto:2.1.2-alpine |
8883 (MQTT/TLS) |
| mqtt-explorer | cloud |
smeagolworms4/mqtt-explorer:browser-1.0.3 |
4000 (UI) |
| telegraf | cloud |
telegraf:1.39-alpine |
— |
| influxdb | cloud |
influxdb:2.9 |
8086 (UI/API) |
| grafana | cloud |
grafana/grafana:13.1.1 |
3000 (UI) |
| ua-cloudaction | cloud |
ghcr.io/opcfoundation/ua-cloudaction:main |
8082 (ClusterIP; HTTPS via ingress) |
| ua-cloudlibrary | cloud |
ghcr.io/opcfoundation/ua-cloudlibrary:latest |
8083 (ClusterIP; HTTPS via ingress) |
| cloudlib-postgres | cloud |
postgres:17.6-alpine |
5432 (ClusterIP only) |
| ua-dataprocessor | cloud |
ghcr.io/opcfoundation/ua-dataprocessor:main |
— |
| i3x4influx | cloud |
ghcr.io/barnstee/i3x4influx:main |
8084 (ClusterIP; HTTPS via ingress) |
| ua-cloudai | cloud |
ghcr.io/barnstee/ua-cloudai:main |
5050 (ClusterIP; HTTPS via ingress) |
| portainer | cloud |
portainer/portainer-ce:2.44.0 |
9443 (HTTPS UI), 9000, 8000 |
What each component does
- ua-edgetranslator — OPC Foundation UA Edge Translator. Connects to southbound assets and translates protocols (LoRaWAN, OCPP, etc.) into an OPC UA information model. Exposes a web UI for configuration.
- ua-cloudpublisher — UA Cloud Publisher. Subscribes to OPC UA data (from the Edge Translator or the simulated line) and publishes it as OPC UA PubSub JSON messages to the MQTT broker. Exposes a web UI for configuration.
- ua-cloudcommander — OPC Foundation UA Cloud Commander, the command & control
Responder. Subscribes to
commands/#forua-action-requestmessages, executes OPC UA operations (Read, HistoricalRead, Write, MethodCall) against on-premises OPC UA servers, and replies on theresponsestopic. No web UI. - mes / assembly / test / packaging — the simulated production line "Munich": four OPC UA servers modelling a factory line (MES shift schedule plus assembly, test and packaging stations), providing live telemetry out of the box. See Simulated Production Line.
- modbus-simulator — a simulated Modbus TCP device on the Munich line, automatically mapped into OPC UA by the Edge Translator via a W3C WoT Thing Description. See Simulated Modbus TCP Device.
- mosquitto — Eclipse Mosquitto MQTT broker carrying the OPC UA PubSub
data/#andmetadatamessages. Configured viamosquitto-confwith a TLS listener on 8883 and username/password authentication (allow_anonymous false) using theIOT_USERNAME/IOT_PASSWORDsupplied at apply time. - mqtt-explorer — MQTT Explorer, a browser-based client for inspecting the
broker: browse the live topic tree, read the OPC UA PubSub payloads on
data/#andmetadata, and publish messages by hand. The broker connection is pre-seeded, so it connects with one click.⚠️ It has no authentication of its own — see Inspecting the Broker with MQTT Explorer. - telegraf — Telegraf agent that consumes the MQTT PubSub messages, parses them
with the
json_v2parser (from thetelegraf-confConfigMap), and writes them to InfluxDB as theopcua_pubsub(data) andopcua_metadata(metadata) measurements. - influxdb — InfluxDB 2.x time-series database storing the telemetry.
Initialized with org
iot, bucketmqtt, and an admin user set to yourIOT_USERNAME. Includes a web UI (Data Explorer / dashboards). - grafana — Grafana dashboarding & alerting UI with a pre-provisioned
InfluxDB data source (Flux, org
iot, bucketmqtt) and three pre-provisioned dashboards ("Production Line OEE", "Modbus Simulator" and "UA Cloud Publisher Diagnostics") — see Pre-Provisioned Grafana Dashboards. - ua-cloudaction — OPC Foundation UA Cloud Action, the command & control
Requestor. Polls a configured InfluxDB field and, when it crosses a threshold,
publishes a
ua-action-request(MethodCall) to thecommandstopic for Cloud Commander to execute — closing the digital feedback loop. Also hosts a status web UI and the OPC UA Web API. - portainer — Portainer CE, a web UI to manage the K3s cluster (workloads,
logs, shells, events). Runs under a
cluster-admin-bound ServiceAccount. - ua-cloudlibrary — a self-hosted UA Cloud Library, the OPC Foundation's store of OPC UA Information Models. Running it locally means models can be resolved without reaching out to the public uacloudlibrary.opcfoundation.org, which matters for air-gapped installations or when storing private models. It can also be used to store EU Digital Product Passports, which is the use case leveraged here. It has a Web UI plus a REST API.
- cloudlib-postgres — PostgreSQL, the Cloud Library's backing store. It
holds both the relational data and the uploaded nodeset files (there is no
separate blob store), so it is the single source of truth for everything you
upload.
ClusterIPonly — never exposed on the node. - ua-dataprocessor — UA Data Processor, a headless worker that reads the OPC UA telemetry back out of InfluxDB and calculates a Product Carbon Footprint (PCF) and a Digital Battery Passport, publishing the results as OPC UA Information Models into the Cloud Library above.
- i3x4influx — an i3X server over InfluxDB. It exposes
the telemetry already in the
mqttbucket through the vendor-neutral i3X REST API, so clients can browse the data as an ISA-95 hierarchy and follow typed relationships instead of writing Flux. See Browsing the Data as a Graph (i3X). - ua-cloudai — an MCP server that makes the plant available to agentic AI applications such as Claude Desktop or VS Code. It does not read the historian itself; it fronts the i3X API and UA Cloud Action's OPC UA Web API and presents them as 14 tools. It is read-only — there is no write, method-call or actuation path. See Asking Questions with AI (MCP).
Configuration resources
| Resource | Kind | Purpose |
|---|---|---|
influxdb-auth |
Secret | Holds the INFLUX_TOKEN used by InfluxDB (admin), Telegraf (write), Grafana (query), and UA Cloud Action (query). Supplied at deploy time via ${INFLUX_TOKEN}. |
cloud-services-tls |
Secret | Certificate and key used by the .NET services to serve HTTPS. Created manually before applying cloud.yaml — see Enabling TLS. |
telegraf-conf |
ConfigMap | Telegraf configuration (MQTT inputs + InfluxDB output). |
mosquitto-conf |
ConfigMap | Mosquitto broker configuration (TLS listener, authentication, persistence). |
ua-cloudpublisher-settings |
ConfigMap | Seeds the Publisher's settings.json (broker connection, topics, metadata) and persistency.json (published nodes for the simulated line) on first start. |
modbus-simulator |
ConfigMap | The Python Modbus TCP simulation server run by the simulated device. |
modbus-thing-description |
ConfigMap | W3C WoT Thing Description seeded into UA Edge Translator so the Modbus device is onboarded as an OPC UA asset at startup. |
grafana-datasources, grafana-dashboard-provider, grafana-dashboards |
ConfigMaps | Provision the InfluxDB data source and the three dashboards (Production Line OEE, Modbus Simulator, UA Cloud Publisher Diagnostics). |
opcua-model-importer |
ConfigMap | Importer script for loading OPC UA Information Models from the UA Cloud Library. |
portainer-sa-clusteradmin / portainer-crb-clusteradmin |
ServiceAccount / ClusterRoleBinding | Grant Portainer in-cluster access to the K3s API server. |
cloudlib-postgres-auth |
Secret | PostgreSQL database name, user and password for the UA Cloud Library. Reuses ${IOT_USERNAME} / ${IOT_PASSWORD}. |
Data flow
Security note: you choose the credentials at deployment time via the
IOT_USERNAME/IOT_PASSWORDvariables (see Apply the Stack Manifests), used consistently across the Edge Translator, Cloud Publisher, Cloud Commander, Mosquitto, and InfluxDB for demo purposes. Mosquitto uses a self-signed TLS certificate generated at pod startup. Change these and use certificates from a trusted CA before any production or exposed deployment.
Raspberry Pi OS ships with the memory cgroup controller disabled, but K3s/containerd requires it. Enable it and reboot before installing K3s:
# NOTE: cmdline.txt must stay a SINGLE line - append, don't add a new line.
sudo sed -i '1 s/$/ cgroup_memory=1 cgroup_enable=memory/' /boot/firmware/cmdline.txt
# On Raspberry Pi OS older than Bookworm the file is /boot/cmdline.txt instead.
sudo reboot nowOnce the device has booted and been updated, install K3s
# Install a single-node K3s cluster (server + agent on the same node)
curl -sfL https://get.k3s.io | sh -
# Verify the node is Ready (may take ~30s)
sudo k3s kubectl get nodesTo use the standard kubectl command and the KUBECONFIG without sudo:
mkdir -p ~/.kube
sudo cp /etc/rancher/k3s/k3s.yaml ~/.kube/config
sudo chown "$(id -u):$(id -g)" ~/.kube/config
export KUBECONFIG=~/.kube/config
echo 'export KUBECONFIG=~/.kube/config' >> ~/.bashrc
kubectl get nodes -AK3s ships with the Traefik ingress controller and a built-in ServiceLB (klipper-lb) load balancer, so the
type: LoadBalancerServices in the manifest are reachable directly on the node's IP address.
-
Download
edge.yamlandcloud.yamlonto the device:curl -fsSLO https://raw.githubusercontent.com/OPCF-Members/Cloud-Initiative-Reference-Solution/main/edge.yaml curl -fsSLO https://raw.githubusercontent.com/OPCF-Members/Cloud-Initiative-Reference-Solution/main/cloud.yaml
-
Create the TLS certificate. Every .NET web UI and API in this solution is reachable from outside the cluster only over HTTPS, through ingresses that read a TLS Secret. Those Secrets are not created by the manifests, and the ingresses will not serve TLS until they exist — so do this before applying:
cd ~ # The subjectAltName must list every name you will browse to, or modern # clients reject the certificate even when the CN matches. IP=$(hostname -I | awk '{print $1}') openssl req -x509 -nodes -days 825 -newkey rsa:2048 \ -keyout tls.key -out tls.crt \ -subj "/CN=$IP" \ -addext "subjectAltName=IP:$IP,DNS:localhost,DNS:translator.plant.local,DNS:publisher.plant.local,DNS:cloudaction.plant.local,DNS:cloudlibrary.plant.local,DNS:i3x.plant.local,DNS:cloudai.plant.local" # The namespaces must exist before Secrets can be created in them. kubectl create namespace cloud --dry-run=client -o yaml | kubectl apply -f - kubectl create namespace edge --dry-run=client -o yaml | kubectl apply -f - # Secrets are namespace-scoped, so the same certificate goes in twice — once # for each ingress. kubectl create secret tls cloud-services-tls -n cloud --cert=tls.crt --key=tls.key kubectl create secret tls edge-services-tls -n edge --cert=tls.crt --key=tls.key # Keep tls.crt to hand out as the trust anchor; the key is now in the cluster. shred -u tls.key
ℹ️ For the full explanation — why TLS is terminated at the ingress, why the hostnames are needed, and how to rotate the certificate — see Enabling TLS.
-
Provide the deployment credentials and InfluxDB token. The manifests reference
${IOT_USERNAME},${IOT_PASSWORD}, and${INFLUX_TOKEN}, so set them and substitute them at apply time:# Choose the shared username/password used by the Edge Translator, # Cloud Publisher, Mosquitto broker, and InfluxDB. # NOTE: InfluxDB requires the password to be at least 8 characters. export IOT_USERNAME="myUsername" export IOT_PASSWORD="ChangeMe123" # Generate a random InfluxDB token (or supply your own) export INFLUX_TOKEN="$(openssl rand -hex 32)" # Substitute ONLY these variables and apply. Restricting the variable list is # important so envsubst does not touch the runtime shell variables (e.g. # $MOSQUITTO_USERNAME) used inside the container start-up commands. # Apply the cloud part first so the broker and database exist for the edge part. envsubst '${IOT_USERNAME} ${IOT_PASSWORD} ${INFLUX_TOKEN}' < cloud.yaml | kubectl apply -f - envsubst '${IOT_USERNAME} ${IOT_PASSWORD} ${INFLUX_TOKEN}' < edge.yaml | kubectl apply -f -
⚠️ Add the hostnames to your hosts file on whichever machine you browse from, or the*.plant.localURLs below will not resolve (replace the IP with your device's):192.168.1.50 translator.plant.local publisher.plant.local cloudaction.plant.local cloudlibrary.plant.local i3x.plant.local cloudai.plant.localThat file is
/etc/hostson Linux and macOS, andC:\Windows\System32\drivers\etc\hostson Windows (edit as Administrator).envsubstis part of thegettextpackage (sudo apt install -y gettext-base). Keep the values you chose — you'll reuseIOT_USERNAME/IOT_PASSWORDto log into the web UIs and the broker, and the generatedINFLUX_TOKENto authenticate Telegraf and log into InfluxDB via the API.Lost the token, or re-applying later? It is stored in the
influxdb-authSecret, so you can read it back out of the cluster rather than generating a new one:export INFLUX_TOKEN="$(kubectl get secret influxdb-auth -n cloud \ -o go-template='{{.data.INFLUX_TOKEN | base64decode}}')"
base64decoderuns inside the Go template, so this needs no externalbase64binary and works the same from Linux, macOS and Windows PowerShell.Always do this before re-applying the manifests. Generating a fresh token only rewrites the Secret — InfluxDB keeps the admin token it was initialised with, so Telegraf, Grafana and UA Cloud Action would all suddenly fail to authenticate against a database that never changed.
Deploying edge and cloud on separate clusters? The manifests reference each other by in-cluster DNS (
mosquitto.cloud.svc.cluster.local,influxdb.cloud.svc.cluster.local, andua-edgetranslator.edge.svc.cluster.local). If the two halves run in different clusters, replace those names with the externally reachable addresses of the remote services before applying. -
Watch the workloads come up (each part lives in its own namespace):
kubectl get pods,svc -n cloud kubectl get pods,svc -n edge # or watch everything at once kubectl get pods -A -wAll pods should reach
Running/Ready, and eachLoadBalancerService should receive anEXTERNAL-IP(the node's IP).If the external IP address for some Kubernetes services shows as
<pending>, use the following command to assign the external IP address of the traefik service: sudo kubectl patch service -p '{"spec": {"type": "LoadBalancer", "externalIPs":[""]}}'.
The ingested OPC UA telemetry is stored by InfluxDB in the mqtt bucket. That
database is backed by a Kubernetes hostPath volume, so the data lives directly
on the Pi's NVMe SSD at:
/influxdb2
Because this is a host directory (not ephemeral pod storage), the telemetry survives pod restarts, redeploys, and reboots.
Related OPC UA telemetry persistence paths are also mapped as hostPath volumes on the Pi:
| Path on the Pi | Component | Contents |
|---|---|---|
/influxdb2 |
InfluxDB | Time-series telemetry, buckets, and InfluxDB config (the primary telemetry store). |
/translator/settings, /translator/nodesets, /translator/pki, /translator/logs |
UA Edge Translator | Configuration, OPC UA nodesets, certificates, and logs. |
/publisher/settings, /publisher/pki, /publisher/logs, /publisher/store |
UA Cloud Publisher | Configuration, certificates, logs, and the Store & Forward message store (queued messages held during broker/connectivity outages). |
/commander/pki, /commander/logs |
UA Cloud Commander | OPC UA client certificates and logs. |
/productionline/munich/<station>/pki, /productionline/munich/<station>/logs |
Simulated production line | OPC UA server certificates and logs for each simulated station (mes, assembly, test, packaging). |
/mosquitto |
Mosquitto | Broker persistence database (mosquitto.db: retained messages and queued messages for persistent sessions). |
/portainer |
Portainer | Portainer database, users, and settings. |
/grafana |
Grafana | Grafana database, users, and user-created dashboards. |
/cloudlib-postgres |
UA Cloud Library (PostgreSQL) | The Cloud Library's entire state: user accounts and every uploaded OPC UA nodeset. This directory is the only thing to back up — and deleting it discards every model you uploaded. |
/cloudlib-dpkeys |
UA Cloud Library | ASP.NET Core Data Protection key ring. These keys encrypt the login cookie, the antiforgery token on every form, and password-reset tokens. Without persisting them, every restart signs all users out and breaks the login form — the antiforgery token becomes undecryptable and the POST is rejected before the password is checked, which looks exactly like a wrong password. |
Note: Keep the
INFLUX_TOKENsafe, to read the telemetry stored in InfluxDB in backup scenarios.If you no longer have it, retrieve it from the cluster — it is held in the
influxdb-authSecret:kubectl get secret influxdb-auth -n cloud -o go-template='{{.data.INFLUX_TOKEN | base64decode}}'; echo
base64decoderuns inside the Go template, so this needs no externalbase64binary and works the same from Linux, macOS and Windows PowerShell. Use the token to query the API directly, for example to list the buckets:TOKEN=$(kubectl get secret influxdb-auth -n cloud -o go-template='{{.data.INFLUX_TOKEN | base64decode}}') kubectl exec -n cloud deploy/influxdb -- influx bucket list --org iot --token "$TOKEN"This is an all-access admin token: it can read and delete every bucket in the org. Treat it as a secret and see Production Hardening Recommendations for scoping it down.
All containers use imagePullPolicy: IfNotPresent so the stack can be
deployed and restarted on an air-gapped network without reaching a registry.
Most images are pinned to an explicit version (grafana/grafana:13.1.1,
influxdb:2.9, …), so bumping one means editing the tag in the manifest and
re-applying — the new tag is not in the local cache, so K3s fetches it.
The following four OPC Foundation and two external components are different. They track a floating
:main tag:
ghcr.io/opcfoundation/ua-edgetranslator:main
ghcr.io/opcfoundation/ua-edgetranslator-drivers:main
ghcr.io/opcfoundation/ua-cloudcommander:main
ghcr.io/opcfoundation/ua-cloudaction:main
ghcr.io/barnstee/ua-cloudpublisher:main
ghcr.io/digitaltwinconsortium/manufacturingontologies:main
Because the tag never changes, IfNotPresent means K3s keeps using the copy it
already has. A kubectl rollout restart will not pick up a new build — it
re-creates the pod from the same cached image.
Updating therefore takes two steps, in this order: scale the workload down so
the image is no longer in use, then delete it. crictl refuses to remove an
image that a running container references, so deleting first silently does
nothing.
# 1. stop the workload so its image becomes unused
kubectl scale deployment/ua-cloudpublisher -n edge --replicas=0
kubectl wait --for=delete pod -l app=ua-cloudpublisher -n edge --timeout=120s
# 2. drop the cached image
sudo k3s crictl --timeout=5m rmi ghcr.io/barnstee/ua-cloudpublisher:main
# 3. start it again - the image is gone, so this pulls the current :main
kubectl scale deployment/ua-cloudpublisher -n edge --replicas=1
kubectl rollout status deployment/ua-cloudpublisher -n edge
sudo k3s crictl rmi --pruneis not a shortcut for this. It removes only images that no running container is using, so every image you actually care about updating is skipped. It is useful for reclaiming disk space after a version bump has left old tags behind, not for refreshing a running component.
To rebuild the whole stack on current images, delete the namespaces first — then nothing is running and a prune clears everything:
kubectl delete namespace cloud edge munich --ignore-not-found
kubectl get ns -w # Ctrl+C once all three have gone
sudo k3s crictl --timeout=5m rmi --prune # now genuinely removes every stack image
envsubst '${IOT_USERNAME} ${IOT_PASSWORD} ${INFLUX_TOKEN}' < cloud.yaml | kubectl apply -f -
envsubst '${IOT_USERNAME} ${IOT_PASSWORD}' < edge.yaml | kubectl apply -f -Updating obviously requires registry access — on an air-gapped node, side-load the new image with
sudo k3s ctr images import <file>.tarinstead, which replaces the cached copy without needing to delete it first.
Deleting the three namespaces stops and removes every workload:
kubectl delete namespace cloud edge munich --ignore-not-found
kubectl get ns -w # Ctrl+C once all three have goneNamespace deletion occasionally stalls on finalizers. If one sits in
Terminatingfor more than a minute or two, inspect what is left withkubectl get all -n <namespace>before forcing anything.
This does not delete your data. Every path in the table under
Where Telemetry Data Is Persisted is a
hostPath on the Pi and survives. Re-applying the manifests brings the stack
back up with the same telemetry, certificates, users and dashboards.
To start genuinely from scratch — new certificates, empty database, fresh admin accounts — also delete the host directories:
sudo rm -rf /mosquitto /influxdb2 /portainer /grafana /cloudlib-postgres /cloudlib-dpkeys
sudo rm -rf /translator /publisher /commander /productionline
⚠️ This is irreversible. It destroys all recorded telemetry, the InfluxDB admin token, the Mosquitto CA and password file, every OPC UA certificate and trust list — the stations, Publisher, Translator and Commander all mint new identities and re-establish trust on the next start — and every nodeset uploaded to the UA Cloud Library, together with its user accounts. Save yourINFLUX_TOKENfirst if you still need to read the old data.
To return the Pi to a plain OS install:
sudo /usr/local/bin/k3s-uninstall.shThat stops the service, removes the binary, the cluster state under
/var/lib/rancher/k3s, and all cached container images. It does not touch
the hostPath directories above (delete them separately, as shown), nor the
cgroup_memory=1 cgroup_enable=memory parameters added to
/boot/firmware/cmdline.txt during
Install K3s — harmless to leave in place, but remove them by hand
if you want the boot configuration back exactly as it was.
So that the stack produces meaningful OPC UA telemetry immediately — without any
physical machines — edge.yaml also deploys a software-only factory simulation.
One production line, named Munich, is deployed into its own munich
namespace. It consists of four OPC UA servers:
| Station | Role | OPC UA endpoint |
|---|---|---|
| mes | Manufacturing Execution System — drives the shift schedule (Morning / Afternoon / Night) for the line. | opc.tcp://mes.munich/ |
| assembly | Assembly station (200 W, 6 s cycle time). | opc.tcp://assembly.munich/ |
| test | Test station (100 W, 6 s cycle time). | opc.tcp://test.munich/ |
| packaging | Packaging station (100 W, 6 s cycle time). | opc.tcp://packaging.munich/ |
Each station simulates a real machine, exposing OPC UA variables such as production status, pressure, energy consumption, and product counts, and it implements OPC UA methods (e.g. opening a pressure relief valve) that the command & control path can invoke.
The stations run in the
munichnamespace on purpose: their in-cluster DNS names (mes.munich,assembly.munich, …) then match the OPC UA application URIs the stations advertise, so the endpoint URLs in the Publisher's configuration resolve without modification.
The line follows a three-shift schedule and is idle outside those shifts, so telemetry pauses during the daily breaks. This matters when reading OEE — see Production Shifts and Choosing the Grafana Time Range.
UA Cloud Publisher is pre-seeded with a published-nodes persistency file
listing the nodes to subscribe to on each station, plus the
Modbus variables that UA Edge Translator maps into OPC UA. Because the seeded
settings.json sets AutoLoadPersistedNodes: true, the Publisher loads this
list on startup and begins publishing OPC UA PubSub messages to Mosquitto right
away — telemetry appears in InfluxDB and Grafana without any manual onboarding.
To guarantee correct start-up order, the Publisher pod runs two init containers
that block until their dependencies are accepting OPC UA connections on port 4840:
wait-for-productionline (all four stations) and
wait-for-edgetranslator (the Edge Translator, which serves the mapped
Modbus asset):
# watch the simulation come up
kubectl get pods -n munich -w
# the Publisher stays in Init: until the line and the translator are ready
kubectl get pods -n edge
kubectl logs -n edge deploy/ua-cloudpublisher -c wait-for-productionline
kubectl logs -n edge deploy/ua-cloudpublisher -c wait-for-edgetranslatorBoth seeded files are only copied if they are not already present, so any changes you later make through the Publisher UI are preserved across restarts.
The Munich line also includes a simulated Modbus TCP device — a "line conditioning unit" — to demonstrate the other half of the story: bringing a non-OPC UA asset into the OPC UA world without writing any code.
It is a small, dependency-free Modbus TCP server (Python standard library only,
so it runs on arm64 and offline) exposing continuously changing registers at
modbus-simulator.munich.svc.cluster.local:502, unit id 1:
| Register | Address | Modbus type | Value |
|---|---|---|---|
| Temperature | Holding 0–1 | float32 | Process temperature (°C) |
| Pressure | Holding 2–3 | float32 | Process pressure (bar) |
| FlowRate | Holding 4–5 | float32 | Coolant flow (l/min) |
| EnergyConsumption | Holding 6–7 | float32 | Cumulative energy (kWh) |
| MotorSpeed | Holding 8 | int16 | Motor speed (rpm) |
| MachineState | Holding 9 | int16 | 0 = stopped, 1 = running, 2 = fault |
| Running | Coil 0 | bool | True while running |
| FaultActive | Coil 1 | bool | True during a high-pressure fault |
The device is described by a W3C WoT Thing Description shipped in the
modbus-thing-description ConfigMap. UA Edge Translator loads every *.jsonld
file in its settings folder at startup and onboards it as an OPC UA asset, so
the Modbus registers appear as browsable, subscribable OPC UA variables on
opc.tcp://<device-ip>:4840 with no manual configuration.
The TD carries the Modbus binding on each property's forms entry:
As with the Publisher, the Edge Translator pod runs two init containers: one
waits for the Modbus simulator to accept connections, and one seeds the Thing
Description into /translator/settings — only if it is not already there, so
assets you add or edit through the Edge Translator UI survive restarts.
# watch the simulator and the translator come up
kubectl get pods -n munich -l app=modbus-simulator
kubectl logs -n edge deploy/ua-edgetranslator -c seed-thing-descriptionsUA Edge Translator registers one OPC UA namespace per onboarded asset, derived
from the Thing Description's name:
http://opcfoundation.org/UA/<td.name>/
Each mapped property becomes a variable in that namespace with a string NodeId equal to the property name. So the simulator's registers are addressable as:
| Property | OPC UA NodeId |
|---|---|
| Temperature | nsu=http://opcfoundation.org/UA/modbus-simulator/;s=Temperature |
| Pressure | nsu=http://opcfoundation.org/UA/modbus-simulator/;s=Pressure |
| FlowRate | nsu=http://opcfoundation.org/UA/modbus-simulator/;s=FlowRate |
| EnergyConsumption | nsu=http://opcfoundation.org/UA/modbus-simulator/;s=EnergyConsumption |
| MotorSpeed | nsu=http://opcfoundation.org/UA/modbus-simulator/;s=MotorSpeed |
| MachineState | nsu=http://opcfoundation.org/UA/modbus-simulator/;s=MachineState |
| Running | nsu=http://opcfoundation.org/UA/modbus-simulator/;s=Running |
| FaultActive | nsu=http://opcfoundation.org/UA/modbus-simulator/;s=FaultActive |
Because each asset is isolated in its own namespace, two devices can expose identically-named properties without colliding.
These NodeIds are already listed in the Publisher's seeded persistency.json
against the Edge Translator endpoint
(opc.tcp://ua-edgetranslator.edge.svc.cluster.local:4840), so the Modbus data
flows all the way through to InfluxDB and Grafana automatically — a non-OPC UA
device published as OPC UA PubSub with zero manual configuration, enabling
fully automatic asset onboarding!
All eight tags are charted out of the box on the pre-provisioned Modbus Simulator dashboard — see Pre-Provisioned Grafana Dashboards.
To onboard a real Modbus (or BACnet, S7, Rockwell, OPC DA, …) device, see Onboarding a Non-OPC UA Device and the additional examples in the UA Edge Translator samples.
Don't want the simulation? Delete the
munichnamespace (kubectl delete namespace munich) and remove thewait-for-productionline/wait-for-modbusinit containers, thepersistency.jsonentries, and themodbus-thing-descriptionConfigMap fromedge.yaml, then onboard your real devices as described in Onboarding an OPC UA Device.
OPC UA is secure by default: a client and a server will only talk to each other once they mutually trust each other's X.509 certificates. Normally that means manually copying certificates into each server's trust list before publishing can start.
The seeded UA Cloud Publisher configuration enables the GDS Server Push
feature ("PushCertsBeforePublishing": true), which automates this entirely.
UA Cloud Publisher acts as a lightweight Global Discovery Server (GDS) and
uses the OPC UA Server Push Configuration interface (IEC 62541-12) to
provision certificates into each OPC UA server it is about to publish from.
Whenever the Publisher is about to process a published-nodes / persistency.json
file (or upload a WoT file to the Edge Translator), it performs the following
against each target OPC UA server:
- Connects to the server's endpoint using administrator credentials — the
ones stored with the endpoint, falling back to the
OPCUA_USERNAME/OPCUA_PASSWORDenvironment variables (i.e. yourIOT_USERNAME/IOT_PASSWORD). - Requests a Certificate Signing Request (CSR) from the server, asking it to
regenerate its private key (rather than reuse the existing one — older
sub-2048-bit keys are rejected by modern servers with
BadCertificatePolicyCheckFailed). - Signs the CSR with the Publisher's own issuer (CA) certificate.
- Pushes the new certificate and the issuer chain back to the server
(
UpdateCertificate). - Adds the server's new certificate to the Publisher's own trust list, so the Publisher keeps trusting the server.
- Pushes the Publisher's trust list to the server (
UpdateTrustList) so the server trusts the Publisher in return. - Applies the changes on the server and disconnects.
The result is a fully automated, mutually trusted, certificate-based OPC UA security relationship — no manual certificate exchange required. This is why the simulated production line starts streaming data as soon as it is up, and why onboarding real OPC UA devices usually needs no manual trust step.
- The Publisher UI's Browse view has a Push Certificate action to trigger a GDS push against the currently connected server on demand.
- The Cert Manager page lets you inspect the Publisher's trust list, download it as a ZIP, and add/remove trusted certificates.
- The behaviour is toggled by "Push new OPC UA certificates to server before WoT
file upload or before processing published nodes files (GDS Server Push
feature)" on the Configuration page (the
PushCertsBeforePublishingsetting).
Requirements & caveats: the target server must implement the OPC UA Server Push Configuration model and the supplied credentials must map to a role allowed to update certificates (typically
SecurityAdmin). Servers that don't support push, or reject the admin credentials, simply log aGDS server push failederror — you then fall back to exchanging certificates manually. Note that pushing replaces the server's certificate with one issued by the Publisher's CA, which is appropriate for this reference deployment but should be reviewed against your PKI policy in production.
Most services in this solution authenticate with HTTP Basic, which sends
reversible credentials on every single request. Over plain HTTP anyone on the
network path can read and replay them, so every .NET web UI and API in this
solution is not exposed on the node at all. They are ClusterIP only, and
the single way in from outside is a TLS-terminating Traefik ingress:
| Service | Namespace | In-cluster (ClusterIP, HTTP) | External (HTTPS via ingress) |
|---|---|---|---|
| UA Edge Translator | edge |
ua-edgetranslator-ui:8080 |
https://translator.plant.local |
| UA Cloud Publisher | edge |
ua-cloudpublisher:8081 |
https://publisher.plant.local |
| UA Cloud Action | cloud |
ua-cloudaction:8082 |
https://cloudaction.plant.local |
| UA Cloud Library | cloud |
ua-cloudlibrary:8083 |
https://cloudlibrary.plant.local |
| i3X for InfluxDB | cloud |
i3x4influx:8084 |
https://i3x.plant.local |
| UA Cloud AI | cloud |
ua-cloudai:5050 |
https://cloudai.plant.local |
Because the HTTP ports are ClusterIP, k3s never binds them on the node IP — so
there is no plain-HTTP port to reach from the LAN, and no way to send credentials
in the clear even by mistake.
ℹ️ There are two Ingresses, one per namespace. A Kubernetes Ingress can only route to Services in its own namespace, so
cloud.yamlandedge.yamleach carry their own. That also keepsedge.yamldeployable on its own, which is the point of the edge/cloud split.
ℹ️ Only HTTP is routed this way. The Edge Translator also listens for OPC UA (
4840), LoRaWAN (5000/5001) and OCPP (19520/19521). Those are raw TCP, not HTTP, so an HTTP ingress cannot carry them and they remain on the node IP. They have their own security: OPC UA uses certificate-based transport security (see Automatic Certificate Provisioning), and LoRaWAN and OCPP offer secure variants on5001and19521. MQTT is already TLS on8883.
ℹ️ Why TLS is terminated at the ingress rather than inside each app. UA Cloud Action, the UA Cloud Library and i3X all call
UseHttpsRedirection()unconditionally. If they bound an HTTPS port themselves, they would answer every plain-HTTP request with a307redirect to HTTPS — including requests from other pods, which do not trust the self-signed certificate. The UA Data Processor's Digital Product Passport uploads would fail withAuthenticationException: UntrustedRoot. Terminating at the ingress keeps the in-cluster paths on clean HTTP while everything external is encrypted.
ℹ️ Why hostnames rather than paths (
/cloudlibrary,/i3x, …). These apps emit root-relative asset URLs —~/css/site.cssrenders as/css/site.css— and none of them callsUsePathBase. Behind a path prefix the HTML loads but every stylesheet and script404s, leaving the UIs unstyled. Giving each app its own hostname keeps its root path intact.
⚠️ Applications behind the ingress must honourX-Forwarded-Proto. Traefik terminates TLS and forwards to the pod over plain HTTP. An app that does not read that header believes the request arrived over HTTP, and anything it derives from the scheme is then wrong:
- Absolute redirects downgrade the connection. ASP.NET Identity builds its login redirect from the observed scheme, so it would send the browser to
http://…/Identity/Account/Login— and the credentials typed there cross the network in the clear.- Per-client rate limits collapse. A limiter partitioned on the connection's remote address sees the ingress pod's IP for every caller, turning a per-client budget into one shared bucket.
- Blazor Server circuits fail. The websocket URI is derived from the request scheme, so the client is told to open
ws://from anhttps://page and the browser blocks it as mixed content.Each .NET app in this solution therefore calls
UseForwardedHeaders()as its first middleware, withKnownIPNetworks/KnownProxiescleared because the ingress pod's address is not known in advance. Clearing that allow-list means the app trusts these headers from any caller, which is only safe because these services areClusterIP-only and unreachable except through the ingress. If you ever expose one directly, restore the allow-list — otherwise a client can spoof its own source IP and apparent scheme.
These names must point at the device. The simplest option is a hosts file entry on each machine that needs access:
192.168.1.50 translator.plant.local publisher.plant.local cloudaction.plant.local cloudlibrary.plant.local i3x.plant.local cloudai.plant.local
That file is /etc/hosts on Linux and macOS, and
C:\Windows\System32\drivers\etc\hosts on Windows (edit as Administrator). Use
your own DNS instead if you have one.
Traefik ships with k3s and is enabled by default, so there is nothing extra to install.
The cloud-services-tls Secret is not created by cloud.yaml. It is
generated as step 2 of Apply the Stack Manifests,
before the manifests are applied — so if you followed the deployment steps you
already have it. For reference, those commands are:
cd ~
# The subjectAltName must list every name you will browse to, or modern clients
# reject the certificate even when the CN matches.
IP=$(hostname -I | awk '{print $1}')
openssl req -x509 -nodes -days 825 -newkey rsa:2048 \
-keyout tls.key -out tls.crt \
-subj "/CN=$IP" \
-addext "subjectAltName=IP:$IP,DNS:localhost,DNS:translator.plant.local,DNS:publisher.plant.local,DNS:cloudaction.plant.local,DNS:cloudlibrary.plant.local,DNS:i3x.plant.local,DNS:cloudai.plant.local"
# Kubernetes Secrets are namespace-scoped, so the same certificate is needed in
# both namespaces - one for each Ingress.
kubectl create secret tls cloud-services-tls -n cloud --cert=tls.crt --key=tls.key
kubectl create secret tls edge-services-tls -n edge --cert=tls.crt --key=tls.keyThen remove the private key from the device, keeping only tls.crt to hand out
as the trust anchor:
shred -u tls.keyBrowsers will warn on first visit because the certificate is self-signed. You can
click through it, but the better fix is to
trust the certificate on your client machines
— dismissing that warning repeatedly trains you to ignore exactly the message
that would appear during a real man-in-the-middle attack. Until you do, pass -k
to command-line clients:
curl -k https://i3x.plant.local/v1/infoTo verify properly instead of skipping the check, use the certificate you kept:
curl --cacert tls.crt https://i3x.plant.local/v1/infoTyping a bare hostname is fine: browsers try http:// first, and the ingress
answers with a permanent redirect to https://. The redirect is issued before
the request reaches the application, so no credentials are ever transmitted
over the cleartext connection. You can see it with:
curl -kIs http://i3x.plant.local/ | head -2
# HTTP/1.1 301 Moved Permanently
# Location: https://i3x.plant.local/Confirm the plain-HTTP ports really are unreachable from the LAN — each of these should fail to connect rather than return data:
curl -sS --max-time 5 http://<device-ip>:8082/ # UA Cloud Action
curl -sS --max-time 5 http://<device-ip>:8083/ # UA Cloud Library
curl -sS --max-time 5 http://<device-ip>:8084/v1/info # i3X
curl -sS --max-time 5 http://<device-ip>:5050/health # UA Cloud AIFrom inside the cluster those same services are still plain HTTP, which is what keeps the internal call paths working:
kubectl run -n cloud probe --rm -it --restart=Never --image=curlimages/curl -- \
curl -sS http://i3x4influx:8084/v1/info -u "$IOT_USERNAME:$IOT_PASSWORD"Until a client trusts the certificate, every browser shows a full-page "Your connection is not private" warning, and command-line and programmatic clients fail outright. Clicking through the browser warning each time is not just irritating — it teaches you to dismiss exactly the warning that would appear during a real attack. Install the certificate once instead.
First copy tls.crt off the device (this is the public certificate, safe to
distribute — the private key never leaves the Pi):
scp pi@<device-ip>:~/tls.crt .ℹ️ If you already ran
shred -u tls.keyand no longer havetls.crteither, fetch it back from the cluster:kubectl get secret cloud-services-tls -n cloud \ -o jsonpath='{.data.tls\.crt}' | base64 -d > tls.crt
Windows — installs for every user and every browser except Firefox. Run in an elevated PowerShell:
Import-Certificate -FilePath .\tls.crt -CertStoreLocation Cert:\LocalMachine\RootmacOS:
sudo security add-trusted-cert -d -r trustRoot \
-k /Library/Keychains/System.keychain tls.crtLinux (Debian/Ubuntu):
sudo cp tls.crt /usr/local/share/ca-certificates/plant-local.crt
sudo update-ca-certificatesFirefox keeps its own trust store and ignores the operating system's. Go to
Settings → Privacy & Security → Certificates → View Certificates → Authorities →
Import, select tls.crt, and tick "Trust this CA to identify websites".
Verify it worked — this should now succeed without -k:
curl https://i3x.plant.local/v1/info -u "$IOT_USERNAME:$IOT_PASSWORD"Node does not use the operating system's certificate store by default — it
ships its own bundled list of certificate authorities. Installing the certificate
as above therefore fixes browsers and curl but not MCP Inspector, which
fails with fetch failed and DEPTH_ZERO_SELF_SIGNED_CERT.
Point Node at the certificate explicitly:
$env:NODE_EXTRA_CA_CERTS = "C:\path\to\tls.crt"
npx @modelcontextprotocol/inspectorexport NODE_EXTRA_CA_CERTS=/path/to/tls.crt
npx @modelcontextprotocol/inspectorThe file must be PEM (it begins -----BEGIN CERTIFICATE-----), which is what
kubectl and openssl produce. A DER-encoded .crt is silently ignored.
For Claude Desktop, set the same variable in the server's env block so the
bridged process inherits it:
"ua-cloudai": {
"command": "npx",
"args": ["-y", "mcp-remote", "https://cloudai.plant.local/mcp", "--header", "..."],
"env": { "NODE_EXTRA_CA_CERTS": "C:\\path\\to\\tls.crt" }
}
⚠️ You will findNODE_TLS_REJECT_UNAUTHORIZED=0suggested for this. It works, but it disables certificate verification for every connection that Node process makes — not just yours. Use it only in a throwaway shell for a one-off test, and never set it permanently or in a service definition.
kubectl delete secret cloud-services-tls -n cloud
kubectl delete secret edge-services-tls -n edge
# ...re-create both as above. Traefik picks the new certificate up automatically,
# because it reads the Secret rather than mounting it into each pod.
⚠️ A self-signed certificate encrypts traffic but does not prove identity. It stops passive eavesdropping on your credentials, which is the main risk on a shared LAN. It does not stop an active attacker who can intercept and present their own certificate, because nothing independently vouches for this one. Issue certificates from your own CA for anything beyond a demonstration.
ℹ️ Not everything is covered. InfluxDB, Grafana, MQTT Explorer and the Portainer HTTP port still serve plain HTTP on the node IP, and each is configured differently. Mosquitto already uses TLS on
8883. The Edge Translator's OPC UA, LoRaWAN and OCPP ports are not HTTP and so cannot go through this ingress at all. See the STRIDE analysis for what that leaves exposed.
The services fall into two groups.
Reached over HTTPS through the ingress. These are ClusterIP services with
no port on the node at all — the hostnames must resolve to the device, so add
them to your hosts file first (see Enabling TLS). The
certificate is self-signed, so your browser will warn on first visit.
| Service | URL | Notes |
|---|---|---|
| UA Edge Translator | https://translator.plant.local |
Configure southbound asset connections and the OPC UA information model. Log in with the IOT_USERNAME / IOT_PASSWORD you set (exposed via the manifest OPCUA_USERNAME / OPCUA_PASSWORD env vars). |
| UA Cloud Publisher | https://publisher.plant.local |
Configure which OPC UA nodes to publish and the MQTT broker target (mosquitto.cloud.svc.cluster.local:8883, TLS). Log in with the IOT_USERNAME / IOT_PASSWORD you set (exposed via the manifest PUBLISHER_USERNAME / PUBLISHER_PASSWORD env vars). |
| UA Cloud Action | https://cloudaction.plant.local |
Status UI for the automated feedback loop (data-source, broker, and Commander connectivity) and OPC UA Web API. Log in with the IOT_USERNAME / IOT_PASSWORD you set (see Automated Feedback Loop with UA Cloud Action). |
| UA Cloud Library | https://cloudlibrary.plant.local |
Web UI for the self-hosted store of OPC UA Information Models and Digital Product Passports — browse, search, upload and download nodesets, and explore the REST API. On first use you must register an account using your IOT_USERNAME and a strong password of your choosing, or the library will appear empty; see First Login. |
| i3X for InfluxDB | https://i3x.plant.local/swagger |
Swagger UI for the i3X REST API over the telemetry in InfluxDB — browse the data as an ISA-95 hierarchy, follow typed relationships, and read current or historical values without writing Flux. The Swagger page itself needs no login (it is exempt from authentication), but Authorize with your IOT_USERNAME / IOT_PASSWORD before calling any endpoint. See Browsing the Data as a Graph (i3X). |
| UA Cloud AI | https://cloudai.plant.local/health |
Not a web UI — this is an MCP endpoint for agentic AI applications, and the URL shown is only a health check that returns {"status":"ok"}. The endpoint itself is /mcp, which answers POST only and returns 405 to a browser. Point Claude Desktop, VS Code or an MCP test client at it and authenticate with your IOT_USERNAME / IOT_PASSWORD. See Asking Questions with AI (MCP). |
Reached directly on the node IP. Replace <device-ip> with the CM5's
address (from ip addr or kubectl get svc). These are still LoadBalancer
services, and apart from Portainer they are not encrypted — they are
third-party components that each configure TLS differently, so they are left as
an exercise; see
Production Hardening Recommendations.
| Service | URL | Notes |
|---|---|---|
| InfluxDB | http://<device-ip>:8086 |
Time-series UI, Data Explorer, and dashboards. Log in with the IOT_USERNAME / IOT_PASSWORD you set (org iot, bucket mqtt). |
| Grafana | http://<device-ip>:3000 |
Dashboards & alerting. Log in with the IOT_USERNAME / IOT_PASSWORD you set. The InfluxDB data source and three dashboards (Production Line OEE, Modbus Simulator, UA Cloud Publisher Diagnostics) are pre-provisioned (see Pre-Provisioned Grafana Dashboards). |
| MQTT Explorer | http://<device-ip>:4000 |
Web UI for the Mosquitto broker — browse the live topic tree, inspect the OPC UA PubSub payloads on data/# and metadata, and publish messages by hand (handy for driving UA Cloud Commander on commands). The broker connection is pre-provisioned — just press Connect; see Inspecting the Broker with MQTT Explorer. |
| Portainer | https://<device-ip>:9443 |
Kubernetes management UI for the K3s cluster. On first access you set the admin password (see Managing the Cluster with Portainer). |
🔒 Why the first group has no port number. Those six are
ClusterIPservices behind a TLS-terminating ingress, so their plain-HTTP ports are not bound on the node IP at all — yourIOT_USERNAME/IOT_PASSWORDcannot be sent in cleartext by accident. See Enabling TLS.
ℹ️ The Edge Translator's protocol ports are unchanged. Only its web UI moved behind the ingress. OPC UA (
4840), LoRaWAN (5000/5001) and OCPP (19520/19521) are raw TCP, not HTTP, so an HTTP ingress cannot carry them — they stay on the node IP. OPC UA has its own transport security (see Automatic Certificate Provisioning), and LoRaWAN and OCPP have secure variants on5001and19521.
Portainer CE provides a web UI to inspect and manage everything running on the
single-node K3s cluster (deployments, pods, logs, container shells, events, and
volumes). It is deployed by cloud.yaml and is pre-wired to manage the
cluster it runs in — no manual endpoint configuration is required (Click on Home -> Live connect after setting the admin password).
How the K3s connection works:
- The manifest creates a
portainer-sa-clusteradminServiceAccount and a ClusterRoleBinding to the built-incluster-adminrole, then runs the Portainer pod under that ServiceAccount. Portainer therefore talks to the K3s API server in-cluster using the mounted ServiceAccount token — it manages the local Kubernetes environment out of the box. - Portainer data (users, settings) is persisted on the Pi at
/portainer.
On a fresh installation Portainer only allows the initial admin account to be created within a few minutes of first start. If you browse to it later than that, the setup page is replaced by:
New Portainer installation — Your Portainer instance timed out for security purposes. To re-enable your Portainer instance, you will need to restart Portainer.
This is a deliberate safeguard: it stops a publicly reachable, unclaimed instance from being taken over by whoever finds it first. Restart the pod to reopen the window, then create the admin account promptly:
kubectl rollout restart -n cloud deployment/portainer
kubectl rollout status -n cloud deployment/portainerReload https://<device-ip>:9443 and the account creation page returns. The
timeout only applies until an admin user exists — once you have created it, normal
logins are not time limited.
Because Portainer's data lives on the node at
/portainer, the admin account survives pod restarts. If you ever need to start completely fresh (for example after losing the password), delete the deployment, remove that directory withsudo rm -rf /portainer, and re-applycloud.yaml.
This matters as soon as you move away from the single-node default:
| Topology | Does Portainer see the edge workloads? |
|---|---|
| One K3s instance (both manifests applied to the same cluster — the default) | Yes. The edge, munich, and cloud namespaces are all in the cluster Portainer runs in, so the in-cluster ServiceAccount covers them. Nothing else to do. |
| Separate edge and cloud clusters | No. A ServiceAccount token is only valid for its own cluster, so a Portainer Server running in the cloud cluster has no visibility of the edge cluster at all. |
For the split topology you need an extra component on the edge: the
Portainer Edge Agent, provided in
portainer-edge-agent.yaml.
The Edge Agent dials outbound to the Portainer Server's tunnel port (8000,
already published by the portainer Service in cloud.yaml), so the edge device
needs no inbound firewall rule and no public IP — the normal situation for an
industrial gateway behind NAT.
[Edge cluster] [Cloud cluster]
portainer-agent --- outbound tunnel :8000 ---> portainer (server)
(edge.yaml workloads) (cloud.yaml workloads)
To connect an edge cluster:
-
In the Portainer UI choose Environments → Add environment → Edge Agent → Kubernetes. Name it and copy the generated Edge ID and Edge key.
-
On the edge cluster, apply the agent manifest with those values:
curl -fsSLO https://raw.githubusercontent.com/OPCF-Members/Cloud-Initiative-Reference-Solution/main/portainer-edge-agent.yaml export PORTAINER_EDGE_ID="<edge id from the UI>" export PORTAINER_EDGE_KEY="<edge key from the UI>" envsubst < portainer-edge-agent.yaml | kubectl apply -f -
-
The environment turns green in the Portainer UI once the tunnel is established, and you can then manage the edge cluster alongside the cloud one.
Alternative: if the edge cluster is reachable from the cloud, you can use the standard (non-Edge) Portainer Agent instead and have the server connect inbound to it on port 9001. The Edge Agent is preferred for industrial deployments precisely because it avoids opening inbound ports on the edge.
Security: the agent manifest binds to
cluster-admin(same as the server) and setsEDGE_INSECURE_POLL=1because the demo server uses a self-signed certificate. Scope the role down and remove that flag once the Portainer Server has a trusted certificate — see Production Hardening Recommendations.
First-time setup:
- Browse to
https://<device-ip>:9443(accept the self-signed certificate warning) within a few minutes of the pod starting.For security, Portainer disables initial admin creation if you don't complete it shortly after startup. If you see a timeout message, restart the pod:
kubectl rollout restart deployment/portainer -n cloud. - Create the admin user and password.
- On the environments page, select the local Kubernetes environment (already connected via the in-cluster ServiceAccount) and click Live connect.
- You can now browse the
defaultnamespace to see the Edge Translator, Cloud Publisher, Cloud Commander, Mosquitto, Telegraf, and InfluxDB workloads, view their logs, exec into containers, and monitor cluster resources.
Mosquitto has no user interface of its own, so the stack deploys MQTT
Explorer as its web UI at http://<device-ip>:4000. Use it to see exactly what
is on the wire between UA Cloud Publisher, Telegraf and UA Cloud Commander.
No setup is required. The Mosquitto connection is pre-provisioned from the
mqtt-explorer-config ConfigMap in cloud.yaml and is already
filled in when the UI first loads — pick Mosquitto (this cluster) and press
Connect.
The seeded connection uses:
| Field | Value |
|---|---|
| Protocol | mqtts:// (TLS) |
| Host | mosquitto.cloud.svc.cluster.local |
| Port | 8883 |
| Username / Password | your IOT_USERNAME / IOT_PASSWORD |
| Validate certificate | off |
| Subscription | # (the whole topic tree) |
Certificate validation must be off: Mosquitto uses a self-signed certificate
generated on first start, which no public CA has signed. This is the same reason
Telegraf sets insecure_skip_verify and UA Cloud Action sets
MQTT_TLS_INSECURE.
Note: connections you edit in the UI are stored in an
emptyDir, so the pod returns to the provisioned connection after a restart. To change the default permanently, edit themqtt-explorer-configConfigMap instead.
Once connected you will see the live topic tree:
data/#— OPC UA PubSub telemetry from UA Cloud Publisher (one message per publishing cycle, carrying theMessages[]array Telegraf parses)metadata— theua-metadatamessages describing eachDataSetWriterId; this is the stream the Grafana panels resolve station names fromcommands/responses— the UA Cloud Commander request/response path
⚠️ MQTT Explorer has no authentication of its own. Unlike the other UIs in this stack it has no login, and anyone who can reach port 4000 can publish to any topic — includingcommands, which UA Cloud Commander will execute against your OPC UA servers. Keep it on a trusted network, place it behind an authenticating reverse proxy, or scale it to zero when it is not needed:kubectl scale deployment/mqtt-explorer -n cloud --replicas=0
The UA Cloud Library is the OPC Foundation's store of OPC UA Information Models, publicly hosted at uacloudlibrary.opcfoundation.org.
In this solution it is self-hosted and used as the store for EU Digital Product Passports (DPPs). A Digital Product Passport is a structured, machine-readable record of what a product is and what it cost the environment to make — its material composition, its carbon footprint, and its end-of-life characteristics — which the EU's Ecodesign for Sustainable Products Regulation progressively makes mandatory for products placed on the EU market. The DPP itself is specified by EN 18222 (data model and unique identifiers) and EN 18223 (system architecture and data exchange), the CEN/CENELEC standards that underpin the regulation — see Interoperability Through Open Standards.
Because a DPP is exactly the kind of structured, versioned, semantically described artefact that OPC UA Information Models already express, the Cloud Library works as a DPP repository without modification: UA Data Processor computes each product's carbon footprint and battery passport from live production telemetry and publishes them here as OPC UA Information Models, one per product. The Cloud Library then provides the storage, versioning, search and retrieval that a DPP repository needs, and its REST API is how downstream consumers — a customer, a recycler, or a regulator — fetch a given product's DPP.
Running your own instance also means DPPs and any proprietary models stay on your own infrastructure rather than in a public library, and that the stack keeps working with no dependency on the public Internet (see the air-gapped notes under Updating the Container Images).
Browse to https://cloudlibrary.plant.local. The UI lets you search and filter the stored
models, inspect their metadata and namespaces, download them, and upload your
own. The same data is available programmatically through a REST API.
ℹ️ This is a different thing from the model importer. The
opcua-model-importerJob pulls a model from a Cloud Library into InfluxDB so queries can resolve node metadata. The Cloud Library itself is the store — here, the DPP store.
Storage. Everything — the relational data and the uploaded models — lives
in the cloudlib-postgres PostgreSQL database backed by the hostPath
/cloudlib-postgres. That directory is the
only thing you need to back up, and removing it (as described under
Uninstalling) discards every Digital Product Passport and model you stored.
The Cloud Library server enables account confirmation only when an email sender API key is configured.
cloud.yaml deliberately does not set EmailSenderAPIKey, so
RequireConfirmedAccount evaluates to false and newly registered users can sign
in immediately. This is what makes the component usable in this reference deployment or on an
isolated network, where there is no outbound email service like SendGrid to deliver a
confirmation link and an unconfirmable account would lock you out of your own
deployment.
⚠️ The consequence is open self-registration. Anyone who can reach:8083can create a working account without proving they control an email address. That is acceptable for a reference deployment on a trusted network and not acceptable on an untrusted one. To restore verification, setEmailSenderAPIKeyfrom a public email service like SendGrid (andRegistrationEmailFrom/RegistrationEmailReplyTo) on theua-cloudlibraryDeployment.
The Cloud Library has no account until you create one. On first use, browse to
https://cloudlibrary.plant.local, choose Register, and sign up with:
| Field | Value |
|---|---|
| Username | exactly the IOT_USERNAME you deployed the solution with |
| Password | a strong password of your own choosing |
Registering under that specific name matters, because the Cloud Library is a multi-tenant service and filters
what you can see by who owns it. UA Data Processor authenticates as
IOT_USERNAME when it uploads, so an account with the same name sees every
Digital Product Passport the processor has produced.
Register under any other name (an email address, a different spelling) and the
library will look empty, even though the DPPs are stored and the uploads
are succeeding.
UA Data Processor closes the loop between raw telemetry and sustainability reporting. It reads the OPC UA data back out of InfluxDB and calculates:
- a Product Carbon Footprint (PCF) — by correlating each product's serial number across the assembly, test and packaging stations, summing the energy each station consumed while that product was inside it, and multiplying by the grid carbon intensity for the production line's location, and
- a Digital Battery Passport — the end-of-line dimensional and quality data for the produced item.
Together these form the Digital Product Passport (DPP) for each item produced. The results are published as OPC UA Information Models into the UA Cloud Library, which acts as the DPP store — so the DPP is generated from real production data rather than assembled by hand after the fact.
It is a headless worker with no web UI, so watch it with:
kubectl logs -f deployment/ua-dataprocessor -n cloudℹ️ WattTime is optional. Real grid carbon intensity comes from the WattTime service. Without
WATTTIME_USER/WATTTIME_PASSWORDthe lookup simply returns an average carbon intensity and the Battery Passport still works. Credentials are commented out incloud.yamlready to be filled in.
Everything up to this point treats the telemetry as time series: Grafana
charts it, the Data Processor queries it with Flux, and both need to know that a
station's status lives in the field Payload_Status_Value of measurement
opcua_pubsub, tagged with a datasetWriterId that must first be looked up in
opcua_metadata. That is precise, but it is storage-specific — the queries only
make sense against this InfluxDB schema.
i3X (Industrial Information Interoperability eXchange) is the specification for the other view: the same data as a connected graph. Clients browse an ISA-95 hierarchy — enterprise → site → area → production line → station — follow typed relationships between objects, and read current or historical values through one API, without knowing or caring which historian sits underneath.
The i3x4influx container serves that API from the telemetry already in the
mqtt bucket. Nothing extra is ingested and no second copy of the data is kept:
it maps i3X calls onto the same opcua_pubsub and opcua_metadata measurements
Telegraf writes and Grafana reads.
Why this matters for interoperability. A client written against i3X (or the OPC UA WebAPI) works against any conforming server. Swap InfluxDB for another historian and the dashboards and queries would have to be rewritten — an i3X client would not. That is the same argument as OPC UA at the edge, applied to the query layer.
The API is at https://i3x.plant.local, and it ships a built-in Swagger UI at
https://i3x.plant.local/swagger — the easiest way to explore it. The Swagger
page loads without credentials, but press Authorize and enter your
IOT_USERNAME / IOT_PASSWORD before invoking anything, or every call returns
401.
Browsing and discovery are GETs; the POST endpoints are bulk operations that
take a JSON body naming the elements to act on:
| Endpoint | Verb | What it does |
|---|---|---|
/swagger |
GET |
Built-in Swagger UI. Exempt from authentication, so the page loads without credentials — use its Authorize button before calling anything |
/v1/info |
GET |
Server information. Also exempt from authentication — useful for checking the service is up |
/v1/objects |
GET |
Browse the ISA-95 hierarchy. Add ?root=true for the top level, or ?typeElementId=… to filter by type |
/v1/namespaces |
GET |
List the OPC UA namespaces present in the data |
/v1/objecttypes |
GET |
List the available object types |
/v1/relationshiptypes |
GET |
List the available relationship types |
/v1/objects/list |
POST |
Bulk look up objects by id (elementIds) — not a browse |
/v1/objects/related |
POST |
Follow typed relationships from one object to others |
/v1/objects/value |
POST |
Read the current value of one or more objects |
/v1/objects/history |
POST |
Read historical values over a time range |
/v1/subscriptions/* |
POST |
Register, list and stream live updates via server-sent events |
Authentication is HTTP Basic with the IOT_USERNAME / IOT_PASSWORD you
deployed with. Start at the root of the hierarchy and walk down:
# is the service alive? (no credentials needed)
curl -s "http://localhost:8084/v1/info"
# the top of the ISA-95 tree
curl -s -u "$IOT_USERNAME:$IOT_PASSWORD" \
"http://localhost:8084/v1/objects?root=true"
# every object, including the variables at the leaves
curl -s -u "$IOT_USERNAME:$IOT_PASSWORD" \
"http://localhost:8084/v1/objects"Take an elementId from that output and use it with the POST endpoints, for
example to read a current value:
curl -s -u "$IOT_USERNAME:$IOT_PASSWORD" \
-X POST -H 'Content-Type: application/json' \
-d '{"elementIds":["<elementId from above>"]}' \
"http://localhost:8084/v1/objects/value"ℹ️ Reading an empty response. The
POSTendpoints are bulk operations driven by the ids in the request body, so sending{}returns an empty result rather than an error — it did exactly what was asked. If a call looks empty, check the body names someelementIds, and useGET /v1/objectsto discover them. Every route is also under/v1: omitting the prefix returns404, andcurlreporting000means no HTTP response at all — the pod is not Ready, so checkkubectl get pods -n cloud -l app=i3x4influxfirst.
ℹ️ Two time ranges control what you see.
INFLUX_BROWSE_RANGE(default-24h) is how far back the server looks when building the hierarchy, andINFLUX_LATEST_RANGE(default-1h) is how far back it looks for a current value. The browse range must comfortably exceed the idle gaps between production shifts, or stations that were quiet all night will simply not appear in the tree. See Production Shifts and Choosing the Grafana Time Range for when the line is actually running.
⚠️ Authentication is mandatory and fails closed. If neither Basic nor OAuth2 is configured the server returns 503 to every request rather than serving data anonymously.cloud.yamlsetsI3X_BASIC_AUTH_USERNAME/I3X_BASIC_AUTH_PASSWORDfor this reason. For production, prefer OAuth2 by settingI3X_OAUTH2_AUTHORITY,I3X_OAUTH2_AUDIENCEandI3X_OAUTH2_ISSUER— Basic auth over plain HTTP sends credentials in the clear on every call.
Everything above assumes a human driving a UI or writing a query. UA Cloud AI removes that assumption: it is an MCP (Model Context Protocol) server, the open standard agentic AI applications use to reach external systems. Point a model at it and you can ask "which station had the lowest OEE last shift, and what was its downtime?" in plain language.
It does not query the historian itself. It fronts the two APIs this solution already exposes and presents them as 14 tools:
| Group | Tools | What it is good for |
|---|---|---|
| Orientation | describe_available_data, check_connectivity |
Explains the two interfaces and which to use; reports each backend separately when something is wrong. |
| i3X (9 tools) | i3x_browse_hierarchy, i3x_get_related_objects, i3x_read_current_values, i3x_read_history, … |
The semantic graph — walking the ISA-95 hierarchy and following typed relationships. |
| OPC UA Web API (3 tools) | opcua_browse_nodes, opcua_read_values, opcua_read_history |
The raw address space — reading specific nodes by NodeId. |
ℹ️ Why an orientation tool? The two APIs overlap but their identifiers are not interchangeable — i3X uses
elementId, OPC UA usesNodeId. A model that guesses will silently pick the wrong one, sodescribe_available_datatells it which interface answers which kind of question before it starts.
The MCP endpoint is not a web page. It speaks JSON-RPC over HTTP POST, so
opening https://cloudai.plant.local/mcp in a browser returns:
405 Method Not Allowed
Allow: POST
That is the correct response, and it is a useful signal: reaching a 405 means
TLS, routing and your credentials all worked, because wrong credentials
return 401 instead.
To check the server in a browser, use the health endpoint instead — it answers
GET and needs no credentials:
https://cloudai.plant.local/health
{"status":"ok"}No extra tooling required. This performs the MCP initialize handshake:
curl -k -u "$IOT_USERNAME:$IOT_PASSWORD" \
-H 'Content-Type: application/json' \
-H 'Accept: application/json, text/event-stream' \
-d '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2024-11-05","capabilities":{},"clientInfo":{"name":"curl","version":"1.0"}}}' \
https://cloudai.plant.local/mcpA response containing "serverInfo":{"name":"UA-CloudAI"...} confirms the server
is working end to end.
Claude Desktop speaks stdio, launching the server itself. Add this to
claude_desktop_config.json (Settings → Developer → Edit Config), using
mcp-remote to bridge to the in-cluster HTTPS endpoint:
{
"mcpServers": {
"ua-cloudai": {
"command": "npx",
"args": [
"-y", "mcp-remote",
"https://cloudai.plant.local/mcp",
"--header", "Authorization:Basic <base64 of IOT_USERNAME:IOT_PASSWORD>"
]
}
}
}Generate the header value with:
printf '%s' "$IOT_USERNAME:$IOT_PASSWORD" | base64
⚠️ This needs Node.js installed (which providesnpx).npxfetches a package from the public npm registry and runs it without installing it permanently — so the machine running Claude Desktop needs internet access, andmcp-remoteis third-party code downloaded at launch. Check withnode --version; if it is missing, install Node.js first.On a locked-down or air-gapped network, avoid
npxentirely by running UA Cloud AI locally in stdio mode instead of bridging to the cluster — see the UA-CloudAI README, which shows aclaude_desktop_config.jsonthat launches the binary directly.
MCP Inspector lists and calls the tools by hand, which is the quickest way to confirm the server works:
npx @modelcontextprotocol/inspectorThis also requires Node.js, and prints a http://localhost:6274 URL to open. Set
Transport to Streamable HTTP, URL to https://cloudai.plant.local/mcp,
and add an Authorization header with the same base64 value as above. You should
see all 14 tools.
ℹ️ MCP Inspector runs on your machine, not the Pi. It therefore needs the
*.plant.localhostnames in its own hosts file, and must trust (or skip) the self-signed certificate. If it cannot connect, use thecurlhandshake above to confirm the server itself is healthy before debugging the client.
⚠️ The certificate is self-signed, hence-kabove. MCP Inspector andmcp-remotewill fail withfetch failed/DEPTH_ZERO_SELF_SIGNED_CERTuntil Node is told to trust it — Node ignores the operating system's certificate store, so installing the certificate system-wide is not enough. See Node.js clients.
🔒 UA Cloud AI is read-only by design. It browses and reads; there is no write, method-call or actuation path. An agent can analyse the plant but cannot change it. Note that this is a narrower boundary than the rest of the solution — UA Cloud Action and UA Cloud Commander can write back to the OPC UA servers, which is what the feedback loop depends on.
Three dashboards are provisioned automatically from the grafana-dashboards
ConfigMap in cloud.yaml and appear under Dashboards in
Grafana without any manual import.
| Dashboard | UID | What it shows |
|---|---|---|
| Production Line OEE | production-line-oee |
Line and per-station OEE, station status, cycle time and product counts for the simulated production line. Has a Station dropdown (assembly, test, packaging). This is also the Grafana home dashboard. |
| Modbus Simulator | modbus-simulator |
All 8 tags of the simulated Modbus TCP device, onboarded through UA Edge Translator (see Simulated Modbus TCP Device). |
| UA Cloud Publisher Diagnostics | publisher-diagnostics |
The Publisher's own health: broker connection, OPC UA session/subscription/monitored-item counts, queue depth, throughput, latency and failure counters. |
Every panel except the line gauge follows the Station dropdown, so the dashboard shows one station at a time:
| Panel | Scope |
|---|---|
| OEE - production line (bottleneck) | the whole line, ignores the dropdown |
| OEE - <station> | selected station |
| Status - <station> | selected station |
| Actual cycle time - <station> | selected station |
| Manufactured / Discarded products, Energy consumption, Pressure | selected station |
Line OEE is the OEE of the slowest station, not an average. On a serial line
(assembly → test → packaging) the stations are coupled: the slowest one
starves everything downstream and blocks everything upstream, so line throughput
is governed by the constraint. Averaging would hide the very station you need to
act on. This matches CalculateOEEForLine() in the
Manufacturing Ontologies
reference, which evaluates each station and then takes summarize min(oee).
The Status panels are stepped line charts rather than smooth ones, because
Status is a discrete enum — interpolating between points would draw the station
passing through states it never occupied:
| Value | State |
|---|---|
| 0 | Ready |
| 1 | WorkInProgress |
| 2 | Done |
| 3 | Discarded |
| 4 | Fault |
Note: the OEE figures here are calculated over the dashboard's time range, not over a fixed shift window. Selecting a range that spans a shift break will under-report OEE — see Production Shifts and Choosing the Grafana Time Range below before reading anything into the numbers.
The simulated line does not run around the clock. The MES station reads
ShiftTimes.csv and holds the line idle outside the configured shifts. The
munich line runs in Europe/Berlin (set via the FactoryTimeZone
environment variable in edge.yaml):
| Shift | Start | End |
|---|---|---|
| Morning | 07:00 | 14:00 |
| Afternoon | 15:00 | 22:00 |
| Night | 23:00 | 06:00 (next day) |
That leaves three daily idle gaps — 06:00-07:00, 14:00-15:00 and 22:00-23:00
— during which the stations report Ready and produce nothing.
The OEE panels compute Availability from the whole selected window:
availability = (windowLength - faultyTime) / windowLength
The simulation has no concept of "planned downtime", so an idle hour is not excluded — it is counted as time the line should have been producing. Selecting a range that includes a shift break therefore reduces OEE, and the more of a break you include, the lower it reads.
Choose a range that sits entirely inside one shift. In Grafana use the time picker's Absolute time range and enter, for example:
| Goal | From | To |
|---|---|---|
| Current morning shift | 07:00 |
14:00 |
| Last hour of production | now-1h |
now (only while a shift is running) |
| A full shift, yesterday | 2026-08-03 07:00:00 |
2026-08-03 14:00:00 |
Grafana renders in your browser's timezone by default. If that is not
Europe/Berlin, the shift boundaries will not fall where the table above says. Set the dashboard timezone explicitly via the time picker's Change time settings → Timezone, or your figures will silently include break time.
now-1h — Grafana's default — is only safe mid-shift. Run it at 14:30 and the
window is entirely inside the afternoon break, so Availability approaches zero
and OEE collapses. That is the calculation working correctly on an idle line, not
a fault.
The Publisher publishes its own health as OPC UA nodes under
nsu=http://opcfoundation.org/UA/CloudPublisher/ (see Diagnostics.cs), so
these values travel the same data/# → Telegraf → InfluxDB path as production
telemetry and need no extra configuration.
When telemetry stops arriving, these panels localise the fault quickly:
| Panel | What it tells you |
|---|---|
| Connected to broker | Whether the MQTT session is up at all. Stepped line, mapped to Connected / Disconnected. |
| OPC UA sessions / subscriptions / monitored items | Whether the Publisher is still attached to the source servers. A drop here means the problem is upstream of MQTT. |
| Internal queue depth | Back-pressure. Sustained growth means the Publisher is reading faster than it can send. |
| Enqueue failures | The queue hit InternalQueueCapacity and data was dropped. Any increase is data loss. |
| Broker messages / second, Monitored item notifications / second | Actual throughput, to compare against the configured publishing interval. |
| Average message size / latency | Broker round-trip health. |
| Send failures, Stored messages left to send | Broker-side trouble; stored messages accumulate while sending fails. |
| Working set | Publisher memory — worth watching on a CM5. |
Note: these counters are cumulative since Publisher start, so they reset on a restart. A flat line is normal: InfluxDB only stores changes, so a counter that stops moving simply stops producing points.
Step-by-step guides live in their own files to keep this README readable:
| Tutorial | What you will do |
|---|---|
| Onboarding an OPC UA Device | Connect UA Cloud Publisher to an OPC UA server and publish its nodes. |
| Onboarding a Non-OPC UA Device | Map a non-OPC UA asset into OPC UA with a W3C WoT Thing Description. |
| Querying Data in the InfluxDB Dashboard | Explore the telemetry with Flux queries and build InfluxDB dashboards. |
| Dashboards with Grafana | Use the pre-provisioned InfluxDB data source and dashboards. |
| Calculating OEE | Compute Availability, Performance, Quality and OEE per station and for the whole line, and chart it in Grafana. |
| Importing an OPC UA Information Model | Load a model from the UA Cloud Library into InfluxDB. |
| Command & Control with UA Cloud Commander | Send OPC UA Actions over MQTT and close the digital feedback loop. |
| Building Custom Apps for the Reference Solution | Use the OPC UA Web API and the UA Web API Starter Kit to build your own applications. |
This section applies the STRIDE threat-modeling framework (Spoofing, Tampering, Repudiation, Information disclosure, Denial of service, Elevation of privilege) to the reference stack. It is intended to help you understand the residual risks of the demo configuration and what to change before an internet-exposed or production deployment.
Important: the reference manifest is optimized for a self-contained, single-node demo. It ships with convenience defaults (shared credentials, a self-signed broker certificate generated at pod start,
LoadBalancerservices bound to the node IP for the third-party components, and permissive TLS verification in Telegraf). The .NET services are reachable only over HTTPS through the ingress, but with a self-signed certificate, and TLS terminates at the ingress so pod-to-pod traffic remains plain HTTP. These are not appropriate for production as-is — see Production Hardening Recommendations.
[Field devices] --(OPC UA / Modbus / LoRaWAN / OCPP / HTTP)--> [UA Edge Translator]
[Simulated line: mes/assembly/test/packaging (OPC UA :4840)] |
[Modbus simulator (Modbus TCP :502, no auth)] --------------------------+
| |
| Boundary A: device <-> edge | (OPC UA server :4840)
v v
[UA Cloud Publisher] --(MQTT/TLS :8883, user/pass)--> [Mosquitto] --(MQTT/TLS)--> [Telegraf] --(HTTP + token)--> [InfluxDB]
^ ^ ^ ^ ^
| [UA Cloud Commander] -+ | (commands/responses) | |
| [UA Cloud Action] --------+--(reads InfluxDB threshold, publishes commands)--------+ |
| [Grafana] -----+ (query token)
| [Model importer Job] ----+ (writes opcua_model)
| [UA Data Processor] --------+ (reads telemetry + metadata)
| [i3X for InfluxDB :8084] ---+ (reads telemetry + metadata)
| |
| (publishes PCF / Battery Passport models) v
| [UA Cloud Library :8083] --> [PostgreSQL :5432, ClusterIP]
|
| [AI agent / MCP client] --(MCP over HTTPS, basic auth)--> [UA Cloud AI] --+--> [i3X :8084]
| Boundary D: AI client <-> plant (read-only) +--> [UA Cloud Action Web API :8082]
|
| Boundary B: operator <-> web UIs (:8080/:8081/:8086/:3000/:9443, basic auth)
| UA Cloud Action, Cloud Library, i3X and UA Cloud AI are ClusterIP-only and reachable
| ONLY through the TLS-terminating ingress (cloudaction|cloudlibrary|i3x|cloudai .plant.local) |
+----------- Boundary C: node/cluster host (K3s + Portainer cluster-admin, hostPath volumes) -----------------------+
Key assets: the telemetry data (in transit and at rest in InfluxDB), the shared
IOT_USERNAME / IOT_PASSWORD credentials, the INFLUX_TOKEN, the UA Cloud
Library credentials used by the import Job, the self-hosted UA Cloud Library's
PostgreSQL database (which holds its user accounts and every stored Digital
Product Passport — a regulatory record whose integrity is the point of keeping
it), the broker's private key, the
Portainer cluster-admin ServiceAccount token (full control of the cluster), and
the K3s node itself (root of trust for all hostPath data). The i3X API
(i3x.plant.local) is a further read path to the same telemetry, so it inherits the value
of the data it exposes. UA Cloud AI is a read path on top of
both APIs, so it inherits the union of what they expose — and because it is
driven in natural language, it lowers the expertise needed to exploit that access.
Note also that Boundary D extends outside the cluster entirely whenever the
MCP client is a hosted AI service, since tool results are sent onward for
inference.
Each category below lists the representative threats in this stack, the mitigations already in place, and the residual risk that remains in the demo configuration. The residual risk is the part to act on: see Production Hardening Recommendations.
Representative threats
- A rogue client impersonates the Publisher or UA Cloud Commander/Action to the broker
- An attacker impersonates a web UI user (Translator, Publisher, Grafana, UA Cloud Action, or Portainer)
- A fake OPC UA server feeds the Publisher
- A forged
ua-action-requesttriggers an OPC UA method - An unauthenticated caller hits the OPC UA Web API
- Anything on the pod network impersonates a Modbus master
- Theft of the Publisher's CA key (
/publisher/pki/issuer/private) lets an attacker mint a trusted certificate for any component - Anyone who can reach the UA Cloud Library UI can self-register a working account and act as a legitimate user.
- An unauthenticated caller queries the i3X API and reads the entire ISA-95 hierarchy and its values
- An unauthenticated caller reaches the MCP endpoint and uses UA Cloud AI as a ready-made interface to the whole plant
- A malicious or compromised MCP client impersonates a legitimate agent, since the server cannot distinguish one Basic-authenticated caller from another
- A caller reaching a .NET service directly spoofs its own source IP and scheme by sending forged
X-Forwarded-For/X-Forwarded-Protoheaders, because those services trust the headers unconditionally (see the residual risks below)
Mitigations already in place
- MQTT broker requires username/password (
allow_anonymous false) - Most web UIs require login (the .NET ones over HTTPS through the ingress; InfluxDB, Grafana and Portainer on the node IP)
- The UA Cloud Action web UI and OPC UA Web API mandate HTTP Basic authentication on every request (no anonymous access)
- OPC UA supports certificate exchange between Publisher/Commander and server
- The Cloud Library requires an account to upload, and its API is authenticated with
ServiceUsername/ServicePassword - The i3X API fails closed: with no Basic or OAuth2 credentials configured it returns
503to every request rather than serving data anonymously - No .NET web UI or API is exposed on the node at all — the Edge Translator UI, Cloud Publisher, UA Cloud Action, the Cloud Library, i3X and UA Cloud AI are all
ClusterIPservices reachable only through a TLS-terminating ingress, so their Basic credentials cannot cross the LAN in cleartext even by misconfiguration - UA Cloud AI compares its inbound credentials in fixed time (
CryptographicOperations.FixedTimeEquals), so they cannot be recovered by timing its responses - UA Cloud AI warns at startup when Basic auth is enabled without TLS, or when it is left unauthenticated entirely
Residual risk / gaps
- Single shared credential set across all components (including Grafana/Portainer admin and the Web API)
- TLS terminates at the ingress and the certificate is self-signed: it encrypts traffic but proves no identity, so an active attacker presenting their own certificate is not prevented, and pod-to-pod traffic behind the ingress is still plain HTTP
- No per-service identities or mutual TLS (mTLS)
- Broker does not authenticate clients by certificate
- Any client that can publish to
commandscan drive Commander - The GDS issuer key is a 12-year self-signed CA stored in a PKCS#12 with an empty password on a
hostPathvolume (see hardening item 10) - The Cloud Library has email verification disabled (
EmailSenderAPIKeyunset), so self-registration is open and accounts are not tied to a provable identity - MQTT Explorer (
:4000) has no authentication of its own, so anyone who can reach it can publish to any topic — includingcommands - Modbus TCP has no authentication whatsoever by protocol design — the simulator (and any real Modbus device) trusts every caller
- The i3X API shares the same
IOT_USERNAME/IOT_PASSWORDas everything else, so it grants no separate identity and a single leaked credential opens it too - UA Cloud AI reuses that same credential pair both inbound and outbound, so one leak exposes the MCP endpoint and, through it, both backends
- The MCP endpoint serves anonymously if
MCP_USERNAME/MCP_PASSWORDare left unset — it warns loudly at startup, but nothing prevents it - The .NET services trust
X-Forwarded-For/X-Forwarded-Protofrom any source. They must, because the ingress pod's address is not known in advance, soKnownIPNetworks/KnownProxiesare cleared. This is safe only because those services areClusterIP-only and unreachable except through the ingress — exposing any of them directly would let a caller forge its apparent client IP and scheme, defeating IP-based rate limiting and any scheme-dependent logic
Representative threats
- Modification of telemetry in transit
- Tampering with
hostPathconfig/cert files on the node - Editing the ConfigMaps
- Altering the imported
opcua_modeldata or the model importer script - A malicious command writing/actuating an OPC UA node via Commander
- Writing Modbus coils/registers on the simulated device
- Forging or altering a stored Digital Product Passport so a product appears to have a lower carbon footprint than it does, or uploading a malicious nodeset that is then trusted as the definition of what a machine reports
Mitigations already in place
- MQTT is carried over TLS (8883)
- Config is delivered via Kubernetes ConfigMaps/Secrets
- Commander/Action send spec-compliant OPC UA PubSub Action envelopes
- The seeded Thing Description and settings are delivered read-only from ConfigMaps
- Cloud Library uploads require an authenticated account
Residual risk / gaps
- Telegraf and UA Cloud Action use TLS verification skip (
insecure_skip_verify/MQTT_TLS_INSECURE=true), so a man-in-the-middle with any cert is accepted hostPathvolumes (/influxdb2,/cloudlib-postgres,/cloudlib-dpkeys,/translator/*,/publisher/*,/commander/*,/productionline/*,/mosquitto,/portainer,/grafana) are writable by anyone with node access- No message signing on payloads
- Commander performs Writes/MethodCalls with no per-action authorization
- Stored Digital Product Passports are not signed or provenance-checked, and because registration is open any account can upload one, so a passport carries no cryptographic proof of origin
- PCF and Battery Passport results are published without a signature, so a consumer, recycler or regulator cannot verify they came from this pipeline
- Modbus traffic is plaintext and unauthenticated, so anything on the pod network can read or write the simulated device's registers
- i3X is a read-only projection, so it cannot alter stored telemetry — and it is now reachable from outside only over HTTPS through the ingress, so responses can no longer be altered in flight by a passive network attacker
- UA Cloud AI cannot tamper with anything by design — it exposes no write, method-call or actuation tool — but it can be misled: it faithfully relays whatever the backends return, so a compromised backend misrepresents the plant to the model
- A model acting on UA Cloud AI's output may drive changes through other paths (an operator acting on its answer, or UA Cloud Action's feedback loop), so its read-only boundary does not by itself make downstream effects safe
- Credentials are injected by
envsubstat apply time with no validation, so an unset or mistyped variable is silently substituted as an empty string or literal text across every manifest. The cluster then runs with credentials nobody intended, and the failure surfaces later as unrelated-looking authentication errors rather than at deploy time
Representative threats
- An operator changes a device mapping, publish set, Grafana dashboard, or issues a command and denies it
- No record of who logged in or who imported a model
- A user uploads a Digital Product Passport to the Cloud Library and denies it
- A disputed Product Carbon Footprint cannot be traced back to the inputs it was derived from, which matters when the passport is presented as a regulatory claim
Mitigations already in place
- Component logs are written to
hostPathlogsdirectories and pod stdout; Portainer records some cluster events; the Cloud Library records the owning account against each upload.
Residual risk / gaps
- No centralized, tamper-evident audit log
- Shared credentials make actions unattributable to an individual
- Command/action requests and model imports are not attributably logged
- Cloud Library accounts are self-registered with unverified email addresses, so the recorded uploader identity is weak evidence
- UA Data Processor does not retain the telemetry window or carbon-intensity figure behind each PCF, so a passport's figures are not independently reproducible
- No log shipping or retention policy
- i3X API queries are not attributably logged, so there is no record of who browsed or exported the production data
- MCP tool calls are not attributably logged either, so there is no record of which agent asked what — and because an AI client may issue many queries per question, this is the component most likely to read broadly across the plant with the least trace
Representative threats
- Sniffing telemetry
- Reading credentials from the manifest
- Exposed dashboards (Grafana, Portainer, UA Cloud Action) on the node IP
- Leaking the UA Cloud Library credentials used by the import Job
- Reading an OPC UA private key — or the Publisher's CA key — off the node (or off the SSD if the device is removed)
- Reading the Cloud Library's PostgreSQL database directly off
/cloudlib-postgres, which exposes every stored Digital Product Passport and all account password hashes - Stealing the Data Protection key ring from
/cloudlib-dpkeys, which would let an attacker forge a valid Cloud Library authentication cookie for any account without ever knowing a password - Inferring production volumes, energy use and product composition from stored passports.
- Reading the whole production hierarchy and its history through the i3X API, which is designed to make exactly that convenient
- Extracting the plant's structure and history through UA Cloud AI in natural language, which lowers the skill needed to do so — no Flux, no OPC UA knowledge and no API familiarity are required
- Leaking plant data to a third-party model provider, since a hosted AI client sends tool results onward for inference
Mitigations already in place
- MQTT is encrypted with TLS
INFLUX_TOKENis stored in a KubernetesSecret- Credentials are supplied at apply time (not committed to git)
- PostgreSQL is
ClusterIPonly, so it is not reachable from outside the cluster - Cloud Library passwords are stored as ASP.NET Identity hashes, not plaintext
- Every .NET web UI and API is reachable only over HTTPS, through the ingress, so credentials and returned data cannot cross the network in cleartext (see Enabling TLS)
- Plain-HTTP requests are redirected to HTTPS before reaching the application, so a mistyped
http://URL cannot leak credentials — the redirect is issued by the ingress, ahead of any authentication challenge - UA Cloud AI caps how much any one tool call returns (
MCP_MAX_RESULTS, default 200), so a single broad request cannot trivially export the whole address space
Residual risk / gaps
- Credentials (including UA Cloud Library and Grafana/Portainer admin) are injected as plain-text env vars (visible via
kubectl describe/exec) - Kubernetes Secrets are base64, not encrypted at rest by default
- Self-signed broker cert offers encryption but no server-identity assurance
- OPC UA private keys, including the GDS issuer (CA) key, are held unprotected in
Directorystores onhostPathvolumes (see hardening item 10) - The PostgreSQL data directory is an unencrypted
hostPathand the database password is the sharedIOT_PASSWORD - Traffic inside the cluster is still plain HTTP — TLS stops at the ingress, so anything able to observe pod-to-pod traffic (or a compromised pod) still sees credentials and data in the clear; mTLS or a service mesh would be needed to close this
- The TLS certificate is self-signed, so it provides encryption but no identity assurance, and users are trained to click through the browser warning
- InfluxDB, Grafana, MQTT Explorer and Portainer's HTTP port are still exposed directly on the node IP without TLS
- The Edge Translator's non-HTTP listeners stay on the node: OPC UA (
4840), LoRaWAN (5000/5001) and OCPP (19520/19521) cannot be carried by an HTTP ingress, so they depend on their own protocol-level security rather than this TLS layer - The i3X Swagger UI and
/v1/infoare exempt from authentication, so anyone who can reach the ingress can enumerate the full API surface and read the server's capabilities before authenticating - UA Cloud AI's
/healthendpoint is likewise unauthenticated, confirming the service exists to an unauthenticated scanner (it returns no plant data) - Tool results leave the cluster entirely when the MCP client is a hosted AI service, which is a disclosure path no amount of in-cluster hardening addresses
Representative threats
- Flooding the broker or web UIs
- Filling the node disk with telemetry or repeated model imports
- A crash loop
- A runaway feedback loop from UA Cloud Action
- Overloading the simulated stations or the Modbus simulator with connections
- Filling the disk by uploading large or numerous passports/nodesets to the Cloud Library
- Exhausting InfluxDB with the Data Processor's repeated multi-day queries.
- Exhausting InfluxDB through the i3X API, whose
/v1/objects/historyand/v1/subscriptions/streamendpoints can each drive repeated backend queries - Amplifying load through UA Cloud AI, where a single natural-language question can fan out into many tool calls and therefore many backend queries — an agent retrying or looping does this without any attacker intent
- Losing the Cloud Library's Data Protection key ring, which locks every user out: the authentication cookie and the login form's antiforgery token both become undecryptable, so logins are rejected before the password is even checked
Mitigations already in place
- Liveness/readiness probes restart unhealthy pods
- Single-replica deployments recover automatically
- The importer is a short-lived Job with
ttlSecondsAfterFinished - UA Cloud Action has a built-in rate limiter that bounds how often it actuates
- The Modbus simulator declares CPU/memory
requests/limits - The Data Processor polls on a fixed interval rather than continuously
- i3X caches metadata (
I3X_METADATA_CACHE_SECONDS) and bounds its browse and latest-value lookups to fixed time ranges rather than scanning the whole bucket - UA Cloud AI caps results per call (
MCP_MAX_RESULTS) and bounds every backend request with a timeout (HTTP_TIMEOUT_SECONDS), so one tool call cannot hang indefinitely - The Cloud Library's Data Protection key ring is persisted to a
hostPathvolume (/cloudlib-dpkeys), so cookies and antiforgery tokens survive restarts and upgrades
Residual risk / gaps
- No rate limiting, quotas, or
resourcesrequests/limits on most pods - Unbounded InfluxDB growth on the local SSD (now including
opcua_modelpoints) - No upload size limit or per-account quota on the Cloud Library, and its PostgreSQL volume has no size cap — filling
/cloudlib-postgresfills the same disk InfluxDB and the broker rely on - A single node is a single point of failure
- The broker persists to
hostPath(/mosquitto), reducing message loss on restart though the single node remains a SPOF - UA Cloud Action's rate limit still needs tuning for your environment
- No rate limit or result cap on the i3X API, so a client may open many concurrent
streamsubscriptions or request unbounded history ranges - UA Cloud AI has no rate limiter, so while each individual call is capped, nothing bounds how many calls an agent makes per second
Representative threats
- Container escape to the node
- A compromised pod reading another component's data via shared host paths
- Using the InfluxDB admin token for full DB control
- Abusing Portainer's
cluster-adminServiceAccount to take over the whole cluster - Using Commander/Action to reach and control OT devices
- A self-registered Cloud Library user escalating to administrative rights over the model store
- A compromised UA Data Processor reusing the admin
INFLUX_TOKENit is given.
Mitigations already in place
- Distinct container images per component
nodeSelectorpins workloads to Linux- The importer Job uses
restartPolicy: Never - The Cloud Library separates ordinary user accounts from the
ServiceUsernameAPI account - UA Cloud AI is read-only: it exposes no write, method-call or actuation tool, so an agent reaching it cannot use it to change the plant — unlike UA Cloud Action and Commander, which deliberately can
Residual risk / gaps
- Containers run with default (often root) user and no
securityContext - No
NetworkPolicyisolation between pods - The InfluxDB token is an all-powerful admin token
- UA Data Processor only ever reads, but is handed the same admin token rather than a read-only one
- Portainer is bound to
cluster-admin, so compromising it compromises the cluster - The Cloud Library's API account shares the single
IOT_PASSWORDused everywhere else, so one leaked credential grants model-store write access - Commander bridges IT→OT with method-call/write capability and no fine-grained authorization
- No RBAC scoping for the workloads
- i3X authorization is all-or-nothing: any caller who authenticates sees the entire hierarchy, with no per-site, per-line or per-tag scoping
- UA Cloud AI inherits that all-or-nothing scope and cannot narrow it, so every agent sees everything both backends expose; its read-only boundary limits what kind of access is possible, not how much
The following changes move the stack from a demo toward a production-grade deployment. Prioritize the items marked (High).
- Use unique, per-service credentials (High). Replace the single shared
IOT_USERNAME/IOT_PASSWORDwith distinct identities for the Translator UI, Publisher UI, broker client, and InfluxDB admin. Store them in a real secrets manager (e.g. HashiCorp Vault, Sealed Secrets, or an external secrets operator) rather than plain-text env vars. - Deploy trusted TLS certificates and enforce verification (High). Replace
the self-signed, pod-generated broker certificate with one from a trusted CA
(e.g. via cert-manager). Remove
insecure_skip_verify = truefrom the Telegraf MQTT inputs and pin the broker CA so man-in-the-middle attacks are prevented. The four .NET services are alreadyClusterIP-only behind a TLS-terminating ingress (see Enabling TLS), but with a self-signed certificate that encrypts without proving identity — replace it with one from your own CA, ideally issued and renewed automatically by cert-manager. Grafana, InfluxDB, MQTT Explorer and Portainer's HTTP port are still exposed directly on the node IP without TLS; move them behind the same ingress. Consider mTLS or a service mesh for pod-to-pod traffic, which is still plain HTTP. - Enable mutual TLS (mTLS) or per-client auth on the broker. Configure Mosquitto to authenticate publishers/subscribers by client certificate in addition to username/password, and use ACLs to restrict which topics each client may publish/subscribe to.
- Scope the InfluxDB token (High). Do not use the all-powerful admin token
for Telegraf. Create a dedicated write-only token limited to the
mqttbucket, and separate read tokens for dashboards. UA Data Processor only ever reads, so give it a read-only token rather than the admin one it currently shares. - Re-enable Cloud Library email verification and close open registration (High).
The self-hosted UA Cloud Library ships with
EmailSenderAPIKeyunset, which disables account confirmation and lets anyone who can reach:8083register a working account (see Registration and the Disabled Email Verification). For production, setEmailSenderAPIKey,RegistrationEmailFromandRegistrationEmailReplyToso accounts are tied to a verified address, and front the UI with an authenticating proxy or SSO. Registration and login are already HTTPS-only from outside the cluster, since the service is reachable only through the ingress. Also give the Cloud Library API its ownServiceUsername/ServicePasswordinstead of reusing the sharedIOT_*credentials, and give its PostgreSQL database a dedicated password. - Restrict network exposure (High). Do not expose
LoadBalancerservices directly on the node IP. Front the web UIs with an authenticating reverse proxy/ingress, place the broker and database on an internal network only, and add KubernetesNetworkPolicyrules so pods can only reach the peers they need. - Secure the i3X API (High). The i3X server exposes the whole production
hierarchy and its history to any caller who authenticates. It is now
ClusterIP-only and reachable externally just over HTTPS through the ingress, but the certificate is self-signed — use a CA-issued one. Switch from Basic auth to OAuth2 by settingI3X_OAUTH2_AUTHORITY,I3X_OAUTH2_AUDIENCEandI3X_OAUTH2_ISSUER, which gives per-client identities and expiring tokens instead of one shared password. SetI3X_CORS_ORIGINSto the specific origins that need browser access rather than leaving it open, and put a rate limit in front of/v1/objects/historyand/v1/subscriptions/stream, neither of which is bounded today. - Control what the AI layer can reach (High). UA Cloud AI turns the plant
into a natural-language query surface, which is useful precisely because it
removes the expertise barrier — and that cuts both ways. Always set
MCP_USERNAME/MCP_PASSWORD(it serves anonymously without them), keep it reachable only through the HTTPS ingress, and give it credentials scoped to only the data an agent should see rather than the sharedIOT_*pair. Be deliberate about where tool results go: a hosted AI client sends them outside your network for inference, so treat that as an export of plant data and check it against your data-handling policy. Add a rate limit in front of the endpoint, since one question can fan out into many backend queries. Its read-only design means an agent cannot actuate anything directly — but do not over-rely on that, because an operator acting on its output can. - Harden the pods. Add a
securityContext(runAsNonRoot: true,readOnlyRootFilesystem: true, drop Linux capabilities,allowPrivilegeEscalation: false) and set CPU/memoryrequests/limitsto contain resource-exhaustion and blast radius. - Protect data at rest. Enable encryption at rest for the node's disk
(
/influxdb2,/cloudlib-postgresand the otherhostPathvolumes) and for Kubernetes Secrets (e.g. a KMS provider or an encrypted etcd). Replace ad-hochostPathvolumes with managedPersistentVolumeClaimswhere possible. Note that/cloudlib-postgresholds the Cloud Library's account password hashes and every nodeset uploaded to it. - Encrypt the OPC UA private keys at rest (High). Every OPC UA component in
this stack holds its application instance certificate in a
Directorycertificate store, so the private key sits unencrypted on the Pi's filesystem under<component>/pki/own/private/*.pfx:
| Path on the Pi | Whose identity |
|---|---|
/publisher/pki/own/private |
UA Cloud Publisher |
/translator/pki/own/private |
UA Edge Translator |
/commander/pki/own/private |
UA Cloud Commander |
/productionline/munich/<station>/pki/own/private |
each simulated station |
These keys are the components' identities. Anyone who can read one can impersonate that component to every OPC UA server that trusts it — and in the Commander's case that means calling methods on your OT devices. They are more sensitive than the telemetry they protect, and unlike the broker certificate they are not regenerated on restart.
⚠️ /publisher/pki/issuer/privateis the most sensitive file in the whole deployment. UA Cloud Publisher acts as a small Certificate Authority for GDS server push: on first start it mints a self-signed CA certificate (SetCAConstraint(), 12-year lifetime) and stores the PKCS#12 there, protected by an empty password. That single file can issue a valid certificate for any OPC UA component in the system, and every station already trusts it. Stealing anownkey impersonates one component; stealing the issuer key lets the holder mint identities at will and be trusted by all of them — and the 12-year lifetime means the exposure does not expire in any useful sense. Treat it as the deployment's root of trust and protect it accordingly.
Mitigate in layers, strongest first:
- Keep them off the plain filesystem. Back the
pkivolumes with an encrypted store rather than a barehostPath— a LUKS-encrypted partition or filesystem-level encryption (e.g.fscrypton ext4) for the directory the volumes bind to, so the keys are unreadable if the SSD is removed from the device. This is the single highest-value step on a physically accessible edge device such as a Pi in a cabinet. - Restrict who can read them. Tighten the directory to the container's own
UID (
chmod 0700), setrunAsNonRootwith a dedicated UID per component, and avoid mounting thepkidirectory into any other pod. Note thathostPathvolumes are readable by anyone with node access, which is one more reason to preferPersistentVolumeClaims(item 9). - Move the CA off the device entirely. The self-signed issuer is a convenience so the demo can provision certificates with no external infrastructure. In production, use a real GDS or an existing enterprise PKI (or a managed CA such as cert-manager with an offline root), so no CA private key is ever stored on an edge node.
- Prefer hardware-backed keys where the platform allows it. The CM5 can be paired with a TPM or secure element, see here for the recommended hardware. Storing the private key there means it never exists in readable form on disk at all. This is the direction OPC UA deployments in regulated environments are expected to take, though it requires a certificate store implementation that supports it.
- Rotate on exposure. Because GDS server push is already wired up, re-issuing a component's certificate is inexpensive — treat any suspected key exposure as a rotation event rather than something to tolerate, and remove the old certificate from every peer's trust list. Note that rotating the issuer is a different matter: every station must be re-provisioned against the new CA, which is why keeping it off the device is preferable to planning to rotate it.
The demo deliberately uses unencrypted
Directorystores so the certificates can be inspected withlsandopensslwhile learning the system. That trade-off is appropriate for a reference deployment and inappropriate for production.
- Add auditing and monitoring. Ship component and access logs to a central, tamper-evident store; enable Kubernetes audit logging; and add alerting on authentication failures, pod restarts, and disk usage.
- Manage capacity and availability. Set InfluxDB retention policies to bound
growth, back up
/influxdb2regularly, and consider multi-node/HA for the broker and database to remove the single-point-of-failure. - Keep software patched. Pin and regularly update the container image versions, apply OS/K3s security updates, and scan images for known vulnerabilities as part of your release process.
- Scope Portainer's cluster access (High). The demo binds Portainer to the
built-in
cluster-adminrole. For production, grant it a least-privilegeRole/ClusterRolelimited to the namespaces and resources operators actually manage, protect its UI behind the ingress, and enforce strong, per-user Portainer accounts (not the shared credentials). - Authorize and throttle the command/control path. Restrict who can publish
to the
commandstopic (broker ACLs) and validate/allow-list the OPC UA methods and nodes UA Cloud Commander may Write/Call. UA Cloud Action includes a built-in rate limiter on its actuation, so a faulty threshold or spoofed value cannot drive OT devices uncontrollably; tune its limit for your environment. Treat the UA Cloud Library import credentials as secrets and restrict the import Job's egress.
