diff --git a/pages/clustering/high-availability/setup-ha-cluster-k8s.mdx b/pages/clustering/high-availability/setup-ha-cluster-k8s.mdx index b5860a5af..2cb06b337 100644 --- a/pages/clustering/high-availability/setup-ha-cluster-k8s.mdx +++ b/pages/clustering/high-availability/setup-ha-cluster-k8s.mdx @@ -1130,16 +1130,17 @@ helm install memgraph-ha memgraph/memgraph-high-availability \ When using an existing Gateway, ensure it has listeners whose **names** match the TCPRoute `sectionName` references. The chart expects listener names in the format -`data-{id}-bolt` for data instances and `coordinator-{id}-bolt` -for coordinators. The listener ports are your choice in this mode, +`data-{id}-bolt` for data instances and a single listener named `coordinators-bolt` +for the entire coordinator tier. The listener ports are your choice in this mode, but keeping them aligned with the chart defaults avoids surprises. For example, the default HA setup (2 data instances, 3 coordinators) needs these listeners: - `data-0-bolt` on port 9000 - `data-1-bolt` on port 9001 -- `coordinator-1-bolt` on port 10001 -- `coordinator-2-bolt` on port 10002 -- `coordinator-3-bolt` on port 10003 +- `coordinators-bolt` on port 7687 (`ports.boltPort`) + +Adding coordinators needs no new listener, while every new data instance needs +one on `dataPortBase + id`. A standalone Gateway manifest with these pre-configured listeners is available in the [Helm charts repository](https://github.com/memgraph/helm-charts/blob/main/examples/gateway/gateway.yaml). @@ -1748,9 +1749,14 @@ To start the HA chart with direct scraping you need: `kube-prometheus-stack` selects it, and it is created in `prometheus.namespace` (falls back to the release namespace). Both must match the `kube-prometheus-stack` Helm release name and namespace you used when installing the stack. +6. **The right scheme** – when instances run with `tls.bolt.enabled: true`, their + monitoring endpoint is served over HTTPS and Prometheus must scrape it with + `prometheus.serviceMonitor.scheme: https` (see + [ServiceMonitor TLS for direct scraping](#servicemonitor-tls-for-direct-scraping)). A minimal values file for the default installation of `kube-prometheus-stack` -(release `kube-prometheus-stack` in namespace `monitoring`) looks like this: +(release `kube-prometheus-stack` in namespace `monitoring`) and instances without +Bolt TLS looks like this: ```yaml # Every coordinator and data instance serves OpenMetrics; the chart exposes their @@ -1764,6 +1770,7 @@ prometheus: enabled: true # in-cluster: kube-prometheus-stack scrapes each instance directly kubePrometheusStackReleaseName: kube-prometheus-stack # must equal the Helm release name interval: 15s + scheme: http # https when tls.bolt.enabled is true on the instances grafanaDashboard: enabled: true # optional: ship the "Memgraph OpenMetrics" dashboard namespace: monitoring # a namespace the Grafana sidecar watches @@ -1802,7 +1809,85 @@ kubectl port-forward -n monitoring svc/kube-prometheus-stack-prometheus 9090:909 Open `http://localhost:9090/targets` and look for the `serviceMonitor/monitoring/-memgraph/0` job; every coordinator and data instance should be listed as `UP`. If the job is missing, the `release` label or -namespace does not match your `kube-prometheus-stack` installation. +namespace does not match your `kube-prometheus-stack` installation. If the targets are +`DOWN` with a TLS or protocol error, the `scheme` does not match whether Bolt TLS is +enabled on the instances. + +#### ServiceMonitor TLS for direct scraping + +When an instance has `tls.bolt.enabled: true`, Memgraph serves its HTTP monitoring +endpoint over HTTPS as well, so Prometheus must scrape it with `https`. Set +`prometheus.serviceMonitor.scheme` and, if needed, a Prometheus Operator +[`TLSConfig`](https://prometheus-operator.dev/docs/api-reference/api/#monitoring.coreos.com/v1.TLSConfig) +through `prometheus.serviceMonitor.tlsConfig`. A complete example with Bolt TLS on +all instances and a self-signed certificate: + +```yaml +scrapeMemgraphDirectly: true + +prometheus: + enabled: false + namespace: monitoring + serviceMonitor: + enabled: true + kubePrometheusStackReleaseName: kube-prometheus-stack + scheme: https + tlsConfig: + insecureSkipVerify: true + +data: + - id: "0" + tls: + bolt: + enabled: true + secretName: bolt-tls-secret + - id: "1" + tls: + bolt: + enabled: true + secretName: bolt-tls-secret + +coordinators: + - id: "1" + tls: + bolt: + enabled: true + secretName: bolt-tls-secret + - id: "2" + tls: + bolt: + enabled: true + secretName: bolt-tls-secret + - id: "3" + tls: + bolt: + enabled: true + secretName: bolt-tls-secret +``` + +The `scheme` applies to every instance selected by the `ServiceMonitor`, so all +coordinators and data instances must expose their metrics endpoint using the same +protocol: either enable `tls.bolt` on all of them and use `https`, or on none of them +and use `http`. `insecureSkipVerify: true` is convenient for self-signed certificates but +not suitable for production; prefer `caFile`/`ca` together with `serverName` instead: + +```yaml +prometheus: + serviceMonitor: + enabled: true + scheme: https + tlsConfig: + ca: + secret: + name: bolt-ca-bundle + key: ca.crt + serverName: memgraph.example.com +``` + +`tlsConfig` is rendered verbatim into the `ServiceMonitor` endpoint, so any field +supported by your Prometheus Operator version can be used. The referenced `Secret` +(`bolt-ca-bundle` above) must exist in the namespace where the Prometheus instance +runs, because the operator mounts it into the Prometheus pod. #### Grafana dashboard for direct scraping @@ -2163,6 +2248,8 @@ and their default values. | `prometheus.serviceMonitor.enabled` | If enabled, a `ServiceMonitor` object will be deployed. | `false` | | `prometheus.serviceMonitor.kubePrometheusStackReleaseName` | The release name under which `kube-prometheus-stack` chart is installed. Rendered as the `release` label on the `ServiceMonitor`, which is how the stack discovers it. | `kube-prometheus-stack` | | `prometheus.serviceMonitor.interval` | How often will Prometheus pull data from Memgraph's Prometheus exporter. | `15s` | +| `prometheus.serviceMonitor.scheme` | Protocol Prometheus uses when scraping Memgraph directly (`http` or `https`). Applies to every instance selected by the `ServiceMonitor`. | `http` | +| `prometheus.serviceMonitor.tlsConfig` | Prometheus Operator `TLSConfig` for direct HTTPS scraping (e.g. `insecureSkipVerify: true`, or `ca`/`serverName` for production). | `{}` | | `prometheus.grafanaDashboard.enabled` | Ship the bundled "Memgraph OpenMetrics" Grafana dashboard as a `ConfigMap` for a Grafana sidecar to auto-load. | `false` | | `prometheus.grafanaDashboard.namespace` | Namespace for the dashboard `ConfigMap`; must be one the Grafana sidecar watches. Defaults to `prometheus.namespace`, else release namespace. | `""` | | `prometheus.grafanaDashboard.label` | Label the Grafana sidecar selects dashboards by. | `grafana_dashboard` |