Skip to content
Merged
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
101 changes: 94 additions & 7 deletions pages/clustering/high-availability/setup-ha-cluster-k8s.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -1130,16 +1130,17 @@ helm install memgraph-ha memgraph/memgraph-high-availability \
<Callout type="warning">
When using an existing Gateway, ensure it has listeners whose **names** match the
TCPRoute `sectionName` references. The chart expects listener names in the format
`data-{id}-bolt` for data instances and `coordinator-{id}-bolt`
for coordinators. The listener ports are your choice in this mode,
`data-{id}-bolt` for data instances and a single listener named `coordinators-bolt`
for the entire coordinator tier. The listener ports are your choice in this mode,
but keeping them aligned with the chart defaults avoids surprises. For example,
the default HA setup (2 data instances, 3 coordinators) needs these listeners:

- `data-0-bolt` on port 9000
- `data-1-bolt` on port 9001
- `coordinator-1-bolt` on port 10001
- `coordinator-2-bolt` on port 10002
- `coordinator-3-bolt` on port 10003
- `coordinators-bolt` on port 7687 (`ports.boltPort`)

Adding coordinators needs no new listener, while every new data instance needs
one on `dataPortBase + id`.

A standalone Gateway manifest with these pre-configured listeners is available in the [Helm charts repository](https://github.com/memgraph/helm-charts/blob/main/examples/gateway/gateway.yaml).
</Callout>
Expand Down Expand Up @@ -1748,9 +1749,14 @@ To start the HA chart with direct scraping you need:
`kube-prometheus-stack` selects it, and it is created in `prometheus.namespace`
(falls back to the release namespace). Both must match the `kube-prometheus-stack`
Helm release name and namespace you used when installing the stack.
6. **The right scheme** – when instances run with `tls.bolt.enabled: true`, their
monitoring endpoint is served over HTTPS and Prometheus must scrape it with
`prometheus.serviceMonitor.scheme: https` (see
[ServiceMonitor TLS for direct scraping](#servicemonitor-tls-for-direct-scraping)).

A minimal values file for the default installation of `kube-prometheus-stack`
(release `kube-prometheus-stack` in namespace `monitoring`) looks like this:
(release `kube-prometheus-stack` in namespace `monitoring`) and instances without
Bolt TLS looks like this:

```yaml
# Every coordinator and data instance serves OpenMetrics; the chart exposes their
Expand All @@ -1764,6 +1770,7 @@ prometheus:
enabled: true # in-cluster: kube-prometheus-stack scrapes each instance directly
kubePrometheusStackReleaseName: kube-prometheus-stack # must equal the Helm release name
interval: 15s
scheme: http # https when tls.bolt.enabled is true on the instances
grafanaDashboard:
enabled: true # optional: ship the "Memgraph OpenMetrics" dashboard
namespace: monitoring # a namespace the Grafana sidecar watches
Expand Down Expand Up @@ -1802,7 +1809,85 @@ kubectl port-forward -n monitoring svc/kube-prometheus-stack-prometheus 9090:909
Open `http://localhost:9090/targets` and look for the
`serviceMonitor/monitoring/<release-name>-memgraph/0` job; every coordinator and data
instance should be listed as `UP`. If the job is missing, the `release` label or
namespace does not match your `kube-prometheus-stack` installation.
namespace does not match your `kube-prometheus-stack` installation. If the targets are
`DOWN` with a TLS or protocol error, the `scheme` does not match whether Bolt TLS is
enabled on the instances.

#### ServiceMonitor TLS for direct scraping

When an instance has `tls.bolt.enabled: true`, Memgraph serves its HTTP monitoring
endpoint over HTTPS as well, so Prometheus must scrape it with `https`. Set
`prometheus.serviceMonitor.scheme` and, if needed, a Prometheus Operator
[`TLSConfig`](https://prometheus-operator.dev/docs/api-reference/api/#monitoring.coreos.com/v1.TLSConfig)
through `prometheus.serviceMonitor.tlsConfig`. A complete example with Bolt TLS on
all instances and a self-signed certificate:

```yaml
scrapeMemgraphDirectly: true

prometheus:
enabled: false
namespace: monitoring
serviceMonitor:
enabled: true
kubePrometheusStackReleaseName: kube-prometheus-stack
scheme: https
tlsConfig:
insecureSkipVerify: true

data:
- id: "0"
tls:
bolt:
enabled: true
secretName: bolt-tls-secret
- id: "1"
tls:
bolt:
enabled: true
secretName: bolt-tls-secret

coordinators:
- id: "1"
tls:
bolt:
enabled: true
secretName: bolt-tls-secret
- id: "2"
tls:
bolt:
enabled: true
secretName: bolt-tls-secret
- id: "3"
tls:
bolt:
enabled: true
secretName: bolt-tls-secret
```

The `scheme` applies to every instance selected by the `ServiceMonitor`, so all
coordinators and data instances must expose their metrics endpoint using the same
protocol: either enable `tls.bolt` on all of them and use `https`, or on none of them
and use `http`. `insecureSkipVerify: true` is convenient for self-signed certificates but
not suitable for production; prefer `caFile`/`ca` together with `serverName` instead:

```yaml
prometheus:
serviceMonitor:
enabled: true
scheme: https
tlsConfig:
ca:
secret:
name: bolt-ca-bundle
key: ca.crt
serverName: memgraph.example.com
```

`tlsConfig` is rendered verbatim into the `ServiceMonitor` endpoint, so any field
supported by your Prometheus Operator version can be used. The referenced `Secret`
(`bolt-ca-bundle` above) must exist in the namespace where the Prometheus instance
runs, because the operator mounts it into the Prometheus pod.

#### Grafana dashboard for direct scraping

Expand Down Expand Up @@ -2163,6 +2248,8 @@ and their default values.
| `prometheus.serviceMonitor.enabled` | If enabled, a `ServiceMonitor` object will be deployed. | `false` |
| `prometheus.serviceMonitor.kubePrometheusStackReleaseName` | The release name under which `kube-prometheus-stack` chart is installed. Rendered as the `release` label on the `ServiceMonitor`, which is how the stack discovers it. | `kube-prometheus-stack` |
| `prometheus.serviceMonitor.interval` | How often will Prometheus pull data from Memgraph's Prometheus exporter. | `15s` |
| `prometheus.serviceMonitor.scheme` | Protocol Prometheus uses when scraping Memgraph directly (`http` or `https`). Applies to every instance selected by the `ServiceMonitor`. | `http` |
| `prometheus.serviceMonitor.tlsConfig` | Prometheus Operator `TLSConfig` for direct HTTPS scraping (e.g. `insecureSkipVerify: true`, or `ca`/`serverName` for production). | `{}` |
| `prometheus.grafanaDashboard.enabled` | Ship the bundled "Memgraph OpenMetrics" Grafana dashboard as a `ConfigMap` for a Grafana sidecar to auto-load. | `false` |
| `prometheus.grafanaDashboard.namespace` | Namespace for the dashboard `ConfigMap`; must be one the Grafana sidecar watches. Defaults to `prometheus.namespace`, else release namespace. | `""` |
| `prometheus.grafanaDashboard.label` | Label the Grafana sidecar selects dashboards by. | `grafana_dashboard` |
Expand Down
Loading