diff --git a/pages/memgraph-zero/memgql/changelog.mdx b/pages/memgraph-zero/memgql/changelog.mdx index 3d22af484..764e5e7ab 100644 --- a/pages/memgraph-zero/memgql/changelog.mdx +++ b/pages/memgraph-zero/memgql/changelog.mdx @@ -7,6 +7,16 @@ description: MemGQL release notes ## MemGQL v0.14.0 - unreleased +### 🍃 New features & Improvements + +- **Amazon Redshift connector.** `ADD CONNECTOR … TYPE redshift` maps + Redshift tables, on provisioned clusters and Serverless workgroups, as a + graph, with reads and writes pushed down as Redshift SQL. Authenticate with a + database password, or with IAM temporary credentials (`redshift+iam://…`) + fetched for your AWS identity on every connection. The endpoint from the AWS + console, or a `jdbc:redshift://` URL, works as the connector URI. See + [Amazon Redshift](/memgraph-zero/memgql/connect/redshift). + ## MemGQL v0.13.0 - October 7th, 2026 ### 🍃 New features & Improvements diff --git a/pages/memgraph-zero/memgql/connect.mdx b/pages/memgraph-zero/memgql/connect.mdx index e1b76f414..f90c73d90 100644 --- a/pages/memgraph-zero/memgql/connect.mdx +++ b/pages/memgraph-zero/memgql/connect.mdx @@ -7,6 +7,7 @@ description: MemGQL connector details and configuration. - [ClickHouse](/memgraph-zero/memgql/connect/clickhouse) - [DuckDB](/memgraph-zero/memgql/connect/duckdb) +- [Amazon Redshift](/memgraph-zero/memgql/connect/redshift) - [Iceberg](/memgraph-zero/memgql/connect/iceberg) - [Microsoft Fabric](/memgraph-zero/memgql/connect/fabric) - [Memgraph](/memgraph-zero/memgql/connect/memgraph) diff --git a/pages/memgraph-zero/memgql/connect/_meta.ts b/pages/memgraph-zero/memgql/connect/_meta.ts index 401a8a8b8..1478ab073 100644 --- a/pages/memgraph-zero/memgql/connect/_meta.ts +++ b/pages/memgraph-zero/memgql/connect/_meta.ts @@ -12,5 +12,6 @@ export default { "mysql": "to MySQL", "sqlserver": "to SQL Server", "pinot": "to Pinot", + "redshift": "to Amazon Redshift", "snowflake": "to Snowflake", } diff --git a/pages/memgraph-zero/memgql/connect/redshift.mdx b/pages/memgraph-zero/memgql/connect/redshift.mdx new file mode 100644 index 000000000..0d2276521 --- /dev/null +++ b/pages/memgraph-zero/memgql/connect/redshift.mdx @@ -0,0 +1,221 @@ +--- +title: Amazon Redshift +description: Connect MemGQL to Amazon Redshift. +--- + +# Amazon Redshift + +The Redshift connector (type `redshift`) translates GQL queries into Redshift +SQL and runs them on a provisioned cluster or a Serverless workgroup. It +requires a [mapping file](/memgraph-zero/memgql/reference#mapping-schema) that +maps graph patterns to relational tables, in the same format as every other SQL +backend. + +Redshift speaks the PostgreSQL wire protocol, so there is no ODBC layer or +Redshift client to install. The connector supports **both reads and writes**, +and authenticates with a database password or with **IAM temporary +credentials**. It is available both in +[`multi` mode](/memgraph-zero/memgql/multiple-graphs) (`ADD CONNECTOR … TYPE +redshift`) and as the standalone `CONNECTOR_TYPE=redshift` mode. + +## 1. Prepare a database user + +Connect to your cluster or workgroup as an admin with any PostgreSQL client or +the Redshift query editor, and create a user for MemGQL: + +```sql +CREATE USER memgql PASSWORD ''; +GRANT USAGE ON SCHEMA public TO memgql; +GRANT SELECT ON ALL TABLES IN SCHEMA public TO memgql; +-- Only if you want writes: +GRANT INSERT, DELETE ON ALL TABLES IN SCHEMA public TO memgql; +``` + +The endpoint is shown on the cluster or workgroup page in the AWS console, for +example `analytics.abc123xyz.us-east-1.redshift.amazonaws.com:5439/dev` +(provisioned) or +`my-workgroup.123456789012.us-east-1.redshift-serverless.amazonaws.com:5439/dev` +(Serverless). + +## 2. Write the mapping + +Save as `mapping.json`, the standard +[mapping format](/memgraph-zero/memgql/reference#mapping-schema). Declare every +property you plan to query. In `multi` mode the mapping is also the routing +schema, so an undeclared property is a routing error, not a passthrough. + +```json +{ + "vertices": [ + { + "label": "Person", + "mappedTableSource": { + "connector": "rs", + "schema": "public", + "table": "persons", + "metaFields": { "id": "id" } + }, + "attributes": [ + { "name": "name" }, + { "name": "age", "type": "Int" } + ] + }, + { + "label": "Company", + "mappedTableSource": { + "connector": "rs", + "schema": "public", + "table": "companies", + "metaFields": { "id": "id" } + }, + "attributes": [ { "name": "name" } ] + } + ], + "edges": [ + { + "label": "KNOWS", + "from": "Person", + "to": "Person", + "mappedTableSource": { + "connector": "rs", + "schema": "public", + "table": "knows", + "metaFields": { "id": "id", "from": "from_id", "to": "to_id" } + } + }, + { + "label": "WORKS_AT", + "from": "Person", + "to": "Company", + "mappedTableSource": { + "connector": "rs", + "schema": "public", + "table": "works_at", + "metaFields": { "id": "id", "from": "person_id", "to": "company_id" } + } + } + ] +} +``` + +## 3. Start MemGQL in multi mode + +```bash +docker run --rm \ + --name memgql \ + --stop-timeout 2 \ + -p 7688:7688 \ + --env CONNECTOR_TYPE=multi \ + -v ./mapping.json:/data/mapping.json \ + memgraph/memgql:latest +``` + +For [IAM authentication](#authentication), also pass your AWS credentials, for +example `--env AWS_ACCESS_KEY_ID --env AWS_SECRET_ACCESS_KEY --env +AWS_SESSION_TOKEN`. + +## 4. Connect and register the backend + +```bash +mgconsole --port 7688 +``` + +Give the endpoint as the connector `URI`. The connector accepts the same URL +shape as the Redshift JDBC driver, so you can paste the endpoint from the +console: + +```gql +ADD CONNECTOR rs TYPE redshift + URI 'redshift://analytics.abc123xyz.us-east-1.redshift.amazonaws.com:5439/dev' + USER 'memgql' PASSWORD ''; +CREATE GRAPH social FROM FILE '/data/mapping.json'; +MATCH (n:Person) RETURN n.name LIMIT 5; +``` + +The port defaults to `5439` and the database to `dev`. TLS is required by +default. The server certificate is verified against the system roots, which +include Amazon's. + +Because the parameters live on the connector, **changing them needs no MemGQL +restart**. To rotate a password, re-register the connector: + +```gql +DROP CONNECTOR rs; +ADD CONNECTOR rs TYPE redshift + URI 'redshift://analytics.abc123xyz.us-east-1.redshift.amazonaws.com:5439/dev' + USER 'memgql' PASSWORD ''; +``` + +## 5. Query + +```gql +MATCH (p:Person) RETURN p.name, p.age; +``` + +```gql +MATCH (p:Person)-[:WORKS_AT]->(c:Company) RETURN p.name, c.name; +``` + +```gql +MATCH (a:Person)((-[:KNOWS]->()){1,3})(b:Person) RETURN DISTINCT b.name; +``` + +## Authentication + +The URI scheme selects the method: + +| Method | URI | +|---|---| +| **Database password** | `redshift://:@:5439/`, or `USER` / `PASSWORD` options | +| **IAM temporary credentials** | `redshift+iam://[@]:5439/` | + +`jdbc:redshift://…` and `jdbc:redshift:iam://…` URLs work too. + +With IAM, MemGQL exchanges your AWS identity for a short-lived database password +on every connection, so there is no password to store or rotate. The identity +comes from the standard AWS credential chain: environment variables, a profile +(`?profile=`), SSO, web identity, or an instance or task role. + +- **Provisioned clusters** call `GetClusterCredentials` when a user is given + (`redshift+iam://memgql@…`), with optional `?AutoCreate=true` and + `?DbGroups=group1,group2`. Without a user they call + `GetClusterCredentialsWithIAM`, which logs in as your IAM identity. +- **Serverless workgroups** call `GetCredentials`, which always logs in as your + IAM identity, so a user in the URI is rejected. + +The cluster or workgroup and the region are read from the AWS endpoint name. +Behind a custom DNS name or a VPC endpoint, add them to the URI: +`?cluster=` or `?workgroup=`, plus `?region=`. + +Every option you omit falls back to its environment variable (`REDSHIFT_URI`, +`REDSHIFT_USER`, `REDSHIFT_PASSWORD`, `REDSHIFT_DATABASE`). The standalone +`CONNECTOR_TYPE=redshift` mode takes all of them from the environment. + +## Dialect notes + +- Table references are `schema.table`. A `mappedTableSource` with a `database` + is spelled `database.schema.table`, for cross-database queries and data + sharing. +- `AVG()` casts its argument to `DOUBLE PRECISION`. Redshift's `AVG` over + integers otherwise truncates. +- Bounded variable-length patterns (`(){1,3}`) run as a recursive CTE with + trail semantics. Redshift limits recursion to 100 levels. +- JSON attributes with a `path` read `VARCHAR` JSON columns with + `JSON_EXTRACT_PATH_TEXT`. +- `DESCRIBE CONNECTOR` lists tables, columns, primary keys, and foreign keys. + Redshift records these constraints but does not enforce them. + +## Known limitations + +- **`collect()`** and **map projections** (`RETURN n {.a, .b}`) are not + available: Redshift has no array type to return them in. +- **Inserting a node and an edge in one statement** requires the node to carry + its id (`INSERT (:Person {id: 7, name: 'Ann'})-[:KNOWS]->(…)`). Redshift + cannot report generated ids back. Without an explicit id, insert the node + first and connect it with `MATCH … INSERT`. +- **Unbounded variable-length paths**, shortest path and graph algorithms are + not pushed down. Use a bounded form or a Cypher backend. +- `SUPER` columns are not mapped yet. + +For connector configuration, see the +[mapping reference](/memgraph-zero/memgql/reference#mapping-schema). diff --git a/pages/memgraph-zero/memgql/features.mdx b/pages/memgraph-zero/memgql/features.mdx index 66774cdf5..0bcbc61ce 100644 --- a/pages/memgraph-zero/memgql/features.mdx +++ b/pages/memgraph-zero/memgql/features.mdx @@ -26,6 +26,7 @@ description: MemGQL Community and Enterprise feature comparison. | [Microsoft Fabric](/memgraph-zero/memgql/connect/fabric) | Yes | Yes | | [MongoDB](/memgraph-zero/memgql/connect/mongodb) | Yes | Yes | | [SAP HANA](/memgraph-zero/memgql/connect/hana) | Yes | Yes | +| [Amazon Redshift](/memgraph-zero/memgql/connect/redshift) | Yes | Yes | | **Multi-Connection Mode** | Yes | Yes | | Max connectors | 2 | Unlimited | | Max simultaneous connections | 2 | Unlimited | diff --git a/pages/memgraph-zero/memgql/reference.mdx b/pages/memgraph-zero/memgql/reference.mdx index a508aee46..d4df4a516 100644 --- a/pages/memgraph-zero/memgql/reference.mdx +++ b/pages/memgraph-zero/memgql/reference.mdx @@ -10,7 +10,7 @@ Memgraph and Neo4j (translation is largely passthrough); "SQL backends" means PostgreSQL, MySQL and DuckDB. **MongoDB** is neither — it translates to aggregation pipelines — so it gets its own column. -SQL Server, SAP HANA, ClickHouse, Apache Iceberg, and Apache Pinot are also supported +SQL Server, SAP HANA, Amazon Redshift, ClickHouse, Apache Iceberg, and Apache Pinot are also supported as connectors but with a narrower verified surface. See each connector's page for the exact list of features each one supports. @@ -202,7 +202,7 @@ apply depends on the type: | Option | Read by | |---------------------------------------------|----------------------------------------------------------------------| | `GRAPH ''` | Memgraph, Neo4j — selects the Cypher database (it is not a mapping) | -| `DATABASE ''` | PostgreSQL, MySQL, SQL Server, Oracle (service name), ClickHouse, MongoDB, Snowflake, Fabric (warehouse/lakehouse item), SAP HANA (tenant database of an MDC system) | +| `DATABASE ''` | PostgreSQL, MySQL, SQL Server, Oracle (service name), ClickHouse, MongoDB, Snowflake, Fabric (warehouse/lakehouse item), SAP HANA (tenant database of an MDC system), Amazon Redshift | | `CATALOG ''` | Iceberg (Trino catalog), Iceberg Direct (warehouse) | | `SCHEMA ''` | Iceberg, Snowflake, Fabric (default `dbo`), SAP HANA (defaults to the connection user's own schema); MongoDB accepts it as a fallback for `DATABASE` | | `WAREHOUSE` / `ROLE` | Snowflake session settings | @@ -297,7 +297,7 @@ Writing a [mapping](/memgraph-zero/memgql/schema-file) shouldn't require out-of-band database access. `DESCRIBE CONNECTOR` reads a source's physical catalog **live through the connector** — tables, columns, types, nullability, primary keys, and single-column foreign keys as `→ table.column` -(PostgreSQL and SAP HANA; a graph backend answers via `SHOW SCHEMA` / +(PostgreSQL, SAP HANA and Amazon Redshift; a graph backend answers via `SHOW SCHEMA` / `REFRESH SCHEMA` instead and says so): ``` @@ -522,6 +522,7 @@ refused with a `Compute limit reached` error; the client can retry. | `pinot` | GQL -> SQL | Apache Pinot | | `mongodb` | GQL -> aggregation pipeline | MongoDB 5.0+ | | `hana` | GQL -> SQL | SAP HANA 2.0 SPS05+, HANA Cloud, HANA Express | +| `redshift` | GQL -> SQL | Amazon Redshift (provisioned and Serverless) | | `multi` | Per-connector | Multiple backends simultaneously | @@ -653,6 +654,21 @@ There is no `CATALOG` level for HANA: the tenant database is chosen at login, not spelled into a qualified table name, so a HANA table reference is at most `schema.table`. See the [SAP HANA connector page](/memgraph-zero/memgql/connect/hana). +#### Amazon Redshift (`redshift`) + +| Variable | Default | Description | +|---------------------|--------------------------|----------------------------------------------------------------| +| `REDSHIFT_URI` | _(required)_ | `redshift://:@:5439/` (password) or `redshift+iam://:5439/` (IAM temporary credentials); `jdbc:redshift://` URLs are accepted | +| `REDSHIFT_USER` | _(from the URI)_ | Database user, when not written into the URI | +| `REDSHIFT_PASSWORD` | _(from the URI)_ | Password, when not written into the URI (password auth only) | +| `REDSHIFT_DATABASE` | `dev` | Database, when not written into the URI | +| `MAPPING_FILE` | _(required)_ | Path to JSON mapping file (same format as Postgres) | + +IAM auth uses the standard AWS credential chain, so the usual `AWS_*` +variables (or an instance or task role) apply. TLS is required unless the URI +says `?sslmode=disable`. See the +[Amazon Redshift connector page](/memgraph-zero/memgql/connect/redshift). + #### Iceberg (`iceberg`) | Variable | Default | Description |