Skip to content

Commit e7ea6ef

Browse files
authored
Add dbt project for the mobile catalog (#1249)
Introduces a dbt project under dbt/ for transformations over the Iceberg tables in the `mobile` catalog, with one model to start: `hotspots.enabled_carriers_inventory`, the current carrier enablement per hotspot derived from `hotspots.enabled_carriers_history`. Lives in this repo rather than its own because most of the catalog is written by the oracles a few directories away, and the schema of each of those tables is a `table_definition()` function rather than a migration -- so a column added in Rust and the model consuming it can land in one reviewed change. The project is public for the same reason the oracles are, and scoped to the `mobile` catalog alone. A model is written to the schema of the data it describes rather than a shared `analytics` bucket, so `enabled_carriers_inventory` sits beside the history it summarises and consumers can source `mobile.<namespace>.<model>`. Two things the model gets right that are worth naming, both documented at length in the source: - The incremental watermark is `received_timestamp`, the clock the oracle assigns on ingest, not the device-supplied `timestamp`. A watermark over a reported clock is pinned by any device reporting a far-future time and then admits nothing ever again, for every hotspot -- and it prunes no partitions, since the table is partitioned on `day(received_timestamp)`. A unit test pins this: reverting the window to `timestamp` fails it. - The lookback is sized in minutes, covering ingest commit skew rather than data lateness. A current-state upsert recomputes nothing, so a days-wide window would re-read and re-upsert most of the fleet on every run. The production cluster runs on Railway with app sleeping, so the image's entrypoint wakes Trino and blocks until it will serve a query before handing over to dbt -- polling /v1/info for `starting=false`, then confirming with SELECT 1 over the real credentials. CI stands up the repo's own Trino/Polaris/Iceberg stack, creates the source table, loads fixtures, and runs the models and tests against it. Rust CI now skips dbt-only changes.
1 parent 0f984e6 commit e7ea6ef

25 files changed

Lines changed: 1513 additions & 0 deletions

‎.github/workflows/CI.yml‎

Lines changed: 5 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -3,6 +3,11 @@ name: CI
33
on:
44
pull_request:
55
branches: ["main"]
6+
# A dbt-only change has no Rust to build; .github/workflows/dbt.yml covers
7+
# dbt/. Should a dbt model ever read an oracle-written table, that workflow
8+
# gains a trigger on the Rust defining it -- see the note in its `paths`.
9+
paths-ignore:
10+
- "dbt/**"
611
push:
712
branches: ["main"]
813
tags: ["*"]

‎.github/workflows/dbt.yml‎

Lines changed: 161 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,161 @@
1+
name: dbt
2+
3+
# NOTE: the third-party actions below are pinned by tag, not by commit SHA as
4+
# the Rust workflows in this repo are. Pin them before merging.
5+
6+
# Triggers on the dbt project and on the local warehouse it is built against.
7+
#
8+
# No Rust paths: the only source this project declares is written by a pipeline
9+
# outside this repo, so there is no `table_definition()` here for CI to create a
10+
# table from. Add the relevant `*/src/iceberg/**` path back alongside the first
11+
# model over an oracle table -- that is what makes a source declaration
12+
# verifiable rather than just written down.
13+
on:
14+
pull_request:
15+
branches: ["main"]
16+
paths:
17+
- "dbt/**"
18+
- "docker-compose.yml"
19+
- "infra/**"
20+
- ".github/workflows/dbt.yml"
21+
push:
22+
branches: ["main"]
23+
tags: ["*"]
24+
workflow_dispatch:
25+
26+
permissions:
27+
contents: read
28+
packages: write
29+
30+
env:
31+
PYTHON_VERSION: "3.12"
32+
33+
jobs:
34+
# Needs no warehouse and no Rust, so it is the first thing in the run to go
35+
# red on a bad model. Jinja and YAML only: `dbt parse` does not look at the SQL
36+
# itself, so a syntax error gets past this and is caught by `dbt run` in the
37+
# `build` job, where Trino parses it for real.
38+
parse:
39+
runs-on: ubuntu-latest
40+
concurrency:
41+
group: ${{ github.workflow }}-${{ github.ref }}-parse
42+
cancel-in-progress: true
43+
defaults:
44+
run:
45+
working-directory: dbt
46+
env:
47+
DBT_PROFILES_DIR: ${{ github.workspace }}/dbt
48+
steps:
49+
- name: Checkout Repository
50+
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6
51+
- name: Set up Python
52+
uses: actions/setup-python@v5
53+
with:
54+
python-version: ${{ env.PYTHON_VERSION }}
55+
cache: pip
56+
cache-dependency-path: dbt/requirements.txt
57+
- name: Install dbt
58+
run: pip install -r requirements.txt
59+
- name: Install dbt packages
60+
# --lock so a drifted package-lock.yml fails rather than being silently
61+
# re-resolved to something the image was never built with.
62+
run: dbt deps --lock && dbt deps
63+
- name: Parse project
64+
run: dbt parse
65+
66+
# Stands up the repo's own Trino/Polaris/Iceberg stack, creates the source
67+
# table from the mirrored DDL, loads fixtures, then builds and tests. Docker
68+
# and Python only -- no Rust toolchain, no cargo cache.
69+
build:
70+
runs-on: ubuntu-latest
71+
concurrency:
72+
group: ${{ github.workflow }}-${{ github.ref }}-build
73+
cancel-in-progress: true
74+
env:
75+
COMPOSE_PROJECT_NAME: oracles
76+
DBT_PROFILES_DIR: ${{ github.workspace }}/dbt
77+
steps:
78+
- name: Checkout Repository
79+
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6
80+
81+
- name: Start Trino, Polaris and RustFS
82+
# `--wait trino` pulls in postgres, rustfs and the polaris bootstrap
83+
# jobs through the compose dependency graph and blocks on their health
84+
# checks, so nothing below needs its own retry loop.
85+
run: docker compose -p "$COMPOSE_PROJECT_NAME" up -d --wait trino
86+
87+
- name: Register the Trino catalog
88+
# Local Trino runs with catalog.store=memory, so the mounted catalog
89+
# properties files are inert and the catalog has to be created at
90+
# runtime. See the script's header.
91+
run: ./dbt/scripts/register_trino_catalog.sh
92+
93+
- name: Create external source mirrors
94+
# Tables written by pipelines outside this repo, so there is no Rust
95+
# definition to create them from. This applies a hand-maintained copy of
96+
# their production DDL -- see the header of the .sql file.
97+
run: ./dbt/scripts/create_external_sources.sh
98+
99+
- name: Load fixtures
100+
# So the data tests run against rows rather than passing trivially on
101+
# empty tables. Model logic is covered by the unit tests, which need no
102+
# data at all and run as part of `dbt build` below.
103+
run: ./dbt/scripts/load_fixtures.sh
104+
105+
- name: Set up Python
106+
uses: actions/setup-python@v5
107+
with:
108+
python-version: ${{ env.PYTHON_VERSION }}
109+
cache: pip
110+
cache-dependency-path: dbt/requirements.txt
111+
- name: Install dbt
112+
working-directory: dbt
113+
run: pip install -r requirements.txt && dbt deps
114+
115+
# `run` then `test`, not `build`. `dbt build` runs a model's unit tests
116+
# before the model itself, and a unit test that supplies `this` needs the
117+
# relation to exist to read its column types -- so on a fresh warehouse
118+
# like this one, `build` errors before it has created anything. See the
119+
# header of models/marts/_unit_tests.yml.
120+
- name: Build models
121+
working-directory: dbt
122+
run: dbt run
123+
- name: Run tests
124+
working-directory: dbt
125+
# Data tests against the fixture rows, unit tests against rows supplied
126+
# inline. Together with the step above: the models compile and execute
127+
# against schemas created from their real definitions, and the logic is
128+
# pinned independently of what is in the warehouse.
129+
run: dbt test
130+
131+
- name: Trino logs on failure
132+
if: failure()
133+
run: docker compose -p "$COMPOSE_PROJECT_NAME" logs --tail 200 trino polaris
134+
135+
# Built and pushed on `dbt-v*` tags only, so shipping models is not coupled to
136+
# the cadence of the Rust binaries' own release tags.
137+
build-image:
138+
if: startsWith(github.ref, 'refs/tags/dbt-v')
139+
needs: [parse, build]
140+
runs-on: ubuntu-latest
141+
concurrency:
142+
group: ${{ github.workflow }}-${{ github.ref }}-build-image
143+
cancel-in-progress: true
144+
steps:
145+
- name: Checkout Repository
146+
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6
147+
- name: Log in to Docker Registry
148+
uses: docker/login-action@b45d80f862d83dbcd57f89517bcf500b2ab88fb2 # v4
149+
with:
150+
registry: ${{ secrets.ZOT_URL }}
151+
username: ${{ secrets.ZOT_USER }}
152+
password: ${{ secrets.ZOT_PASSWORD }}
153+
- name: Set up Docker Buildx
154+
uses: docker/setup-buildx-action@4d04d5d9486b7bd6fa91e7baf45bbb4f8b9deedd # v4
155+
- name: Build and Push dbt Image
156+
uses: docker/build-push-action@d08e5c354a6adb9ed34480a06d141179aa583294 # v7
157+
with:
158+
context: dbt
159+
file: dbt/Dockerfile
160+
push: true
161+
tags: ${{ secrets.ZOT_URL }}/oracles/dbt:${{ github.ref_name }}

‎dbt/.gitignore‎

Lines changed: 11 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,11 @@
1+
target/
2+
dbt_packages/
3+
logs/
4+
.user.yml
5+
.venv/
6+
7+
# The repo root ignores *.json globally (../.gitignore), which would silently
8+
# drop any dbt fixture or seed in JSON. Re-include them here. The build
9+
# directories above are ignored as directories, so git never descends into them
10+
# and this cannot pull in artefacts.
11+
!*.json

‎dbt/Dockerfile‎

Lines changed: 49 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,49 @@
1+
# Runtime image for the mobile_dbt project.
2+
#
3+
# Separate from the repo-root Dockerfile, which builds Rust binaries and takes a
4+
# PACKAGE build-arg. Nothing is shared between them, so they are two images with
5+
# two cache keys -- see the `dbt-*` jobs in .github/workflows/dbt.yml.
6+
#
7+
# Python 3.12: see the note in requirements.txt about 3.14.
8+
FROM python:3.12-slim AS base
9+
10+
ENV PYTHONDONTWRITEBYTECODE=1 \
11+
PYTHONUNBUFFERED=1 \
12+
DBT_PROFILES_DIR=/dbt \
13+
# dbt writes logs/ and target/ relative to the project dir; keep them off
14+
# the read-only project copy so the image can run with a read-only root fs.
15+
DBT_LOG_PATH=/tmp/dbt/logs \
16+
DBT_TARGET_PATH=/tmp/dbt/target
17+
18+
WORKDIR /dbt
19+
20+
# Dependencies first: requirements change far less often than models, so a model
21+
# edit reuses this layer.
22+
COPY requirements.txt ./
23+
RUN pip install --no-cache-dir -r requirements.txt
24+
25+
# `dbt deps` is resolved at build time from the committed lock file, so the
26+
# image never reaches out to the package hub at run time.
27+
COPY packages.yml package-lock.yml ./
28+
COPY dbt_project.yml profiles.yml ./
29+
RUN dbt deps
30+
31+
COPY macros/ ./macros/
32+
COPY models/ ./models/
33+
COPY tests/ ./tests/
34+
35+
# Only the connection probe -- the rest of scripts/ drives docker compose and
36+
# is for local development, not for inside the image.
37+
COPY scripts/wait_for_trino.py ./scripts/
38+
COPY docker-entrypoint.sh ./
39+
40+
# Fail the build rather than the first run if the project does not parse. Needs
41+
# no warehouse.
42+
RUN dbt parse
43+
44+
# The entrypoint waits for Trino to be awake and serving before handing over to
45+
# dbt; see its header. No default target and no default command: the profile
46+
# takes its target from DBT_TARGET, every connection value comes from the
47+
# environment (see profiles.yml), and no arguments means the full
48+
# `dbt run` + `dbt test` cycle.
49+
ENTRYPOINT ["/dbt/docker-entrypoint.sh"]

0 commit comments

Comments
 (0)