adaptive_export: reliable dx-steered pem-direct capture (chunk/end_time, breaker, dc_snoop filter, DaemonSet) - #92
adaptive_export: reliable dx-steered pem-direct capture (chunk/end_time, breaker, dc_snoop filter, DaemonSet)#92ConstanzeTU wants to merge 186 commits into
Conversation
|
Important Review skippedAuto reviews are disabled on base/target branches other than the default branch. Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: Organization UI Review profile: ASSERTIVE Plan: Pro Plus Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
…e fix) Root cause of the flaky dx-steered capture (dc_snoop/http erratically 0 while light tables always land): OrderExportAll fans out ~20 tables concurrently, each OrderQuery issued ONE unbounded PxL query over the full ~600s control window against the single node-local PEM (pem-direct). QueryFor only set start_time, so every query re-scanned [sliceStart, now] and post-filtered — the heavy tables materialize huge result sets on a saturated PEM and lose the fixed 180s deadline race, dropping out; the cheap tables (redis/conn/stack) return instantly and survive. Reconcile fingerprint: the same dc_snoop query returns 2459 rows in isolation but 0 + 1 err under the fan-out. Fix (durable — removes the data-volume↔deadline coupling, not just tunes it): - pxl.QueryFor: bound the PEM source scan on BOTH sides. Emit a relative end_time (floored toward now so nothing real is clipped; the exact upper bound stays enforced by the df.time_ < sliceEnd nanos post-filter) whenever sliceEnd is in the past. Live-edge slices keep scanning to now (no end_time), preserving prior behavior for the most-recent window. - controller.OrderQuery: walk the capture window in OrderChunk-sized sub-windows (default 60s, env ADAPTIVE_ORDER_CHUNK_SEC), each a both-sides bounded query, so no single query re-materializes the whole window. captureSpan adaptively halves any chunk that still fails with a transient (deadline/overload) error down to orderMinChunk (1s); non-transient errors (missing dark table) surface immediately without wasteful splitting. Overlapping/retried spans dedupe in the ReplacingMergeTree evidence tables, so re-pulls are idempotent. One aggregated reconcile row per table (not per chunk). Chunks run sequentially per table, so OrderExportAll's per-table concurrency is unchanged while each table now issues cheap bounded queries instead of one firehose — reliable capture without needing the global inflight throttle set. Tests: queryfor end_time present for past windows / absent at the live edge; OrderQuery chunking, single aggregated reconcile row, adaptive subdivision on transient error, no-split on non-transient error, termination at min-chunk.
… (dc_snoop) The dx-steered OrderExportAll path applied only a partial comm denylist and NO namespace filter to the node-scoped dark-vector tables — unlike the shipped cron preset (script/presets dc_snoop.pxl __DC_SNOOP_EXCLUSION__, built from presets.go defaultExcludeNamespaces + defaultExcludeComms). So every dc_snoop capture drowned in infra dcache churn: on a real k3s node a single window returned ~54k rows dominated by ConfigReloader/iptables/CNI(host-local,bridge,flannel,loopback)/host daemons(systemd-udevd,dbus-daemon,tailscaled)/kubevuln — burying the salient attack specimens (whoami/cat/getent reading /etc/shadow + the SA token). - Extend darkExcludeCommsDefault with the host/CNI/node daemons that were leaking (systemd-udevd, host-local, bridge, flannel, loopback, bandwidth, dbus-daemon, mount, umount, tailscaled, grpc_health_pro, kubevuln, opm, kube-proxy, …). - Add darkExcludeNamespacesDefault + darkNamespaceExclusion(), applied in the IsDarkVector branch AFTER PodEnrichPxL resolves df.namespace, dropping infra namespaces (pl, kube-system, clickhouse, …). Blank-namespace transient rows survive (each `!=` is true for ''), so the attack's short-lived children — which resolve blank — are never dropped. Overridable via DC_SNOOP_EXCLUDE_NAMESPACES. Kept in sync with script/presets.go. Tests: infra namespaces + host/CNI comms dropped; df.namespace never pinned to the alert pod (node-scoped); env override replaces the default list.
… depth cap) Live RCA on aeprod54: the chunk fix is correct in isolation (pem unit suite — dc_snoop 54k, redis/conn/stack written per-chunk) but UNSAFE under the dx steering firehose. dx does generic collect-per-alert, so OrderExportAll (20 tables) fires on every noisy pl system pod continuously; all land on the ONE node-local PEM (pem-direct) → it saturates → 100% DeadlineExceeded. captureSpan then split every timeout into two narrower retries, amplifying a busy PEM into a query storm where nothing completes (observed: "0 ordered pixie rows written" across the whole run; draining dx + restarting AE → pem-direct instantly serves again). Make subdivision safe: - Circuit-breaker: orderTimeoutStreak (atomic) counts CONSECUTIVE transient failures; any success resets it. Above orderBreakerTrip (8) captureSpan stops subdividing — a saturated PEM must not be flooded with retries. It still splits a genuinely-oversized window on a healthy PEM (the reset keeps that path live). - Depth cap: maxOrderSplitDepth (3) bounds one chunk to ≤2^3 leaf queries even if it keeps timing out (was ~64 splitting 60s→1s). Tests: a 10-chunk all-timeout window stays <60 queries (ungated ≈640); a single transient failure still recovers (breaker resets on success, no latch). NOTE (deployment, not code): the firehose root also needs dx steering scoped so it doesn't fire 20-table captures on every noisy pl/system-pod alert — tracked separately for dx-agent.
Live RCA (aeprod55): every dx-steered capture in the e2e returned 0 rows, and the reconcile showed why — all 36 ordered captures had ~512ns-wide windows (width_s=0), so they matched no pixie rows. /export/start already reaches back controlExportLookback, but a control client that keys the /query window on a single finding's event_time sends lo≈hi (a sub-microsecond span). That passes the lo<hi validation yet captures nothing. handleQuery now widens any window narrower than minControlQueryWindow (5s) to controlExportLookback ending at hi — a point-in-time referral still captures the evidence leading up to it. hi is preserved; comfortably-wide windows pass through unchanged. Isolated /query probes (proper windows) already proved the capture path works — dc_snoop 54k→16k filtered, redis/conn/stack per-chunk; this makes the dx-driven path robust to degenerate windows too. Tests: a 512ns window is widened to >=5s (hi preserved); a 120s window is untouched. NOTE (dx-agent): dx should send a real window (or use /export/start) rather than a point window per finding — tracked separately. This is the AE-side safety net.
The bootstrap manifest was a replicas:0 Deployment with minimal env (EXPORT_MODE= auto, no pem-direct, no throttle) — it never ran and could not do node-local pem-direct. Replace it with the working config that the e2e RCA validated: - DaemonSet (one-per-node) so each pod queries its OWN node's vizier-pem at HOST_IP:50305 (pem-direct: node-local, desync-immune). - dx-steered: EXPORT_MODE=never + CONTROL_ADDR=:9100 + the control Service (internalTrafficPolicy:Local so dx reaches its co-located AE). - PEM-protection: ADAPTIVE_MAX_INFLIGHT_QUERIES_GLOBAL=4 and ADAPTIVE_ORDER_CHUNK_SEC =600 (one query per table, no window pre-chunking) so the AE never saturates the single node-local PEM it shares with dx. See RCA_ae_capture_20260803. Secret still seeded per-cluster (unchanged).
…efault; trim comments - queryfor.go: add darkExcludeCommSubstrings (kworker/ksoftirqd/rcu_/… — kernel threads with variable suffixes exact-match misses) applied via px.logicalNot( px.contains); add pause + systemd-logind exact. Workload comms (redis-*) untouched. - controller.go: defaultOrderChunk 60s -> 600s (one query per table; pre-chunking 10x-amplified queries on the single node-local PEM). - Strip verbose comments across queryfor.go/controller.go/server.go + the AE manifest. Test: kernel-thread substrings dropped, workload comms kept, pause dropped.
Deploys the dx-daemon DaemonSet + Service into honey and mirrors the pl->honey secrets (jwt-signing-key, cluster-id, cloud-addr, api-key, clickhouse http-url) via a before-hook, replacing the hand-applied manifest used in the e2e. Deploy with: skaffold deploy -f k8s/vizier/dx/skaffold.yaml CH http-url defaults to the soc clickhouse Service; override with DX_CH_HTTP_URL.
Replaces the imperative seed-secret + patch-cloud-addr + sed-image +
kubectl-apply sequence with a single skaffold module:
skaffold deploy -f k8s/vizier/adaptive_export/skaffold.yaml
- kustomize overlay reuses bootstrap/adaptive_export_{role,deployment}
and pins the image via images: (ghcr aeprod tag) instead of sed.
- before-hook patches PL_CLOUD_ADDR :443 and seeds
pl-adaptive-export-secrets ONLY when PIXIE_API_KEY/PX_API_KEY is set,
never clobbering an existing secret with an empty key.
- LoadRestrictionsNone so the overlay can reuse the bootstrap manifests
in place (no duplication/drift).
Pairs with the dx-daemon skaffold (k8s/vizier/dx). Bump the AE image by
editing newTag in kustomization.yaml.
…aths
The AE/dx skaffold configs lived inside their overlay dirs with kustomize
paths: [.], which skaffold resolves against the shell CWD (repo root), not
the config-file dir -> 'unable to find kustomization.yaml in /.../pixie'.
Match the repo convention instead (skaffold/skaffold_vizier.yaml et al.):
skaffold configs live in skaffold/ and reference overlays by repo-root-
relative kustomize paths. Overlays stay in k8s/vizier/{adaptive_export,dx}.
skaffold deploy -f skaffold/skaffold_adaptive_export.yaml
skaffold deploy -f skaffold/skaffold_dx.yaml # run from repo root
- dx overlay gains a kustomization.yaml (was rawYaml).
- both validated with 'skaffold render' from repo root (image overrides +
RBAC/DaemonSet/Service resolve).
00d654e to
2e05ed2
Compare
|
@build-agent — (reposting here, channel moved) please build the AE image from branch Fixes the |
Builds for
|
|
@build-agent — Please build the AE image from branch |
|
@build-agent — being specific, please build the AE (adaptive_export) image from this exact ref, NOT the #92 PR head:
This ref = #94 ppid/pcomm dc_snoop capture + PR#95 streaming dark-vector fix. Verify:
The last build (aeprod59) was #92 head |
|
Re-built from TTL image (ready now, expires ~24h from 2026-08-12 11:25 UTC): GHCR image — tag One note for rig 6a7be0ac: |
scanner.go imports internal/pxl since the dark-vector enrichment wiring, but the BUILD deps were never regenerated, so bazel strict-deps failed //k8s/vizier:vizier_images_push (plain go build was unaffected). Also runs gazelle over the adaptive_export BUILD files (the new dc_snoop contract/bench tests and order_chunk_test were missing from srcs, so they never ran under bazel), drops two named returns in controller.go and renames deadlineErr to errDeadline. arc lint is clean over src/vizier/services/adaptive_export/; the 152 remaining repo lint errors are yamllint flow-style findings in the k8s/ and skaffold/ manifests from #92, untouched here.
|
GHCR build update:
The #95 wiring added the
The already-posted TTL image Not touched: 152 yamllint findings (flow-style braces/indentation) in |
|
✅ Green — run 31594196141 (build-release, create-github-release, update-gh-artifacts-manifest all success). From Superseded, do not deploy: |
|
@build-agent — please rebuild the AE (adaptive_export) image from this exact ref (supersedes aeprod61/
This adds the OOM fix on top of the BUILD.bazel fix. aeprod61 fixed the Verify before building:
No new cross-package imports vs |
|
@build-agent — ping (no reply/CI run in ~17min). Please build BOTH images from
Post both refs when ready. This is the OOM-fixed AE (shared |
|
TTL image ready — built from Expires ~2026-08-13 15:40 UTC. GHCR Superseded, do not deploy: |
|
@build-agent — one more turn of the loop. Please build BOTH images from
What changed: aeprod62/TTL Verify: |
|
TTL image ready — Expires ~2026-08-13 15:47 UTC. GHCR |
|
✅ GHCR green — run 31613928161, all jobs success. Same commit as the TTL image above ( Tag ledger: aeprod63 = current. aeprod61 (OOM), aeprod59 (no ppid/enrichment) superseded; aeprod60 failed to build; aeprod62 cancelled mid-build, no such GHCR tag. |
|
@build-agent — this is a DX build (entlein/dx repo, NOT the AE/pixie image). Posting here since this is the channel you watch.
Please post BOTH:
(The entlein release-tag CI is out of GitHub-hosted Actions minutes, so it queues forever — that is why I need you to build it.) This = deployed rc2 + one fix: |
|
DX build answered on entlein/dx#136 — TTL |
…orest evictions) as explicit zeros beside trivial
…ition + 30-day TTL; migration rebuilds on partition drift
…ies the schema once on UNKNOWN_TABLE and retries; kssync forgets its cache)
…nd write), drilled on rig 6aa4409e
…_node_addrs DDL removed (no scheduled work; KPIs are PxL aggregates over the typed views)
…), managed_by from label or annotation with kind user|learned and mirror_version published by the views (0.16.25-rc2), recovery on failed reads with the engine migration inside recovery (0.16.26)
… unreachable store (ErrSpooled = deferred, /rows answers 202), declared columns added before the views at boot and in recovery, spool copies the body
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014GDT6HWiFmRxmUaGFSKjPY
…e views) Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014GDT6HWiFmRxmUaGFSKjPY
…ofiles, dx_profiles__compare) Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014GDT6HWiFmRxmUaGFSKjPY
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014GDT6HWiFmRxmUaGFSKjPY
… 4; compare matches like node-agent) Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014GDT6HWiFmRxmUaGFSKjPY
…-shaped orders; zero-row stream log) Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014GDT6HWiFmRxmUaGFSKjPY
…s tables) Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014GDT6HWiFmRxmUaGFSKjPY
…by the probe on every path; pxtrace names itself) Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014GDT6HWiFmRxmUaGFSKjPY
…ger read) Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014GDT6HWiFmRxmUaGFSKjPY
… offender pid in lineage) Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014GDT6HWiFmRxmUaGFSKjPY
… T3) Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014GDT6HWiFmRxmUaGFSKjPY
…queue depth, no silent drop; AE 5xx deferred) Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014GDT6HWiFmRxmUaGFSKjPY
…ent at zero) Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014GDT6HWiFmRxmUaGFSKjPY
…es held and re-sent in order) Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014GDT6HWiFmRxmUaGFSKjPY
… collect, attachment caps; every bound's discard series present at boot) Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014GDT6HWiFmRxmUaGFSKjPY
… and collects the delta; per-order caps; merged seed on the order; reopen counter) Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014GDT6HWiFmRxmUaGFSKjPY
…ws under flood on edge4) Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014GDT6HWiFmRxmUaGFSKjPY
…merge extends and collects the delta; merged seed on the order; zero-instant findings refused and counted) Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014GDT6HWiFmRxmUaGFSKjPY
…; AddMissingConstraints at boot; records-shape guards) Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014GDT6HWiFmRxmUaGFSKjPY
…HECK as ADD COLUMN and crash-looped the leader) Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014GDT6HWiFmRxmUaGFSKjPY
…; column parser skips CONSTRAINT/INDEX/PROJECTION) Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014GDT6HWiFmRxmUaGFSKjPY
… and a counter) Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014GDT6HWiFmRxmUaGFSKjPY
… as well as rows; per-table merge delta spans) Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014GDT6HWiFmRxmUaGFSKjPY
Summary: Makes the dx-steered adaptive-export capture reliable on a node-local PEM and turns the bootstrap into a pem-direct DaemonSet. Bounded, chunked source scans with a subdivision circuit breaker; infra-noise filter for the node-scoped dark tables; widened near-zero /query windows; pod enrichment in the streaming scanner.
Test Plan:
bazel test //src/vizier/services/adaptive_export/...; deployed on a k3s stack with kubescape, dx and this build; verifiedredis_events,dc_snoop,stack_trace,conn_statsanddns_eventsrows in ClickHouse, deduplicated by ReplacingMergeTree.Type of change: /kind bugfix