These modules ship with security-conservative defaults. You should have to explicitly opt out of safety, not opt in.
| Surface | Default | Knob |
|---|---|---|
| Root volume / OS disk | Encrypted with cloud-managed CMK | enable_customer_managed_key = true switches to KMS / Key Vault key created by the module |
| Data volume / data disk | Encrypted | Same as above |
| RDS / Azure DB for PostgreSQL storage | Encrypted | Customer-managed KMS/Key Vault key supported |
| RDS / Azure DB backups | Encrypted | n/a (forced on) |
| In-transit — client to LB | TLS 1.2+ required | min_tls_version (default TLS_1_2) |
| In-transit — LB to VM | AWS: TLS to the instance on 443. Azure: the Standard LB is L4 passthrough; with enable_application_gateway = true the gateway terminates TLS and the backend hop defaults to HTTP inside the customer vnet |
Azure: appgw_backend_protocol = "Https" for end-to-end TLS — requires appgw_backend_root_cert_pem and appgw_backend_host_header, because App Gateway v2 validates the backend certificate |
| In-transit — VM to DB | TLS required (rds.force_ssl=1; require_secure_transport=ON on Azure) |
Forced on. In db_mode = "external", external_db_sslmode accepts only require/verify-ca/verify-full |
- Security groups / NSGs default to deny all inbound. Required ports (443 for the LB, 5432 between VMs and DB) are opened only to the CIDRs you pass in
allowed_cidrs. Wide-open0.0.0.0/0requiresallow_internet_ingress = trueand emits a deprecation warning. - SSH is not exposed by default. For break-glass access,
enable_management_access = truewires up AWS SSM Session Manager (IAM-gated, viaAmazonSSMManagedInstanceCore) on AWS, and theAADSSHLoginForLinuxextension on Azure — Entra-authenticated, RBAC-gated SSH throughaz ssh vmwith no public IP. Azure Bastion is not provisioned by the modules; the extension is the lighter-weight equivalent, and operators still need theVirtual Machine Administrator LoginorVirtual Machine User Loginrole granted to them. - IMDSv2 is required on every EC2 launch (
http_tokens = "required"). - Public IPs are off by default. The LB has a public DNS name; the VMs sit in private subnets.
- VMs are tagged for AWS Inspector / Azure Defender for Servers auto-enrollment when those services are enabled at the account level.
The Azure modules set resource_provider_registrations = "none" on the azurerm
provider. This is deliberate, and it changes what your subscription must have in place
before the first apply.
Terraform's default ("legacy") attempts to register roughly seventy resource providers
on every apply — including many nothing here touches, such as Microsoft.BotService,
Microsoft.HealthcareApis and Microsoft.DataFactory. Registration is a
subscription-scoped write (*/register/action). A service principal scoped to a
resource group, or any least-privilege operator role, does not hold it, so the apply dies
on a wall of AuthorizationFailed 403s before creating a single resource. Granting
subscription-wide write purely to satisfy a registration sweep is the wrong trade.
With registration off, a subscription owner registers the providers the modules actually use, once per subscription:
for rp in Microsoft.Compute Microsoft.DBforPostgreSQL Microsoft.Insights \
Microsoft.KeyVault Microsoft.ManagedIdentity Microsoft.MarketplaceOrdering \
Microsoft.Network Microsoft.OperationalInsights Microsoft.Storage; do
az provider register --namespace "$rp" --wait
doneAdd Microsoft.Cache if you set enable_redis = true, and Microsoft.Network already
covers Application Gateway.
Most subscriptions that have ever deployed a VM already have Microsoft.Compute,
Microsoft.Network and Microsoft.Storage registered; Microsoft.DBforPostgreSQL is the
one most often missing. Check before you plan:
az provider show --namespace Microsoft.DBforPostgreSQL --query registrationState -o tsvIf a required provider is unregistered, the failure surfaces from whichever resource is
created first and reads as API version 20XX-XX-XX was not found for Microsoft.Foo — a
message that points at the API version rather than the real cause. The CI smoke in
hailbytes-sat asserts registration state up front for exactly this reason, and names the
missing provider.
- Each module creates a least-privilege instance profile / managed identity scoped to:
- Read its own marketplace product code (for subscription verification)
- Read/write its own CloudWatch / Azure Monitor namespace
- Read its own KMS / Key Vault keys
- Read its own secrets (DB connection string) from Secrets Manager / Key Vault
- Nothing else
- No wildcard
*resources in attached policies. - No long-lived access keys. Instance profiles / managed identities only.
- DB master passwords are generated via
random_password(32 chars, full symbol set) and stored in AWS Secrets Manager or Azure Key Vault, never in plaintext outputs or Terraform state diff logs (sensitive = true). - Customer-supplied admin credentials are accepted only via
sensitivevariables and are pushed to Secrets Manager / Key Vault on first apply, not stored in state alongside the resources.
- AWS: VPC Flow Logs are enabled by default (
enable_flow_logs = true), to a KMS-encrypted CloudWatch log group. ALB access logs are opt-in (enable_alb_access_logging) to a versioned, lifecycled S3 bucket. - Azure: VNet flow logs are enabled by default (
enable_flow_logs = trueonnetwork/azure), to a Storage Account with shared-key auth disabled, retained 180 days (flow_log_retention_days). Retention must exceed 90 days to satisfy CKV_AZURE_12 and is validated in the variable; it drives Storage Account cost roughly linearly, so lower it only alongside a documented waiver. Implemented as VNet flow logs rather than NSG flow logs, since NSG flow logs are being retired. - Azure: diagnostic settings for the load balancer, database, cache and (when enabled) Application Gateway go to a Log Analytics workspace by default (
enable_diagnostics = true); passlog_analytics_workspace_idto use your landing zone's central workspace instead. Note a Standard Load Balancer is Layer 4 and has no request-level access log —ApplicationGatewayAccessLogis the closest analogue of AWS ALB access logs, and it only exists whenenable_application_gateway = true. - CloudTrail / Azure Activity Log are not managed by these modules — those are account-level concerns and should be owned by your landing-zone tooling.
The HailBytes marketplace image is the source of truth for in-image patches. When HailBytes publishes a new image version, you:
- Pull the latest module version (or bump the pinned image version in your tfvars).
- Run the pre-patch backup via the SSM Run Command (AWS) or Azure Run Command document the module provisions. This produces a tamper-evident bundle in the immutable backup bucket / container — DB dump + uploads + manifest with the encryption-key fingerprint stamped in.
terraform planshows the launch template / VMSS image reference changing.terraform applytriggers instance refresh (ASG) or rolling upgrade (VMSS) — zero-downtime inha-hot-hotandunlimited-scaletiers, ~2 min downtime insingle-vm. On AWS autoscale, CloudWatch tripwires (5xx-rate, unhealthy-host count) auto-rollback the refresh. On Azure autoscale,automatic_instance_repair+ rolling-upgrademax_unhealthy_*pauses the upgrade.- Post-patch verify: curl
module.<name>.schema_version_endpoint(the module emits this output) and confirm the encryption-key fingerprint hasn't drifted.
Do not run apt/yum updates inside the running VM. The image is immutable; replace it.
Full runbook with audit pointers (no HailBytes admin access, no phone-home, customer-initiated only): docs/PATCHING_AND_MIGRATION.md.
Security issues in these Terraform modules: open a private security advisory in this repo's GitHub Security tab.
Security issues in the HailBytes software itself (inside the marketplace VM image): see SECURITY.md in the relevant product repo (hailbytes-asm, hailbytes-sat).