CUDly supports deployment via Terraform across three cloud providers (AWS, Azure, GCP). A helper script (scripts/tf-deploy.sh) simplifies common operations.
- Quick Start
- Platform Comparison
- Terraform Deployment
- AWS Deployment Details
- Azure Deployment Details
- GCP Deployment Details
- Deploying Code Updates
- Fargate (Side-by-Side with Lambda)
- CDN Architecture
- Cost Estimates
- Monitoring
- Maintenance
- Troubleshooting
# Deploy to AWS dev environment
./scripts/tf-deploy.sh aws dev
# Plan only (dry run)
./scripts/tf-deploy.sh aws dev plan
# Deploy to other providers/environments
./scripts/tf-deploy.sh azure dev
./scripts/tf-deploy.sh gcp dev
./scripts/tf-deploy.sh aws prodThe script uses profile-based tfvars from terraform/profiles/<provider>/<profile>.tfvars.
cd terraform/environments/aws
cp dev.tfvars.example dev.tfvars # edit with your values
terraform init -backend-config=backends/dev.tfbackend
terraform plan -var-file=dev.tfvars
terraform apply -var-file=dev.tfvarsTerraform automatically handles: Docker image build/push (via build module), frontend build/deploy to CDN, CDN cache invalidation, and admin user creation.
- Docker with buildx support
- Terraform >= 1.6.0
- Go 1.26.9+
- Cloud CLI configured:
aws,az, orgcloud
| Provider | Serverless | Containers/Kubernetes | Status |
|---|---|---|---|
| AWS | Lambda | Fargate (ECS) | Fully implemented |
| Azure | Container Apps | AKS | Container Apps implemented |
| GCP | Cloud Run | GKE | Cloud Run implemented |
| Lambda | Fargate | Container Apps | Cloud Run | |
|---|---|---|---|---|
| Timeout | 15 min max | Unlimited | Unlimited | 60 min max |
| Memory | 128-10240 MB | 512-30720 MB | 0.5-4 GB | 128 MB-32 GB |
| CPU | Tied to memory | 256-4096 units | 0.25-2 vCPU | 1-8 vCPU |
| Scaling | Auto (1000 concurrent) | Auto (1-N tasks) | Auto (0-N) | Auto (0-1000) |
| Cold Start | ~1-2s | No (always warm) | ~1-2s | ~0.5-1s |
| Cost (idle) | $0 | ~$30/mo (1 task) | $0 (scale to zero) | $0 (scale to zero) |
| Load Balancer | Not needed | Required (ALB) | Built-in | Built-in |
terraform/
├── environments/
│ ├── aws/ # main.tf, variables.tf, outputs.tf, backend.tf,
│ │ # networking.tf, database.tf, compute.tf, frontend.tf,
│ │ # secrets.tf, build.tf, ses.tf, route53.tf, acm.tf,
│ │ # dev.tfvars.example, backends/
│ ├── azure/ # similar structure
│ └── gcp/ # similar structure
├── modules/
│ ├── build/ # Docker build (docker-build.tf)
│ ├── compute/
│ │ ├── aws/lambda/
│ │ ├── aws/fargate/
│ │ ├── aws/cleanup-lambda/
│ │ ├── azure/container-apps/
│ │ └── gcp/cloud-run/
│ ├── database/ # aws/ (Aurora), azure/ (Flexible Server), gcp/ (Cloud SQL)
│ ├── frontend/ # aws/ (CloudFront+S3), azure/ (CDN+Blob), gcp/ (Cloud CDN+GCS)
│ ├── monitoring/ # aws/, azure/, gcp/
│ ├── networking/ # aws/ (VPC), azure/ (VNet), gcp/ (VPC)
│ ├── registry/ # aws/ (ECR), azure/ (ACR), gcp/ (Artifact Registry)
│ └── secrets/ # aws/ (Secrets Manager), azure/ (Key Vault), gcp/ (Secret Manager)
└── profiles/
├── aws/ # dev.tfvars, prod.tfvars, fargate-dev.tfvars, *.example
├── azure/ # dev.tfvars.example
└── gcp/ # dev.tfvars.example
cd terraform/environments/aws
terraform init -backend-config=backends/dev.tfbackend
terraform plan -var-file=../../profiles/aws/dev.tfvars
terraform apply -var-file=../../profiles/aws/dev.tfvars
terraform output
terraform destroy -var-file=../../profiles/aws/dev.tfvarsSee terraform/environments/aws/dev.tfvars.example for a complete reference. Key variables:
project_name = "cudly"
environment = "dev"
stack_name = "cudly-dev"
region = "us-east-1"
aws_profile = "default"
# Compute platform: "lambda" or "fargate"
compute_platform = "lambda"
# Lambda configuration
lambda_architecture = "arm64"
lambda_memory_size = 512
lambda_timeout = 60
# Database configuration (Aurora Serverless v2)
database_engine_version = "16.4"
database_name = "cudly"
database_username = "cudly"
database_min_capacity = 0.5
database_max_capacity = 2.0
database_backup_retention_days = 7
# Admin user
admin_email = "admin@example.com"
# Frontend (optional)
enable_frontend_build = true
frontend_price_class = "PriceClass_100"Local (dev): State stored in terraform.tfstate (default).
Remote (production): Configure backend using backends/*.tfbackend files:
# Initialize with a specific backend config
terraform init -backend-config=backends/prod.tfbackendExample backend config (backends/prod.tfbackend):
bucket = "cudly-terraform-state-prod"
key = "prod/terraform.tfstate"
region = "us-east-1"
encrypt = true
use_lockfile = true- Compute: Lambda (ARM64, container image) with Function URL, or Fargate (ECS) with ALB
- Database: Aurora Serverless v2 PostgreSQL 16.4 (0.5-2.0 ACU)
- Network: VPC (10.0.0.0/16) with IPv6 dual-stack, private/public subnets, no NAT Gateway
- Proxy: RDS Proxy for Lambda connection pooling
- Frontend: CloudFront (dual-origin: S3 for static, Lambda/ALB for API) + S3
- Secrets: Secrets Manager (DB password, JWT secret, session secret)
- Monitoring: CloudWatch log groups, alarms, EventBridge scheduled tasks
# Deploy to dev
./scripts/tf-deploy.sh aws dev
# Plan only
./scripts/tf-deploy.sh aws dev plan
# Deploy to staging/prod
./scripts/tf-deploy.sh aws prodcd terraform/environments/aws
FUNCTION_URL=$(terraform output -raw lambda_function_url)
curl "$FUNCTION_URL/health"
# Expected: {"status":"healthy","version":"...","timestamp":"...","checks":{"config_store":{"status":"healthy"},"auth_store":{"status":"healthy"}}}
# Monitor logs
aws logs tail /aws/lambda/cudly-dev-api --follow
aws logs tail /aws/lambda/cudly-dev-api --filter-pattern "ERROR"cd terraform/environments/aws
S3_BUCKET=$(terraform output -raw frontend_bucket)
CF_DIST_ID=$(terraform output -raw cloudfront_distribution_id)
# Build and upload
cd ../../../../frontend && npm install && npm run build
aws s3 sync dist/ "s3://${S3_BUCKET}/" --delete \
--cache-control "public,max-age=3600"
aws cloudfront create-invalidation --distribution-id "$CF_DIST_ID" \
--paths "/*"- Compute: Azure Container Apps (serverless containers)
- Database: Azure PostgreSQL Flexible Server
- Frontend: Azure Blob Storage (static website) + Azure CDN
- Secrets: Azure Key Vault
- Monitoring: Azure Monitor alerts
# Deploy with script
./scripts/tf-deploy.sh azure devaz storage blob upload-batch \
--account-name cudlyfrontendprod \
--destination '$web' --source dist/ --overwrite
az cdn endpoint purge \
--resource-group cudly-prod-rg \
--profile-name cudly-cdn-profile \
--name cudly-cdn-endpoint --content-paths "/*"- Compute: Cloud Run service (serverless containers)
- Database: Cloud SQL PostgreSQL
- Frontend: Cloud Storage + Global HTTPS Load Balancer + Cloud CDN
- Secrets: Secret Manager
- Monitoring: Cloud Monitoring alerts
- Optional: Cloud Armor (WAF/DDoS protection)
# Deploy with script
./scripts/tf-deploy.sh gcp devgsutil -m rsync -r -d -x ".*\.map$" dist/ gs://cudly-frontend-prod/
gsutil -m setmeta -h "Cache-Control:public, max-age=31536000, immutable" \
gs://cudly-frontend-prod/js/**
gcloud compute url-maps invalidate-cdn-cache cudly-url-map --path "/*"The explicit "Compare with Archera" feature is off by default. It is enabled only when all three settings are set; with none set, the platform makes no Archera request.
| Terraform variable (environment) | Runtime setting | Meaning |
|---|---|---|
archera_org_id |
ARCHERA_ORG_ID |
Archera organization UUID |
archera_plan_id |
ARCHERA_PLAN_ID |
Archera commitment plan UUID (an Archera ID, not a CUDly plan ID) |
archera_api_key_secret_arn (AWS), archera_api_key_secret_id (GCP), archera_api_key_secret_name (Azure) |
ARCHERA_API_KEY_SECRET |
Reference to the secret holding the Archera API key |
Create the secret yourself, outside Terraform, so the key never enters Terraform state; Terraform only passes the reference. The platform resolves the key lazily on the first comparison request through the configured secret provider.
Runtime access is scoped to that one secret:
- AWS (Lambda and Fargate):
secretsmanager:GetSecretValueon the secret ARN only. - GCP (Cloud Run and GKE): a per-secret
roles/secretmanager.secretAccessorbinding on that secret only. - Azure (Container Apps and AKS): no new role assignment. The secret lives in the platform Key Vault, where the runtime identity already holds the vault-wide Key Vault Secrets User role. For a secret in a different vault, add a secret-scoped role assignment for the runtime identity on that secret; narrowing the vault-wide grant is a separate owner decision.
Required disclosures, shown verbatim wherever Archera is offered (pkg/common ArcheraNonGatingDisclosure and ArcheraSponsorshipDisclosure):
This is entirely optional. CUDly's purchase and management features work fully without Archera.
For full disclosure, Archera sponsors CUDly's Open Source development from a fraction of their insurance premiums.
The GCP secret variable takes the short secret ID, not the projects/<number>/secrets/<id> resource name: the per-secret binding sets the project, and the full name makes the binding be replaced on every apply.
The comparison shows Archera's own figures as returned by its API. Amounts carry no currency (Archera provides none), premiums are already included in totals, and an unknown value is shown as unknown, never as zero. Products Archera does not document as supported are shown as "not known to be supported by Archera". Nothing in this view is a guarantee or a purchase; the existing Archera disclosures stay in place.
When infrastructure already exists and you only need to update application code:
./scripts/tf-deploy.sh aws dev # dev
./scripts/tf-deploy.sh aws prod # production
./scripts/tf-deploy.sh aws dev plan # dry runThe script: initializes Terraform if needed, runs terraform apply with the profile-specific tfvars, and shows outputs on success.
# 1. Build image
GIT_COMMIT=$(git rev-parse --short HEAD)
AWS_ACCOUNT_ID=$(aws sts get-caller-identity --query Account --output text)
IMAGE_URI="${AWS_ACCOUNT_ID}.dkr.ecr.us-east-1.amazonaws.com/cudly:${GIT_COMMIT}"
docker build --platform linux/arm64 --build-arg VERSION="$GIT_COMMIT" -t "$IMAGE_URI" .
# 2. Push to ECR
aws ecr get-login-password --region us-east-1 | \
docker login --username AWS --password-stdin ${AWS_ACCOUNT_ID}.dkr.ecr.us-east-1.amazonaws.com
docker push "$IMAGE_URI"
# 3. Force replace Lambda
cd terraform/environments/aws
terraform apply -replace="module.compute_lambda[0].aws_lambda_function.main" -auto-approve# .github/workflows/deploy.yml
name: Deploy to AWS
on:
push:
branches: [main]
jobs:
deploy:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: aws-actions/configure-aws-credentials@v4
with:
aws-access-key-id: ${{ secrets.AWS_ACCESS_KEY_ID }}
aws-secret-access-key: ${{ secrets.AWS_SECRET_ACCESS_KEY }}
aws-region: us-east-1
- run: ./scripts/tf-deploy.sh aws devYou can run Fargate alongside an existing Lambda deployment for comparison testing. Both connect to the same Aurora database. The terraform setup uses a single directory with different .tfvars and .tfbackend files per environment.
Lambda (dev) |
Fargate (fargate-dev) |
|
|---|---|---|
| VPC | 10.0.0.0/16 | 10.1.0.0/16 (separate) |
| Compute | Lambda + Function URL | ECS Fargate + ALB |
| Database | Aurora via RDS Proxy | Aurora via direct endpoint |
| Frontend | (none yet) | CloudFront + S3 |
cd terraform/environments/aws
# Initialize with fargate-dev backend
terraform init -backend-config=backends/fargate-dev.tfbackend
# Plan and apply with fargate-dev vars
terraform plan -var-file=fargate-dev.tfvars
terraform apply -var-file=fargate-dev.tfvars # ~10-15 minutes
# Get URLs
terraform output fargate_api_url
terraform output frontend_urlcd terraform/environments/aws
# Lambda (1-2s cold start, then fast)
time curl $(terraform output -raw lambda_function_url)/health
# Fargate (consistent ~100-200ms)
time curl $(terraform output -raw fargate_api_url)/healthcompute_platform = "fargate"
fargate_cpu = 512 # 256, 512, 1024, 2048, 4096
fargate_memory = 1024
fargate_desired_count = 2
fargate_min_capacity = 1
fargate_max_capacity = 10
fargate_enable_https = false
fargate_enable_execute_command = false # ECS Exec for debuggingChange compute_platform in your tfvars and redeploy. The frontend module automatically adapts the API endpoint (Function URL vs ALB DNS). Update custom domain DNS if applicable.
# Destroy only Fargate (leaves Lambda and database intact)
cd terraform/environments/aws
terraform destroy -var-file=fargate-dev.tfvarsThe frontend is served through CDN with dual-origin routing:
User Request
|
CDN (CloudFront / Azure CDN / Cloud CDN)
|
+-- /api/* --> Backend (Lambda / Container Apps / Cloud Run)
+-- /* --> Static Files (S3 / Blob Storage / Cloud Storage)
Static assets (JS, CSS, images): Cached 1 year (content-hashed filenames). Compression enabled.
HTML files: No cache (no-cache, no-store, must-revalidate).
API requests (/api/*): No caching. All headers, cookies, and query strings forwarded.
The frontend uses relative paths (/api) by default. Since the CDN proxies /api/* to the backend, requests are same-origin and CORS is not needed.
AWS CloudFront: Origin Access Control (OAC) for S3, CloudFront Function for security headers (HSTS, X-Frame-Options), custom error pages for SPA routing, optional WAF, X-CloudFront-Secret header for origin verification.
Azure CDN: Static website hosting via Blob Storage, Standard CDN or Front Door, managed SSL certificates, delivery rules for URL rewriting.
GCP Cloud CDN: Global HTTPS Load Balancer, Cloud Armor for WAF/DDoS protection, automatic managed SSL certificates.
CF_DIST_ID=$(terraform output -raw cloudfront_distribution_id)
# Check origins (should show S3 + Lambda/ALB)
aws cloudfront get-distribution --id "$CF_DIST_ID" \
--query 'Distribution.DistributionConfig.Origins' --output json
# Test path routing
CF_URL=$(terraform output -raw frontend_url)
curl -I "$CF_URL/" # X-Cache: Hit from cloudfront (static)
curl "$CF_URL/api/health" # X-Cache: Miss from cloudfront (API)| Resource | Monthly Cost |
|---|---|
| Lambda | ~$0.20 |
| Aurora Serverless v2 (0.5 ACU) | ~$43.80 |
| RDS Proxy | ~$10.95 |
| CloudFront + S3 | ~$1 |
| Secrets Manager | ~$1 |
| Total | ~$57/month |
| Resource | Monthly Cost |
|---|---|
| Fargate (0.25 vCPU, 0.5GB x 2) | ~$21.90 |
| ALB | ~$16.20 |
| Aurora Serverless v2 (0.5 ACU) | ~$43.80 |
| CloudFront + S3 | ~$1 |
| Total | ~$83/month |
| Platform | Estimated Monthly Cost |
|---|---|
| AWS Lambda | ~$57 |
| GCP Cloud Run | ~$13-27 |
| Azure Container Apps | ~$18-32 |
- ARM64 Lambda/Fargate: 20% cheaper than x86
- Aurora scales to 0.5 ACU when idle
- Lambda/Cloud Run/Container Apps scale to zero
- IPv6 dual-stack eliminates NAT Gateway costs
- CloudFront PriceClass_100: US/Europe only for lower costs
- Fargate Spot: up to 70% discount for non-critical workloads
aws logs tail /aws/lambda/cudly-dev-api --follow
aws logs tail /aws/lambda/cudly-dev-api --filter-pattern "ERROR" --since 10m# Service status
aws ecs describe-services --cluster cudly-dev-fargate --services cudly-dev-fargate \
--query 'services[0].{Status:status,Running:runningCount,Desired:desiredCount}'
# Logs
aws logs tail /ecs/cudly-dev-fargate --follow
# ECS Exec (if enabled)
aws ecs execute-command --cluster cudly-dev-fargate --task <task-id> \
--container app --interactive --command "/bin/sh"az containerapp logs show --name cudly-dev --resource-group cudly-rggcloud run services logs read cudly-dev --region us-central1Each cloud provider has lifecycle policies configured via Terraform (terraform/modules/registry/{aws,azure,gcp}/):
- AWS ECR: Keep last 10 tagged images, delete untagged after 7 days, vulnerability scanning on push
- GCP Artifact Registry: Cleanup policies for untagged and old images
- Azure ACR: ACR tasks for automated cleanup
Automated cleanup for expired sessions and completed purchase executions, running on a daily schedule:
- AWS: Lambda function triggered by EventBridge (
terraform/modules/compute/aws/cleanup-lambda/) - Azure: Azure Function App with timer trigger
- GCP: Cloud Function with Cloud Scheduler
Alternatively, use pg_cron for database-native scheduling.
docker ps # is Docker running?
docker system df # disk space
go build ./cmd/server # syntax errors?
go mod tidy # dependency issues?Terraform may show "No changes" if the image tag didn't change. Force replace:
terraform apply -replace="module.compute_lambda[0].aws_lambda_function.main" -auto-approve# Check Lambda is in VPC
aws lambda get-function-configuration --function-name cudly-dev-api --query 'VpcConfig'
# Check RDS Proxy
aws rds describe-db-proxies --db-proxy-name cudly-dev-proxy
# Check database cluster
aws rds describe-db-clusters --db-cluster-identifier cudly-dev-postgresDeclare every AWS account the deployment will call sts:AssumeRole against. Nothing is reachable
cross-account until you do.
| Deployment shape | Setting |
|---|---|
Terraform (terraform/environments/aws) |
cross_account_target_account_ids = ["111111111111", ...] in your tfvars |
CloudFormation (cloudformation/stacks/CUDly) |
CrossAccountTargetAccountIds stack parameter, comma-separated |
Both render an aws:ResourceAccount condition onto the grant. An account that is not listed is
denied by IAM, not merely by the app's account selection. Leaving the setting empty creates no
cross-account grant at all, which is the intended fail-closed default rather than an oversight.
For accounts using bastion auth mode, list the bastion's account ID rather than the target's.
Only the first hop runs on the deployment's own identity; the bastion assumes into the target on its
own identity policy, which this setting does not govern.
Upgrading an existing multi-account deployment: the grant used to be scoped by role name only (
arn:aws:iam::*:role/CUDly*), which matched a CUDly role in every AWS account rather than in yours (#1636). List your linked accounts in the same change that picks up this version. On Terraform the apply removesaws_iam_role_policy.cross_account_stsand exits 0; on CloudFormation the stack update drops the statement. Either way the first symptom otherwise is a runtimeAccessDeniedduring collection, not a failed deploy.
Multi-account support requires an AES-256-GCM encryption key for stored cloud account credentials. Terraform creates the key secret automatically (see specs/multi-account-execution/iac.md).
New environment variables (AWS Lambda):
| Variable | Source | Purpose |
|---|---|---|
CREDENTIAL_ENCRYPTION_KEY_SECRET_ARN |
Terraform → Lambda env | ARN of Secrets Manager secret holding the 32-byte AES key |
CREDENTIAL_ENCRYPTION_KEY |
Direct (local dev only) | 64-char hex key; bypasses Secrets Manager |
CUDLY_MAX_ACCOUNT_PARALLELISM |
Terraform → Lambda env | Fan-out goroutine cap for parallel account execution (default: 10) |
Migration 000011 (000011_cloud_accounts) adds:
cloud_accounts— central registry for all managed accounts (AWS/Azure/GCP)account_credentials— encrypted credential materialaccount_service_overrides— sparse per-account service config overridesplan_accounts— M2M join: which accounts a purchase plan targetscloud_account_idFK column onpurchase_executions,purchase_history,savings_snapshots,ri_exchange_history
Migration 000012 (000012_global_config_fields) adds auto_collect, collection_schedule, and notification_days_before columns to the global_config table.
Migration 000013 (000013_add_running_status) adds a running status value to the purchase_executions status enum.
Migration 000014 (000014_azure_auth_mode) adds azure_auth_mode column to cloud_accounts. Supported values: client_secret, managed_identity, workload_identity_federation.
Migration 000015 (000015_gcp_auth_mode) adds gcp_auth_mode column to cloud_accounts. Supported values: service_account_key, application_default, workload_identity_federation.
Migration 000016 (000016_aws_wif) adds aws_web_identity_token_file column to cloud_accounts, used for AWS Workload Identity Federation (token file path on the host platform).
If DB_AUTO_MIGRATE=true (the default), migration 000011 runs automatically on Lambda cold start. To run manually:
The URL scheme is
pgx5://, notpostgres://. golang-migrate selects its driver by scheme, andmigrateis built here with-tags pgx5so it does not linklib/pq(issue #1849). Apostgres://URL fails withunknown driver postgres.The password is percent-encoded with
jq -sRr @uribecause it may contain@,:,/,%,#,?or spaces, which break the URL when left raw.
DB_PASSWORD=$(aws secretsmanager get-secret-value --secret-id cudly-dev-db-password-* --query SecretString --output text | jq -r .password)
RDS_ENDPOINT=$(cd terraform/environments/aws && terraform output -raw database_proxy_endpoint)
ENCODED_PASSWORD=$(printf '%s' "$DB_PASSWORD" | jq -sRr @uri)
migrate -path internal/database/postgres/migrations \
-database "pgx5://cudly:${ENCODED_PASSWORD}@${RDS_ENDPOINT}:5432/cudly?sslmode=require" upCheck: DB_AUTO_MIGRATE=true, DB_MIGRATIONS_PATH=/app/internal/database/postgres/migrations, correct credentials. Run manually:
DB_PASSWORD=$(aws secretsmanager get-secret-value --secret-id cudly-dev-db-password-* --query SecretString --output text | jq -r .password)
RDS_ENDPOINT=$(cd terraform/environments/aws && terraform output -raw database_proxy_endpoint)
ENCODED_PASSWORD=$(printf '%s' "$DB_PASSWORD" | jq -sRr @uri)
migrate -path internal/database/postgres/migrations \
-database "pgx5://cudly:${ENCODED_PASSWORD}@${RDS_ENDPOINT}:5432/cudly?sslmode=require" upterraform force-unlock <LOCK_ID>S3 bucket is empty. Deploy frontend or add a placeholder:
S3_BUCKET=$(terraform output -raw frontend_bucket)
echo "<h1>CUDly</h1>" | aws s3 cp - "s3://${S3_BUCKET}/index.html"aws ecs describe-services --cluster "$CLUSTER" --services "$SERVICE" --query 'services[0].events[:5]'
# Common: image pull errors, resource limits, health check failuresTest ALB directly to isolate:
ALB_DNS=$(terraform output -raw fargate_alb_dns_name)
curl "http://${ALB_DNS}/health"
# If ALB works but CloudFront doesn't: check origin settings, custom header, security groupsDeploy the previous git commit:
git log --oneline -5
git checkout <previous-commit>
./scripts/tf-deploy.sh aws dev
git checkout main