This file tracks outstanding limitations and things that require operator action. Resolved items are moved to the Resolved section at the bottom.
Until issue #1545 was fixed, arm/CUDly-CrossSubscription/template.json
declared two things that reached past the subscription being onboarded:
- A second role assignment carrying
"scope": "/providers/Microsoft.Capacity", in addition to the intended subscription-scope assignment. Written as an absolute path, that scope denotes tenant-wide reservation orders: an assignment there covers every reservation order in the Azure AD tenant, including subscriptions the customer never onboarded. /providers/Microsoft.Capacityin the custom role definition'sassignableScopes, which declares the role eligible to be assigned at tenant scope by anyone who can create role assignments there.
What (1) actually produced is worth recording, because it is not what the
template appears to say. Running az deployment sub validate against the
pre-fix template resolves that assignment to
/subscriptions/<subId>/providers/providers/Microsoft.Capacity/providers/Microsoft.Authorization/roleAssignments/<guid>
Note the doubled providers/providers. In a subscription-scoped deployment
ARM appends the scope value beneath the subscription rather than treating it
as an absolute tenant path, so the most likely apply-time outcome is a
malformed target that fails or lands somewhere meaningless, not a clean
tenant-wide grant. Whether it ever resolved to a real tenant-scope assignment
(for example when deployed by a principal holding tenant-root authority) has
not been established, and cannot be without applying the template to a live
tenant.
Treat it as possibly live rather than assuming either way: verify, then revoke if present. (2) is a real widening regardless of how (1) resolved.
The template no longer creates that assignment, and /providers/Microsoft.Capacity
has been dropped from the role definition's assignableScopes. This matches
terraform/modules/iam/azure/cudly-reservation-role, whose
include_capacity_provider_scope flag has always defaulted to false, and
iac/federation/azure-target/terraform, which has only ever assigned at
subscription scope.
ARM deployments are incremental: removing the resource from the template does not revoke anything it previously created. Anyone who deployed the template before this fix keeps whatever it granted until they delete it by hand. Redeploying alone is not sufficient.
Remediation, in this order:
# 1. Check what the pre-fix template actually left behind. This takes TWO
# queries, because neither one alone can see both shapes described above.
# Project the assignment id in each: it is what step 2 deletes by.
#
# 1a. The tenant-level provider path. `--all` cannot reach this: the CLI
# documents it as "show all assignments under the current subscription",
# and /providers/Microsoft.Capacity sits outside any subscription, so a
# surviving tenant-wide grant would not appear in 1b at all. Nor is it a
# parent scope of the subscription, so --include-inherited does not
# surface it either. Query the scope directly. Anything returned here is
# over-broad by definition, so there is no filter to get wrong.
# Reading at this scope needs tenant-level rights (User Access
# Administrator at tenant root, or Global Administrator with elevated
# access); an authorization error here is NOT an all-clear -- re-run it
# with a principal that can read the scope.
az role assignment list \
--assignee <SP-object-id> \
--scope "/providers/Microsoft.Capacity" \
--query "[].{id:id, scope:scope, role:roleDefinitionName}" \
-o table
# 1b. The subscription and below, which is where the malformed
# doubled-providers target shown above would land. Run once per onboarded
# subscription (`az account set --subscription <subId>` between runs).
# Filtered with `grep -i`, not a JMESPath `--query "[?contains(...)]"`:
# JMESPath's contains() is case-sensitive, while ARM provider namespaces
# are not, so a row stored as /providers/microsoft.capacity satisfies the
# grant and silently fails the filter. Do not reintroduce a
# case-sensitive path match here.
# Projected as a JMESPath list, not a hash, so the tsv column order is
# fixed by the query rather than by key ordering: id, scope, role. Step 2
# deletes by the FIRST column.
az role assignment list \
--assignee <SP-object-id> \
--all \
--query "[].[id, scope, roleDefinitionName]" \
-o tsv | grep -i 'microsoft\.capacity'
# 2. Revoke anything step 1 listed, FIRST, before redeploying.
# Delete by --ids, not by --scope: the pre-fix template could produce the
# malformed doubled-providers scope shown above, and `az role assignment
# delete --scope /providers/Microsoft.Capacity` rejects that as an invalid
# scope, leaving the row listed but undeletable. The id always works.
az role assignment delete --ids <id-from-step-1> [<id> ...]
# 3. Then redeploy the corrected template to narrow assignableScopes.
az deployment sub create \
--location eastus \
--template-file arm/CUDly-CrossSubscription/template.json \
--parameters servicePrincipalObjectId=<SP-object-id> \
--name CUDly-CrossSubscription \
--no-promptBoth step-1 queries returning nothing is a good outcome, and the expected one
if the malformed target described above simply failed to apply. Only 1a and 1b
together are a clean result: 1b alone cannot see a tenant-level grant, and 1a
alone cannot see the malformed subscription-relative one. Step 3 is still
required either way: it is what removes the tenant entry from
assignableScopes.
The order matters: Azure refuses to remove an assignable scope from a role definition while assignments still exist at that scope, so redeploying before step 1 can fail on the role-definition update.
Purchases are unaffected by the narrower grant. Azure authorises a reservation
purchase against the subscription named in the request body's billingScopeId,
not against the tenant-level Microsoft.Capacity provider path, which is why
the Terraform onboarding path has always worked with subscription scope alone.
If a deployment ever does need a tenant-wide grant, it must be applied manually
as an explicitly consented step (the az role assignment create mirror of
step 1) and must not be reintroduced into the template.
scripts/check-azure-role-parity.sh fails CI if it is.
The built-in "Reservation Purchaser" role (f7b75c60-3036-4b75-91c3-6b41c27c1689)
does not include Microsoft.Capacity/calculateprice/action,
Microsoft.Capacity/reservationorders/write, or
Microsoft.BillingBenefits/savingsPlanOrderAliases/write. Without these,
the live purchase API returns 403.
arm/CUDly-CrossSubscription/template.json has been updated (fix/731-arm-roles)
to add a custom role "CUDly Reservation and Savings Plan Purchaser" that enumerates
all required actions explicitly. Existing tenants who applied the ARM template before
this fix MUST re-deploy it:
az deployment sub create \
--location eastus \
--template-file arm/CUDly-CrossSubscription/template.json \
--parameters servicePrincipalObjectId=<SP-object-id> \
--name CUDly-CrossSubscription \
--no-promptUntil re-deployed, PurchaseCommitment and ValidateOffering for savings plans
will continue to return 403.
-
ARM template role-definition scoping +
Reservation Readertenant gap: Resolved.arm/CUDly-CrossSubscription/template.jsonnow uses unscoped global role-definition paths (/providers/Microsoft.Authorization/roleDefinitions/{id}) and drops the fragileReservation Readerassignment in favour of aReservation Purchaserassignment at/providers/Microsoft.Capacityscope (a superset available in every tenant). Operators who previously applied the buggy template may need to clean up the orphaned subscription-scopedReservation Readerassignment manually withaz role assignment delete --assignee <sp-object-id> --role "Reservation Reader" --scope /subscriptions/<subId>. Superseded by issue #1545: the/providers/Microsoft.Capacityassignment described here was a workaround forReservation Readernot existing in every tenant. The custom role introduced by issue #731 removed that need, but the tenant-wide assignment was left behind and silently granted access across the whole tenant. It has since been removed; do not reintroduce it. See the #1545 entry under Outstanding for the revocation steps existing deployments still need. -
Azure ACS SMTP credential generation requires manual portal step: Microsoft's API gap remains (no REST endpoint generates ACS SMTP credentials; the Portal is the only supported path). The ergonomic gap is closed:
scripts/azure-smtp-setup.shprints a pre-filled checklist with the direct Azure Portal URL plus the exactaz keyvault secret setcommands for this deployment. Thesmtp_setup_instructionsTerraform output surfaces the command to run at the end ofterraform apply. Seespecs/azure-smtp-setup.mdfor the runbook and troubleshooting.
- Helm LoadBalancer IP may not be available on first apply:
Resolved.
terraform/modules/compute/azure/aks/main.tfnow emits atime_sleep.wait_for_lb_ip(5-minute create_duration) between thehelm_release.nginx_ingressand thekubernetes_servicedata source read, covering Azure's typical 2–5 minute LB provisioning window. First-apply no longer requires a follow-up run in the common case; thetry()fallback on the output still handles the rare beyond-budget provisioning tail.
-
Same-family-only recommendations: Fully resolved for the allowlisted family groups — advisory names in the first pass (commit edc8d7838), real offering IDs +
EffectiveMonthlyCostranking in the follow-up (commit 0347b3111).pkg/exchange.ReshapeRecommendationnow carries anAlternativeTargets []OfferingOptionfield (renamed from the earlierAlternativeTargetInstanceTypes []string— note for anyone auditing JSON payloads: the response key changed fromalternative_target_instance_typestoalternative_targets).providers/aws/services/ec2/client.go's newFindConvertibleOfferingsbatches all candidate instance types into ONEDescribeReservedInstancesOfferingscall per reshape page load (≤4 API calls for a diverse fleet; 1 for a homogeneous one) and ranks by monthly cost.pkg/exchange.AnalyzeReshapingWithOfferingscomposes the base analyzer with offering enrichment; the auto-exchange pipeline still uses the plainAnalyzeReshaping(no pricing needed) so automated behaviour is unchanged. Allowlist covers general-purposem5/m6i/m7g, compute-optimisedc5/c6i/c7g, memory-optimisedr5/r6i/r7g, burstablet3/t3a/t4g. Specialty (p*/g*/x*/hpc*) and legacy-generation (m4/c4/r3) families are deliberately out of the allowlist — see the follow-up below.The reshape-recommendations dashboard page renders the alternatives as a new "Alternatives" column with per-instance
$X.XX/mocost chips (commit 97fc2597d); when the user clicks "Exchange" from a reshape row, the modal receives the rec'salternative_targetsand shows a matching cost chip next to each target-offering input plus a live-updating running total (sum(chip.cost × row.count)). End-to-end coverage is exercised by the handler integration test atinternal/api/handler_ri_exchange_integration_test.go(build-tagintegration, commit da762067c) which wires a real Postgres through the reshape handler with mocked AWS clients via newly- added factory injection points on the Handler struct (reshapeEC2Factory/reshapeRecsFactory, both nil-safe so prod behaviour is unchanged). -
Multi-target exchange: Fully resolved — backend (commit 5eb274690) and frontend (commit 2ff1ebe89).
pkg/exchange.ExchangeQuoteRequestandExchangeExecuteRequestaccept aTargets []TargetConfigslice; legacyTargetOfferingID/TargetCountfields are retained as a single-target alias so existing callers keep working. The HTTP API gains an optionaltargets[]array on the quote + execute bodies; when present it wins over the legacy singleton fields. Spend-cap semantics: AWS returns a single aggregatedPaymentDueacross all targets, somax_payment_due_usdnaturally functions as a TOTAL cap for multi-target requests. Dashboard modal gained add/remove target rows: the modal posts the singleton shape when exactly one row is present (preserving existing wire format) and poststargets[]when ≥2 rows are present. With commit 97fc2597d the modal also shows per-row cost chips (when the caller suppliesalternativeTargets) and a running total that updates live as the user edits offering-type / count inputs. -
Utilization caching: Resolved with a Postgres-backed TTL cache plus stale-while-revalidate on non-Lambda runtimes. Migration
000031_ri_utilization_cacheaddsri_utilization_cache (region, lookback_days, payload, fetched_at).internal/api/handler_ri_exchange.goroutes bothgetRIUtilizationandgetReshapeRecommendationsthrough the cache wrapper (internal/api/ri_utilization_cache.go) so one Cost Explorer call per TTL window serves every warm and cold Lambda container. Two TTL knobs:CUDLY_RI_UTILIZATION_CACHE_TTL(default15m, soft-freshness window) andCUDLY_RI_UTILIZATION_CACHE_STALE_TTL(default30m, hard expiry). On non-Lambda, reads in[soft, hard)serve the stale row and kick a singleflight-guarded background refresh (golang.org/x/sync/singleflight); reads pasthardforce a synchronous refetch. Lambda runtimes always synchronously refetch on any staleness — background goroutines aren't safe there (containers freeze between invocations). Errors are never cached — a transient CE 5xx cannot lock the dashboard out for the full TTL. Observability:logging.Infofon SWR kick and hard-expiry paths;logging.Debugfon the Lambda-skip branch. See the Config section ofspecs/recommendations-cache.md. End-to-end Postgres integration test atinternal/api/ri_utilization_cache_integration_test.go(build-tagintegration).
- Migration 000027 non-idempotent on fresh DBs: Resolved.
internal/database/postgres/migrations/000027_savings_snapshots_pk.up.sqlnow runsALTER TABLE savings_snapshots DROP CONSTRAINT IF EXISTS savings_snapshots_pkey;before the existing DELETE CTE + ADD CONSTRAINT sequence. The guard makes the migration safe on fresh containers (where 000018 already added the PK) without changing behaviour on production DBs where 000027 was the first to add the PK. Theinternal/api/ri_utilization_cache_integration_test.gobootstrap now uses the standardmigrations.RunMigrationspath instead of the earlier table-create workaround.
- t.Parallel() adoption (partial): Resolved for three audit-safe
packages —
pkg/exchange/{auto,exchange,reshape}_test.go,providers/aws/services/ec2/client_test.go, andinternal/api/validation_test.go. Remaining packages haven't been audited per-file and keep their sequential execution — see the follow-up below.
-
Cross-family RI recommendations for specialty + legacy families— RESOLVED. ExtendedpeerFamilyGroupsinpkg/exchange/reshape.gowith specialty (p3/p4d/p5,g4dn/g5,hpc6a/hpc6id/hpc7g) and legacy-generation (m4/m5,c4/c5,r3/r4/r5) groups. Added a localpassesDollarUnitsCheck(srcNF, srcMonthlyCost, srcCurrency, target)pre-filter applied infillAlternativesFromOfferings: a target survives only iftarget.NF × target.EffectiveMonthlyCost >= src.NF × src.MonthlyCost(with an explicit currency-equality guard that's a no-op when either side is empty). The check approximates AWS's runtime two-parallel-≥-checks rule using the already-computedEffectiveMonthlyCost(which folds upfront amortisation + recurring + usage), so no per-pairGetReservedInstancesExchangeQuoteAPI calls are needed — false positives are caught by the existingauto.goIsValidExchange=falseskip path at execution time.OfferingOptiongainedNormalizationFactor+CurrencyCodefields populated byFindConvertibleOfferings;ConvertibleRIgainedCurrencyCode+RecurringHourlyAmountpopulated byListConvertibleReservedInstances;RIInfogainedMonthlyCost+CurrencyCodepopulated by both API and server handlers via a newmonthlyCostFromConvertibleRIhelper using AWS's canonical(FixedPrice/hours_per_term + UsagePrice + recurring_hourly) × 730formula. Follow-up: make the family allowlist obsolete by sourcing cross-family candidates from CUDly's already-cached Cost Explorer RI purchase recommendations (data we already collect) instead of a hardcoded family list or a per-rec offering API enumeration — seeknown_issues/24_exchange_offering_cache.mdfor the full design. -
t.Parallel() adoption for remaining packages: Adoption is complete only for
pkg/exchange/,providers/aws/services/ec2/, andinternal/api/validation_test.go. Other packages need a per-test-file audit for shared state before parallelizing:internal/api/(other test files besidesvalidation_test.go) use handler fixtures and shared mocks; not race-safe without review.internal/config/*_test.gointegration tests share a Postgres container and cannot naively parallelize.internal/server/app_test.gouses package-level vars (runMigrations,migrationsTimeout) that are not race-safe.- Any test file using
os.Setenv/t.Setenvfor process-wide state needs verification that the variable scope is per-test.
Expected incremental speedup is meaningful but each package needs its own small audit commit; scheduled as ad-hoc cleanup rather than a single sweeping change.
-
Migration 000027 non-idempotent on fresh DBs: Integration tests that spin up a fresh Postgres via
testcontainers-gocan't run the full migration set — migration 000027 (savings_snapshots_pk) tries toADD PRIMARY KEYthat migration 000018 already added, failing with "multiple primary keys for table". Production DBs aren't affected because they were already in the "duplicate rows needing dedup" state that 000027 was written to fix. Fix: make the ADD CONSTRAINT idempotent (e.g. DROP CONSTRAINT IF EXISTS first, or wrap in a conditional PL/pgSQL block) without changing the behaviour on already-migrated databases. Tracked separately because it requires careful review against real prod migration history. Commit2d8f1e2baworks around it by bypassing migrations entirely for the cache integration test (creates only theri_utilization_cachetable directly). -
GCP account
serene-bazaar-666deploy SA missingcompute.regions.list: Visible in production Lambda logs (2026-04-21T16:28:22Zand onward):[ERROR] GCP account GCP serene-bazaar-666 (serene-bazaar-666): get recommendations: failed to get regions: failed to list regions: googleapi: Error 403: Required 'compute.regions.list' permission for 'projects/serene-bazaar-666'The deploy service account that CUDly impersonates for that project doesn't have
roles/compute.viewer(or a custom role that includescompute.regions.list). Two paths to fix:- Operator action (preferred): grant the GCP service account
roles/compute.vieweron the project (or a narrower custom role containingcompute.regions.list+compute.zones.listif least- privilege matters). - Code-side mitigation: the GCP region-fetch already short-circuits
on errors but every fetch attempt logs as
[ERROR]. The collector could downgrade to[WARN]for permission errors specifically (so the operator notices once but the noise stops) — tracked as a follow-up inknown_issues/22_scheduler.mdunder the silent- failure entry.
The collector's account-failure-swallow bug masks this entirely: the GCP provider is reported as successful even when this account fails, so the operator only sees the issue if they tail logs.
- Operator action (preferred): grant the GCP service account
-
Per-plan-type SP split: caveats exposed in plans/recommendations views: The migration to four per-plan-type Savings Plans cards (Compute / EC2 Instance / SageMaker / Database) replaces the umbrella
(aws, savings-plans)ServiceConfig row with four per-plan-type rows and rewritespurchase_plans.servicesJSONB keys atomically (migration 000040). Two pre-existing UX limitations are now visible with the split and are tracked here as follow-ups, not blockers:-
Multi-SP purchase-plan summary shows only one plan type.
frontend/src/plans.ts:231renders a plan summary by reading the FIRST entry fromplan.services(a JSONB-derived map). A purchase plan that targets multiple SP plan types (e.g., Compute + SageMaker) will list only one — whichever sorts first — in its summary card. Pre-split this was hidden because the singleaws:savings-planskey always rendered as "Savings Plans"; post- split the same plan now has four keys and only one displays. Fix is plans.ts-only: render a comma-separated list or a count badge when multiple SP plan types are present in the same plan. Out of scope for the issue #22 follow-up PR. -
Bulk-buy-from-Recommendations no longer sees "all SP types" rows. The bulk-buy modal in recommendations.ts groups recommendations by
(provider, service). Pre-split, every SP recommendation sharedservice: "savings-plans", so a Compute SP rec and a SageMaker SP rec landed in the same bucket and could be bought in one click. Post-split, each plan type is its own service, so an operator who used to bulk-buy SP must now bulk-buy four times (once per plan type). Fix is a UI-side aggregator that groups byIsSavingsPlan(rec.service)for the bulk-buy view only, leaving the underlying service distinction intact for the per-card save path. Out of scope for the issue #22 follow-up PR.
-
-
OpenSearch RI tagging: best-effort, may be rejected by AWS: Implemented in
providers/aws/services/opensearch/client.go. The client now resolves the caller's AWS account ID via STS (cached on first tag call), constructs an ARN (arn:aws:es:<region>:<account>:reserved-instance/<id>), and callsopensearch:AddTagspost-purchase. AWS documentation only explicitly supports AddTags on domain/data-source/application ARNs, so the call MAY be rejected with aValidationException. When that happens,retry.ErrPermanentshort-circuits the retry budget and the failure is logged at WARN — the purchase still succeeds. If AWS extends AddTags to cover reserved-instance ARNs (or CUDly switches to ResourceGroupsTaggingAPI if that ever adds the resource type), the code will start working with no change. Source is also persisted inpurchase_history.sourcefor DB-side reconciliation.