diff --git a/src/content/blog/aws-resource-control-policies-rcp.md b/src/content/blog/aws-resource-control-policies-rcp.md index cb96577..c7039f3 100644 --- a/src/content/blog/aws-resource-control-policies-rcp.md +++ b/src/content/blog/aws-resource-control-policies-rcp.md @@ -2,7 +2,7 @@ title: "AWS Resource Control Policies (RCPs) vs SCPs" description: "AWS resource control policies are Organizations resource-based guardrails: they cap what any caller—including principals outside the org—can do to member-account resources. SCP vs RCP vs IAM, a public-S3 deny, sandbox OU tests." pubDate: 2026-08-27 -updatedDate: 2026-08-27 +updatedDate: 2026-10-02 author: OpenSourceOM Team tags: - AWS @@ -61,7 +61,7 @@ They **do not** apply to: - **AWS managed KMS keys** - `kms:RetireGrant` (special-cased) -When you enable RCPs, AWS attaches `RCPFullAWSAccess` (allow all through the RCP layer) to root, OUs, and accounts. Removing that without a replacement is how you freeze S3. +When you enable RCPs, AWS attaches `RCPFullAWSAccess` (allow all through the RCP layer) to root, OUs, and accounts, and does not let you detach it. A custom RCP Allow therefore does not narrow the account. Narrowing is a Deny, and that Deny is decided before any Allow — [SCP vs RCP evaluation order](/blog/aws-scp-rcp-evaluation-order/). ## SCP vs RCP vs IAM @@ -130,7 +130,7 @@ Remaining public-or-cross-account edges after RCP belong on [attack path analysi ## Checklist -- [ ] Policy type `RESOURCE_CONTROL_POLICY` enabled; `RCPFullAWSAccess` left in place +- [ ] Policy type `RESOURCE_CONTROL_POLICY` enabled; custom statements are Deny, because `RCPFullAWSAccess` cannot be detached - [ ] First Deny attached to a sandbox OU/account, not the org root - [ ] External-account GetObject test **fails**; in-org app role GetObject **succeeds** - [ ] `aws:PrincipalIsAWSService` exception verified with Trail/Config diff --git a/src/content/blog/aws-scp-rcp-evaluation-order.md b/src/content/blog/aws-scp-rcp-evaluation-order.md new file mode 100644 index 0000000..dd77db3 --- /dev/null +++ b/src/content/blog/aws-scp-rcp-evaluation-order.md @@ -0,0 +1,117 @@ +--- +title: "SCP vs RCP Evaluation Order" +description: "AWS checks every explicit Deny before RCP and SCP Allows. RCPFullAWSAccess cannot be detached, so an RCP narrows access only with Deny — an identity Allow never skips that." +pubDate: 2026-10-02 +updatedDate: 2026-10-02 +author: OpenSourceOM Team +tags: + - AWS + - Organizations + - SCP + - RCP + - IAM +focusKeyword: SCP RCP evaluation order +faq: + - question: Does AWS evaluate SCPs before RCPs? + answer: >- + No. The enforcement code first looks for an explicit Deny in every + applicable policy, including both SCPs and RCPs. Only if none match + does it require an Allow in RCPs, then an Allow in SCPs, then + identity and resource policies. An Allow in an SCP does not skip + an RCP Deny. + - question: Can an RCP Allow replace RCPFullAWSAccess and narrow the account? + answer: >- + No. When you enable RCPs, AWS attaches RCPFullAWSAccess and does not + let you detach it, so the RCP Allow step always succeeds. A second + RCP that only Allows s3:GetObject does not remove the pass-through + Allow. Narrowing is a Deny statement, and that Deny is decided in + the first step. + - question: Why does a bucket-account SCP fail to stop another account? + answer: >- + SCPs apply to principals in the account where the SCP is attached. + A caller in a different account is not that principal. The bucket + account's RCP still applies, because RCPs apply to the resource. + A bucket policy Principal of that other account does not skip the + RCP Deny. +--- + +An identity policy Allows `s3:GetObject`. The bucket policy Allows the same role. SCP `FullAWSAccess` is attached. The call still returns `AccessDenied`, and the CloudTrail `errorMessage` mentions an Organizations policy. The statement that matched is an RCP `Deny`. It was decided **before** any of those Allows were consulted. + +**SCP vs RCP evaluation order** is that sequence. Which policy type attaches to principals vs resources is [AWS resource control policies](/blog/aws-resource-control-policies-rcp/). This page is only the order the enforcement code walks. Official sequence: [How AWS evaluates requests to allow or deny](https://docs.aws.amazon.com/IAM/latest/UserGuide/reference_policies_evaluation-logic_policy-eval-denyallow.html). + +## The sequence + +For a request evaluated in an account that is an Organizations member: + +1. **Explicit Deny.** Every applicable policy is scanned for a `Deny` that matches the request: SCPs, RCPs, identity-based policies, resource-based policies, permissions boundaries, session policies. One match is a final Deny. Evaluation stops. +2. **RCP Allow.** If RCPs are enabled, some RCP statement must Allow the action. `RCPFullAWSAccess` is attached to the root, every OU, and every account, and AWS does not let you detach it, so this step finds an Allow. A custom RCP Allow is redundant with that pass-through. +3. **SCP Allow.** Some SCP statement must Allow the action. `FullAWSAccess` is the usual pass-through. Unlike `RCPFullAWSAccess`, the SCP managed policy **can** be detached. If nothing else Allows, this step is a final Deny. +4. **Resource-based policy, then identity-based policy.** In the same account, an Allow in either is enough for most services. IAM role trust policies and KMS key policies still need their own Allow. An implicit deny in the identity policy does not cancel a resource-policy Allow granted directly to the IAM user or to the role-session ARN. +5. **Permissions boundary.** When the Allow came from an identity policy, the boundary must Allow the action too. A resource-policy Allow granted directly to the IAM user or the role-session ARN is not limited by a missing Allow in the boundary. An explicit Deny in the boundary already finished the request in step 1. +6. **Session policy.** If the caller passed a session policy and it does not Allow the action, final Deny. No session policy means this step Allows. The same direct resource-policy grant to a role-session ARN is not limited by an implicit deny in the session policy. + +``` +request + → any explicit Deny? stop, Deny + → RCP Allow present? else Deny (RCPFullAWSAccess always is) + → SCP Allow present? else Deny + → identity or resource Allow + → boundary Allow + → session Allow + → Allow +``` + +There is no “closest policy wins,” and a child OU Allow does not outrank a parent Deny. The parent Deny already matched in step 1. + +## What an RCP can actually change + +Because step 2 always sees `RCPFullAWSAccess`, an RCP changes the decision only by matching a **Deny in step 1**. + +| What you attach | What the order does with it | +| --- | --- | +| RCP `Deny` on `s3:*` unless `aws:PrincipalOrgID` matches | Step 1. Final Deny for callers the condition catches. Later Allows are not read. | +| RCP `Allow` of only `s3:GetObject` | Step 2 still passes via `RCPFullAWSAccess`. Other S3 actions stay allowed if IAM allows them. | +| SCP `Deny` of `iam:CreateUser` | Step 1. Final Deny. An AdministratorAccess identity policy is not consulted. | +| SCP `Allow` of only `ec2:*` and `s3:*`, with `FullAWSAccess` detached | Step 3. Actions outside that Allow are a final Deny. | +| Identity `Allow` of `s3:*` | Reached only if steps 1–3 did not already Deny. | + +A sandbox test that “proves” the RCP Allow list works is usually proving that `RCPFullAWSAccess` is still attached. Detach is not available. The test that matters is an explicit Deny from a principal the condition should catch, and a success from a principal it should not. + +The condition operators on that Deny are evaluated inside step 1, not after IAM. `StringNotEqualsIfExists` does not match when the key is absent, so that Deny statement does not apply and evaluation continues. That is a different bug from order; the statement shape is in the [RCP vs SCP](/blog/aws-resource-control-policies-rcp/) page. A Deny that **does** match is never undone by an Allow later in the list. + +## External callers + +SCPs apply to **principals in the account where the SCP is attached**. RCPs apply to **resources in the account where the RCP is attached**. + +A caller in account B, reading a bucket in account A: + +- Account A’s SCPs do not apply to B’s principal. An SCP Deny in A is invisible to this request. +- Account A’s RCPs do apply. A matching RCP Deny is step 1 in A and is final. +- Account A’s bucket policy is the resource-based policy. Its Allow does not skip the RCP Deny. +- Account B still evaluates B’s own SCPs and B’s identity policy on the way out. B’s SCP Allow does not skip A’s RCP Deny. + +Same-account evaluation and cross-account evaluation are different flows. The cross-account diagram is [Cross-account policy evaluation](https://docs.aws.amazon.com/IAM/latest/UserGuide/reference_policies_evaluation-logic-cross-account.html). The practical rule on this page: a resource policy `Principal` of another account is not evidence that the resource account’s SCP ran, and it is not evidence that the RCP did not. + +Service-linked roles skip SCPs and RCPs. A Deny that “should” have caught CloudTrail and did not is often that skip, not a reordering. The service-principal condition belongs on the Deny statement itself. + +## Two tests + +Run both in a sandbox OU before the Deny moves anywhere else. There is no RCP audit mode. + +1. **In-org role that IAM allows.** `s3:GetObject` succeeds. If it fails, the Deny condition is matching principals you meant to keep (org id typo, missing service-principal exception). +2. **Principal outside the org** whose bucket policy Allows them. `s3:GetObject` is `AccessDenied`. If it succeeds, the RCP Deny did not match. The bucket account’s SCP was never going to be the control that failed. + +CloudTrail in the bucket account shows the denied call. The identity policy in the caller account will still look correct. That is what step 1 looks like from the outside. + +Failure mode: debugging for a day inside the role’s permission boundary because the boundary is step 5 and the RCP already finished at step 1. Read `errorMessage` for the Organizations policy id before editing IAM. + +## Checklist + +- [ ] RCP statements that are meant to constrain are `Deny`. An RCP `Allow` list is not a ceiling while `RCPFullAWSAccess` is attached +- [ ] SCP `FullAWSAccess` detached only when another SCP Allow covers every action you intend to keep +- [ ] External-account `GetObject` test fails; in-org app role `GetObject` succeeds +- [ ] Parent OU Deny tested from a child account. A child Allow did not change the result +- [ ] Service-linked roles and `aws:PrincipalIsAWSService` checked against CloudTrail delivery, not assumed from the order +- [ ] Access Analyzer or the IAM policy simulator used with the Organizations policies included. A simulator run that omits RCPs will show Allow + +**Related:** [AWS resource control policies](/blog/aws-resource-control-policies-rcp/) · [AWS security best practices](/blog/aws-security-best-practices-2026/) · [Attack path analysis](/blog/attack-path-analysis-cloud-security/) diff --git a/src/content/blog/azure-blueprint-locks-deny-assignments.md b/src/content/blog/azure-blueprint-locks-deny-assignments.md new file mode 100644 index 0000000..c21d0db --- /dev/null +++ b/src/content/blog/azure-blueprint-locks-deny-assignments.md @@ -0,0 +1,106 @@ +--- +title: "Azure Blueprint Locks Are Deny Assignments" +description: "Blueprint Read Only and Do Not Delete locks are deny assignments. Microsoft removes those denies on 31 January 2027. Map them onto a deployment stack before then." +pubDate: 2026-10-02 +updatedDate: 2026-10-02 +author: OpenSourceOM Team +tags: + - Azure + - Blueprints + - deny assignments + - deployment stacks + - landing zone +focusKeyword: Azure Blueprint lock deny assignment +faq: + - question: What happens to blueprint locks on 31 January 2027? + answer: >- + Azure Blueprints is retired that day. The API stops, and blueprint + locks — the deny assignments — are removed. Resources the blueprint + deployed stay in place and keep ordinary RBAC. Subscription Owners + can then delete or edit what the lock was blocking. Policy assignments + are a different object and are not removed by this retirement. + - question: Which deployment stack deny setting replaces Read Only? + answer: >- + denyWriteAndDelete. Do Not Delete maps to denyDelete. A template spec + stores the template and creates no deny assignment. The lock exists + only if you deploy that template with a stack and set deny settings. + - question: Can I delete the blueprint deny assignment from IAM? + answer: >- + Not while the blueprint assignment still owns it. Change the lock + mode or remove the assignment after a deployment stack is already + enforcing the same deny. Deleting the blueprint first drops the + protection for the whole window until the stack exists. +--- + +The hub firewall’s IAM blade shows a deny assignment the platform team did not create with `az role assignment create`. The name and description point at a blueprint assignment, lock mode `AllResourcesDoNotDelete`. That deny is the lock. It is not an Azure Policy `Deny`, and it is not on the role-assignment list. + +**Azure Blueprint locks** are deny assignments created by the blueprint assignment. How deny assignments differ from Policy, and how landing zones use them on purpose, is [landing-zone deny assignments](/blog/azure-landing-zone-deny-assignments/). This page is the blueprint lock itself, and the date Microsoft removes it. + +Retirement schedule: [Azure Blueprints retirement](https://learn.microsoft.com/en-us/azure/governance/blueprints/blueprint-retirement). Lock behavior: [resource locking](https://learn.microsoft.com/en-us/azure/governance/blueprints/concepts/resource-locking). + +## What is already on the clock + +| Date | What changes | +| --- | --- | +| 31 July 2026 | New blueprint definitions and versions can no longer be created. This date has passed. | +| 31 October 2026 | Existing definitions can no longer be modified. New blueprint assignments can no longer be created. | +| 31 December 2026 | Existing assignments can no longer be modified. | +| 31 January 2027 | The API stops. Definitions and assignments disappear from the portal. **Blueprint deny assignments are removed.** Deployed resources remain. | + +A subscription Owner who cannot delete the hub today can delete it on 1 February 2027 if the only control was the blueprint lock. Policy assignments the blueprint also deployed are separate resources; this retirement does not delete them. Anything you were enforcing only with the lock — not with Policy, not with a management lock, not with a deployment stack — goes away that day even if you never press delete. + +Export definitions you still need before 31 January 2027. After retirement they are not recoverable from the service. + +## Lock mode to deny setting + +| Blueprint lock | What the deny assignment blocks | Deployment stack `--deny-settings-mode` | +| --- | --- | --- | +| None | Nothing from the blueprint | `none` | +| Do Not Delete (`AllResourcesDoNotDelete`) | Delete | `denyDelete` | +| Read Only (`AllResourcesReadOnly`) | Write and delete | `denyWriteAndDelete` | + +The blueprint assignment’s identity is excluded so the blueprint can still update its own resources. Everyone else, including Owner, hits `AuthorizationFailed`. `az role assignment list` does not show this object. List deny assignments: + +```bash +az rest --method GET \ + --url "https://management.azure.com/subscriptions/${SUB}/providers/Microsoft.Authorization/denyAssignments?api-version=2022-04-01" +``` + +Match `properties.description` / `denyAssignmentName` to the blueprint assignment. Portal: Subscription → Access control (IAM) → Deny assignments. “Created by” is the tell. A `CanNotDelete` resource lock is a different blade and survives blueprint retirement; do not confuse the two when you inventory what is actually protecting the hub. + +A template spec (`Microsoft.Resources/templateSpecs`) stores a versioned template. It does not create a deny assignment. Moving the blueprint JSON into a template spec and stopping there drops the lock on the day you remove the blueprint, not on 31 January 2027. + +## Replace the lock before you remove the blueprint + +Put the stack at the **parent** of the resources — management group for a subscription-scoped set, subscription for a resource-group set — so the people who have Owner on the workload cannot delete the stack and take the deny with it. Microsoft’s migration path is deployment stacks, not a second blueprint. + +```bash +az stack sub create \ + --name platform-hub \ + --location eastus \ + --subscription "${SUB}" \ + --template-file hub.bicep \ + --action-on-unmanage detachAll \ + --deny-settings-mode denyDelete \ + --deny-settings-excluded-principals "${PLATFORM_GROUP_OBJECT_ID}" +``` + +`denyDelete` is the Do Not Delete equivalent. Use `denyWriteAndDelete` for Read Only. Excluded principals are Entra object ids, at most five; a platform group is the usual one, because removing a person should be a group edit rather than a stack update. `--action-on-unmanage detachAll` means deleting the stack later detaches the resources instead of deleting them. `deleteAll` deletes the managed resources. On a hub, that flag is the outage. + +Confirm the stack’s deny assignment is listed and that a non-excluded Owner still cannot delete a protected resource. Then remove the blueprint assignment. The other order — delete the blueprint, then author the stack — is a window where Owner works. Do that in a sandbox subscription, not on the hub. + +Deployment stacks and who is allowed to edit `denySettings` are easy to widen: subscription Owner on the stack’s scope can change the mode. Parent scope is the control. Break-glass for the stack is the same problem as any other deny assignment: an excluded principal you can still operate, tested before you need it. The landing-zone page covers that debug loop. + +Failure mode: two denies on the same firewall, blueprint and stack, and an operator deletes the blueprint assignment thinking it is unused. The stack deny should remain. Check the deny list after the delete and confirm one assignment is still there. Failure mode: `deny-settings-excluded-principals` is a user who has left, stack updates start failing, and someone sets the mode to `none` to unblock CI. That is the lock being removed on purpose. + +## Checklist + +- [ ] Every production blueprint assignment’s lock mode written down (None vs Do Not Delete vs Read Only) +- [ ] Deny assignments correlated to those assignments; resource locks and Policy Deny counted separately +- [ ] Replacement stack exists at parent scope with `denyDelete` or `denyWriteAndDelete` **before** the blueprint assignment is removed +- [ ] `--action-on-unmanage detachAll` on stacks that must not delete the hub +- [ ] Excluded principal is a platform group object id, and a non-member Owner still cannot delete +- [ ] Definitions exported before 31 January 2027 +- [ ] A calendar reminder that 31 October 2026 stops new assignments and definition edits, and 31 January 2027 removes the denies whether or not you migrated + +**Related:** [Landing-zone deny assignments](/blog/azure-landing-zone-deny-assignments/) · [Azure CSPM implementation](/blog/azure-cspm-implementation-guide/) · [CIEM explained](/blog/ciem-explained-for-cloud-teams/) diff --git a/src/content/blog/azure-landing-zone-deny-assignments.md b/src/content/blog/azure-landing-zone-deny-assignments.md index 5354328..22df904 100644 --- a/src/content/blog/azure-landing-zone-deny-assignments.md +++ b/src/content/blog/azure-landing-zone-deny-assignments.md @@ -2,7 +2,7 @@ title: "Azure Landing Zone Deny Assignments That Block Owners" description: "CAF-style deny assignments vs Policy Deny, platform vs app landing zones, break-glass, and debugging “I am Owner but cannot delete.”" pubDate: 2026-08-27 -updatedDate: 2026-08-27 +updatedDate: 2026-10-02 author: OpenSourceOM Team tags: - Azure @@ -91,6 +91,7 @@ Patterns that work: - **Deployment stack** at the platform subscription with `denySettings.mode = denyDelete` (or `denyWriteAndDelete`) on the stack’s resources. Azure creates a deny assignment owned by the stack. - **Managed application** for a marketplace/platform offering; the publisher identity is excluded, everyone else is denied on the managed RG. - **ALZ / CAF accelerator** artifacts that drop deny assignments so subscription Owners cannot remove diagnostic settings or move the subscription out of the MG (depending on version—read **your** deployed JSON, not a 2022 blog post). +- **Blueprint locks** (`AllResourcesDoNotDelete`, `AllResourcesReadOnly`). Those denies are removed when Blueprints retires on 31 January 2027. Replacing them with a deployment stack is [blueprint locks](/blog/azure-blueprint-locks-deny-assignments/). Patterns that hurt: @@ -148,4 +149,4 @@ When the deny is working as designed, the answer to the ticket is: move the work - [ ] Policy Deny still used for public IPs / locations on app MGs ([Azure CSPM](/blog/azure-cspm-implementation-guide/)) - [ ] Locks vs deny assignments distinguished in the debug tree -**Related:** [Azure CSPM implementation](/blog/azure-cspm-implementation-guide/) · [CIEM explained](/blog/ciem-explained-for-cloud-teams/) +**Related:** [Blueprint locks](/blog/azure-blueprint-locks-deny-assignments/) · [Azure CSPM implementation](/blog/azure-cspm-implementation-guide/) · [CIEM explained](/blog/ciem-explained-for-cloud-teams/) diff --git a/src/content/blog/eks-pod-identity-vs-irsa.md b/src/content/blog/eks-pod-identity-vs-irsa.md index c36486e..703ebbd 100644 --- a/src/content/blog/eks-pod-identity-vs-irsa.md +++ b/src/content/blog/eks-pod-identity-vs-irsa.md @@ -2,7 +2,7 @@ title: "EKS Pod Identity vs IRSA" description: "EKS Pod Identity binds an IAM role to a namespace and service account. Trust policy, leftover IRSA annotations, and an agent that fails closed." pubDate: 2026-09-25 -updatedDate: 2026-09-25 +updatedDate: 2026-10-02 author: OpenSourceOM Team tags: - EKS @@ -78,7 +78,7 @@ Behavior and the trust-policy shape are documented in [EKS Pod Identity](https:/ The association is what EKS will attempt. The request-tag conditions are what stop a second association from reusing this role: EKS sets `kubernetes-namespace` and `kubernetes-service-account` on the assume. `aws:SourceArn` stops a different cluster. Without the tags, any service account in `payments` that someone can associate will get the role. Treat `eks:CreatePodIdentityAssociation` and `eks:DeletePodIdentityAssociation` like `iam:PassRole`: cluster-admin and CI only, not every namespace developer. Session tags can be turned off on an association; if you disable them, these request-tag conditions stop matching and the assume fails. Leave tags on. -IRSA’s trust looks nothing like this. It is `AssumeRoleWithWebIdentity` against `oidc.eks..amazonaws.com/id/` with `sub` = `system:serviceaccount:payments:payments-api`. Reusing that document for Pod Identity fails closed. Leaving it in place beside the new trust is how one role stays assumable two ways. +IRSA’s trust looks nothing like this. It is `AssumeRoleWithWebIdentity` against `oidc.eks..amazonaws.com/id/` with `sub` = `system:serviceaccount:payments:payments-api` and `:aud` = `sts.amazonaws.com`. What that `aud` condition does and does not pin is [IRSA trust policy aud](/blog/irsa-trust-policy-aud/). Reusing the IRSA document for Pod Identity fails closed. Leaving it in place beside the new trust is how one role stays assumable two ways. ## What IRSA still is diff --git a/src/content/blog/irsa-trust-policy-aud.md b/src/content/blog/irsa-trust-policy-aud.md new file mode 100644 index 0000000..3adb3c3 --- /dev/null +++ b/src/content/blog/irsa-trust-policy-aud.md @@ -0,0 +1,136 @@ +--- +title: "IRSA Trust Policy aud Mistakes" +description: "The IRSA aud claim is sts.amazonaws.com and does not name the cluster. Wrong condition keys, a second audience on the OIDC provider, and a trust policy that only checks sub." +pubDate: 2026-10-02 +updatedDate: 2026-10-02 +author: OpenSourceOM Team +tags: + - EKS + - IRSA + - IAM + - OIDC + - Kubernetes +focusKeyword: IRSA trust policy aud +faq: + - question: What should the IRSA trust policy aud condition equal? + answer: >- + The condition key is the cluster OIDC issuer host plus :aud, and the + value is sts.amazonaws.com unless the service account sets + eks.amazonaws.com/audience to something else. The same value on every + cluster is normal. The cluster id lives in the issuer host, not in aud. + - question: Why does AssumeRoleWithWebIdentity return Incorrect token audience? + answer: >- + STS compared the token aud to the IAM OIDC provider ClientIDList and + it was not listed. That check happens before the role trust policy. + Adding :aud to the trust policy does not fix a provider whose client + id list is missing sts.amazonaws.com. + - question: Is a trust policy that only checks sub safe if ClientIDList is sts.amazonaws.com? + answer: >- + Only while that list stays a single audience. STS rejects tokens whose + aud is not in ClientIDList. The day a second audience is added to the + provider, every role that forgot :aud will accept it. Pin :aud with + StringEquals on the role anyway. +--- + +`AssumeRoleWithWebIdentity` failed. The token in the pod has `"aud": "sts.amazonaws.com"` and `"sub": "system:serviceaccount:payments:payments-api"`. The trust policy also says `aud`. The condition key is the bare string `aud`, so IAM never compares it to the token. That role is not assumable, and the line that looks like a security control is not one. + +**IRSA trust policy aud** is the `:aud` condition on `sts:AssumeRoleWithWebIdentity`. How the projected token is mounted is [projected service account tokens](/blog/kubernetes-projected-service-account-tokens/). Pod Identity uses a different principal and does not use this claim — [EKS Pod Identity vs IRSA](/blog/eks-pod-identity-vs-irsa/). This page is the `aud` mistakes on the IRSA document. + +Issuer setup and the trust shape: [IAM roles for service accounts](https://docs.aws.amazon.com/eks/latest/userguide/iam-roles-for-service-accounts.html). + +## The condition that matches + +```json +{ + "Effect": "Allow", + "Principal": { + "Federated": "arn:aws:iam::123456789012:oidc-provider/oidc.eks.us-east-1.amazonaws.com/id/EXAMPLED539D4633E53DE1B71EXAMPLE" + }, + "Action": "sts:AssumeRoleWithWebIdentity", + "Condition": { + "StringEquals": { + "oidc.eks.us-east-1.amazonaws.com/id/EXAMPLED539D4633E53DE1B71EXAMPLE:aud": "sts.amazonaws.com", + "oidc.eks.us-east-1.amazonaws.com/id/EXAMPLED539D4633E53DE1B71EXAMPLE:sub": "system:serviceaccount:payments:payments-api" + } + } +} +``` + +The key prefix is the issuer host with no `https://`. It is the same string as `cluster.identity.oidc.issuer` after that prefix is stripped. `:aud` and `:sub` are claims on that provider. A key named `aud` is a different key, and it is not in this request, so `StringEquals` on it does not match. The API error is `AccessDenied` / not authorized to perform `sts:AssumeRoleWithWebIdentity`, which is a trust-policy miss. + +`InvalidIdentityToken` / `Incorrect token audience` is earlier. STS compares the token’s `aud` to the OIDC provider’s `ClientIDList` before the role document runs. + +```bash +aws iam get-open-id-connect-provider \ + --open-id-connect-provider-arn "$PROVIDER_ARN" \ + --query 'ClientIDList' +``` + +That list has to include `sts.amazonaws.com` for a default IRSA token. Putting `:aud` in the trust policy does not add it here. + +## aud does not name the cluster + +Every default IRSA token carries `aud` of `sts.amazonaws.com`. Cluster A and cluster B both do. The cluster is the id in the issuer: + +```text +https://oidc.eks.us-east-1.amazonaws.com/id/EXAMPLED539D4633E53DE1B71EXAMPLE +``` + +A copied trust policy that still points `Principal.Federated` and the condition prefix at cluster A’s issuer will not match a token whose `iss` is cluster B. Fixing only the `:aud` value changes nothing, because the value was already `sts.amazonaws.com`. + +The loose copy is the dangerous one: two `Federated` principals, `:aud` set, and `:sub` of `system:serviceaccount:*:*` under `StringLike`. Any service account in either cluster can assume the role. The `aud` line is identical to a locked-down role, so a review that only checks “aud is sts.amazonaws.com” passes it. `:sub` is the service account. `:aud` is the intended consumer of the token. The issuer host is the cluster. + +## ClientIDList is the first aud check + +STS rejects a JWT whose `aud` is not in the provider `ClientIDList`, with `Incorrect token audience`. The in-cluster token mounted for the API server has `aud` of `https://kubernetes.default.svc` and the same `iss` and `sub`. With a provider that lists only `sts.amazonaws.com`, that API token cannot assume the role even if the trust policy forgot `:aud`. + +That protection disappears when someone adds a second audience to the provider: + +- `eks.amazonaws.com/audience` set on a service account to a private string, and that string added to `ClientIDList` so the pod can assume one role +- a vendor’s audience added to the same provider +- `https://kubernetes.default.svc` added while debugging “Incorrect token audience” + +Every role on that provider whose trust policy checks only `:sub` now accepts tokens minted for the new audience. `StringEquals` on `:aud` is what keeps those roles on `sts.amazonaws.com`. `StringEqualsIfExists` is the wrong operator: a token that omits `aud` makes the condition pass. Use `StringEquals`. + +Do not add `https://kubernetes.default.svc` to `ClientIDList`. That is the API-server audience. Combined with a trust policy that has no `:aud` key, the default projected token becomes a cloud credential. + +## Custom audience annotation + +The webhook reads `eks.amazonaws.com/audience` on the service account. Omitted, the projected token’s `aud` is `sts.amazonaws.com`. + +```yaml +metadata: + annotations: + eks.amazonaws.com/role-arn: arn:aws:iam::123456789012:role/payments-api + eks.amazonaws.com/audience: sts.amazonaws.com +``` + +If you set a different audience, three places have to carry the same string: + +1. The annotation (what the webhook asks the API server to mint) +2. The provider `ClientIDList` (what STS will accept at all) +3. The trust policy `:aud` value (what this role will accept) + +Annotation changed, trust policy still `sts.amazonaws.com`: `AccessDenied` on the role. Annotation changed, `ClientIDList` not updated: `Incorrect token audience`. Annotation left at the default and the trust policy updated to a custom string: `AccessDenied` again. Decode the file the SDK is actually sending before editing the third place. + +```bash +cut -d. -f2 < "$AWS_WEB_IDENTITY_TOKEN_FILE" \ + | tr '_-' '/+' | base64 -d 2>/dev/null \ + | jq '{iss,aud,sub}' +``` + +`iss` must be your cluster’s issuer. `aud` must be the trust policy value and a member of `ClientIDList`. `sub` must be `system:serviceaccount::` with that namespace and name, not a wildcard you meant to tighten later. + +A second statement on the same role that checks only `:sub`, left in place “for Terraform,” accepts every audience the provider lists. Delete that statement. Pod Identity’s trust (`pods.eks.amazonaws.com`, `sts:AssumeRole`, `sts:TagSession`) is not an `aud` statement; leaving the IRSA statement beside it keeps the OIDC door open after you add an association. + +## Checklist + +- [ ] `:aud` key is `:aud`, not `aud` +- [ ] `:aud` value is `StringEquals` `sts.amazonaws.com`, or the custom audience on all three of annotation, `ClientIDList`, and trust policy +- [ ] `:sub` is one `system:serviceaccount:namespace:name`, not `StringLike` `*:*` +- [ ] `Federated` principal and the condition prefix are this cluster’s issuer id +- [ ] `ClientIDList` includes `sts.amazonaws.com` and does not include `https://kubernetes.default.svc` +- [ ] No second statement on the role allows `AssumeRoleWithWebIdentity` without `:aud` +- [ ] Decoded `AWS_WEB_IDENTITY_TOKEN_FILE` matches `iss`, `aud`, and `sub` before the trust policy is rewritten again + +**Related:** [EKS Pod Identity vs IRSA](/blog/eks-pod-identity-vs-irsa/) · [Projected service account tokens](/blog/kubernetes-projected-service-account-tokens/) · [GitHub Actions OIDC to AWS](/blog/github-actions-oidc-aws-iam/) diff --git a/src/content/blog/kubernetes-validating-admission-policy.md b/src/content/blog/kubernetes-validating-admission-policy.md index b746424..0a8b166 100644 --- a/src/content/blog/kubernetes-validating-admission-policy.md +++ b/src/content/blog/kubernetes-validating-admission-policy.md @@ -2,7 +2,7 @@ title: "Kubernetes ValidatingAdmissionPolicy: CEL vs Webhooks" description: "In-process CEL for deny-only checks; keep Kyverno or Gatekeeper for mutation. FailurePolicy Fail vs Ignore, parameter resources, and 1.30+ GA." pubDate: 2026-08-27 -updatedDate: 2026-09-22 +updatedDate: 2026-10-02 author: OpenSourceOM Team tags: - Kubernetes @@ -108,16 +108,16 @@ Docs: [ValidatingAdmissionPolicy](https://kubernetes.io/docs/reference/access-au | Need | VAP (CEL) | Kyverno / Gatekeeper | | --- | --- | --- | | Deny privileged / hostPath / latest | Yes, in-process | Yes, extra hop | -| Mutate (add RuntimeDefault, drop caps) | **No** (use MutatingAdmissionPolicy in 1.32+ or a webhook) | Yes | +| Mutate (add RuntimeDefault, drop caps) | `MutatingAdmissionPolicy` is GA in 1.36; older clusters still need a webhook | Yes | | Generate NetworkPolicy from a label | No | Kyverno generate | | Query other resources (“does the SA exist?”) | Limited (params, not cluster walk) | Gatekeeper referential | | Cosign verifyImages | No | Kyverno / custom | | Policy language your team already writes | CEL | YAML/Rego | | Availability | API server (always) | Your Deployment | -Keep Gatekeeper/Kyverno when you already have fifty Rego constraints and no budget to rewrite, or when you **mutate**. Do not add a webhook for `privileged == false`. That check belongs in VAP so a webhook outage cannot fail open unless you configured it that way. +Keep Gatekeeper/Kyverno when you already have fifty Rego constraints and no budget to rewrite, or when you **mutate** on a cluster older than 1.36. Do not add a webhook for `privileged == false`. That check belongs in VAP so a webhook outage cannot fail open unless you configured it that way. -MutatingAdmissionPolicy (CEL mutations) is the native follow-on. Until it is everywhere you run, a small Kyverno install for *mutations only* plus VAP for *denies* is a coherent split. One engine that does both is also fine; two engines that both deny the same pod is how you debug for a week. +Which engine owns generate, image verify, background reports, and apply-time rejection of Deployments is [Kyverno vs ValidatingAdmissionPolicy](/blog/kyverno-vs-validating-admission-policy/). `MutatingAdmissionPolicy` is GA in 1.36. On older clusters, a small Kyverno install for mutations plus VAP for denies is the split. Two engines that both deny the same pod is how you debug for a week. Latency: webhooks add network RTT on every Pod create. VAP runs in the apiserver process. That is the operational argument, not a religion. diff --git a/src/content/blog/kyverno-vs-validating-admission-policy.md b/src/content/blog/kyverno-vs-validating-admission-policy.md new file mode 100644 index 0000000..33536d7 --- /dev/null +++ b/src/content/blog/kyverno-vs-validating-admission-policy.md @@ -0,0 +1,108 @@ +--- +title: "Kyverno vs ValidatingAdmissionPolicy" +description: "ValidatingAdmissionPolicy denies object shape in-process. Kyverno still generates objects, verifies images, and reports on resources that already exist. Do not run both on the same rule." +pubDate: 2026-10-02 +updatedDate: 2026-10-02 +author: OpenSourceOM Team +tags: + - Kubernetes + - Kyverno + - ValidatingAdmissionPolicy + - admission control + - policy as code +focusKeyword: Kyverno vs ValidatingAdmissionPolicy +faq: + - question: Does ValidatingAdmissionPolicy replace Kyverno? + answer: >- + For deny-only checks on the object in the request, yes. CEL runs in + the API server, so a down policy controller cannot fail that check + open. Kyverno remains the engine for generating objects, verifying + image signatures, background scans of resources that already exist, + and mutation on clusters older than 1.36. + - question: Why does kubectl apply of a Deployment succeed when pods are forbidden? + answer: >- + A ValidatingAdmissionPolicy that matches only pods runs when the + ReplicaSet creates the pod, not when the Deployment is admitted. + The apply succeeds and the new pods never start. Kyverno pod-controller + autogen rejects the Deployment. Generating a ValidatingAdmissionPolicy + from a Kyverno policy turns that autogen off. + - question: Is MutatingAdmissionPolicy a reason to remove Kyverno? + answer: >- + On a cluster that is actually 1.36 or newer, in-process mutation is + GA and on by default. Many clusters are older; there the mutator is + still a webhook. Generation, image verification, and background + reports are not mutation, and CEL admission does not do them. +--- + +`kubectl apply` of the Deployment returns success. The ReplicaSet then sits at 0 ready, and the event says the pod was denied by a `ValidatingAdmissionPolicy`. The policy matches `pods`. The Deployment object was never in that match. A Kyverno rule with pod-controller autogen would have rejected the apply. + +**Kyverno vs ValidatingAdmissionPolicy** is which engine owns which check. CEL syntax, `failurePolicy`, and parameter resources are [ValidatingAdmissionPolicy](/blog/kubernetes-validating-admission-policy/). Image signatures are [image provenance and SLSA](/blog/kubernetes-image-provenance-slsa/). This page is the split. + +Kyverno’s own matrix: [ValidatingPolicy](https://kyverno.io/docs/policy-types/validating-policy/). Native mutation: [MutatingAdmissionPolicy](https://kubernetes.io/docs/reference/access-authn-authz/mutating-admission-policy/), stable in Kubernetes 1.36. + +## What each one can decide + +| Need | ValidatingAdmissionPolicy | Kyverno | +| --- | --- | --- | +| Deny privileged, hostPath, hostNetwork, floating tags | Yes, in the API server | Yes, if you still webhook it | +| Reject a Deployment whose pod template is bad, at apply time | Only if you also match that controller and write CEL against `spec.template` | Pod-controller autogen | +| Scan pods that already exist | No. Admission does not replay | Background scan, PolicyReport | +| Create a NetworkPolicy when a Namespace appears | No | Generate rules | +| Verify Cosign / Sigstore at admission | No. CEL does not call a registry | `verifyImages` | +| Mutate (default seccomp, drop caps, labels) | `MutatingAdmissionPolicy`, GA in 1.36 | Webhook on older clusters | +| Namespaced exception object | No. Bypass is editing the cluster-scoped policy or its binding | `PolicyException` | + +A deny you can express as “this object’s fields are wrong” belongs in `ValidatingAdmissionPolicy` with `failurePolicy: Fail` and a binding `validationActions: ["Deny"]`. The API server evaluates it when the policy-controller Deployment is unschedulable. A webhook with `failurePolicy: Ignore` does the opposite during that outage: the privileged pod is admitted. + +## Autogen and generated VAPs are alternatives + +Kyverno can mint a `ValidatingAdmissionPolicy` from a `ValidatingPolicy` (`spec.autogen.validatingAdmissionPolicy.enabled`). It can also expand a pod rule to Deployments, Jobs, CronJobs, and StatefulSets (`spec.autogen.podControllers`). + +Those two switches are mutually exclusive. Pod-controller autogen means Kyverno does **not** generate the `ValidatingAdmissionPolicy`, and the policy status says why: a `ValidatingAdmissionPolicy` is evaluated by the API server and cannot carry the controller-shaped copies Kyverno would have written. Turn on VAP generation and you lose apply-time rejection of the Deployment unless the CEL you wrote already matches those kinds. + +Matching only `pods` is still a correct deny. It is a bad developer experience and a bad audit signal: Git shows a green Deployment, the cluster has no new pods, and the old pods keep running because nothing re-admits them. If the control has to fail the pipeline, match the controllers too or keep autogen and do not generate a VAP for that rule. + +```yaml +# VAP: this denies the pod the ReplicaSet tries to create. +# The Deployment apply already succeeded. +spec: + matchConstraints: + resourceRules: + - apiGroups: [""] + apiVersions: ["v1"] + operations: ["CREATE", "UPDATE"] + resources: ["pods"] +``` + +Write a second `resourceRules` entry for `apps/v1` `deployments` only if the expression reads `object.spec.template.spec`, not `object.spec`. One expression copied from the pod rule onto a Deployment does not see containers and can error. With `failurePolicy: Fail` that error denies the Deployment, including ones that were fine. That is a different outage from the one you meant to build. + +## What stays on Kyverno + +**Resources that already exist.** Installing a VAP does not list privileged pods. They remain until something updates them. Kyverno’s background scan writes a PolicyReport for the current objects. A program that only checks “admission would deny” will call the cluster clean while those pods run. Keep the scan, or accept that the VAP’s enforce date is the date each workload is rolled. + +**Generate.** A namespace label that should produce a default-deny NetworkPolicy is a second API call. CEL admission can reject the Namespace. It cannot create the NetworkPolicy. That rule stays a Kyverno generate rule. + +**Signatures.** “Digest is pinned” is a VAP. “This digest was signed by our key” needs a witness outside the object. That is Kyverno or another verifier, at admission, on the digest. The provenance page is that control. Do not encode it as a CEL prefix check and call it verification. + +**Mutation below 1.36.** `MutatingAdmissionPolicy` and `MutatingAdmissionPolicyBinding` are `admissionregistration.k8s.io/v1` and on by default in Kubernetes 1.36 (April 2026). `kubectl version` on the cluster you are changing is the fact that matters. EKS and other managed offerings lag upstream. On 1.35 and older the mutator is still a webhook. Deleting Kyverno the week upstream went GA, without checking the cluster, removes the mutation and leaves the pods without the seccomp profile you thought was in-process. + +## One owner per rule + +Two denies for privileged pods — a VAP and a Kyverno webhook — fail in different orders and with different text. The webhook’s `failurePolicy: Ignore` then becomes a hole next to a VAP that was supposed to be the control, or a second 403 you debug for a week when both are `Fail`. Pick one. + +Kyverno `PolicyException` is namespaced. Whoever can create that object in the namespace can exempt a pod from the Kyverno rule. RBAC on `policyexceptions` is the bypass. A VAP has no exception CRD. The bypass is `update` on the cluster-scoped `ValidatingAdmissionPolicy` or its binding, which is a cluster-admin verb, not a namespace developer verb. Moving a rule from Kyverno to a VAP changes who can waive it. Do that on purpose. Namespace developer RBAC is [Kubernetes RBAC](/blog/kubernetes-rbac-security-best-practices/). + +Leave Kyverno installed for generate, verify, background reports, and mutation you have not moved. Stop sending it the denies CEL already enforces. A webhook outage should not be on the path of `privileged == false`. + +## Checklist + +- [ ] `kubectl version` recorded. Mutation is in-process only on 1.36+ +- [ ] Shape denies (privileged, hostPath, hostNetwork, digest) are VAP with `failurePolicy: Fail` and binding `Deny` +- [ ] Those same rules are not also enforced by a Kyverno webhook +- [ ] Deployments are rejected at apply, or you have accepted ReplicaSet failures and written that down +- [ ] `autogen.podControllers` and `autogen.validatingAdmissionPolicy` are not both set on one policy +- [ ] Background PolicyReports still cover objects that predate the VAP +- [ ] Generate rules and `verifyImages` still have a running Kyverno (or another engine that can do those jobs) +- [ ] `PolicyException` RBAC reviewed before calling a Kyverno validate rule a control + +**Related:** [ValidatingAdmissionPolicy](/blog/kubernetes-validating-admission-policy/) · [Image provenance and SLSA](/blog/kubernetes-image-provenance-slsa/) · [Kubernetes RBAC](/blog/kubernetes-rbac-security-best-practices/)