Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
91 changes: 91 additions & 0 deletions .claude/skills/planning-prebid-aws/SKILL.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,91 @@
---
name: planning-prebid-aws
disable-model-invocation: true
description: Plan Prebid Server Go deployments on AWS through an interactive requirements interview. Use when choosing deployment architecture or AWS services, or generating Terraform, runtime configuration, deployment scripts, and runbooks for a new or revised PBS deployment. Stops before cloud changes and live traffic.
---

# Planning Prebid Server Go on AWS

Produce an approved deployment design and locally checked implementation files. Prebid Server Go is fixed; ask for its release and image source, not its implementation language.

## Authority

This workflow permits planning and, after design approval, file generation and safe local checks. Cloud provisioning, credential writes, application deployment, production traffic changes, and teardown require a separate execution workflow with explicit authorization. Generated scripts and CI jobs must remain inactive. Ask before authenticated AWS inspection or state access, naming the account, role, regions, and read scope. Keep real credentials out of the conversation and generated files.

## 1. Establish the baseline

Inspect repository instructions, Git status, existing infrastructure, runtime files, and supplied documents. Preserve unrelated changes. Identify the target repository and deployment directory before proposing edits.

For read-only discovery from any existing `trusted-server.toml`, and for designing the operator commands and PBS config/secret delivery, read [configuration and secrets](references/configuration-and-secrets.md). Select the authoritative config source before inferring requirements; examples and disabled integrations are not active deployment inputs.

Record requirements in one decision record, initially in the conversation, then in the approved deployment plan:

| Requirement | Value | Status | Evidence or decision owner | Blocks |
| -------------------- | ----------------------- | ---------------------------------- | -------------------------- | ----------------------------------------- |
| One row per decision | Answer or open question | Confirmed, proposed, or unresolved | Source or person | Design, generation, live rollout, or none |

Treat supplied examples as evidence of intent, not accepted requirements for this deployment. Reconcile conflicting inputs with the user. Inspect existing answers before asking again.

Done when the current deployment, requested outcome, reusable resources, selected Trusted Server config source or its absence, and unanswered decisions are identified.

## 2. Interview by decision

Ask two to four related questions per turn. Start with questions that change the architecture; explain the consequence of unfamiliar choices. Offer a recommendation the user can accept rather than requiring AWS expertise.

Cover these topics, skipping already confirmed answers:

- Purpose: disposable test, live pilot, or ongoing production; retained provider and migration scope.
- Availability: acceptable interruption, host/AZ/region failure tolerance, maintenance windows, recovery time, and loss tolerance for required data.
- Workload: absolute peak eligible auction QPS, allocation, regional mix, bidder fan-out, payload sizes, burst duration, and caller latency budget.
- Operations: fixed headroom versus automatic scaling, budget ceiling, owner and backup, existing CI and AWS platform, regions, DNS, network restrictions, and bidder IP allowlists.
- Security: permitted callers/publishers, public or private access, data residency, retention, and the privacy policy owner.
- Auction behavior: caller, bidder set, formats, consent and identity, account settings, stored requests, and cache dependencies. Read [Prebid Go requirements](references/prebid-go.md) before resolving these inputs.
- Outbound bidder connectivity: for each bidder and region, choose public/NAT egress, an internal RTB Fabric link, or an RTB Fabric outbound external link. If RTB Fabric is a candidate, read [RTB Fabric connectivity](references/rtb-fabric.md) before asking the conditional questions. Record partner participation, gateway and link ownership, PBS endpoint mapping, regional support, quotas, timeout, cost, fallback, and monitoring.

When RTB Fabric is a candidate, ask these questions in related groups:

1. Which bidders participate in RTB Fabric, in which regions, and which partner provides each responder gateway ID? Who accepts and owns each link?
2. For each participating bidder, should PBS use an internal Fabric link or an outbound external link? What remains on NAT egress, and what is the explicit behavior when a Fabric link is pending, unavailable, or over quota?
3. Which pinned PBS Go release and adapter configuration own the endpoint mapping? Who approves link creation, partner acceptance, configuration rollout, and endpoint changes?
4. What peak transactions per second, payload sizes, bidder deadlines, regional failover load, and monthly volume should size each link and compare Fabric cost with NAT?

Translate "production ready" into measurable availability, security, capacity, and recovery requirements. Scaling and availability are separate decisions. Offer a measurement plan for unknown traffic rather than inventing capacity.

Done when every topic is confirmed, explicitly inapplicable, or recorded as an unresolved blocker. For an RTB Fabric branch, every participating bidder and region has a selected path, partner owner, endpoint mapping, quota and timeout check, cost assumption, fallback, and monitoring owner. Continue a provisional design around unknowns, but pause affected file generation until architecture-changing decisions are approved.

## 3. Recommend and obtain approval

Read [architecture decisions](references/architecture.md). If the interview selected RTB Fabric, also read [RTB Fabric connectivity](references/rtb-fabric.md). For Terraform state/authentication decisions and HCL generation, read [Terraform guidance](references/terraform.md). For a two-region standalone-host pilot, consult [the worked example](examples/two-region-pilot.md); its values remain conditional.

Present one recommended design and only alternatives that resolve a real tradeoff. Include:

- A top-to-bottom Mermaid diagram and the failure boundaries.
- Each selected AWS service, its requirement, and whether to reuse or create it.
- Capacity assumptions, cost drivers and estimate date, accepted limitations, and blockers.
- Runtime, infrastructure, secret, and traffic-control ownership.
- Use the experimental `ts prebid server` commands for supported local checks, secret writes, and EC2 status. Record the selected YAML/secret delivery path and unsupported operations explicitly. Deployment, rollback, and other runtime support need a separately approved implementation; avoid competing wrappers for implemented commands.
- Target files and checks, with cloud-dependent checks separated from local checks.

Ask the user to approve the architecture, assumptions, target files, and accepted limitations. Approval to generate files is not approval to execute them. Reopen approval if later findings change topology, cost commitments, or ownership.

Done when the user explicitly approves the design and file scope. If blockers remain, agree on a bounded draft and label it incomplete rather than producing apparently deployable infrastructure.

## 4. Generate the approved files

Read [file generation and validation](references/file-generation.md). Follow existing repository conventions and generate only artifacts used by the selected design. Keep the decision record in the deployment plan; reference it from the runbook rather than repeating it.

Read the [PBS CLI usage and descriptor schema](../../../crates/trusted-server-cli/README.md) before generating inputs consumed by `ts prebid server`. Its current descriptor supports EC2/Compose only. Keep other architecture choices available, but mark their CLI integration deferred rather than generating unsupported fields.

Verify version-specific PBS fields and adapter bindings against the selected release. For an approved RTB Fabric branch, reread [RTB Fabric connectivity](references/rtb-fabric.md) and keep partner acceptance, link activation, and any unsupported CLI integration visible as separate work. Verify AWS/Terraform behavior and pricing against current primary documentation. Record source links, versions, and verification dates in the deployment plan. Unavailable evidence remains a named blocker; do not invent image digests, configuration keys, prices, or benchmark results.

Done when every approved artifact exists, has a named owner and check, and every unresolved input is visible and prevents unsafe use where applicable. Every generated operator command must have documented inputs, access requirements, output, failure behavior, and a recovery action; proposing command names alone is not implementation.

## 5. Validate and hand off

Run the applicable safe checks in the loaded generation references and inspect the final diff. Report changed paths, exact commands, results, checks not run, and remaining blockers. Separate these states:

- Draft: unresolved generation inputs or incomplete artifacts.
- Locally checked: applicable local checks passed; cloud and integration behavior remain unverified.
- Ready for deployment review: file scope complete, blockers for generation cleared, local evidence recorded, and cloud plan/live checks listed for the authorized operator.

This skill never establishes production readiness from generated files alone. End with the next approval or evidence needed, not a provisioning command executed on the user's behalf.
74 changes: 74 additions & 0 deletions .claude/skills/planning-prebid-aws/examples/two-region-pilot.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,74 @@
# Worked example: two-region pilot

This is a conditional design derived from a supplied pilot plan, not a default architecture or evidence of provisioned capacity. Use it only when the user accepts its limitations. Operational rules live in the skill references.

## Example requirements

Assume the user has confirmed:

- Prebid Server Go on AWS with Terraform.
- One deployment in `us-east-1` and one in `us-west-2`.
- A retained provider and a caller-controlled trial targeting 5% of eligible traffic.
- A fixed approved bidder set, versioned configuration, and credential rotation.
- Host maintenance interruptions, manual host replacement, and no autoscaling for this bounded pilot.

Still resolve absolute peak traffic, regional distribution, bidder fan-out, caller type, domain ownership, inventory/cache dependencies, secret permissions, budget, recovery expectations, and numeric rollout gates. The 5% value supplies no instance-size evidence.

## Conditional recommendation

```mermaid
flowchart TD
A["Existing caller"] --> S["Caller experiment allocation"]
S -->|"95% of eligible traffic"| P["Existing provider"]
S -->|"5% of eligible traffic"| R["Selected regional routing"]
R --> E["East EC2: Caddy and PBS Go"]
R --> W["West EC2: Caddy and PBS Go"]
E --> B["Approved bidder endpoints"]
W --> B
```

Per region, propose a public subnet, internet gateway, EC2 host with encrypted storage and Elastic IP, Compose, Caddy, SSM access, regional secret access, release artifacts, and telemetry. Select actual instance sizes and images only after workload and compatibility evidence.

Prefer existing caller-side regional routing when suitable. Otherwise evaluate shared-hostname Route 53 latency routing with independent regional health checks and a verified DNS-challenge certificate setup. Implement only the selected branch.

Use one pilot Terraform root with two provider aliases and explicit regional module calls, unless ownership requires separate roots. Keep backend bootstrap independent. Generate Compose/Caddy and host deployment tools only after approving this profile.

No ALB, NAT gateway, ECS, distributed database, or cache is justified solely by the pilot's two regions. Add or replace components when a confirmed requirement demands it. Exclude inventory with unmet dependencies explicitly rather than presenting a reduced-function trial as equivalent.

## Approval and evidence gates

The user must accept one-host-per-region outages and maintenance behavior. If regional failover concentrates the pilot on one node, prove that capacity or agree on reducing allocation. Automatic replacement, uninterrupted deployment, or tighter recovery requirements invalidate this profile and require another architecture decision.

File generation starts after architecture-changing questions and target paths are approved. Live-traffic gates remain deferred to authorized execution: zero-allocation deployment checks, internal inventory validation, then agreed 1% and 5% observation windows. A tested caller kill switch and measured refresh bound are required before the live pilot.

## Configuration and operator walkthroughs

The experimental `ts prebid server` CLI implements local inspection/checks, secret value writes, and EC2 infrastructure status. The release, runtime delivery, and rollback scenarios below remain acceptance criteria for a future approved implementation, not executed deployment evidence. Consult the [current command contract](../references/configuration-and-secrets.md#operator-command-contract) before documenting an invocation.

| Input or task | Expected behavior |
| ------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Several `trusted-server.toml` files, including a disabled example | Ask which source/environment is authoritative; preserve every source file and report uncertainty about remote overrides |
| The config lists server bidders, client-side bidders, and bundle adapters | Classify them separately; verify candidate server adapters and report host-secret needs as required, not needed, or unresolved without activating bidders |
| An operator changes a PBS timeout | Check and render regional YAML, prepare an immutable release, preview and approve deployment; no manual environment-file edits or Terraform apply |
| A binding conflicts with YAML or a required secret key is missing | Reject conflicting local inputs or stop the authorized deployment preflight; preserve working capacity and never print credential values |
| A credential contains quotes, dollars, or newlines | Prompt without echo or accept file/stdin, validate and safely encode it; no secret-value command argument or log output |
| East consumes a rotated credential while West still runs the old version | Report actual regional versions and pending work, verify replication and replacement, then coordinate partner revocation; never report global success from East alone |
| An operator requests rollback after the old credential was revoked | Refuse known-incompatible recovery and explain the credential action needed; a Git release rollback cannot restore bidder validity |
| A command resolves the wrong AWS account or an ambiguous deployment | Stop before mutation and require corrected target selection; local inspect/check stay credential-free |
| ECS is selected instead of Compose | Generate one concrete YAML delivery mechanism and task secret references behind the same operator commands; do not generate unused host loaders |

## Walkthrough assertions

Use these contrasts when reviewing the skill. They are expected behavior, not executed deployment tests. Terraform cases exercise the rules in `references/terraform.md`.

| Input | Expected planning behavior |
| ------------------------------------------------------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| This pilot, with outage acceptance but unknown QPS | Preserve the supplied answers; ask about workload and other blockers; propose this profile provisionally without inventing instance capacity |
| Production requires AZ survival, uninterrupted releases, and automatic scaling | Reopen topology; evaluate multi-AZ ECS/ALB or EC2 Auto Scaling/ALB against the existing platform; define minimum/failover capacity and scaling evidence; generate only the selected runtime's files |
| Pilot adds video with independent regional caches | Resolve cache write/retrieval and failure behavior before generating affected infrastructure, or obtain approval to exclude that inventory |
| User approves file generation but has supplied no AWS execution authorization | Generate inactive tools and run safe local checks; leave cloud actions and traffic changes unexecuted |
| A requirement is unknown or a validation tool is unavailable | Record a blocker or not-run check; preserve the distinction between a draft and locally checked files |
| A proposed local test omits `command` and uses a real AWS provider | Do not run it: the default is apply. Generate an explicitly commanded, isolated mocked test or defer it as an authorized cloud integration test |
| A plan-mode suite mocks East but retains a real West provider alias | Reject the credential-free claim; inspect all provider mappings and setup dependencies, mock the remaining external provider, and select exact reviewed test files |
| The existing S3 backend uses DynamoDB locking | Preserve it while proposing a version/permissions/recovery-aware migration; no backend replacement, lock removal, or state migration during generation |
| An upstream CI example uses stored AWS keys and applies after merge | Generate the approved short-lived identity and protected saved-plan review procedure as inactive tooling; upstream examples grant no execution authority |
Loading
Loading