Deploy an OpenAI, Anthropic, Cohere & Ollama compatible AI gateway on AWS. ECS Fargate infrastructure with auto-scaling, private subnets, KMS encryption and least-privilege IAM by default; HTTPS, WAF, API key authentication and CloudWatch alarms are opt-in inputs — see What this minimal configuration deploys and What gets provisioned, and when.
🌐 Documentation · 🚀 Start 14-Day Free Trial · 💻 GitHub Repository
- Subscribe to stdapi.ai on AWS Marketplace (14-day free trial included)
- Install Terraform or OpenTofu >= 1.9 — see Requirements for exact version constraints
- Configure AWS credentials with IAM permissions to create VPC, ECS, Application Auto Scaling, ALB, S3, KMS, IAM, SSM Parameter Store, SQS, DynamoDB and CloudWatch resources — plus ACM, Route 53 and WAFv2 for a deployment that serves HTTPS on its own domain (
alb_domain_name) or enables the WAF (alb_waf_enabled) - Tenant API keys (
tenants) and the shared model cache (model_cache_shared) also needkms:CreateGrantandkms:DescribeKeyon the deployment's KMS key — see Shared DynamoDB table. Asking for a tenant key rotation (tenant_key_rotation_days, or atenantsentry declaringkey_generation) additionally needssecretsmanager:CreateSecret,DescribeSecret,UpdateSecret,TagResource,GetResourcePolicyandDeleteSecretfor the one secret per tenant the module then creates — see Tenant key secrets
Deployable in any AWS region with ECS Fargate support.
module "stdapi_ai" {
source = "stdapi-ai/stdapi-ai/aws"
version = "~> 1.0"
}What this minimal configuration deploys:
- ECS Fargate service (auto-scaling, private subnets)
- Dedicated VPC with private app subnets, NAT gateways for outbound AWS access, and the free S3 gateway endpoint
- S3 bucket for generated and temporary files (KMS-encrypted)
- IAM roles (least privilege) and CloudWatch logs
What it does NOT deploy (you add these explicitly for production):
- ❌ No public endpoint — the service is only reachable from inside the VPC until you enable the ALB
- ❌ No HTTPS / custom domain
- ❌ No WAF
- ❌ No API key authentication
Production-ready next step — add a public HTTPS endpoint with WAF and API key:
module "stdapi_ai" {
source = "stdapi-ai/stdapi-ai/aws"
version = "~> 1.0"
# Public HTTPS endpoint on a custom domain (ACM cert auto-issued via Route 53)
alb_enabled = true
alb_public = true
alb_domain_name = "api.example.com"
alb_route53_zone_name = "example.com"
# Protection
alb_waf_enabled = true
alb_waf_rate_limit = 2000
alb_waf_block_anonymous_ips = true
# Authentication — module generates a secure key and exposes it as a sensitive output
api_key_create = true
}For ready-to-deploy variants (single-region, EU/US multi-region, Open WebUI), see the samples repository. For deeper patterns (BYO VPC / ALB / Route 53 / S3, manual ECS, cost-optimized), see the advanced deployment guide.
Every input has a working default, so a deployment sets only what it means to change. These sixteen are the ones a first deployment actually decides; everything else is listed under Inputs and can be left alone.
| Input | What it does | Set it when |
|---|---|---|
alb_enabled, alb_public |
Puts an Application Load Balancer in front of the service, in public subnets when alb_public |
The API must be reachable from outside the VPC. Both default to false, so the service answers only from inside it |
alb_domain_name, alb_route53_zone_name |
Serves HTTPS on your own name, with an ACM certificate issued and validated through that Route 53 zone | You enable the ALB — without a name the listener is plain HTTP |
alb_waf_enabled, alb_waf_rate_limit, alb_waf_block_anonymous_ips |
AWS WAF web ACL in front of the ALB, with per-IP rate limiting and anonymous-network blocking | The endpoint is public. The WAF defaults to false, and rate limiting stays off until you give it a number |
api_key_create |
Generates an API key and returns it as the sensitive api_key output |
Anyone but you can reach the endpoint. Default false: the server authenticates nobody, and a public ALB fails the plan without an authentication method |
aws_cognito_user_pool_id, aws_cognito_client_ids |
Accepts Amazon Cognito user pool tokens, validated per request, instead of a shared key | Your callers already have an identity provider, or you want per-user tokens rather than one key |
tenants |
One API key per tenant, each with its own model and endpoint scopes | Several teams or customers share one deployment |
aws_bedrock_regions |
The regions models are served from and quota is drawn from | You want more models, more quota or a specific data residency than the deployment region alone gives you |
cpu, memory |
Task size, 0.25 vCPU and 512 MiB by default | You serve audio, video or inline input files: those hold bytes in memory, and 512 MiB is OOM-killed under load rather than answering slowly |
autoscaling_min_capacity, autoscaling_max_capacity |
Task count floor and ceiling | You need to bound the bill or guarantee headroom. Defaults: one task per availability zone, ceiling five times the floor |
autoscaling_spot_percent |
Share of the capacity above the floor served by Fargate Spot, at roughly a 70% discount | Interruptible capacity is acceptable for the peaks. Default 0 |
availability_zones_count |
Caps how many availability zones the dedicated VPC spans | You want to bound the NAT gateway bill or the task floor — both scale per zone. Defaults to every zone in the region |
nat_gateways_allowed |
Chooses between NAT gateways (private tasks, billed hourly) and public application subnets for internet egress | The NAT bill matters more than keeping the tasks unaddressable. Default true |
subnet_ids |
Deploys into subnets you already own instead of a dedicated VPC | You have an existing network. Default [], and the module builds the VPC |
sns_topic_arn |
The topic the CloudWatch alarms notify; setting it is what turns the alarms on | You want the alarms to reach someone. Default: no alarms |
deletion_protection |
Deletion protection on the resources that support it | The deployment is production |
tags |
Your own tags on every resource the module creates | Cost allocation, ownership or environment tagging. Default {} |
Production-ready infrastructure following AWS Well-Architected Framework:
- 🚀 Serverless Compute — ECS Fargate with intelligent auto-scaling (0.25-16 vCPU, CPU/Memory/Request-based)
- ⚖️ Load Balancing — Application Load Balancer with HTTPS/TLS, configurable idle timeout for long operations
- 🌐 Networking — Dedicated VPC with private app subnets; AWS access via NAT gateways or interface VPC endpoints; optional public subnets for the ALB; IPv4/IPv6 support
- 🔒 Security — WAF with rate limiting & IP filtering, KMS encryption, IAM roles with least privilege
- 🪪 Authentication — a shared API key (
api_key_create, or one you supply) or Amazon Cognito user pool tokens (aws_cognito_user_pool_id+aws_cognito_client_ids), validated on every request;authentication_modepins which of the two the deployment accepts, andoauth_resource_identifierpublishes the OAuth 2.0 protected-resource metadata that the 401 challenge points an agent at - 🔑 Multi-Tenancy — Declarative per-tenant API keys (
tenants) with model and endpoint scopes and per-minute request and token limits; secrets are minted server-side and never enter Terraform state — delivered once through an SSM SecureString parameter, or, once a rotation is asked for (tenant_key_rotation_days, or a tenant'skey_generation), stored and rotated as one AWS Secrets Manager secret per tenant - 🗄️ Shared State — DynamoDB table holding what outlives a task — the tenant records, and the shared Bedrock model list (
model_cache_shared) a fleet discovers once instead of once per server — created automatically with the first feature that needs it; nothing needs configuring, butaws_dynamodb_tableandaws_dynamodb_regionpoint at a table you manage yourself instead - 🧠 Your Own Models — Amazon SageMaker AI endpoints published as chat models (
aws_sagemaker_endpoints); the module creates no endpoint, it only grants the task role the permission to invoke the ones you name - 📈 Usage & Costs API — OpenAI-compatible
/v1/organization/usage/*and/v1/organization/costs(usage_api), answered from the CloudWatch metricscloudwatch_metricspublishes - 🦙 Ollama Compatibility — Ollama's
/api/*routes served alongside the OpenAI ones, relocatable under a prefix withollama_routes_prefix - 🎙️ WebRTC Media Mode — Opt-in Realtime API WebRTC transport (
realtime_webrtc_media_enabled): public task IP, UDP media ingress, STUN/TURN configuration — single-task by design - 📊 Monitoring — Container Insights dashboards, optional CloudWatch alarms, VPC Flow Logs, request/response logging
- 💾 Storage — S3 buckets with encryption, versioning, lifecycle policies, multi-region support
- 💰 Cost Optimization — Fargate Spot support (~70% discount), scheduled auto-scaling, resource right-sizing
- 🏷️ Tagging —
tagsapplies your cost-allocation, ownership or environment tags to every resource the module creates
Internet
│ egress only when required (see table)
┌──────────▼──────────┐
│ WAF (optional) │ alb_waf_enabled
└──────────┬──────────┘
│ inbound only when the ALB is enabled
┌──────────▼──────────┐
│ ALB (optional) │ alb_enabled / alb_public
│ HTTPS / HTTP │ (public subnets only if alb_public)
└──────────┬──────────┘
│
┌───────────────────────┼───────────────────────┐
│ VPC (dedicated, or bring-your-own subnet_ids) │
│ │ │
│ ┌──────────▼──────────┐ S3 gateway ┌──────────────┐
│ │ ECS Fargate │ endpoint │ S3 Bucket │
│ │ ┌─────────────┐ │ (always, free) ──▶│ (+ regional │
│ │ │ stdapi.ai │ │ │ buckets, │
│ │ │ Container │ │ │ KMS-encr.) │
│ │ └─────────────┘ │ └──────────────┘
│ │ app subnet (private) │
│ └──────────┬──────────┘ │
│ egress to Bedrock, Polly, Transcribe, … │
│ uses exactly ONE of (mutually exclusive): │
│ • NAT gateways (default) │
│ • Interface VPC endpoints (no-internet) │
└───────────────────────┬───────────────────────┘
│
┌──────────▼──────────┐
│ CloudWatch │
└─────────────────────┘
| Component | Created when |
|---|---|
| Dedicated VPC, private app subnets, ECS Fargate service, KMS-encrypted S3 bucket(s), S3 gateway endpoint, CloudWatch logs | Always — except: passing your own subnet_ids skips VPC/subnet/endpoint creation entirely, and passing your own aws_s3_bucket skips bucket creation. A DynamoDB gateway endpoint joins the S3 one as soon as the module creates the shared table; both are gateway endpoints, so they are free and are created on the feature alone, never weighed against a monthly cost |
| NAT gateways (private internet egress) — one per availability zone, billed hourly, the module's largest fixed monthly cost | Default. Created whenever the app needs internet: AWS Marketplace auto-subscribe is on (aws_bedrock_marketplace_auto_subscribe, enabled unless explicitly set to false) or any AWS service runs outside the deployment region (e.g. multi-region aws_bedrock_regions). Set nat_gateways_allowed = false to instead make the app subnets public (cheaper, less isolated). |
Interface VPC endpoints — Amazon Bedrock, Amazon Polly, Amazon Transcribe, Amazon Comprehend, Amazon Translate, CloudWatch Logs, SSM, ECR, Marketplace metering (Secrets Manager only with api_key_secretsmanager_secret or the per-tenant key secrets, Amazon SQS only with the indexing queue in the deployment region, s3vectors only with the Vector Stores API in the deployment region) |
Only when the app needs no internet egress: aws_bedrock_marketplace_auto_subscribe = false and every AWS service is in the deployment region and vpc_endpoints_allowed = true (default). Replaces the NAT path — the two are never created together. |
| Public subnets | Only with a public ALB (alb_enabled = true and alb_public = true) |
| ALB + HTTPS listener / ACM certificate | alb_enabled = true (HTTPS when alb_domain_name / alb_certificate_arn is set; auto ACM + Route 53 via alb_domain_name). Without an ALB the service is only reachable from inside the VPC. alb_public = true requires an authentication method — an API key source, aws_cognito_user_pool_id, or tenants — and fails the plan without one: the server authenticates nobody by default, and a public ALB would put that on the internet. ACM cannot validate into a private hosted zone, so alb_route53_zone_private = true leaves an HTTP-only listener unless you supply alb_certificate_arn; the plan warns. |
| WAF (rate limiting, IP filtering) | alb_waf_enabled = true (requires alb_enabled) |
| API key authentication | One of api_key_create, api_key, api_key_ssm_parameter, api_key_secretsmanager_secret |
| Amazon Cognito / OAuth 2.0 authentication | aws_cognito_user_pool_id + aws_cognito_client_ids. The module creates no pool: it points the server at yours, which then validates the bearer token's signature, issuer, expiry, application and scopes on every request — an alternative to the API key, and enough on its own to satisfy the alb_public precondition. authentication_mode pins which methods the deployment accepts, and oauth_resource_identifier publishes the protected-resource metadata that the 401 challenge points an agent at. See the getting_started_cognito sample. |
| Realtime API signing key (stored as an SSM parameter, like the API key) | Generated when realtime_client_secret_key is unset and no API key is configured: the server otherwise signs the Realtime API ephemeral client secrets with a per-process random value that no other task can verify. Set realtime_client_secret_key to bring your own; with an API key configured, the server derives it from that key. |
| WebRTC media path (UDP security-group ingress and egress, public task IP, STUN/TURN configuration) | realtime_webrtc_media_enabled = true. The mode needs nat_gateways_allowed = false (the public task IP is the media's inbound path) and a single task (autoscaling_min_capacity = autoscaling_max_capacity = 1, because a call lives in the memory of the task that answered it) — it sets all three itself when you leave them unset, and fails the plan only if you set a contradicting value. Opens the UDP media port range (default 32768-60999) to realtime_webrtc_ingress_ipv4_cidrs/ipv6_cidrs (default: the internet — narrow it when the callers are known), the matching ephemeral range outbound to the same networks, and the STUN and TURN ports the task itself reaches out on. By default the task learns its public address from Google's public STUN server (realtime_webrtc_stun_server, default stun:stun.l.google.com:19302), over an egress rule open to 0.0.0.0/0 on that one port. That is a third-party endpoint outside your account, so point realtime_webrtc_stun_server at a STUN server you operate to keep that flow inside your own perimeter. A security group is only half of the path, so the same flows are written on the application subnets' network ACL, which is stateless and evaluated first: a module-created VPC needs nothing extra. Those subnets reach IPv6 through an egress-only gateway, so an IPv6 caller connects on the candidate pair the task itself initiates. Bringing your own subnet_ids means bringing your own network ACLs — allow the media UDP range inbound and the ephemeral range (1024-65535) in both directions there yourself, since the module cannot write rules in a VPC it did not create. The optional TURN password rides an SSM-backed ECS secret. |
| S3 vector bucket + its KMS key (Vector Stores API) | aws_s3_vectors_bucket_create = true, or pass your own with aws_s3_vectors_bucket. Single-region (aws_s3_vectors_region, default the deployment region): a vector bucket is never handed to a model, so it needs no per-region twin. The API is disabled while neither is set. |
| Amazon SQS indexing queue + its dead-letter queue (durable vector store indexing) | Default, but only once the Vector Stores API is enabled — there is nothing to index without it, and the server refuses the setting on its own. Makes an indexing job outlive the task that accepted it: whichever task is still running finishes it, so a deployment or a scale-in settles the file as completed instead of failed. Set aws_sqs_vector_store_queue_create = false to keep indexing inside the task that accepted the request, or pass your own standard queue with aws_sqs_vector_store_queue_url. Both queues are encrypted with the deployment's KMS key and carry no queue policy. |
| Amazon Bedrock batch service role (Batch API) | aws_bedrock_batch_role_create = true, or pass your own with aws_bedrock_batch_role_arn. The API is disabled while neither is set. |
| Per-end-user role + its inline policy (per-end-user cost attribution) | aws_bedrock_user_role_create = true, or pass your own with aws_bedrock_user_role_arn. All model usage is reported under the task role while neither is set. The created role trusts this deployment's task role alone and may only invoke models and read, from this deployment's own S3 buckets, the media an invocation references by S3 URI; activate the aws_bedrock_user_role_tag_key session tag as an 'IAM principal' cost allocation tag to see the split. |
| Shared DynamoDB table (server state that outlives a task) | Automatically, with the first feature that needs it — one entry in tenants, tenant_api_keys = true, or model_cache_shared = true. Nothing needs configuring, but aws_dynamodb_table and aws_dynamodb_region point at a table you manage yourself instead. Created in the deployment region, encrypted with the deployment's KMS key, on-demand billed so it costs what it is used for. Deletion protection follows what the table holds, not deletion_protection: on whenever tenant API keys are enabled — tenants non-empty, or tenant_api_keys = true — since the tenant secret hashes exist nowhere else; off when the table holds only the shared models list (a cache the next discovery sweep rebuilds). See Destroying a deployment with tenant API keys. |
| Tenant records (one DynamoDB item + one key identifier per tenant) | One entry in tenants, which also creates the table above. Terraform owns the record — identity, model and endpoint scopes, the disabled flag, the optional tenant aws_role_arn, the optional key_generation, and the optional requests_per_minute and tokens_per_minute, which replace tenant_rate_limit_requests_per_minute and tenant_rate_limit_tokens_per_minute for that tenant — and never the key secret: the server mints it and, unless a rotation is asked for (next row), delivers it once through an SSM SecureString under the prefix in the tenant_keys output. Declaring a role also grants the task role sts:AssumeRole on exactly those roles. |
| Tenant key secrets (one AWS Secrets Manager secret per tenant, on the deployment's KMS key) | tenant_key_rotation_days set, or any tenants entry declaring key_generation — either asks for a rotation, and a rotated key needs a durable place to be published that a one-shot SSM delivery is not. Secrets Manager bills per secret per month, so this is the one part of the bill that follows the tenant count. The keys are then stored as each secret's current version instead of delivered through SSM, and rotated in place: on the schedule, and once per key_generation increase; the superseded key stays readable as AWSPREVIOUS and keeps working for tenant_key_rotation_overlap_seconds (server default 7 days). Terraform owns the container — name, tags, encryption key, lifecycle — and the server owns the versions; the task role gains secretsmanager:CreateSecret, DescribeSecret, GetSecretValue, PutSecretValue and UpdateSecretVersionStage on the deployment's prefix alone, never DeleteSecret. The tenant_keys output names each secret; grant a tenant secretsmanager:GetSecretValue on its own and kms:Decrypt on the deployment's key (via Secrets Manager) to let it re-read its key itself. That grant is an identity policy in this deployment's own account: the module writes no resource policy on the secret and takes no extra principal for the KMS key policy, so a tenant reading from its own AWS account needs both written outside the module — kms_key_id brings a key whose policy you own, the secret policy has to be added out of band — or the key handed to it by an operator, as with the SSM delivery. |
| Amazon SageMaker AI endpoint permissions (your own endpoints published as chat models) | One entry in aws_sagemaker_endpoints. No endpoint is created: the task role gains sagemaker:InvokeEndpoint on exactly the endpoint ARNs you name, plus sagemaker:CallWithBearerToken on * — the action that mints the short-term key the OpenAI-compatible route authenticates with, which AWS defines with no resource-level scope. Endpoint hours are billed by SageMaker AI, not by the token. |
Usage & costs API (/v1/organization/usage/*, /v1/organization/costs) |
usage_api = true, which turns on cloudwatch_metrics (it answers from those metrics) and cost_tracking (what puts a cost against the usage) unless either is set explicitly, and adds cloudwatch:GetMetricData and cloudwatch:ListMetrics. Every query is billed by CloudWatch per metric read and is outside its free tier. cloudwatch_metrics_user_dimension = true is what allows grouping by user_id, at one custom metric series per user, model and metric name. |
| Shared Bedrock model list (one discovery pass per fleet) | model_cache_shared = true, which creates the shared DynamoDB table for it. One server refreshes the list and publishes it while the others read it, so a starting task is ready without a discovery pass of its own. Billed on the published list, roughly $0.60 to $2 per month at the default model_cache_seconds. |
| CloudWatch alarms (ECS service alarms, plus one on error/critical log lines) | Setting sns_topic_arn, which is what the alarms notify. alarms_enabled overrides that either way: true without a topic creates up to five alarms with no action, false with a topic creates nothing. |
| VPC Flow Logs | vpc_flow_log_enabled (default true) |
Ready-to-deploy Terraform examples live in the stdapi.ai samples repository.
Deployment shapes — the gateway on its own:
| Example | What it deploys |
|---|---|
| getting_started_production | Single-region production deployment with HTTPS, WAF, auto-scaling |
| getting_started_production_gdpr | Multi-region EU deployment (4 regions) for GDPR data residency |
| getting_started_production_us | Multi-region US deployment (3 regions) for high availability |
| getting_started_cognito | Amazon Cognito user pool authentication instead of a shared API key |
Application integrations — an application in front of a single-region gateway:
| Example | What it deploys |
|---|---|
| getting_started_openwebui | Full Open WebUI chat platform stack (Aurora PostgreSQL + Valkey + SearXNG + stdapi.ai) |
| getting_started_lobehub | LobeHub chat UI (Postgres on EFS + Valkey + S3) |
| getting_started_n8n | n8n workflow automation over the full route surface (Aurora PostgreSQL + Valkey + EFS) |
| getting_started_hermes | Hermes Agent's autonomous agent loop on Bedrock (EFS) |
| getting_started_openclaw | OpenClaw personal-assistant and coding-agent gateway (EFS) |
| getting_started_home_assistant | Home Assistant Assist voice through Amazon Transcribe and Polly (EFS + RDS PostgreSQL Multi-AZ) |
| getting_started_docling | Docling Serve document conversion, the ingestion stage of a RAG pipeline |
| getting_started_ragflow | RAGFlow document Q&A over your own corpus (Amazon OpenSearch + Aurora PostgreSQL + Valkey + S3) |
For integration against existing infrastructure and non-Terraform deployments, see the advanced deployment guide.
stdapi.ai is dual-licensed: AGPL-3.0-or-later for the free community container image, or a commercial license obtained by subscribing on AWS Marketplace. This module is commercial-only by construction — it deploys the Marketplace ECR image, so an active Marketplace subscription is required. To run the AGPL community image instead, see the local deployment guide.
The Marketplace license is metered at $0.10 per container-hour, with a 14-day free trial on the license. autoscaling_min_capacity defaults to one task per availability zone and availability_zones_count defaults to all AZs in the region, so a default deployment in a 3-AZ region runs 3 tasks — about $216/month in license (720 h × $0.10 × 3), and about $432/month in a 6-AZ region such as us-east-1. Those are the floor, not the ceiling: autoscaling_cpu_target_percent defaults to 70, so the fleet scales out under CPU pressure up to autoscaling_max_capacity, five times the minimum unless you set it. Set availability_zones_count, autoscaling_min_capacity and autoscaling_max_capacity explicitly to control this, or autoscaling_cpu_target_percent = null to hold the fleet at a fixed size.
Only the license is covered by the trial. AWS resources this module creates (Fargate, ALB, NAT gateways, KMS, CloudWatch, S3, SQS, DynamoDB, and one Secrets Manager secret per tenant once a key rotation is asked for) and Amazon Bedrock inference are billed by AWS from the first hour, with no markup. Of those, the indexing queues are the one resource with no standing charge at all: Amazon SQS bills per request, an idle queue costs nothing, and indexing a file is a handful of requests. See the cost management guide and the licensing guide.
| Resource | Description |
|---|---|
| Getting Started | Deployment examples and first API call |
| Advanced Deployment | VPC integration, multi-region, cost optimization |
| Configuration | All environment variables and module parameters |
| API Reference | OpenAI & Anthropic compatible API documentation |
| Use Cases | Open WebUI, n8n, coding assistants, and more |
| Features | Full product capabilities |
| Cost Management | License metering, AWS resource costs, and per-request cost estimation |
| Resilience & Failover | Multi-region routing, retry scope, and what does not fail over |
| Licensing | AGPL-3.0 community edition vs the Marketplace commercial license |
| Compliance | Data residency, region allow-lists, encryption, and outbound paths |
| IAM Permissions | Task-role permissions the running gateway needs, for custom deployments and policy auditing |
stdapi.ai is an AWS Qualified Software solution, verified against AWS technical and security requirements for AWS Marketplace.
Controls are grouped below by the deployment feature that gates them, not by which internal module implements them — the baseline section always applies; the rest only come into play once you enable the corresponding input. Only controls whose resolution is worth calling out are listed; N/A controls for resource types this module never creates (EFS, Classic Load Balancers, ECS task sets, Windows containers, Route 53 hosted zones/health checks) are omitted entirely.
Severity: 🔴 Critical · 🟠 High · 🟡 Medium · 🔵 Low
Always created, regardless of subnet_ids, alb_enabled, or any other toggle. Key inputs: tags = local.apn_tags (never null) is applied to every resource below, cloudwatch_logs_retention_in_days defaults 365, the main container sets read_only_root_filesystem = true and user = "65532:65532", secrets are passed via secrets never environment, and no security_group_rules_ingress/security_group_connect_ingress (internal wiring — not module variables) is passed (the only extra ingress rule references the ALB's security group, never a CIDR).
| Control | Severity | Title | Status | Notes |
|---|---|---|---|---|
| KMS.1 / KMS.2 | 🟡 Medium | IAM policies/inline policies should not allow decryption on all KMS keys | ✅ Pass | The application is only ever allowed to decrypt data using the specific encryption keys this deployment creates or is given — never every key in the account. |
| KMS.3 | 🔴 Critical | KMS keys should not be deleted unintentionally | ✅ Pass | No deletion_window_in_days override — AWS's 30-day maximum applies. |
| KMS.4 | 🟡 Medium | KMS key rotation should be enabled | ✅ Pass | Hardcoded for every key, including the per-Bedrock-region ones. |
| KMS.5 | 🔴 Critical | KMS keys should not be publicly accessible | ✅ Pass | Every policy statement scopes a specific AWS service principal with an ArnLike/StringEquals condition — none is a wildcard principal. |
| ECS.3 / ECS.4 / ECS.9 / ECS.10 / ECS.17 / ECS.18 | 🟠 High / 🟡 Medium | Various ECS controls (host PID namespace, non-privileged, logging config, Fargate platform version, host network mode, EFS in-transit encryption) | ✅ Pass | Unconditional defaults — not overridden. |
| ECS.14 | 🔵 Low | ECS clusters should be tagged | ✅ Pass | The cluster always receives a Name tag regardless of tags. |
| ECS.5 | 🟠 High | Task definitions should use read-only root filesystems | ✅ Pass | read_only_root_filesystem = true on the main container. |
| ECS.8 | 🟠 High | Secrets should not be passed as container environment variables | ✅ Pass | The API key is passed via secrets. |
| ECS.12 | 🟡 Medium | ECS clusters should use Container Insights | ✅ Pass | Defaults to "enabled". |
| ECS.13 / ECS.15 | 🔵 Low | Service / task definition should be tagged | ✅ Pass | Non-null tags is applied directly, with no fallback. |
| ECS.20 | 🟡 Medium | Task definitions should configure non-root users for Linux containers | ✅ Pass | user = "65532:65532", matching the Chainguard python:latest base image's actual default non-root user (nonroot, uid/gid 65532). |
| EC2.13 / EC2.14 / EC2.18 / EC2.53 / EC2.54 | 🟠 High | ECS service security group should not allow unrestricted/admin-port ingress | ✅ Pass | No CIDR-based ingress rule exists. |
| EC2.19 | 🔴 Critical | ECS service security group should not allow unrestricted access to high-risk ports | ✅ Pass | Same reasoning as above. |
| EC2.43 | 🔵 Low | ECS service security group should be tagged | ✅ Pass | Unconditional. |
| CloudWatch.16 | 🟡 Medium | CloudWatch log groups should be retained for a specified time period | ✅ Pass | cloudwatch_logs_retention_in_days defaults 365, applied to every ECS log group including Container Insights. |
| CloudWatch.17 | 🟠 High | CloudWatch alarm actions should be activated | ✅ Pass | Unconditional whenever alarms exist (see Other options below). |
| IAM.1 | 🟠 High | IAM policies should not allow full "*" administrative privileges | ✅ Pass | The ECS execution/task role policies and the two aggregated policies, aws_iam_policy.server and aws_iam_policy.server_services, use no wildcard actions. |
| IAM.21 | 🔵 Low | IAM customer managed policies should not allow wildcard actions for services | ✅ Pass | Same statements as IAM.1 — no wildcard (service:*) actions. |
| S3.2 / S3.3 | 🔴 Critical | S3 buckets should block public read/write access | ✅ Pass | aws_s3_bucket_public_access_block sets all four flags to true for every bucket (main and regional). |
| S3.8 | 🟠 High | S3 buckets should block public access (account/bucket combined check) | ✅ Pass | Same configuration as S3.2/S3.3. |
| S3.5 | 🟡 Medium | S3 buckets should require requests to use SSL | ✅ Pass | Bucket policy denies all s3:* actions when aws:SecureTransport is false. |
| S3.6 | 🟠 High | S3 bucket policies should restrict access to other AWS accounts | ✅ Pass | The only statement is the TLS-enforcement Deny; no cross-account Allow. |
| S3.9 | 🟡 Medium | S3 buckets should have server access logging enabled | The main bucket logs to a shared SSE-S3-encrypted logs bucket (also used for ALB access logs); each regional bucket logs to its own per-Region logs bucket, since S3 access log destinations must stay in the source bucket's Region. The logs buckets themselves have no destination: pointing one at itself is refused by S3, and pointing two at each other makes each delivery generate another. | |
| S3.10 / S3.13 | 🟡 Medium / 🔵 Low | S3 buckets should have lifecycle configurations | ✅ Pass | Every bucket gets an unconditional lifecycle configuration (tmp cleanup, files expiration, intelligent-tiering). |
| S3.11 | 🟡 Medium | S3 buckets should have event notifications enabled | ✅ Pass | eventbridge = true is set unconditionally — zero-config, no targets/rules required. |
| S3.12 | 🟡 Medium | ACLs should not be used to manage access to S3 buckets | ✅ Pass | New buckets default to BucketOwnerEnforced (ACLs disabled). |
| S3.14 | 🔵 Low | S3 buckets should have versioning enabled | ✅ Pass | status = "Enabled" unconditionally on every bucket. |
| S3.15 | 🟡 Medium | S3 buckets should have Object Lock enabled | ⬜ N/A | Object Lock (WORM immutability) doesn't fit this bucket's purpose — it's temporary storage with active expiration rules (1-day tmp cleanup, 30-day Files API expiration), the opposite of what Object Lock is for. |
| S3.17 | 🟡 Medium | S3 buckets should be encrypted at rest with AWS KMS keys | SSE-KMS with a dedicated customer-managed key, bucket_key_enabled = true. The logs buckets are SSE-S3: S3 server access logging and ALB access logging both refuse a customer-managed key as their destination, so this is the strongest encryption those buckets can carry. |
|
| S3.20 | 🔵 Low | S3 buckets should have MFA delete enabled | ⬜ N/A (exempt) | AWS's own control text exempts buckets with a lifecycle configuration — every bucket here always has one. |
| S3.22 / S3.23 | 🟡 Medium | S3 buckets should log object-level read/write events | ⬜ N/A | Account-level control requiring an org-wide multi-Region CloudTrail trail — outside this module's scope. |
Thanks to tags always being non-null, ECS.13/ECS.15 (tagged, above) actually pass — leaving tags unset would fail them; ECS.20 also passes thanks to the explicit user.
Key inputs: flow-log retention follows cloudwatch_logs_retention_in_days (default 365); internet access is enabled by default (internal wiring — driven by aws_bedrock_marketplace_auto_subscribe, whose auto-subscribe default requires internet access for the AWS Marketplace API, and by any cross-region service usage); nat_gateways_allowed defaults true; the interface endpoint set (internal wiring — not a module variable) always includes s3, ssm, logs, ecr.api, ecr.dkr.
| Control | Severity | Title | Status | Notes |
|---|---|---|---|---|
| EC2.2 | 🟠 High | VPC default security groups should restrict all traffic | ✅ Pass | Unconditional. |
| EC2.6 | 🟡 Medium | VPC flow logging should be enabled in all VPCs | ✅ Pass | vpc_flow_log_enabled defaults true. |
| EC2.21 | 🟡 Medium | Network ACLs should not allow ingress from 0.0.0.0/0 to port 22/3389 | ❌ Fail (accepted) | The stateless network ACLs allow the ephemeral range for return traffic, which spans 3389. Unreachable: the security groups admit only 80/443 from the CIDRs you name, then only port 8000 from the load balancer. Details. |
| EC2.53 / EC2.54 / EC2.13 / EC2.14 | 🟠 High | VPC default security group should not allow ingress from 0.0.0.0/0 to remote administration ports | ✅ Pass | Unconditional. |
| EC2.12 | 🔵 Low | Unused EIPs should be removed | ✅ Pass | Unconditional. |
| EC2.37 / EC2.39 / EC2.40 / EC2.41 / EC2.42 / EC2.43 / EC2.44 / EC2.46 / EC2.174 | 🔵 Low | Various VPC resources should be tagged | ✅ Pass | A Name tag is always merged in regardless of tags. |
| EC2.48 | 🔵 Low | VPC flow logs should be tagged | ✅ Pass | Non-null tags is applied directly, with no fallback. |
| IAM.24 | 🔵 Low | IAM roles should be tagged | ✅ Pass | Same reasoning as EC2.48, for the flow log's IAM role. |
| CloudWatch.16 | 🟡 Medium | CloudWatch log groups should be retained for a specified time period | ✅ Pass | Flow-log retention follows cloudwatch_logs_retention_in_days (default 365). |
Sub-variant — NAT gateways (default) vs. VPC interface endpoints (no internet egress)
| Control | Severity | Title | Status | Notes |
|---|---|---|---|---|
| EC2.15 | 🟡 Medium | EC2 subnets should not automatically assign public IP addresses | ✅ Pass | No subnet assigns addresses on launch, in either architecture. With nat_gateways_allowed = false the app subnets become public, but the task still takes its address from assign_public_ip on its own network configuration rather than from the subnet. |
| ECS.2 | 🟠 High | Services should not have public IP addresses assigned automatically | ✅ Pass by default | Same trigger as EC2.15 — the ECS service only gets a public IP when nat_gateways_allowed = false, which you either set yourself or get from realtime_webrtc_media_enabled = true (see WebRTC media mode, which always fails this control). |
Sub-variant — compliance VPC endpoints (compliance_vpc_endpoints_enabled = true)
| Control | Severity | Title | Status | Notes |
|---|---|---|---|---|
| EC2.55 / EC2.56 / EC2.57 / EC2.58 / EC2.60 | 🟡 Medium | VPC should be configured with an interface endpoint for ECR API / Docker Registry / SSM / SSM Incident Manager Contacts / SSM Incident Manager | The endpoint set (internal wiring vpc_endpoints_services) already requests ecr.api/ecr.dkr/ssm, but that only takes effect when there's no direct internet route — and by default there is one. Set compliance_vpc_endpoints_enabled = true to force these 5 endpoints regardless of internet posture — except when the app subnets would end up public, leaving no private subnet to place an endpoint in: with realtime_webrtc_media_enabled, or with nat_gateways_allowed = false on its own while internet access is otherwise required (e.g. the default AWS Marketplace auto-subscribe). Both are refused at plan time rather than silently dropping the endpoint. |
Thanks to tags always being non-null, EC2.48 and IAM.24 (tagged, above) actually pass — leaving tags unset would fail them. One gap remains at default settings: EC2.55/56/57/58/60 are silently ineffective, because internet access is required for AWS Marketplace auto-subscribe — set compliance_vpc_endpoints_enabled = true to close it.
| Control | Severity | Title | Status | Notes |
|---|---|---|---|---|
| ELB.1 | 🟡 Medium | ALB should redirect all HTTP requests to HTTPS | The HTTP listener only redirects when a certificate exists. Set alb_certificate_arn or alb_domain_name (with a resolvable Route 53 zone) to get a certificate and enable the redirect. |
|
| ELB.4 | 🟡 Medium | ALB should be configured to drop invalid HTTP headers | ✅ Pass | drop_invalid_header_fields = true is hardcoded. |
| ELB.5 | 🟡 Medium | ALB should have logging enabled | alb_access_logging_enabled defaults true — a dedicated SSE-S3-encrypted bucket is created and wired to the load balancer's access_logs block. |
|
| ELB.6 | 🟡 Medium | ALB should have deletion protection enabled | Set deletion_protection = true (default false). |
|
| ELB.12 | 🟡 Medium | ALB should use defensive or strictest desync mitigation mode | ✅ Pass | Not set explicitly, but AWS's own default (defensive) satisfies the control. |
| ELB.13 | 🟡 Medium | ALB should span multiple Availability Zones | Subnets use all available AZs by default (availability_zones_count = null) — always ≥2 in practice. Fails only if availability_zones_count is explicitly set to 1. |
|
| ELB.17 | 🟡 Medium | ALB listeners should use recommended security policies | alb_ssl_policy defaults to ELBSecurityPolicy-TLS13-1-2-Res-PQ-2025-09, one of AWS's recommended policies. Only applies once the HTTPS listener exists (see ELB.1). |
|
| ELB.18 | 🟡 Medium | ALB listeners should be configured with a secure listener protocol | ❌ Fail | The HTTP listener (port 80) always exists; there's no exemption for a redirect-only listener. Inherent to offering both HTTP and HTTPS. |
| ELB.21 | 🟡 Medium | ELB target groups should have health check configured with encrypted protocol | ❌ Fail | The health check uses the default HTTP protocol; no variable exposes HTTPS health checks. Requires a code change to pass. |
| ELB.22 | 🟡 Medium | ELB target groups should use encrypted transport protocol | ❌ Fail | The target group forwards to the ECS task over plain HTTP on the container port (TLS terminates at the ALB, not re-established to the backend). Requires a code change to pass. |
| ACM.1 | 🟡 Medium | ACM certificates should be renewed after a specified time period | ✅ Pass | DNS validation (validation_method = "DNS"), which ACM renews automatically. |
| ACM.2 | 🟠 High | RSA certificates managed by ACM should use a key length of at least 2,048 bits | ✅ Pass | key_algorithm isn't set, so ACM uses its default RSA_2048. |
| ACM.3 | 🔵 Low | ACM certificates should be tagged | ✅ Pass | Tagged via non-null tags plus a Name tag. |
Sub-variant — WAF (alb_waf_enabled = true, requires alb_enabled = true)
| Control | Severity | Title | Status | Notes |
|---|---|---|---|---|
| ELB.16 | 🟡 Medium | ALB should be associated with a WAF web ACL | Set alb_waf_enabled = true to pass. |
|
WAFV2.1 (AWS WAF WAF.10) |
🟡 Medium | AWS WAF web ACLs should have at least one rule or rule group | ✅ Pass once enabled | Three AWS managed rule groups are always attached — never empty. |
WAFV2.2 (AWS WAF WAF.11) |
🔵 Low | AWS WAF web ACL logging should be enabled | ✅ Pass once enabled | alb_waf_logging_enabled defaults true. |
With alb_enabled = false (default), none of the ALB/WAF/ACM controls above apply — no load balancer exists. Once enabled, ELB.4 and ELB.5 pass out of the box. ELB.18, ELB.21, ELB.22 fail unconditionally — closing them requires re-architecting to terminate TLS on the backend, not just adding a variable — while ELB.1, ELB.6, ELB.16 fail by default until their respective variables are set.
Client IP trust (X-Forwarded-For) — not a Security Hub control, but a hardening applied automatically for ALB deployments. When the ALB is enabled together with log_client_ip = true, the module enables ENABLE_PROXY_HEADERS so the real client IP (not the ALB's) is recorded in request logs and OpenTelemetry spans. To keep that value trustworthy, it also pins PROXY_TRUSTED_HOSTS to the ALB's own subnet CIDRs (IPv4 and IPv6): the server honors X-Forwarded-* only when the immediate peer is the ALB. Because an ALB appends to X-Forwarded-For rather than replacing it, without this restriction a client could prepend a forged entry and poison the recorded client IP; pinning the trust to the ALB subnets makes the real appended address authoritative instead. This is defense in depth on top of the ECS security group, which already allows ingress only from the ALB's security group. Set proxy_trusted_hosts explicitly to override — for example when fronting the ALB with an additional proxy such as CloudFront, set it to that proxy's egress range. On an IPv6-enabled VPC the container binds a dual-stack socket (GRANIAN_HOST=::, needed because service discovery publishes an AAAA record per task), and the kernel then reports an IPv4 peer in IPv4-mapped form such as ::ffff:10.0.1.5, which belongs to no IPv4 network. The module therefore adds the matching ::ffff: range for every IPv4 entry — including entries you set through proxy_trusted_hosts yourself — so the ALB stays trusted. Write entries in their natural address family and let the module handle the mapping; a hand-rolled PROXY_TRUSTED_HOSTS on a dual-stack listener must cover both forms itself.
Independent toggles that aren't required to pass any control above. Only the WebRTC media mode changes the status of a control listed earlier, and its table says which.
WebRTC media mode (realtime_webrtc_media_enabled = true)
| Control | Severity | Title | Status | Notes |
|---|---|---|---|---|
| EC2.18 | 🟠 High | Security groups should only allow unrestricted incoming traffic for authorized ports | Enabling the mode opens the UDP media port range to realtime_webrtc_ingress_ipv4_cidrs (default 0.0.0.0/0), which this control flags unless the allowed ports are customized. WebRTC media is inbound UDP on ephemeral ports to the task itself — there is no narrower shape that still serves arbitrary browsers. Narrow the ingress CIDRs to your callers' networks where they are known, or accept and document the finding. |
|
| EC2.21 | 🟡 Medium | Network ACLs should not allow ingress from 0.0.0.0/0 to port 22/3389 | ❌ Fail (accepted, unchanged) | The mode adds the media UDP range and the STUN/TURN reply range to the application subnets' network ACL, for 0.0.0.0/0 and ::/0: a network ACL takes no narrower source, and the security group in front is what restricts the callers. The control already fails on the ephemeral return-traffic range (see its row in the main table above), and the added entries cover neither 22 nor 3389 unless you move realtime_webrtc_udp_port_range over them. |
| ECS.2 | 🟠 High | Services should not have public IP addresses assigned automatically | Always fails while the mode is on. The mode derives nat_gateways_allowed = false, which is what gives the ECS service the public IP the media arrives on; the ECS.2 row of the NAT sub-variant above passes only because that value is otherwise unset. Terminating media on the task means addressing the task publicly — there is no configuration of this mode that passes. Accept and document the finding, or keep the mode off. |
CloudWatch Alarms (created by sns_topic_arn, which they notify; alarms_enabled overrides either way)
| Control | Severity | Title | Status | Notes |
|---|---|---|---|---|
| CloudWatch.15 | 🟠 High | CloudWatch alarms should have specified actions configured | alarms_enabled defaults to whether sns_topic_arn is set, so setting the topic is what creates the alarms and every one of them attaches it: the control passes. alarms_enabled = true with no topic creates the same alarms with no action and fails it; alarms_enabled = false suppresses them all and leaves it N/A. |
Enabling it creates up to five alarms:
- High memory usage — ECS service
MemoryUtilization> 90% for 4 of 5 one-minute periods. - Unhealthy containers — ECS service
HealthCheckFailed> 0. - CPU anomaly detection — CloudWatch anomaly-detection band around
CPUUtilization; fires when usage exceeds the expected upper bound. - Max autoscaling capacity reached — Container Insights
DesiredTaskCount>=autoscaling_max_capacity(only created if min/max capacity differ and Container Insights is enabled). - Application error/critical logs — a log metric filter counts
error/ERROR/critical/CRITICALmatches in the app's CloudWatch log group; fires when any appear within a 60-second period.
None of them requires sns_topic_arn: each attaches the topic as its action when one is set, and is created without any action when there is none. So alarms_enabled = true with no topic bills you for every alarm created — four when autoscaling_min_capacity equals autoscaling_max_capacity, since the max-capacity alarm is then not created, and the anomaly-detection one at the higher anomaly-detection rate — and notifies nobody.
Durable vector store indexing (aws_sqs_vector_store_queue_create = true, the default, effective once the Vector Stores API is enabled)
Two queues are created — the indexing queue and its dead-letter queue — and both are treated identically: a dead-letter queue is a queue, and every control below is evaluated on each of them.
| Control | Severity | Title | Status | Notes |
|---|---|---|---|---|
| SQS.1 | 🟡 Medium | Amazon SQS queues should be encrypted at rest | ✅ Pass | kms_master_key_id is set to the deployment's customer-managed key on both queues, so each is SSE-KMS. The control accepts SSE-SQS too; the CMK is used to match the encryption posture of every other resource here. |
| SQS.2 | 🔵 Low | SQS queues should be tagged | ✅ Pass | Both queues get the non-null tags plus a Name tag, so a tag key always exists. Fails only if you configure requiredTagKeys with keys this module does not set. |
| SQS.3 | 🔴 Critical | SQS queue access policies should not allow public access | ✅ Pass | Neither queue has an access policy at all: the task role reaches the queue through its identity policy, scoped to that one queue ARN. A control that inspects the resource policy cannot find a wildcard principal where there is no resource policy. |
The dead-letter queue additionally carries a RedriveAllowPolicy restricted to this deployment's own queue (redrivePermission = "byQueue"), so no other queue in the account can redrive into it. That is not a Security Hub control — left unset, Amazon SQS would allow every queue of the account.
Shared DynamoDB table (created with the first feature that needs it: an entry in tenants, tenant_api_keys = true, or model_cache_shared = true)
| Control | Severity | Title | Status | Notes |
|---|---|---|---|---|
| DynamoDB.1 | 🟡 Medium | DynamoDB tables should automatically scale capacity with demand | ✅ Pass | billing_mode = "PAY_PER_REQUEST" (on-demand), unconditionally. |
| DynamoDB.2 | 🟡 Medium | DynamoDB tables should have point-in-time recovery enabled | ✅ Pass | point_in_time_recovery { enabled = true }, unconditionally. |
| DynamoDB.3 / DynamoDB.7 | 🟡 Medium | DynamoDB Accelerator (DAX) clusters should be encrypted at rest / in transit | ⬜ N/A | This module never creates a DAX cluster. |
| DynamoDB.4 | 🟡 Medium | DynamoDB tables should be present in a backup plan | ❌ Fail | This module does not manage AWS Backup, so the table is never assigned to a backup plan. Point-in-time recovery (DynamoDB.2, always on) is the compensating control: it restores to any point in the last 35 days, but it is not an AWS Backup plan and does not satisfy this control. Assign the table to a backup plan yourself if you need this control to pass. |
| DynamoDB.5 | 🔵 Low | DynamoDB tables should be tagged | ✅ Pass | The table gets the non-null tags plus a Name tag, so a tag key always exists. |
| DynamoDB.6 | 🟡 Medium | DynamoDB tables should have deletion protection enabled | deletion_protection_enabled is derived from what the table holds, not from deletion_protection. With tenant API keys enabled — tenants declared, or tenant_api_keys = true — the table holds the tenant secret hashes and salts the server minted — they exist nowhere else, and losing them invalidates every tenant credential with no way back — so protection is on and the control passes. Holding only the shared models list, the table is a cache: a manifest, compressed shards and a 120-second lease, all TTL'd and all rebuilt by the next discovery sweep. There is nothing there to protect and a protected cache is only a destroy that cannot finish, so protection is off and the control reports a finding on a table whose entire contents are disposable. |
The table is encrypted with this deployment's own KMS key, the same one the S3 buckets, the log groups and the queues use — there is nothing to configure and no AWS owned key fallback. The ECS task role is granted kms:Decrypt on that key through DynamoDB alone (kms:ViaService), since DynamoDB decrypts the table's key as the caller; the principal running Terraform does need kms:CreateGrant and kms:DescribeKey on that key (see Prerequisites). An administrator principal already has both; a narrowly scoped deployment role may not, and fails at apply with an explicit KMS authorization error. There is deliberately no fallback to the AWS owned key: silently downgrading the table to weaker encryption is worse than failing where you can see it.
While tenant API keys are enabled — tenants has at least one entry, or tenant_api_keys = true — the shared DynamoDB table carries deletion_protection_enabled = true and AWS refuses DeleteTable on it. tofu destroy therefore fails partway through, and so does any plan that would drop the table — emptying tenants on its own does exactly that, since the table then has no feature left to exist for. The teardown is two applies, and the first has to keep the table alive while clearing its protection:
# 1. Drop the tenant records and clear the protection, keeping the table itself.
# Every tenant API key stops working here. Add -var 'tenant_api_keys=false'
# when the deployment sets tenant_api_keys: the protection follows that
# override, so emptying tenants alone does not clear it.
tofu apply -var 'tenants={}' -var 'model_cache_shared=true'
# 2. Destroy the rest, table included.
tofu destroymodel_cache_shared = true is what holds the table through step 1: the table's protection follows whether tenant API keys are enabled, which tenants alone decides unless tenant_api_keys overrides it, so emptying tenants clears it, and keeping a second reason for the table to exist is what turns that same apply into an in-place update instead of a destroy AWS refuses (ValidationException: Resource cannot be deleted as it is currently protected against deletion). Set both values in your own configuration rather than on the command line if your root module does not expose them as variables. The second apply then removes the table for good, and skipping straight to it leaves a half-destroyed deployment.
Step 1 also drops each tenant's Secrets Manager secret where the keys are stored there. Secrets Manager only schedules a deletion: each secret stays in the account for the provider's default 30-day recovery window, still billed and with its name reserved against a re-creation under the same tenant. Delete one immediately — to free the name, or to stop the charge — with aws secretsmanager delete-secret --secret-id <secret> --force-delete-without-recovery, which cannot be undone.
A deployment already stuck — the table protected, its apply failing — is unblocked with one API call, after which the two steps above (or tofu destroy alone, if tenants is already empty in the configuration) complete:
aws dynamodb update-table --table-name <table> --no-deletion-protection-enabledA deployment that never enabled tenant API keys needs none of this: tofu destroy completes on its own.
Tenant key secrets (created with tenant_key_rotation_days, or a tenants entry declaring key_generation)
| Control | Severity | Title | Status | Notes |
|---|---|---|---|---|
| SecretsManager.1 | 🟡 Medium | Secrets Manager secrets should have automatic rotation enabled | ❌ Fail (justified) | The control passes only for a secret rotated by a Lambda function Secrets Manager invokes, and these secrets are rotated by the gateway itself through the native staging labels — on tenant_key_rotation_days, and on demand per key_generation — with RotationEnabled never set. A rotation function would have to reproduce the gateway's salted hashing in a second runtime holding write access to the credential table, would leave every deployment without this module (Docker, CloudFormation) unrotated, and would add a Lambda runtime to patch; the finding is the cost of not doing that. Suppress it for the deployment's prefix, or accept it as documented. |
| SecretsManager.2 | 🟡 Medium | Secrets Manager secrets configured with automatic rotation should rotate successfully | ⬜ N/A | Evaluated only for secrets with automatic rotation configured, which these never have. |
| SecretsManager.3 | 🟡 Medium | Remove unused Secrets Manager secrets | Measures the last read of each secret. A tenant that re-reads its secret at least every 90 days passes; one that keeps its key locally and never re-reads is expected to fail, and so is a rotation schedule longer than 90 days (the server never reads a version back on the steady state). Nothing here is unused: a flagged secret is still the tenant's live credential. | |
| SecretsManager.4 | 🟡 Medium | Secrets Manager secrets should be rotated within a specified number of days | tenant_key_rotation_days ≤ 90) |
The control measures LastChangedDate, which every rotation's version write moves. A schedule of 90 days or less keeps every secret within the control's default window; keys rotated on demand only, or on a longer schedule, report a finding. |
| SecretsManager.5 | 🔵 Low | Secrets Manager secrets should be tagged | ✅ Pass | Each secret gets the non-null tags. The server creates a secret itself only where none exists — a deployment declaring tenants outside this module — and never tags one. |
Amazon GuardDuty Runtime Monitoring endpoint (guardduty_vpc_endpoint_enabled = true, dedicated VPC only) — not a Security Hub control. Enforced regardless of internet posture, except when the app subnets would end up public, leaving no private subnet for the endpoint's network interface: with realtime_webrtc_media_enabled, or with nat_gateways_allowed = false on its own while internet access is otherwise required (e.g. the default AWS Marketplace auto-subscribe). Both are refused at plan time rather than silently dropping the endpoint.
Route 53 Resolver DNS Firewall (dns_firewall_enabled = true, dedicated VPC only) — not a Security Hub control. Blocks/alerts on DNS queries to known-malicious domains (AWS Managed Domain Lists, plus DGA/DNS-tunneling detection via dns_firewall_advanced_enabled). Complements application-level SSRF protection against malicious-URL injection through user-supplied URL/file reference fields. Has no effect (and cannot be enabled) when subnet_ids is set.
All the options above are off by default and never required to pass a control in the sections higher up.
| Name | Version |
|---|---|
| terraform | >= 1.9.0 |
| aws | >= 6.27.0 |
| random | >= 3.0.0 |
| Name | Version |
|---|---|
| aws | >= 6.27.0 |
| random | >= 3.0.0 |
| terraform | n/a |
| Name | Source | Version |
|---|---|---|
| kms_key | JGoutin/kms-key/aws | ~> 1.2 |
| regional_kms | JGoutin/kms-key/aws | ~> 1.2 |
| server | JGoutin/ecs-fargate/aws | >= 1.4.4, < 2.0.0 |
| vectors_kms | JGoutin/kms-key/aws | ~> 1.2 |
| vpc | JGoutin/vpc/aws | ~> 1.6 |
| Name | Description | Type | Default | Required |
|---|---|---|---|---|
| ai_response_timeout | Maximum time in seconds to wait for an AI model to complete a response. Applies to both streaming and non-streaming requests. The default of 600 seconds accommodates models with extended reasoning. Increase for long-running requests (e.g., large document analysis); decrease to fail fast on unexpectedly slow responses. Default to 600. | number |
null |
no |
| alarms_enabled | Enable CloudWatch alarms. Default to whether sns_topic_arn is set: an alarm with no topic to notify is one nobody sees, and a topic configured for alarms that do not exist is a setting that does nothing. Set explicitly to override either way. | bool |
null |
no |
| alb_access_logging_enabled | If true, enable ALB access logging to a dedicated S3 bucket. Security Hub: ELB.5 (Application Load Balancers should have logging enabled) — default true = pass; only relevant when var.alb_enabled is true. | bool |
true |
no |
| alb_certificate_arn | Existing ACM certificate ARN to attach to the HTTPS listener. When specified, takes precedence over certificate_create. If not specified and certificate_create is true, a certificate will be created automatically. | string |
null |
no |
| alb_certificate_create | If true, create an ACM certificate and validate it via DNS. Only used when certificate_arn is not specified. Requires route53_zone_id, domain_name, and route53_zone_private=false. | bool |
true |
no |
| alb_domain_name | Primary domain name for the application (e.g., api.example.com). Creates Route 53 A record and ACM certificate. If route53_zone_id is not specified, automatically looks up the most specific parent domain zone. | string |
null |
no |
| alb_enabled | If true, create an Application Load Balancer for the ECS service. Cannot be used with external subnets (subnet_ids). | bool |
false |
no |
| alb_idle_timeout | The time in seconds that the connection is allowed to be idle. Range: 1-4000 seconds. Default to 3600 (1 hour) to support slow LLM responses and long-running operations like AWS Transcribe. | number |
3600 |
no |
| alb_ingress_ipv4_cidrs | List of IPv4 CIDR blocks allowed to access the ALB. Default to ['0.0.0.0/0'] for public access. | list(string) |
[ |
no |
| alb_ingress_ipv6_cidrs | List of IPv6 CIDR blocks allowed to access the ALB. Default to ['::/0'] for public access. | list(string) |
[ |
no |
| alb_public | If true, create a public (internet-facing) ALB with dedicated public subnets. If false, create a private (internal) ALB using app subnets. | bool |
false |
no |
| alb_route53_zone_id | Route 53 hosted zone ID for DNS records. If not specified, automatically infers the zone from the parent domain of domain_name (e.g., 'api.example.com' → 'example.com', 'api.sandbox.example.com' → 'sandbox.example.com'). | string |
null |
no |
| alb_route53_zone_name | Route 53 hosted zone name for DNS records (e.g., 'example.com'). Alternative to route53_zone_id - module will look up the zone ID automatically. If specified with domain_name, creates DNS records and ACM certificate. | string |
null |
no |
| alb_route53_zone_private | If true, the Route 53 zone is private. If false, it's public. Used when looking up the zone by name. | bool |
false |
no |
| alb_ssl_policy | SSL/TLS security policy for the ALB HTTPS listener. Defaults to the AWS-recommended post-quantum policy. See https://docs.aws.amazon.com/elasticloadbalancing/latest/application/describe-ssl-policies.html | string |
"ELBSecurityPolicy-TLS13-1-2-Res-PQ-2025-09" |
no |
| alb_waf_block_anonymous_ips | If true, block requests from anonymous IP addresses (VPNs, proxies, Tor exit nodes). | bool |
false |
no |
| alb_waf_enabled | If true, create a WAF WebACL and associate it with the ALB (requires alb_enabled=true). | bool |
false |
no |
| alb_waf_logging_enabled | If true, enable WAF logging to CloudWatch Logs. | bool |
true |
no |
| alb_waf_rate_limit | Maximum number of requests allowed from a single IP address in a 5-minute period. If null, rate limiting is disabled. | number |
null |
no |
| anthropic_beta_allowlist | Additional anthropic_beta flags to allow beyond the built-in defaults. Comma-separated string. Merged with the built-in set of Bedrock-supported flags. Only effective when anthropic_beta_filter is true. | string |
null |
no |
| anthropic_beta_filter | Enable filtering of unsupported anthropic_beta flags for Anthropic Claude models. When enabled, flags not in the allowlist are silently removed to prevent Bedrock ValidationException errors. Default to true. | bool |
null |
no |
| anthropic_routes_prefix | Anthropic API compatible routes prefix. Default to '/anthropic'. | string |
null |
no |
| api_key | API key for client authentication. When specified, all API requests must include this key. Mutually exclusive with api_key_create, api_key_ssm_parameter, and api_key_secretsmanager_secret. | string |
null |
no |
| api_key_create | If true, generate and return an API key using the 'api_key' output. When specified, all API requests must include this key. Mutually exclusive with api_key, api_key_ssm_parameter, and api_key_secretsmanager_secret. | bool |
false |
no |
| api_key_secretsmanager_key | Key name within the AWS Secrets Manager secret containing the API key. Only used when api_key_secretsmanager_secret is specified. | string |
null |
no |
| api_key_secretsmanager_secret | AWS Secrets Manager secret name containing the API key. Mutually exclusive with api_key_create, api_key, and api_key_ssm_parameter. When using this option, you must create an IAM policy granting secretsmanager:GetSecretValue permission and pass the policy ARN to var.ecs_task_role_policy_arns. | string |
null |
no |
| api_key_ssm_parameter | AWS Systems Manager Parameter Store parameter name containing the API key. Mutually exclusive with api_key_create, api_key, and api_key_secretsmanager_secret. When using this option, you must create an IAM policy granting ssm:GetParameter permission and pass the policy ARN to var.ecs_task_role_policy_arns. | string |
null |
no |
| authentication_mode | Which client authentication methods this deployment accepts: 'any' for every method that is configured, 'api_key' for the API key only, or 'cognito' for Amazon Cognito user pool tokens only. The value asserts the intended security posture: the server fails to start when the selected method is not configured, and when a method that would be ignored is configured anyway, so a credential is never accepted or silently refused by accident. Default to 'any'. | string |
null |
no |
| autoscaling_alb_target_requests_per_target | Target number of ALB requests per ECS task for auto-scaling. If null or ALB not enabled, request-based scaling is disabled. This tracks the load the gateway actually carries more closely than CPU does. No policy is created when the minimum and maximum capacities are equal. | number |
null |
no |
| autoscaling_cpu_target_percent | Target CPU utilization percentage for auto-scaling. If null, CPU-based scaling is disabled. The gateway spends most of its time awaiting AWS responses, so this mainly answers the audio and video transcoding that does use the CPU; set autoscaling_alb_target_requests_per_target as well to scale on request volume. No policy is created when the minimum and maximum capacities are equal. | number |
70 |
no |
| autoscaling_max_capacity | Maximum number of ECS tasks for auto-scaling. If null, defaults to five times the minimum capacity. | number |
null |
no |
| autoscaling_memory_target_percent | Target memory utilization percentage for auto-scaling. If null, memory-based scaling is disabled. No policy is created when the minimum and maximum capacities are equal. | number |
null |
no |
| autoscaling_min_capacity | Minimum number of ECS tasks. If not specified, defaults to the number of availability zones. | number |
null |
no |
| autoscaling_scale_in_cooldown | Time in seconds after a scale-in activity completes before another scale-in can start. If null, uses AWS default. | number |
null |
no |
| autoscaling_scale_out_cooldown | Time in seconds after a scale-out activity completes before another scale-out can start. If null, uses AWS default. | number |
null |
no |
| autoscaling_schedule_start | Schedule to start the service if stopped. Format: cron(fields) or at(yyyy-mm-ddThh:mm:ss) in UTC. | string |
null |
no |
| autoscaling_schedule_stop | Schedule to stop/pause the service (scale to 0). Format: cron(fields) or at(yyyy-mm-ddThh:mm:ss) in UTC. | string |
null |
no |
| autoscaling_spot_on_demand_min_capacity | Minimum number of on-demand tasks when autoscaling_spot_percent is enabled. If not specified, defaults to autoscaling_min_capacity. | number |
null |
no |
| autoscaling_spot_percent | Percent of capacity over the minimum capacity to run with Fargate Spot (~70% cost discount). Set to 100 to use only Spot instances. Set to 0 to disable Spot instances. | number |
0 |
no |
| availability_zones_count | Maximum count of availability zones to provision with the dedicated VPC. Default to all available availability zones. | number |
null |
no |
| aws_adaptive_retry | Enable adaptive retry mode for all AWS service calls. When enabled, the client dynamically adjusts its retry behavior based on observed error rates, slowing down when a service appears congested. Default to false. | bool |
null |
no |
| aws_bedrock_allow_application_inference_profile_arn | If True, allow users to pass application inference profile ARNs directly as model IDs. Application inference profiles are custom routing configurations for specific use cases. When disabled, only standard model IDs and configured profiles are accepted. | bool |
null |
no |
| aws_bedrock_allow_cross_region_inference_profile_arn | If True, allow users to pass cross-region inference profile ARNs directly as model IDs. Cross-region inference profiles enable routing to multiple regions for better availability. When disabled, only standard model IDs and configured profiles are accepted. | bool |
null |
no |
| aws_bedrock_allow_external_web_access_override | If true, allow clients to override aws_bedrock_external_web_access per request with the web search tool's 'external_web_access' field. When false, a request that sets a different value is rejected. Default to false. | bool |
null |
no |
| aws_bedrock_allow_guardrail_override | Allow users to override the global guardrail configuration at request level using headers (X-Amzn-Bedrock-GuardrailIdentifier, X-Amzn-Bedrock-GuardrailVersion, X-Amzn-Bedrock-Trace). When disabled and a global guardrail is configured, request headers are ignored for security. Defaults to false for security. | bool |
null |
no |
| aws_bedrock_allow_mantle_project_override | If true, allow clients to override the configured Amazon Bedrock Mantle project per request via the 'OpenAI-Project' / 'anthropic-workspace' header. Default to false. | bool |
null |
no |
| aws_bedrock_allow_marketplace_endpoint_arn | If true, allow users to pass the ARN of an Amazon Bedrock Marketplace model endpoint directly as a model ID. Cost-bearing: a caller who can name any endpoint ARN can direct traffic at instances you are already paying for. Default to false. | bool |
null |
no |
| aws_bedrock_allow_prompt_arn | If true, allow users to reference an Amazon Bedrock Prompt Management prompt ARN in the OpenAI Responses API 'prompt.id' parameter, for example 'arn:aws:bedrock:us-east-1:123456789012:prompt/ABCDE12345:1'. The prompt template is rendered by Amazon Bedrock and its variables are filled from 'prompt.variables'. Setting it to true also grants the task role bedrock:GetPrompt and bedrock:RenderPrompt on every prompt of the account. Default to false, which rejects any 'prompt' parameter with a 400 error. | bool |
null |
no |
| aws_bedrock_allow_prompt_router_arn | If True, allow users to pass prompt router ARNs directly as model IDs. Prompt routers enable dynamic model selection based on prompt characteristics. When disabled, only standard model IDs and configured profiles are accepted. | bool |
null |
no |
| aws_bedrock_allow_service_tier_override | Allow users to select the service tier at request level, through the 'service_tier' request parameter or the X-Amzn-Bedrock-Service-Tier header. When disabled, a request cannot change the tier configured for the model by default_model_service_tiers or by the model alias it names. A model with no configured tier still honors the request in either case. Defaults to true. | bool |
null |
no |
| aws_bedrock_batch_role_arn | ARN of an existing IAM service role Amazon Bedrock assumes to run batch inference jobs. Its trust policy must allow 'bedrock.amazonaws.com' to assume it, and it must be able to read and write every bucket the server may use, under aws_s3_batches_prefix; this module grants the task role 'iam:PassRole' on this ARN alone, for Amazon Bedrock only. When specified, takes precedence over aws_bedrock_batch_role_create. Default to none, meaning a role is created when aws_bedrock_batch_role_create is true, and the Batch API is disabled otherwise. | string |
null |
no |
| aws_bedrock_batch_role_create | If true, create the IAM service role Amazon Bedrock assumes to run batch inference jobs, allowed to read the submitted requests and write the results under aws_s3_batches_prefix in the module-managed buckets, and to invoke foundation models and inference profiles. Only used when aws_bedrock_batch_role_arn is not specified. When aws_bedrock_batch_role_arn is specified, this value is ignored. Default to false (Batch API disabled). | bool |
false |
no |
| aws_bedrock_cross_region_inference | If true, allow cross region inference to be used. Default to true. | bool |
null |
no |
| aws_bedrock_cross_region_inference_global | If True, allow 'global' cross region inference that can route requests to any region, worldwide. Default to true. | bool |
null |
no |
| aws_bedrock_deprecated_model_fallback | If true, requests that use a deprecated model ID are transparently retried with the recommended replacement model instead of returning a 404 error. Disable if you want deprecated model IDs to fail explicitly so clients are forced to migrate. Default to true. | bool |
null |
no |
| aws_bedrock_deprecated_models | Additional deprecated model ID mappings, merged with the built-in deprecation registry at startup. User-provided entries take precedence over built-in ones. Keys are deprecated model IDs, values are the recommended replacement model IDs. Example: { "my-old-model-v1" = "my-new-model-v2" } |
map(string) |
null |
no |
| aws_bedrock_external_web_access | If true, let the built-in web search tool reach the public web instead of answering from the Amazon Bedrock web index and cache. Requires the 'bedrock-websearch:ExternalWebAccess' IAM permission, granted by this module when enabled. Default to false. | bool |
null |
no |
| aws_bedrock_guardrail_checks_pii_entities | PII entity types to detect on the Moderations API when it classifies with inline Amazon Bedrock guardrail checks, for instance ["EMAIL", "PHONE"]. Billed as a check of its own. Empty or null detects none. Broad types such as ADDRESS, NAME and URL match ordinary prose. Has no effect on classifications served by a guardrail resource, which applies the sensitive information policy its own configuration defines. | list(string) |
null |
no |
| aws_bedrock_guardrail_checks_prompt_attack | Detect prompt attacks (jailbreaks, prompt injection, prompt leakage) on the Moderations API when it classifies with inline Amazon Bedrock guardrail checks. Billed as a check of its own. Has no effect on classifications served by a guardrail resource, which applies the prompt attack filter its own configuration defines. | bool |
null |
no |
| aws_bedrock_guardrail_identifier | Amazon Bedrock Guardrails ID. | string |
null |
no |
| aws_bedrock_guardrail_scope_turns | Number of trailing user turns an Amazon Bedrock guardrail evaluates on the chat routes. Lowers the bill on long conversations by no longer evaluating the history the client replays, and lowers detection with it. Only the text of a scoped turn is submitted, so images and other non-text content stop being evaluated at any value; model output is always evaluated in full. Defaults to null: the whole conversation is evaluated on every request. | number |
null |
no |
| aws_bedrock_guardrail_trace | Amazon Bedrock Guardrails trace setting: disabled, enabled, or enabled_full. | string |
null |
no |
| aws_bedrock_guardrail_version | Amazon Bedrock Guardrails version. | string |
null |
no |
| aws_bedrock_knowledge_base_ids | Allowlist of Amazon Bedrock knowledge bases served through the Vector Stores API. Each allowlisted knowledge base is addressed as the vector store vs_kb_<knowledgeBaseId> on every /v1/vector_stores endpoint and is listed alongside the stores the server owns: searching runs against it, attaching a file ingests a document, listing and reading files report its documents back, and deleting a file deletes the document.Write each entry as <knowledgeBaseId>, or as <knowledgeBaseId>/<dataSourceId> when the knowledge base has more than one data source; with a single data source the server resolves it itself. For example ["ABCDE12345", "FGHIJ67890/KLMNO13579"]. Each knowledge base must live in the first aws_bedrock_regions region, which is the region this module grants access to it in.The knowledge base stays yours: this module never creates or deletes one, and the task role is granted no action that would reshape it, only bedrock:Retrieve and the document actions of its data source, on the ARN of each listed knowledge base.Default to an empty list, which grants no permission on any knowledge base and makes none of them addressable: a vs_kb_... identifier is then answered exactly as an unknown vector store is. |
list(string) |
[] |
no |
| aws_bedrock_legacy | If true, allow legacy Bedrock models to be used. Default to false. | bool |
null |
no |
| aws_bedrock_mantle_enabled | If true (application default), expose models served by the Amazon Bedrock Mantle endpoint (OpenAI GPT, xAI Grok, Google Gemma, and more) in addition to the classic Bedrock Converse models. Set to false to disable Mantle. When enabled but Mantle is unreachable or the region lacks the service, Mantle models are simply not listed. | bool |
null |
no |
| aws_bedrock_mantle_endpoint_url | Override the Amazon Bedrock Mantle endpoint URL template, with '{region}' substituted for the target region. Point it at a VPC endpoint or an inspection proxy you already operate. Default to 'https://bedrock-mantle.{region}.api.aws'. | string |
null |
no |
| aws_bedrock_mantle_preferred_models | Model IDs (or ID prefixes) served by Amazon Bedrock Mantle even when also available on the classic bedrock-runtime endpoint. Cannot be combined with Bedrock Guardrails, which Mantle does not apply. Left unset, the server's own default applies: the OpenAI GPT-5.6 and GPT-6 families, which serve the OpenAI server tools only on Mantle and are therefore billed at the In-Region rate, 1.10x the cross-region one. An explicit list replaces that default entirely rather than adding to it: repeat 'openai.gpt-5.6' and 'openai.gpt-6' to keep it. Set to an empty list to route every model through bedrock-runtime instead. | list(string) |
null |
no |
| aws_bedrock_mantle_project | Default Amazon Bedrock Mantle project (workspace) ID used to attribute Mantle inference requests for cost tracking and observability. A bare project ID such as 'proj_abc123' or 'default' (not an ARN). Default to none. | string |
null |
no |
| aws_bedrock_mantle_regions | List of AWS regions used for Amazon Bedrock Mantle, in failover priority order. Default to var.aws_bedrock_regions. | list(string) |
null |
no |
| aws_bedrock_mantle_service_header | If true, honor the 'x-stdapi-service: bedrock-mantle' request header to route a dual-homed model through Bedrock Mantle for that request. Cannot be combined with Bedrock Guardrails. Default to false. | bool |
null |
no |
| aws_bedrock_marketplace_auto_subscribe | If true, allow the server to automatically subscribe to new models in the AWS Marketplace. Default to true. | bool |
null |
no |
| aws_bedrock_marketplace_endpoint_regions | Regions searched for Amazon Bedrock Marketplace model endpoints. Every region listed must also appear in var.aws_bedrock_regions. Default to every var.aws_bedrock_regions region. | list(string) |
null |
no |
| aws_bedrock_marketplace_endpoints_enabled | If true, publish the Amazon Bedrock Marketplace model endpoints deployed in this account and serve them like any other chat model. The module never creates an endpoint: it only grants the server the permissions to discover and invoke the ones you deployed. Disabled by default. A Marketplace model endpoint runs on dedicated instances and is billed by the instance-hour for as long as it exists, whether or not it is called. |
bool |
null |
no |
| aws_bedrock_max_retries | Maximum number of retries for Bedrock invocations. When region routing is enabled, retries cycle through all available regions. Default to 9. | number |
null |
no |
| aws_bedrock_model_arn_mapping | Map standard model IDs to custom inference profile or prompt router ARNs. This allows server administrators to override the default cross-region inference profiles with custom application inference profiles, cross-region inference profiles, or prompt routers. Supported ARN types: - Cross-region inference profile: arn:aws:bedrock:REGION:ACCOUNT:inference-profile/ID - Application inference profile: arn:aws:bedrock:REGION:ACCOUNT:application-inference-profile/ID - Prompt router: arn:aws:bedrock:REGION:ACCOUNT:default-prompt-router/ID Example: { "anthropic.claude-3-5-sonnet-20241022-v2:0" = "arn:aws:bedrock:us-east-1:123456789012:application-inference-profile/my-custom-profile" "anthropic.claude-haiku-4-5-20251001-v1:0" = "arn:aws:bedrock:us-east-1:123456789012:default-prompt-router/my-router" } |
map(string) |
null |
no |
| aws_bedrock_model_region_restrict | Restrict a model to specific region(s) only. Can be used when a model provides important features only in certain regions. Keys are Bedrock model IDs (or prefixes), values are ordered lists of allowed regions. When set, the model will only be available in the listed regions (intersected with the regions where it is actually available). Example: { "amazon.nova-pro-v1:0" = ["us-east-1"] } Use case: Nova grounding is only available in us-east-1, so restricting nova-pro to us-east-1 ensures grounding always works. |
map(list(string)) |
null |
no |
| aws_bedrock_region_routing | Automatic region routing strategy for Bedrock invocations. Distributes requests across configured regions to handle quota limits and regional unavailability. Strategies: 'disabled' (no routing), 'ordered' (try regions in configured order, default), 'lowest_latency' (prefer region with lowest measured latency), 'round_robin' (distribute evenly, incompatible with prompt caching). Requires at least 2 regions in aws_bedrock_regions. | string |
null |
no |
| aws_bedrock_region_routing_max_quota_backoff_seconds | Hard ceiling in seconds on the exponential quota backoff for a single region. Quota backoff doubles on each consecutive error; this value caps how high it can grow. Only effective when aws_bedrock_region_routing is not 'disabled'. Default to 3600 (1 hour). | number |
null |
no |
| aws_bedrock_region_routing_quota_backoff_seconds | Seconds to avoid a region after receiving a quota/throttling error. Only effective when aws_bedrock_region_routing is not 'disabled'. Default to 60. | number |
null |
no |
| aws_bedrock_region_routing_quota_stale_factor | Multiplier applied to the max quota backoff to compute the stale-error threshold. If the most recent quota error for a region is older than (max_quota_backoff * factor) seconds, the consecutive-error counter is reset. Only effective when aws_bedrock_region_routing is not 'disabled'. Default to 2. | number |
null |
no |
| aws_bedrock_region_routing_unavailable_backoff_seconds | Seconds to avoid a region after receiving an unavailability error. Only effective when aws_bedrock_region_routing is not 'disabled'. Default to 30. | number |
null |
no |
| aws_bedrock_regions | List of AWS regions where Bedrock AI models are available. Default to the current region. | list(string) |
null |
no |
| aws_bedrock_session_encryption_key_arn | KMS key ARN encrypting the AWS Bedrock sessions that back stored responses and chat completions (store=true). Default to the AWS-managed key. | string |
null |
no |
| aws_bedrock_user_role_arn | ARN of an existing IAM role the server assumes once per end user, so AWS reports Amazon Bedrock model usage per end user in Cost Explorer and the Cost and Usage Report. The role's trust policy must allow this module's task role to call both 'sts:AssumeRole' and 'sts:TagSession' on it; this module grants the task role those two actions on this ARN. Only model invocations carry the end user session: usage served through Amazon Bedrock Mantle is attributed by project (aws_bedrock_mantle_project), and every other service keeps the task role's identity. When specified, takes precedence over aws_bedrock_user_role_create. Default to none, meaning a role is created when aws_bedrock_user_role_create is true, and all usage is reported under the task role otherwise. | string |
null |
no |
| aws_bedrock_user_role_create | If true, create the IAM role the server assumes once per end user, so AWS reports Amazon Bedrock model usage per end user in Cost Explorer and the Cost and Usage Report. Reporting also requires activating the session tag key named by aws_bedrock_user_role_tag_key as a cost allocation tag of type 'IAM principal', in the Billing console. The created role trusts this module's task role alone, and may only invoke models and built-in tools, apply guardrails, run the built-in web search (reaching the public web only when aws_bedrock_external_web_access is enabled), invoke Marketplace model endpoints through Amazon Bedrock when aws_bedrock_marketplace_endpoints_enabled is set, and read from this deployment's own S3 buckets the media an invocation references by S3 URI instead of carrying inline, which Amazon Bedrock reads under the end user's session. Only model invocations carry the end user session: usage served through Amazon Bedrock Mantle is attributed by project (aws_bedrock_mantle_project), and speech, transcription, translation, video and every other service keeps the task role's identity. Its ARN is returned as the 'bedrock_user_role_arn' output. Only used when aws_bedrock_user_role_arn is not specified. When aws_bedrock_user_role_arn is specified, this value is ignored. Default to false (all usage reported under the task role). | bool |
false |
no |
| aws_bedrock_user_role_require_identity | If true, reject a model request that identifies no end user instead of running it under the server's own identity. Requires aws_bedrock_user_role_create or aws_bedrock_user_role_arn. Default to false. | bool |
null |
no |
| aws_bedrock_user_role_session_duration | Lifetime in seconds of a per-end-user role session, from 900 to 3600. The ceiling is imposed by AWS: a role session obtained from another role session cannot last longer than one hour. Default to 3600. | number |
null |
no |
| aws_bedrock_user_role_tag_key | Session tag key carrying the end user identity on per-end-user role sessions. Activate it as a cost allocation tag of type 'IAM principal' to group Bedrock costs per end user, and test it in IAM policies as 'aws:PrincipalTag/'. Default to 'user'. | string |
null |
no |
| aws_cognito_accept_id_token | If true, accept Amazon Cognito identity tokens in addition to access tokens. Identity tokens describe the signed-in user rather than granting API access, and carry no scopes; enable only for clients that cannot obtain an access token. Default to false. | bool |
null |
no |
| aws_cognito_client_ids | Amazon Cognito user pool application client IDs whose tokens are accepted, as a comma-separated list. A token issued to any other application is rejected. Required when aws_cognito_user_pool_id is specified. | string |
null |
no |
| aws_cognito_issuer_type | Issuer configuration of the Amazon Cognito user pool, which decides the issuer URL its tokens carry: 'original' for 'https://cognito-idp..amazonaws.com/', or 'updated' for 'https://issuer-cognito-idp..amazonaws.com/', available on the Essentials and Plus pool tiers. Tokens whose issuer does not match are rejected, so this must match the pool's own setting. Default to 'original'. | string |
null |
no |
| aws_cognito_required_scopes | OAuth 2.0 scopes a token must all carry to be accepted, as a comma-separated list. Custom scopes exist only on tokens obtained from the user pool's OAuth 2.0 token endpoint, which requires a resource server and a pool domain; tokens obtained by signing in with a username and password carry only 'aws.cognito.signin.user.admin', so requiring a custom scope rejects them. Default to none (any scope set is accepted). | string |
null |
no |
| aws_cognito_user_pool_id | Identifier of the Amazon Cognito user pool whose tokens authenticate clients, for example 'eu-west-3_a1b2c3d4e'. Clients send a pool access token in the 'Authorization: Bearer ' header; its signature, issuer, expiry, application and scopes are validated on every request. The identifier is prefixed by the pool's AWS Region, which is where the signing keys are read from. Requires aws_cognito_client_ids. Default to none (user pool authentication disabled). | string |
null |
no |
| aws_comprehend_region | AWS region for Comprehend language detection service. Default to every var.aws_bedrock_regions region as a failover candidate, or the current region. | string |
null |
no |
| aws_connect_timeout | Timeout in seconds for establishing a connection to an AWS service endpoint. Keeping this value short allows fast failover to another region when a connection cannot be established. Default to 5. | number |
null |
no |
| aws_dynamodb_region | AWS region holding the DynamoDB table. Embedded in the task role's DynamoDB grant, so a table in another region only works when this setting names it. A table this module creates is encrypted with this deployment's own regional KMS key, which cannot encrypt a table in another region, so this must match the deployment's region while the module creates the table; only a table supplied through aws_dynamodb_table is free to live in whatever region this setting names. A DynamoDB table has no built-in cross-region replication, so this setting has no failover, unlike aws_bedrock_regions. Default to the region this module is deployed in. | string |
null |
no |
| aws_dynamodb_table | Existing DynamoDB table name backing the tenant records and the shared models list, in place of the table this module creates automatically. Must carry a string hash key 'pk', a string range key 'sk', and TTL enabled on the 'expires_at' attribute -- see the aws_dynamodb_table.main resource in dynamodb.tf for the schema. When specified, no table is created, and the task role's DynamoDB grant is scoped to this table's ARN, composed from this name and aws_dynamodb_region (or the deployment's own region) rather than read back from a Terraform-managed resource. Default to none, meaning a table is created automatically the first time tenants or model_cache_shared needs one. | string |
null |
no |
| aws_failover_max_retries | Maximum SDK retry attempts per candidate region for the multi-region failover services (Polly, Transcribe, Translate, Comprehend). Only applied when the service has several candidate regions (no dedicated region setting configured). Default to 2. | number |
null |
no |
| aws_max_pool_connections | Maximum number of concurrent HTTP connections per AWS service client. Each AWS service client (per region) maintains its own connection pool up to this limit. Increase if you observe connection pool exhaustion under high concurrency. Default to 50. | number |
null |
no |
| aws_polly_region | AWS region for Polly text-to-speech service. Default to every var.aws_bedrock_regions region as a failover candidate, or the current region. | string |
null |
no |
| aws_s3_accelerate | Enable S3 Transfer Acceleration for presigned URLs. Default to false. | bool |
null |
no |
| aws_s3_accepted_buckets | S3 buckets that the application has read access to, mapped to their region. These buckets can be used as input S3 data sources, and S3 HTTP URLs (including presigned URLs) for these buckets will be automatically converted to S3 URIs for direct access. Keys are bucket names, values are AWS region identifiers. Example: { "my-data-bucket" = "us-east-1", "my-eu-bucket" = "eu-west-1" } If not specified, only the application's own S3 buckets (aws_s3_bucket and aws_s3_regional_buckets) are recognized for S3 URI conversion. |
map(string) |
null |
no |
| aws_s3_accepted_buckets_kms_key_arn | List of KMS key ARNs used to encrypt the accepted S3 buckets (var.aws_s3_accepted_buckets). Required to grant the server permissions to decrypt objects from KMS-encrypted accepted buckets. | list(string) |
null |
no |
| aws_s3_batches_prefix | S3 prefix (folder path) for the Batch API's own data — the submitted requests, the results and the batch records themselves. Each batch stores its data under a folder of its own below this prefix, in the bucket that served it, and the batch service role is granted access to this prefix alone. Must not be the bucket root. Default to 'batches/'. | string |
null |
no |
| aws_s3_bucket | Existing S3 bucket name for storing generated files and application data. When specified, takes precedence over aws_s3_bucket_create. If not specified and aws_s3_bucket_create is true, a bucket will be created automatically. | string |
null |
no |
| aws_s3_bucket_create | If true, create an S3 bucket for the application. Only used when aws_s3_bucket is not specified. When aws_s3_bucket is specified, this value is ignored. | bool |
true |
no |
| aws_s3_buckets_kms_keys_arns | List of KMS key ARNs used to encrypt user-provided regional S3 buckets specified in aws_s3_regional_buckets.Required to grant the server permissions to access encrypted regional buckets. When using aws_s3_regional_buckets_create = true (default), KMS keys are created automatically and do not need to be specified here. |
list(string) |
[] |
no |
| aws_s3_files_prefix | S3 prefix (folder path) for Files API objects. Default to 'files/'. | string |
null |
no |
| aws_s3_regional_buckets | By default (aws_s3_regional_buckets_create = true), buckets are created automatically for every region in aws_bedrock_regions not listed here. Use this variable only to point to existing buckets you manage yourself.Keys are AWS region identifiers, values are bucket names. Example: { "us-east-1" = "my-bucket-us-east-1", "us-west-2" = "my-bucket-us-west-2" } Required for Bedrock operations with multimodal input or document processing. |
map(string) |
null |
no |
| aws_s3_regional_buckets_create | If true (default), create regional S3 buckets and per-region KMS keys for every region in aws_bedrock_regionsnot already present as a key of aws_s3_regional_buckets and not equal to the provider's primary region.Has no effect while aws_bedrock_regions is unset: with no second region there is no region to create a bucket for.Set to false to disable automatic creation (for example, if you manage these buckets out-of-band). |
bool |
true |
no |
| aws_s3_tmp_prefix | S3 prefix (folder path) for temporary files used during job processing. Default to 'tmp/'. | string |
null |
no |
| aws_s3_vector_stores_prefix | S3 prefix (folder path) in the general purpose bucket for the Vector Stores API's own records — the stores, their attached files and their file batches. Default to 'vector_stores/'. | string |
null |
no |
| aws_s3_vectors_bucket | Existing Amazon S3 vector bucket name backing the Vector Stores API. A vector bucket is a distinct resource type from a general purpose bucket. When specified, takes precedence over aws_s3_vectors_bucket_create. Default to none, meaning a bucket is created when aws_s3_vectors_bucket_create is true, and the Vector Stores API is disabled otherwise. | string |
null |
no |
| aws_s3_vectors_bucket_create | If true, create an S3 vector bucket backing the Vector Stores API. Only used when aws_s3_vectors_bucket is not specified. When aws_s3_vectors_bucket is specified, this value is ignored. Default to false (Vector Stores API disabled). | bool |
false |
no |
| aws_s3_vectors_kms_key_arn | KMS key ARN encrypting the vector bucket specified in aws_s3_vectors_bucket. Required to grant the server permission to use an SSE-KMS encrypted vector bucket. When using aws_s3_vectors_bucket_create = true, a key is created automatically and does not need to be specified here. | string |
null |
no |
| aws_s3_vectors_region | AWS region holding the vector bucket. A vector bucket is a regional resource and its indexes are only reachable in that region, so this setting has no failover. Default to the region this module is deployed in. Amazon S3 Vectors is not available in every region: see AWS Regions, endpoints, and quotas for S3 Vectors. |
string |
null |
no |
| aws_s3_videos_expires_after | Retention period in seconds for generated videos. When set, Video.expires_at is reported, expired downloads return 404, and a matching S3 Lifecycle expiration rule is created on the module-managed buckets. Default to no expiry. | number |
null |
no |
| aws_s3_videos_prefix | S3 prefix (folder path) for videos generated through the Videos API. Default to 'videos/'. | string |
null |
no |
| aws_sagemaker_endpoint_url | Override for the Amazon SageMaker AI runtime endpoint URL, with '{region}' substituted per region. Only needed to reach the runtime through a VPC endpoint or an inspection proxy; the correct host for every AWS partition is resolved automatically otherwise. | string |
null |
no |
| aws_sagemaker_endpoints | Amazon SageMaker AI endpoints published as chat models, keyed by the model ID clients name them with. The module never creates an endpoint: it only grants the server the permissions to invoke the ones you name. The endpoint's container must serve the OpenAI Chat Completions API, which the SageMaker AI vLLM and SGLang containers do. An endpoint is billed by the instance-hour rather than by the token, so the tokens it serves are reported without a cost; one configured to scale to zero costs nothing while idle, at the price of a slow first request. Fields: "endpoint" (name, never the ARN) and "region" are required; "inference_component" is required for a component-hosted endpoint, which a scale-to-zero one always is; "name", "provider" and "input_modalities" are optional catalogue metadata. Example: { "my-qwen3": { endpoint = "my-endpoint" region = "us-east-1" inference_component = "my-inference-component" name = "Qwen3 1.7B" provider = "Qwen" } } |
map(object({ |
null |
no |
| aws_sagemaker_warmup_timeout | Seconds a request may wait for an Amazon SageMaker AI endpoint that has scaled to zero to provision capacity again. The request that finds the endpoint cold is what makes SageMaker AI scale it back up, so the server holds the connection and retries instead of failing, and concurrent callers share one wait. Set to 0 to disable the wait and fail immediately. Cannot exceed var.ai_response_timeout. Default to 600. |
number |
null |
no |
| aws_sqs_vector_store_queue_create | If true, create the Amazon SQS queue, and its dead-letter queue, that make vector store indexing durable: a file attached to a store keeps being indexed by another task when the task that accepted it stops, instead of being reported as failed. Only used when the Vector Stores API is enabled and aws_sqs_vector_store_queue_url is not specified. When aws_sqs_vector_store_queue_url is specified, this value is ignored. Default to true. | bool |
true |
no |
| aws_sqs_vector_store_queue_kms_key_arn | KMS key ARN encrypting the queue specified in aws_sqs_vector_store_queue_url. Required to grant the server permission to use an SSE-KMS encrypted queue; leave unset for a queue encrypted with the Amazon SQS managed key. When using aws_sqs_vector_store_queue_create = true, the deployment key is used automatically and does not need to be specified here. | string |
null |
no |
| aws_sqs_vector_store_queue_url | URL of an existing Amazon SQS queue carrying the vector store indexing jobs. Must be a standard queue, never FIFO, and must have a dead-letter queue: a file the server cannot index is settled as failed once its deliveries run out, and its message is kept in the dead-letter queue. The queue is reached in the single region its URL names, and only ever carries identifiers, never file content. Requires the Vector Stores API to be enabled: the server refuses to start with a queue and nothing to index. When specified, takes precedence over aws_sqs_vector_store_queue_create. Default to none, meaning a queue is created when aws_sqs_vector_store_queue_create is true and the Vector Stores API is enabled, and indexing runs in the task that accepted the request otherwise. | string |
null |
no |
| aws_transcribe_output_encryption_key_arn | KMS key ARN encrypting the transcription job output written to aws_transcribe_s3_bucket. The key must be usable from every region a transcription job can be served from; this module grants the task role kms:GenerateDataKey and kms:Decrypt on it, and the key policy must allow the same actions. Default to the bucket's own default encryption. | string |
null |
no |
| aws_transcribe_region | AWS region for Transcribe speech-to-text service. Default to every var.aws_bedrock_regions region as a failover candidate, or the current region. | string |
null |
no |
| aws_transcribe_s3_bucket | AWS S3 bucket name for temporary file storage during transcription. Defaults to aws_s3_bucket if not specified. | string |
null |
no |
| aws_transcribe_stream_languages | Languages a streamed transcription (stream=true) picks between when the request names none, as two or more language codes, for example ['en-US', 'es-US', 'fr-FR']. A streamed transcription starts before the recording has been fully read, which requires knowing which languages to expect. Default to none, meaning a request naming no language is transcribed once the whole recording has been read, and its language detected. | list(string) |
null |
no |
| aws_translate_region | AWS region for Translate text translation service. Default to every var.aws_bedrock_regions region as a failover candidate, or the current region. | string |
null |
no |
| chat_completions_reasoning_field | Field carrying a reasoning model's thinking text on '/v1/chat/completions', which the OpenAI API itself does not return, so vendors differ: 'reasoning_content' is the DeepSeek spelling most clients read, 'reasoning' is the one OpenRouter and vLLM use, and 'none' emits neither and keeps responses strictly OpenAI-shaped. Default to 'reasoning_content'. | string |
null |
no |
| cloudwatch_logs_retention_in_days | Cloudwatch logs retention in days. Applies to every log group this module and its child modules create, including the Container Insights performance log group. Security Hub: CloudWatch.16 (CloudWatch log groups should be retained for a specified time period) requires at least 365 days by default — default 365 = pass; lowering it fails this control. | number |
365 |
no |
| cloudwatch_metrics | If True, emit per-request AWS-billed usage as CloudWatch Embedded Metric Format (EMF) log lines. Left unset, usage_api turns it on, since the usage API answers from these metrics. Default to false. | bool |
null |
no |
| cloudwatch_metrics_namespace | CloudWatch namespace for the emitted usage metrics. Default to 'stdapi'. | string |
null |
no |
| cloudwatch_metrics_region | AWS region the usage API reads the published metrics from. Defaults to the region the server runs in, which is where its logs are ingested and therefore where the metrics exist. Only needs setting when the logs are shipped to another region. | string |
null |
no |
| cloudwatch_metrics_user_dimension | If True, also publish the authenticated caller as a 'User' dimension on the usage metrics, which is what lets the usage API group by user_id. Stores one CloudWatch metric series per user, model and metric name, each billed as a custom metric, so its cost follows the size of your user population. Requires cloudwatch_metrics and usage_api. Default to false. | bool |
null |
no |
| cohere_routes_prefix | Cohere API compatible routes prefix. Default to '/cohere'. | string |
null |
no |
| compliance_vpc_endpoints_enabled | If true, add the interface VPC endpoints for ECR API, ECR Docker Registry, Systems Manager, SSM Incident Manager Contacts and SSM Incident Manager. Enable only if you have high compliance requirements — each interface endpoint adds cost. Security Hub: EC2.55/EC2.56/EC2.57/EC2.58/EC2.60 — default false = fail; set to true to pass. | bool |
false |
no |
| container_insight | Container insight configuration. Valid values: 'enhanced', 'enabled', 'disabled'. Default to 'enabled'. Security Hub: ECS.12 (ECS clusters should use Container Insights) — default 'enabled' = pass; setting 'disabled' fails this control. | string |
"enabled" |
no |
| cors_allow_origins | List of origins allowed to make cross-origin requests (CORS). Use ['*'] to allow all origins. Default to no CORS headers. | list(string) |
null |
no |
| cost_price_overrides | Unit price overrides for models not covered by the AWS Price List API, as a map of model IDs to dimension-name/price maps. | map(map(number)) |
null |
no |
| cost_tracking | Enable per-request cost estimation from AWS Price List values (adds the pricing:GetProducts permission). Reported costs are an estimate from published prices, not your actual AWS bill; use cost_price_overrides for models the Price List API does not cover. Left unset, usage_api turns it on, since the costs endpoint reports nothing without it. Default to false. | bool |
null |
no |
| cpu | ECS task CPU count. Valid values: 0.25, 0.5, 1, 2, 4, 8 & 16. Paired with var.memory: each CPU value constrains which memory values Fargate accepts, see the ECS documentation. Default of 0.25 vCPU is suitable for common use cases (text generation, embeddings). Increase for intensive workloads (multimodal requests, large LLM models). | number |
0.25 |
no |
| cpu_architecture | CPU architecture. Valid values: 'X86_64' or 'ARM64'. | string |
"ARM64" |
no |
| default_model_params | Default inference parameters applied to specific models automatically. JSON string format. | string |
null |
no |
| default_model_service_tiers | Default service tier applied to specific models automatically when no explicit tier is provided (default, flex, priority, reserved). JSON string format, e.g. {"amazon.nova-pro-v1:0": "flex"}. | string |
null |
no |
| default_tts_language | Default text-to-speech language to use if not specified in the request. Default to language autodetection. | string |
null |
no |
| default_tts_model | Default text-to-speech model to use if not specified in the request. Default to 'amazon.polly-standard'. | string |
null |
no |
| deletion_protection | If true, enable deletion protection on eligible resources. One resource ignores it: the shared DynamoDB table derives its own from what it holds. With tenants declared it holds the tenant secret hashes and salts, which exist nowhere else, so it is always protected regardless of this setting and destroying the deployment takes two applies: first one with tenants emptied and model_cache_shared = true, which deletes the tenant records and clears the protection while keeping the table, then the destroy itself. Emptying tenants on its own would drop the table in that same apply, which AWS refuses while the protection is still on. Holding only the shared models list (model_cache_shared) it is a pure cache the next discovery sweep rebuilds, so it is never protected and destroys cleanly. | bool |
false |
no |
| dns_firewall_action | Action taken by DNS Firewall when a query matches a domain from var.dns_firewall_managed_domain_list_ids, and (if var.dns_firewall_advanced_enabled) a DNS Firewall Advanced threat detection. Valid values: 'ALLOW', 'BLOCK', 'ALERT'. 'ALLOW' isn't valid for DNS Firewall Advanced rules, so it's treated as 'BLOCK' for those only. Ignored if var.dns_firewall_enabled is false. | string |
"BLOCK" |
no |
| dns_firewall_advanced_confidence_threshold | Confidence threshold for DNS Firewall Advanced rules. Valid values: 'LOW', 'MEDIUM', 'HIGH'. Lower thresholds catch more threats at the cost of more false positives. Ignored if var.dns_firewall_advanced_enabled is false. | string |
"HIGH" |
no |
| dns_firewall_advanced_enabled | If true, add Route 53 Resolver DNS Firewall Advanced rules (additional cost) blocking DNS queries identified as domain generation algorithm (DGA) or DNS tunneling activity, on top of any managed-domain-list rules. Ignored if var.dns_firewall_enabled is false. | bool |
false |
no |
| dns_firewall_enabled | If true, create a Route 53 Resolver DNS Firewall rule group and associate it with the dedicated VPC, blocking/alerting on DNS queries per var.dns_firewall_managed_domain_list_ids and var.dns_firewall_advanced_enabled. Helps mitigate malicious-URL injection via user-supplied URL/file references (images, documents, audio) by blocking outbound DNS resolution to known-malicious domains, in addition to the application's own SSRF protection. Only supported for the dedicated VPC this module creates; cannot be enabled when using external subnets (var.subnet_ids). Not mapped to a Security Hub control; default false = feature not created. | bool |
false |
no |
| dns_firewall_managed_domain_list_ids | Map of AWS Managed Domain List name to ID (e.g. { "AWSManagedDomainsAggregateThreatList" = "rslvr-fdl-..." }) to block/alert on via var.dns_firewall_action. Defaults (null) to the Aggregate Threat List ID built into the underlying VPC module for the current region, covering commercial regions enabled by default — no AWS CLI call or extra permissions required. For a region not covered by that default, look up the ID with 'aws route53resolver list-firewall-domain-lists' and pass it explicitly. Ignored if var.dns_firewall_enabled is false. Set to {} to skip managed-list rules while still using var.dns_firewall_advanced_enabled. | map(string) |
null |
no |
| dns_firewall_priority | Processing priority for the DNS Firewall rule group association within the VPC (lower is processed first). Must be unique among all rule group associations on the same VPC, including ones created outside this module. Ignored if var.dns_firewall_enabled is false. | number |
101 |
no |
| drop_unsupported_system_prompt | If true, system prompts are silently dropped when models don't support them. If false, an error is returned when a system prompt is passed to a model that doesn't support system prompts (e.g., mistral.mistral-7b models). Default: true for backward compatibility. | bool |
null |
no |
| ecs_task_role_policy_arns | List of IAM policy ARNs to attach to the ECS task role. Use this to grant additional permissions to the ECS task, such as access to SSM parameters or Secrets Manager secrets specified in api_key_ssm_parameter or api_key_secretsmanager_secret. | list(string) |
[] |
no |
| enable_docs | Enable interactive API documentation UI at /docs. Default to false. | bool |
null |
no |
| enable_gzip | Enable GZip compression middleware for HTTP responses. Disabled by default. | bool |
null |
no |
| enable_mcp_sse | Enable the MCP (Model Context Protocol) server using Server-Sent Events (SSE) transport. When enabled, exposes MCP endpoints at /sse. Maintained for backwards compatibility with older MCP clients; prefer enable_mcp_streamable_http for new deployments. Default to false. | bool |
null |
no |
| enable_mcp_streamable_http | Enable the MCP (Model Context Protocol) server using Streamable HTTP transport. When enabled, exposes an MCP-compatible endpoint at /mcp. This is the recommended MCP transport. Default to false. | bool |
null |
no |
| enable_openapi_json | Enable OpenAPI JSON schema endpoint at /openapi.json. Default to false. | bool |
null |
no |
| enable_proxy_headers | Enable ProxyHeadersMiddleware to trust X-Forwarded-* headers from reverse proxies. Automatically enabled when var.alb_enabled is true and var.log_client_ip is true. | bool |
null |
no |
| enable_redoc | Enable ReDoc API documentation UI at /redoc. Default to false. | bool |
null |
no |
| extra_model_params_denylist | Additional parameter names to strip from the 'extra model parameters' passthrough, as a comma-separated list. Merged with the built-in default denylist of client-control parameters (such as 'drop_params', 'api_key' or 'custom_llm_provider') that some OpenAI-SDK-based clients leak into extra_body and that are never legitimate Bedrock model parameters. Only effective when extra_model_params_drop_all is false. Default to the built-in denylist alone. Example: 'x_internal_debug_flag,x_proxy_trace_id' | string |
null |
no |
| extra_model_params_drop_all | If true, disable the 'extra model parameters' passthrough entirely: no undeclared request field is ever forwarded to Amazon Bedrock as a provider-specific inference parameter, on any route that supports it. Overrides extra_model_params_denylist, which no longer matters once nothing is forwarded. Default to false, which keeps the passthrough, filtered by the built-in default denylist and extra_model_params_denylist. | bool |
null |
no |
| guardduty_vpc_endpoint_enabled | If true, add the interface VPC endpoint required by GuardDuty Runtime Monitoring. Only relevant if you use GuardDuty Runtime Monitoring on resources in this VPC — leave false otherwise. Recommended whenever Runtime Monitoring is enabled, even with GuardDuty's automated agent configuration, since managing it here ensures correct subnet placement. Not mapped to a Security Hub control; default false = endpoint not created. | bool |
false |
no |
| image_generation_model | Default model ID for image generation (e.g. 'amazon.nova-canvas-v1:0'). Required unless the client or the LLM specifies a model per call. | string |
null |
no |
| kms_key_id | If specified, directly use this KMS key instead of creating a dedicated one for the application. | string |
null |
no |
| log_client_ip | If True, log the client IP address for each request and add it to OpenTelemetry spans. Default to false. | bool |
null |
no |
| log_level | Minimum logging level to output: info, warning, error, critical, or disabled. Default to info. | string |
null |
no |
| log_request_params | If True, add requests and responses parameters to logs. Should not be enabled in production. Default to false. | bool |
null |
no |
| max_concurrent_input_downloads | Maximum number of input files fetched or resolved concurrently within a single request, bounding outbound downloads against socket/memory exhaustion and SSRF amplification. Default to 8. | number |
null |
no |
| max_input_file_size | Maximum size in bytes of an inline input file loaded into memory (base64, data URI, or a downloaded HTTP(S)/S3 source). Requests exceeding it are rejected with HTTP 413 before the content is fully decoded. Default to 0 (no limit). | number |
null |
no |
| mcp_exclude_tools | Comma-separated list of MCP tool names to hide from MCP clients. All other tools remain exposed. When mcp_include_tools is also specified, these values are removed from it. Example: 'openai_files_delete,anthropic_files_delete' | string |
null |
no |
| mcp_include_tools | Comma-separated list of MCP tool names to expose exclusively. Only the listed tools will be available to MCP clients; all others are hidden. When both mcp_include_tools and mcp_exclude_tools are specified, mcp_exclude_tools values are removed from mcp_include_tools. Example: 'openai_chat_completion,openai_embedding,openai_model_list' | string |
null |
no |
| mcp_stateless_http | Serve the MCP Streamable HTTP transport in stateless mode. Each request is then handled by a fresh transport that keeps no session state, so any client may call /mcp without initializing a session first and any task may serve any request. Required by hosts that provide their own session isolation and inject an 'Mcp-Session-Id' header the server never issued. Left unset, enable_mcp_streamable_http is turned on, since this is a mode of that transport. Default to false. | bool |
null |
no |
| memory | ECS task memory (MiB). Valid values depend on the var.cpu value, see the ECS documentation. The default of 512 MiB covers text generation and embeddings, where the task holds little more than the request in flight. It is not enough for the paths that hold bytes in memory: audio and video through the ffmpeg pipeline, inline input files up to max_input_file_size, and max_concurrent_input_downloads of them fetched at once. Raise it to 1024 or beyond before using those, or the task is OOM-killed under load rather than answering slowly — which surfaces to callers as a 502 or 504 from the load balancer, not as an error from stdapi.ai. | number |
512 |
no |
| model_aliases | Map of model aliases to actual model IDs. Allows users to reference models using custom alias names. This is merged with default system aliases at startup. User-provided aliases take precedence over system defaults. An alias maps either to a model ID, or to an object carrying that model plus the configuration to apply to requests naming the alias: "service_tier", "guardrail_id" with "guardrail_version" (and optionally "guardrail_trace"), "metadata" and "extra_params". Those values override the equivalent server-wide configuration, and a value sent with the request still wins unless its override variable (aws_bedrock_allow_guardrail_override, aws_bedrock_allow_service_tier_override) is disabled. Example: { "my-tts": "amazon.polly-neural", "my-stt": "amazon.transcribe", "my-chat": { "model": "amazon.nova-lite-v1:0", "service_tier": "flex", "metadata": { "team": "research" }, "extra_params": { "temperature": 0.2 } } } |
any |
null |
no |
| model_cache_max_stale_seconds | Maximum age in seconds the cached Bedrock models list may reach while its refresh keeps failing. Below it an expired list is served while the refresh runs; at or beyond it the next request waits for a successful refresh instead, so a deployment whose refreshes fail silently cannot serve an arbitrarily old list. Set to 0 to always wait for a refresh once the list has expired. Default to the application default (86400, 24 hours). | number |
null |
no |
| model_cache_seconds | Age in seconds at which the cached Bedrock models list is refreshed. Once reached, the next request needing the list is answered from the cached one and the refresh runs in the background, so no request waits for it. Default to the application default (900, 15 minutes). | number |
null |
no |
| model_cache_shared | If true, share one Bedrock models list between the deployment's servers through the shared DynamoDB table, which this module then creates for it. One server refreshes the list and publishes it while the others read it, so a fleet performs one discovery pass per model_cache_seconds instead of one per server and a starting server is ready without a discovery pass of its own. Billed on the published list and on the table it lives in, roughly $0.60 to $2 per month at the default interval. Default to the application default (false). | bool |
null |
no |
| name_prefix | Prefix to add to all created resources names. Every created name is "<name_prefix>-<8 hexadecimal characters>-", so a long prefix reaches the length limit of the name it is longest in: 22 characters is the most any region accepts, and 13 when alb_enabled is true. Both assume the shortest region name AWS publishes, nine characters, and a longer one lowers them at different rates: the 22 by two characters per extra character (the bound is 40 minus twice the region name length, because the VPC flow log role name repeats the region), the 13 by one (22 minus the region name length). | string |
"stdapiai" |
no |
| nat_gateways_allowed | If true, NAT gateways give the application its internet access. One NAT gateway is created per availability zone and each is billed hourly whether or not traffic flows, which makes this the module's largest fixed monthly cost; set availability_zones_count to bound it. If false and internet access is required, the application subnets are public instead: no hourly charge, and the tasks are addressed directly. Default to true, except with realtime_webrtc_media_enabled, which needs the task publicly addressed and turns it off for you. | bool |
null |
no |
| oauth_authorization_servers | Issuer URLs of the OAuth 2.0 authorization servers that issue tokens for the API, as a comma-separated list, published in the protected resource metadata. Leave it unset when aws_cognito_user_pool_id is specified: the pool's own issuer is published, resolved for the partition the pool lives in. Setting it explicitly is for a deployment that accepts tokens from another authorization server, and the list must still name the configured pool's issuer. Required only when no user pool is configured and oauth_resource_identifier is. | string |
null |
no |
| oauth_resource_identifier | Public URL clients use to reach the API, for example 'https://api.example.com', which is normally 'https://' followed by alb_domain_name. Setting it publishes an OAuth 2.0 protected resource metadata document at '/.well-known/oauth-protected-resource' and puts that address in the challenge every 401 response carries, so an AI agent can discover where to obtain a token. Must be the exact origin clients dial: scheme and host, no path, no trailing slash. Requires either aws_cognito_user_pool_id, whose issuer is then published, or an explicit oauth_authorization_servers. If not specified, nothing is published. | string |
null |
no |
| oauth_scopes_supported | OAuth 2.0 scopes a token needs to call the API, as a comma-separated list, advertised in the protected resource metadata and in the 401 challenge. Requires oauth_resource_identifier. If not specified, aws_cognito_required_scopes is advertised, so the scopes a token needs are declared in one place. | string |
null |
no |
| ollama_routes_prefix | Ollama API compatible routes prefix. Default to the root, where Ollama clients expect '/api/*'. | string |
null |
no |
| openai_routes_prefix | OpenAI API compatible routes prefix. Default to the root, where OpenAI clients expect '/v1/*'. | string |
null |
no |
| otel_enabled | Enable OpenTelemetry distributed tracing. Default to false. | bool |
null |
no |
| otel_exporter_endpoint | OpenTelemetry traces export endpoint URL. | string |
null |
no |
| otel_sample_rate | OpenTelemetry trace sampling rate (0.0 to 1.0). | number |
null |
no |
| otel_service_name | Service name identifier for OpenTelemetry traces. Default to 'stdapi.ai'. | string |
null |
no |
| proxy_trusted_hosts | Trusted proxy hosts/IPs (CIDRs) whose X-Forwarded-* headers are honored when proxy headers are enabled. Restrict to your reverse proxy's IP range so direct clients cannot forge their source IP. Write entries in their natural address family: on an IPv6-enabled VPC the server binds a dual-stack socket and sees IPv4 peers in IPv4-mapped form, and the module adds the matching '::ffff:' range for every IPv4 entry automatically. When null and proxy headers are auto-enabled (var.alb_enabled and var.log_client_ip both true), defaults to the ALB subnet CIDRs so only the ALB is trusted; otherwise the server default ('*') applies. | list(string) |
null |
no |
| realtime_allow_session_override | Allow a client connecting to the Realtime API with an ephemeral client secret to override the session configuration that secret carries. When disabled, the model, the instructions and the output token cap minted into the secret are final: a 'model' query parameter naming another model is refused, and a session.update changing one of them answers an error. Default to true, which is the upstream behavior; disable it in a multi-tenant deployment where the secret is the only thing constraining an untrusted client. | bool |
null |
no |
| realtime_client_secret_key | Secret the ephemeral client secrets of the Realtime API are signed with. Any value works as long as every task of the deployment shares it: a secret minted by one task is verified by whichever one the client's WebSocket reaches. Passed to the container as an ECS secret stored in AWS Systems Manager Parameter Store, never as a plain environment variable. When not specified, the server derives the key from the configured API key. When no API key is configured either (api_key, api_key_create, api_key_ssm_parameter and api_key_secretsmanager_secret all unset), the module generates a key and stores it the same way, because the server would otherwise fall back to a per-process random value, under which a client secret minted by one task is rejected by every other task and by any task replacing it after a deployment. |
string |
null |
no |
| realtime_webrtc_allow_private_candidates | Accept the ICE candidates a WebRTC caller offers on addresses that are not globally routable: private (RFC 1918), shared (RFC 6598), loopback and link-local ones. They are dropped by default and an offer left with no candidate at all is refused, which is what keeps a caller holding nothing but an ephemeral client secret from aiming the task's UDP connectivity checks at addresses inside the deployment's own VPC. Enable it only where the callers legitimately share that network, such as a same-VPC or on-premises deployment -- including one where realtime_webrtc_enabled reports WebRTC available with realtime_webrtc_media_enabled off, because every caller already reaches the task on its private address. Hostname and mDNS ('.local') candidates are dropped either way: resolving one is itself a lookup on the deployment's network. Requires realtime_webrtc_enabled, which follows realtime_webrtc_media_enabled unless overridden. Default to the application default (false). | bool |
null |
no |
| realtime_webrtc_enabled | If set, overrides whether the Realtime API reports WebRTC as available to callers, in place of following realtime_webrtc_media_enabled. The infrastructure the media path needs -- the public task, the security group and network ACL rules, the pinned autoscaling capacity -- stays governed by realtime_webrtc_media_enabled alone, so setting this to true without it advertises WebRTC with no public address and no opened UDP ports: only callers already inside the VPC, reaching the task on its private address with realtime_webrtc_allow_private_candidates, can then connect. Default to none, meaning realtime_webrtc_media_enabled alone decides. | bool |
null |
no |
| realtime_webrtc_ingress_ipv4_cidrs | IPv4 CIDR blocks allowed to send WebRTC call media to the task. Defaults to the whole internet, which is what serving arbitrary browsers requires; narrow it to your callers' networks whenever they are known. | list(string) |
[ |
no |
| realtime_webrtc_ingress_ipv6_cidrs | IPv6 CIDR blocks allowed to send WebRTC call media to the task, applied when the VPC has IPv6 enabled. Defaults to the whole internet; narrow it to your callers' networks whenever they are known. | list(string) |
[ |
no |
| realtime_webrtc_media_enabled | If true, enables the Realtime API's WebRTC transport (POST /v1/realtime/calls) with the media path terminated by the server task: the task gets a public IP, its security group opens the WebRTC UDP media port range to realtime_webrtc_ingress_ipv4_cidrs/ipv6_cidrs inbound and the callers' own ephemeral range outbound, plus the ports of the STUN and TURN servers the task reaches out to, the application subnets' network ACL carries those same flows, and the server is configured with the STUN/TURN settings below. Requires nat_gateways_allowed = false (the task must be publicly addressed for inbound media) and a single task (autoscaling_min_capacity = autoscaling_max_capacity = 1), because calls live in the answering task's memory and the media path cannot drain. On operator-supplied subnet_ids the network ACLs are yours: allow the media UDP port range inbound and the ephemeral range (1024-65535) in both directions there, since the module cannot write rules in a VPC it did not create. Off by default: this opens inbound UDP from the internet and pins the service to one task. | bool |
false |
no |
| realtime_webrtc_stun_server | STUN server the server queries to discover the public address it advertises to WebRTC callers, as a STUN URI. Required behind the 1:1 NAT of a public ECS task; any public STUN server works and learns nothing but the deployment's public address. | string |
"stun:stun.l.google.com:19302" |
no |
| realtime_webrtc_turn_password | Long-term credential password of realtime_webrtc_turn_server, delivered to the task as an ECS secret backed by an encrypted SSM parameter. | string |
null |
no |
| realtime_webrtc_turn_server | Operator-run TURN relay advertised to WebRTC callers whose networks block UDP, as a TURN URI (e.g. 'turn:turn.example.com:3478?transport=udp'). AWS offers no managed TURN; requires realtime_webrtc_turn_username and realtime_webrtc_turn_password. | string |
null |
no |
| realtime_webrtc_turn_username | Long-term credential username of realtime_webrtc_turn_server. | string |
null |
no |
| realtime_webrtc_udp_port_range | UDP port range opened for WebRTC call media. Media sockets are bound by the OS from its ephemeral range, so the opened range must cover it; the default matches the Linux kernel default (net.ipv4.ip_local_port_range). | object({ |
{ |
no |
| security_group_id | If specified and 'subnet_ids' is specified, use this security group instead of creating a new one giving access to internet and AWS services. | string |
null |
no |
| service_discovery_dns_name | DNS name for service discovery. By default, uses the service name. Only if service_discovery_dns_namespace_id is specified. | string |
null |
no |
| service_discovery_dns_namespace_id | If specified, enable Service discovery on the ECS service and attach it to this Cloud Map namespace. | string |
null |
no |
| shutdown_drain_timeout | Maximum time in seconds the server waits, once asked to stop, for background work that requests started and did not wait for: temporary file cleanups, vector store file indexing, and the release of live audio sessions. Work still running when the wait ends is cancelled and counted as a warning in the server's stop log event. This wait is best effort, not a delivery guarantee: a container runtime sends SIGKILL a fixed delay after the stop signal. This module raises the task's own stop timeout to match, so the wait is not cut short here; a deployment that does not use this module must raise it itself, because the default on Amazon ECS is 30 seconds. Values above 110 are capped, since Fargate accepts at most 120. Set to 0 to stop as fast as possible, cancelling background work immediately. Default to 10. | number |
null |
no |
| sns_topic_arn | SNS topic ARN notified by the CloudWatch alarms. Setting it is what turns alarms on (see alarms_enabled): the ECS service alarms of the underlying module -- high memory, unhealthy containers, CPU anomaly, autoscaling at its ceiling -- plus one alarm created here on ERROR and CRITICAL log lines. Default to none. | string |
null |
no |
| ssrf_protection_block_private_networks | Enable SSRF protection by blocking requests to private/local networks. When enabled, the server will reject requests to RFC 1918 private addresses (10.0.0.0/8, 172.16.0.0/12, 192.168.0.0/16), loopback, link-local, reserved, and multicast addresses. Default to true. | bool |
null |
no |
| strict_input_validation | If True, raise error on extra fields in input request. Default to false. | bool |
null |
no |
| subnet_ids | If specified, directly use theses subnets instead of creating a dedicated VPC. | list(string) |
[] |
no |
| tags | Tags applied to every resource this module creates, on top of the tags it sets itself. Use these for cost allocation, ownership or environment. The 'aws-apn-id' tag is reserved: it attributes the deployment to the AWS Marketplace product and is merged last, so an entry of that key here is ignored. | map(string) |
{} |
no |
| tenant_api_keys | If set, overrides whether the server validates tenant API keys, in place of deriving it from tenants being non-empty. Set to true to validate against a shared DynamoDB table populated outside this module, for example one named in aws_dynamodb_table; Terraform still only writes DynamoDB table items, and grants the sts:AssumeRole permission, for tenants actually declared here. It cannot be set to false while tenants declares any. Default to none, meaning tenants alone decides. | bool |
null |
no |
| tenant_aws_credentials | If set, overrides whether the server reports tenant AWS credentials as enabled, in place of deriving it from any tenants entry declaring aws_role_arn. The sts:AssumeRole grant stays scoped to exactly the roles tenants actually declares: forcing this to true with none declared reports the feature as enabled with no role for it to assume. Default to none, meaning tenants alone decides. | bool |
null |
no |
| tenant_key_cache_seconds | Seconds each server instance caches a tenant API key validation before re-reading the shared table. This is the revocation window: a key revoked, disabled or re-scoped in 'tenants' keeps its previous decision for up to this long per instance, traded against the table reads a shorter window costs. 0 reads the table on every request. Only applied while tenant API keys are enabled. Default to 60. | number |
null |
no |
| tenant_key_rotation_days | Rotate every tenant API key once it is this many days old, counted from its mint or its last rotation. Setting it stores the tenant keys in AWS Secrets Manager -- one secret per tenant, named in the tenant_keys output, encrypted with this deployment's own KMS key -- instead of delivering each key once through SSM Parameter Store: a rotated key becomes its secret's current version (AWSCURRENT), the superseded one stays readable as AWSPREVIOUS and keeps working for tenant_key_rotation_overlap_seconds, and a tenant granted secretsmanager:GetSecretValue on its own secret and kms:Decrypt on this deployment's key through Secrets Manager -- a principal of this deployment's own account, since the module writes no resource policy on the secret -- re-reads its key without an operator in the loop. 90 or less keeps the secrets within the periodic-rotation window AWS Security Hub checks. Unsetting it again destroys every tenant's secret and reverts delivery to one-shot Parameter Store, unless a tenants entry still declares key_generation: the store is selected by these two triggers alone, and removing the last one takes it down with every key stored in it. Only applied while tenant API keys are enabled. Default to none: keys are delivered once through Parameter Store and only rotated on demand, through a tenants entry's key_generation. | number |
null |
no |
| tenant_key_rotation_overlap_seconds | Seconds a rotated tenant API key keeps working after its replacement was stored, so a client that has not re-read its secret yet is not locked out; 0 refuses the superseded key as soon as the new one is stored. Only the last superseded key is kept, so the grace ends at the next rotation whatever this is set to, and it must be shorter than tenant_key_rotation_days. Setting disabled on its tenants entry refuses both keys within tenant_key_cache_seconds, whatever this value. Only applied while tenant keys are stored in AWS Secrets Manager, that is with tenant_key_rotation_days set or a tenants entry declaring key_generation. Default to the server's own default of 604800 (7 days). | number |
null |
no |
| tenant_key_ssm_kms_key_id | ARN of the KMS key encrypting the SSM parameter tenant API keys are delivered through, in place of this deployment's own key. Only applied while tenant API keys are enabled and delivered through Parameter Store, that is without tenant_key_rotation_days or a tenants entry declaring key_generation; the task role's KMS delivery grant is scoped to it, and the key's own policy must additionally allow the task role kms:Encrypt, kms:GenerateDataKey and kms:Decrypt. Must be a key ARN, not a key id or an alias: an IAM policy resource takes nothing else. Default to this deployment's own KMS key. | string |
null |
no |
| tenant_key_ssm_parameter_prefix | SSM Parameter Store path prefix the server delivers each tenant's minted API key under, in place of the prefix this module derives from its own name. Only applied while tenant API keys are enabled and delivered through Parameter Store, that is without tenant_key_rotation_days or a tenants entry declaring key_generation; the task role's SSM and KMS delivery grants are scoped to it, so a custom prefix is never wider than the derived one. Default to '//tenant-keys'. | string |
null |
no |
| tenant_rate_limit_requests_per_minute | Requests each tenant API key may make per minute, unless its tenants entry declares its own requests_per_minute. Minutes are fixed windows shared by every task through the DynamoDB table, and a request over the limit answers 429 with a retry-after naming the seconds left in the minute -- on a Realtime connection, an error event after the WebSocket upgrade instead, since the upgrade succeeds before the limit is checked. The value is a ceiling, not an exact allowance: each task reserves request slots ahead in batches of up to an eighth of the limit and never returns before the minute ends what it does not use, so a tenant spread over many tasks can be refused somewhat below its limit -- size the limit with headroom for the number of tasks its traffic reaches. Only tenant keys are limited; the deployment API key and Amazon Cognito tokens are not. Only applied while tenant API keys are enabled, and the dynamodb:UpdateItem grant the counters need follows the limits declared through this module: a limit written into the table by other tooling needs that grant widened by hand, or one of these defaults set. Default to none: no request limit, except for the tenants entries that declare one. | number |
null |
no |
| tenant_rate_limit_tokens_per_minute | Tokens each tenant API key may bill per minute -- input, cache-write and output tokens; cached reads are free, and a batch job's tokens are never counted -- unless its tenants entry declares its own tokens_per_minute. A request is admitted on an estimate and reconciled from what the model actually billed, so a burst of requests larger than the estimate overshoots the limit and the next ones answer 429 until the minute ends -- on a Realtime connection, an error event after the WebSocket upgrade instead. The estimate is learned per task: until a request of the key has billed on a task, each request the key holds in flight there counts as an eighth of the limit, so a freshly started task -- every one, after a deployment or a scale-out -- admits about eight concurrent requests of that key whatever their real size and refuses the next one, and afterwards each counts as the key's mean tokens per billed request on that task. A token limit alone also caps the requests a key holds in flight at 64 per task; declaring tenant_rate_limit_requests_per_minute alongside it replaces that ceiling with the request limit. Only applied while tenant API keys are enabled. Default to none: no token limit, except for the tenants entries that declare one. | number |
null |
no |
| tenants | Per-tenant API keys, one entry per tenant keyed by the tenant's name. Terraform owns each tenant's record — its identity and every field below — in the shared DynamoDB table this module creates for the first tenant declared, and never owns the key secret itself. Default to no tenants, which leaves tenant API keys disabled. Fields, all optional: "models_allow" and "models_deny" scope the models the tenant may name; "endpoints_allow" and "endpoints_deny" scope the routes, as glob patterns against route path templates such as '/v1/chat/completions'; "disabled" suspends the tenant's keys without deleting the record; "aws_role_arn" is an IAM role of the tenant's own AWS account its model invocations then run under, on the tenant's own Amazon Bedrock quota and bill; "key_generation" rotates the tenant's key on demand; "requests_per_minute" and "tokens_per_minute" cap what the tenant's key may do per minute, overriding tenant_rate_limit_requests_per_minute and tenant_rate_limit_tokens_per_minute for that tenant. An absent list restricts nothing; an empty list allows nothing; deny wins over allow. Example: { "acme" = { models_allow = ["anthropic.claude-"] endpoints_deny = ["/v1/images/"] key_generation = 2 requests_per_minute = 600 } "globex" = { aws_role_arn = "arn:aws:iam::123456789012:role/globex-stdapi" } } The key secret never enters Terraform state: the server mints it and, unless a rotation is asked for (below), delivers it once through the SSM parameter named in the tenant_keys output. That parameter is a SecureString encrypted with this deployment's own KMS key, so reading the key also takes kms:Decrypt on that key and not merely ssm:GetParameter on the path: retrieve it and delete it as soon as it appears. Declaring 'key_generation' on any tenant, like setting tenant_key_rotation_days, stores every tenant's key in an AWS Secrets Manager secret of its own instead (named in the tenant_keys output, on this deployment's KMS key), which is where a rotated key is published: the server rotates the tenant's key once whenever the value exceeds the generation it recorded at the previous rotation, so raising it -- 1, then 2, then 3 -- is a declarative, idempotent request for one rotation. The superseded key keeps working for tenant_key_rotation_overlap_seconds. Removing the last rotation trigger reverses the move: with tenant_key_rotation_days unset, clearing the one 'key_generation' a tenant still declares -- tidying away a value raised once and forgotten -- deselects the store, destroys every tenant's secret and reverts delivery to one-shot Parameter Store, so leave the value in place rather than removing it. Secrets Manager only schedules that deletion, with the recovery window and the --force-delete-without-recovery escape the README's destroy section describes. Declaring 'aws_role_arn' enables tenant AWS credentials on the server, grants the task role 'sts:AssumeRole' on exactly the declared roles, and cannot be combined with Amazon Bedrock Guardrails; the tenant must condition its role's trust policy on the ExternalId the server mints, read from the tenant's 'secret#' record. |
map(object({ |
{} |
no |
| timezone | Timezone for request date & time (IANA timezone identifier). Default to UTC. | string |
null |
no |
| tokens_estimation | Deprecated and ignored since stdapi.ai v1.14.0: token estimation has been removed and only real AWS-billed usage is reported. Accepted, and passed nowhere, so a configuration that still sets it keeps applying. Scheduled for removal in the next major version. | bool |
null |
no |
| tokens_estimation_default_encoding | Deprecated and ignored since stdapi.ai v1.14.0: token estimation has been removed. Accepted, and passed nowhere, so a configuration that still sets it keeps applying. Scheduled for removal in the next major version. | string |
null |
no |
| trusted_hosts | List of trusted host header values for Host header validation. Supports wildcard subdomains. Disabled by default. | list(string) |
null |
no |
| usage_api | Serve the organization usage and costs endpoints (/v1/organization/usage/*, /v1/organization/costs) from the metrics cloudwatch_metrics publishes (adds the cloudwatch:GetMetricData and cloudwatch:ListMetrics permissions). Every query is billed by CloudWatch per metric read and is excluded from its free tier, and enabling this also stores the usage metrics under additional dimensions. Enabling it turns on cloudwatch_metrics and cost_tracking unless either is set explicitly, since the endpoints report nothing without them. Default to false. | bool |
null |
no |
| usage_api_admin_scopes | OAuth 2.0 scopes an Amazon Cognito token must all carry to read the organization usage and costs endpoints, as a comma-separated list. With no scope named, no token is accepted and only the deployment's own API key may read them; a tenant API key is never accepted. Default to none. | string |
null |
no |
| usage_api_cache_ttl | Seconds an answered usage API query is reused for, so a client polling faster than the bucket width is not billed for a query that cannot have changed. Set to 0 to disable. Default to 60. | number |
null |
no |
| usage_api_max_metrics | Maximum number of metric series one usage API query may read; a query matching more is refused rather than billed. Default to 500, which is also the CloudWatch per-request maximum. | number |
null |
no |
| usage_api_max_range_days | Maximum span, in days, between start_time and end_time on a usage API query. Default to 92. | number |
null |
no |
| vector_store_chunk_overlap_tokens | Default number of tokens shared between consecutive chunks, for files indexed into a vector store without an explicit chunking_strategy. Must not exceed half of vector_store_chunk_size_tokens, which the server checks on startup. Default to 400. | number |
null |
no |
| vector_store_chunk_size_tokens | Default chunk size, in tokens, for files indexed into a vector store without an explicit chunking_strategy; a request's own chunking_strategy always wins. Tokens are approximated from the text length, and a chunk is additionally capped by what the embedding model accepts in one input. Default to 800. | number |
null |
no |
| vector_store_embedding_model | Model used to embed the files indexed into a vector store, and the queries searched against them. The model is frozen on each store when it is created, so changing it only affects stores created afterwards; existing stores keep answering with the model they were created with. Default to 'amazon.titan-embed-text-v2:0'. | string |
null |
no |
| version_to_deploy | Container image version tag from AWS Marketplace. Defaults to the server version this module release was built and tested against, which is what makes a given module version reproducible; there is no 'latest' resolution. Raise it to take a newer server without changing module version, or lower it to roll back. A '-arm64' or '-amd64' suffix is appended automatically based on var.cpu_architecture, so the value must not include an architecture suffix. | string |
"1.19.1" |
no |
| vpc_cidr | CIDR block for the dedicated VPC. | string |
"10.0.0.0/16" |
no |
| vpc_endpoints_allowed | If true, VPC endpoints interfaces are privileged to give AWS services access to the application if no internet access is required. VPC endpoint Gateway are always provisioned. Disable only if cost is privileged over security. | bool |
true |
no |
| vpc_flow_log_enabled | If true, enable VPC flow log. Disable only if cost is privileged over security. | bool |
true |
no |
| Name | Description |
|---|---|
| alb_arn | ARN of the Application Load Balancer (only if ALB is enabled). |
| alb_dns_name | DNS name of the Application Load Balancer (only if ALB is enabled). |
| alb_security_group_id | Security group ID of the Application Load Balancer (only if ALB is enabled). |
| alb_waf_web_acl_arn | ARN of the WAF WebACL (only if WAF is enabled). |
| alb_waf_web_acl_id | ID of the WAF WebACL (only if WAF is enabled). |
| alb_zone_id | Zone ID of the Application Load Balancer (only if ALB is enabled). |
| api_key | Returns API key value from var.api_key or var.api_key_create. API key values from var.api_key_ssm_parameter or var.api_key_secretsmanager_secret are not returned. |
| application_url | Application URL (uses domain name if configured, otherwise ALB DNS name). |
| aws_s3_tmp_prefix | S3 prefix (folder path) for temporary files used during job processing. To pass to compagnon module. |
| bedrock_batch_role_arn | ARN of the IAM service role Amazon Bedrock assumes to run batch inference jobs, or null when the Batch API is disabled. |
| bedrock_user_role_arn | ARN of the IAM role the server assumes once per end user, whether created by this module or supplied through aws_bedrock_user_role_arn, or null when per-end-user cost attribution is disabled. Activate the session tag key named by aws_bedrock_user_role_tag_key as a cost allocation tag of type 'IAM principal' to group Amazon Bedrock costs per end user. |
| bucket_arn | Configuration S3 bucket ARN. |
| bucket_id | Configuration S3 bucket ID. |
| cloudwatch_log_groups_names | CloudWatch log group names for each container in the server. |
| cluster_name | ECS cluster name. |
| deletion_protection | If true, enable deletion protection on eligible resources. To pass to compagnon module. |
| kms_key_arn | KMS key ARN. |
| kms_key_id | KMS key ID. |
| kms_policy_documents_json | KMS policy documents to add to the policy of the key specified via var.kms_key_id. |
| name_prefix | Name prefix for resources. To pass to compagnon module. |
| port | Container port exposed by the application. |
| regional_buckets | Map of region → bucket name (user-provided + auto-created). |
| security_group_id | Security group ID for the ECS server service. |
| service_discovery_service_name | Service discovery service name for the server (only if service discovery is enabled). |
| service_name | ECS service name. |
| subnet_ids | Subnets IDs where the ECS service is deployed. |
| tenant_keys | Per tenant of var.tenants: the public key ID, and where the server publishes the minted API key within a minute of starting or reconciling. By default that is the SSM SecureString parameter it is delivered through once (ssm_parameter): retrieve it with 'aws ssm get-parameter --name <ssm_parameter> --with-decryption --query Parameter.Value --output text', hand it to the tenant, then delete the parameter -- the copy it holds is the only one, and it is the only thing between a reader of this deployment's KMS key and a working tenant credential. With tenant_key_rotation_days set, or a tenants entry declaring key_generation, it is instead the AWS Secrets Manager secret holding the key as its current version (secret_name, secret_arn): read it with 'aws secretsmanager get-secret-value --secret-id <secret_name> --query SecretString --output text', and again after a rotation, when the superseded key is the AWSPREVIOUS version; grant the tenant secretsmanager:GetSecretValue on secret_arn, and kms:Decrypt on this deployment's KMS key through Secrets Manager, to let it re-read its own key -- an identity policy in this deployment's own account, since the module writes no resource policy on the secret nor on the KMS key. Whichever is not in use is null. Both are encrypted with that key, so reading either takes kms:Decrypt on it as well as the read permission itself. The key itself never enters Terraform state. |
| vector_store_queue_arn | ARN of the Amazon SQS queue carrying the vector store indexing jobs, or null when durable indexing is disabled. |
| vector_store_queue_url | URL of the Amazon SQS queue carrying the vector store indexing jobs, or null when durable indexing is disabled. |
| vectors_bucket_name | S3 vector bucket name backing the Vector Stores API, or null when it is disabled. |
| vectors_region | Region holding the S3 vector bucket, or null when the Vector Stores API is disabled. |