Every AI feature your team ships is also a decision nobody signed off on. SHIP is an AI reviewer that reads every pull request the moment it opens, catches the ones that quietly cross a real legal line, and only ever interrupts a human when it's actually found something. Ask it what's going on out loud and it tells you in one line, then puts the actual findings on whatever screen is nearby. Never in your ear.
Built for Build, Ship, Shape: the Amazon Developer Hackathon (Alexa+ track, AWS Builder mini-challenge). The core review engine started life as an entry to the Agents for Humans Hackathon (Strands Agents SDK). See Built across two hackathons for exactly what changed during this submission window.
Say a five-person startup adds a feature this week: an AI that reads a loan application and suggests whether to approve it. It works, it ships, everyone moves on. Then months later someone realizes the AI's prompt included the applicant's Social Security number in plain text, or that a zip code was quietly swaying who got approved, or that the AI's opinion had quietly become the actual decision, with no person ever looking at it. Nobody meant for any of that to happen. It's just what happens when a small team ships fast and nobody's job is to catch it.
That's not a hypothetical, and it's not a someday problem. GDPR's security obligations (Article 32, the exact rule this project's PII detectors enforce) have been binding, actively enforced law since 2018, with real fines reaching tens of millions of euros for exactly this kind of raw personal data reaching somewhere it shouldn't. Credit-scoring AI is separately named, in the EU AI Act's own Annex III, as a high-risk use case subject to human-oversight and bias-examination duties once its compliance timeline lands. That timeline is currently December 2027, pushed back from a mid-2026 extension, and worth watching precisely because it has already moved once. A small team's exposure here is real today under GDPR and growing on a clock that's still ticking. The team that shipped the feature has no compliance person, no legal review queue, and no time to build one.
The obvious fixes both fail. Ship blind and hope nothing surfaces. Or slow every single pull request down for a human to review by hand. That defeats the entire point of moving fast with AI in the first place.
SHIP is the third option. It reads the diff the moment a PR opens, works out whether something's actually wrong instead of just checking whether a risky-looking word shows up, explains what it found in plain English, points to the exact rule it breaks, and drafts the fix. A person only ever gets pulled in when it's found something real. Every other PR ships exactly as fast as it always would have. You don't even have to open a laptop to ask. Say "what's up" out loud and SHIP tells you in one sentence whether anything needs you. Never the finding itself, which only ever renders on a screen (see Ask, don't read).
- Detector's every citation is grounded in retrieved regulation text (GDPR, EU AI Act, OWASP), never a model's unaided recollection. It must call its retrieval tool before judging anything.
- Screener's keyword match is a signal, not a verdict. Detector independently judges whether it's a real violation or a false positive, verified against genuine violations and deliberate look-alikes designed specifically to trip a naive pattern match.
- The stricter categories carry an explicit evidence bar in the prompt itself. A protected characteristic must connect to an actual scoring operation (an arithmetic adjustment, a conditional) rather than merely appear in the same function, so a field that's present isn't confused with a field that's actually driving the decision.
- Triage only ever proposes an action: log-and-continue, or freeze. It never resolves anything by itself.
- Gate (the review console, gated behind its own credential separate from the webhook's) is the only place a frozen alert gets resolved. Approve/Reject is a one-way, guarded transition. A retry or a duplicate webhook delivery cannot silently re-open or overwrite a decision a human already made.
The pipeline underneath is mechanically generic enough to flag other things too, but SHIP deliberately stays scoped to AI-feature risk (raw data reaching a model, an AI output driving a decision with no checkpoint, a protected characteristic feeding a score, an agent granted an unscoped dangerous capability). That scope is the product. Diluting it into general-purpose static analysis would trade away the one thing that differentiates this from tools that already exist.
A webhook redelivery or a retried job can never create a duplicate alert, and can never silently reopen or overwrite a decision a human already made. Every flagged issue in a PR is reviewed independently and in parallel, so one slow or unlucky issue can never crowd out the review of another in the same PR (see Architecture below). Every property on this list is verified against the real, deployed system, not asserted from a passing test suite alone.
- Screener: a fast, free regex/AST pre-filter. Runs on every commit; if nothing matches, the PR passes in milliseconds and never costs a model call. A match doesn't mean a violation. It means "worth a real look."
- Detector: a Strands Agent, RAG-grounded against real, sourced regulation text. Runs only on the fragments Screener actually flagged, one fragment at a time, and returns a structured verdict: matched or not, which category, a 1–10 risk score, a plain-English explanation, the exact citation, and a draft remediation patch.
- Triage: routes the verdict. Below the threshold, it's logged and
nothing else happens. At or above it, the build freezes: an alert is
created, pending human review, and the PR's own commit status turns
red (
ship/compliance, promotable to a required check in branch protection. This is what actually blocks the merge button, not just a dashboard entry). The threshold is per-category, not one number for everything. A confirmed violation that breaks a required safety guarantee (unmasked PII reaching an external service, an automated decision with no human checkpoint at all) is held to a lower bar than one that's more a matter of degree. - Gate: the review console. Every frozen alert shows the file, the category, the risk score, the plain-English summary, the exact regulatory citation, and the suggested patch. A human clicks Approve or Reject; that decision is recorded as the one-way resolution of the alert, posted back to the PR as a comment, and folded into the recomputed commit status. Resolving the last blocking finding is what turns the check green. (Actually pushing the approved patch back to the PR via the GitHub API is a scoped-out next step, not yet wired in. Today, a human still applies the fix themselves once they've reviewed it here.) Gate also has a history view of every past disposition with the reason a human gave, and a connected-repos view. Connecting a repository is a form submission here, not a redeploy (see Tech stack).
Each is independently verified against real, deliberately adversarial test cases: genuine violations and deliberate false-positive look-alikes alike.
| ID | What it catches | Grounded in |
|---|---|---|
| PIIE-001 | Raw, direct PII (SSN, account number, full profile) reaching an external sink with no masking | GDPR Article 32 |
| PIIE-002 | The same kind of raw PII, written to a log stream | GDPR Article 32 |
| PIIE-003 | Raw PII stored in a cache/session store with no encryption | GDPR Article 32 |
| TLGP-002 | An AI-produced decision applied as final with no human checkpoint anywhere in the fragment | EU AI Act Article 14 |
| ALBP-001 | A protected characteristic (or a clear proxy) directly driving a scoring calculation | EU AI Act Article 10 + Annex III §5(b) |
| TLGP-001 | A dangerous capability (shell exec, unscoped DB write) granted to an AI agent with no gate | OWASP Top 10 for LLM Apps, LLM06:2025 |
SHIP is deliberately scoped to what a single PR diff can actually prove. See the architecture doc for the reasoning behind that boundary, and what's on the roadmap next.
The engine is demonstrated against a fork of MicroPyramid/micro-finance (MIT-licensed, a real Django lending app), used purely as a realistic third-party target and kept as a fully separate repository from this submission. Nothing from it is incorporated here:
- PR #1 plants an AI-assisted underwriting function sending an applicant's full raw profile to an external LLM.
- PR #2 plants one genuine violation per remaining detector across three new files, interleaved with four deliberate false-positive look-alikes (a non-agent backup job, a display-only profile field, a properly-hashed log call, a non-PII cache write).
Run directly against real Bedrock, both PRs together: 9 for 9. Every genuine violation correctly caught with an accurate citation, every look-alike correctly dismissed with real reasoning for why, not a coin-flip.
Nobody wants a voice assistant reading a two-minute monologue of PII findings and article citations aloud. Voice is good at exactly one thing here: an ambient, hands-free check for whether anything's wrong and where to look. It's bad at everything after that. Relay, SHIP's MCP server, hard-caps every spoken response to one short sentence with no line breaks. A finding list cannot fit in that space, so the attempt fails loudly instead of narrating. Citations, file paths, and code only ever reach a screen.
"Alexa, what's up?" "Three findings, one blocking. Want it on a screen?" "Show me on the TV."
The actual findings then appear, live, on whatever device just answered
to that name. That last step is a real push, not a shared-tab trick. Any
device with a browser (a TV's browser, an iPad, a laptop, even a smart
fridge's) can open ship-display.html, name
itself once, and sit idle with no polling until Relay pushes a finding to
it by name over an open WebSocket connection. A real compliance event
happens on the order of weeks, not seconds. A display that polled for it
every few seconds would spend nearly all of that traffic finding nothing
changed. An idle connection costs nothing until there's actually
something to say.
The one thing voice is never allowed to do: resolve a finding. SHIP exists to stop AI systems from making consequential decisions with no human accountably in the loop. That's the exact pattern its own TLGP-002 detector flags in other people's code. A version of SHIP that let someone clear a blocking GDPR finding by saying "approve it" to a speaker would be committing that same violation in its own interface. Try it:
"Approve it." "That needs a written reason on the record. Opening it on your screen."
There is no tool in Relay's surface that can perform an approval. The refusal is a missing capability, enforced server-side, not a prompt asking the model to decline.
Alexa+'s own MCP toolkit requires a live account relationship with an
Amazon Solutions Architect before its CLI/device path will connect at
all. That requirement isn't documented anywhere until you're already
mid-setup (see FRICTION_LOG.md for exactly where and
how it surfaced). The hackathon's own rules anticipate exactly this gap
and name a first-class alternative: a simulated Alexa+ experience in a
web app, source included.
Try it live.
Every response above comes from the real, deployed Relay endpoint over
Streamable HTTP, not a mock.
Full diagram and component-by-component detail: pandayv.github.io/ship-engine (source).
The one piece worth calling out here: each flagged issue in a PR is reviewed as its own independent, retryable job, queued and picked up by an independently-scaling reviewer function, rather than one sequential pass through the whole PR. A fragment that fails outright retries automatically and lands in a dead-letter queue after repeated failure instead of vanishing. A fragment that's just slow, or that hits AWS's own request-rate limit, backs off and retries on its own without blocking anything else. A PR with several flagged issues takes about as long as its slowest single issue, not the sum of all of them, and nothing is silently dropped.
- Agent framework: Strands Agents SDK
- Models: Amazon Nova Lite as the primary Bedrock backend, chosen on
measured evidence rather than preference. Benchmarked against Claude
Haiku 4.5, Nova Pro, Qwen3-235B, GLM-5, and DeepSeek-V3.2 on the real
adversarial fixtures below: all six caught every genuine violation and
dismissed every look-alike, but per-model request-per-minute quota is
what actually bounds review throughput on this account (10/min for
Claude models vs. 200/min for Nova Lite). See the benchmark table in
src/agents/detector.py. Google Gemini serves as a credit-exhaustion fallback and a local Ollama model as a fully offline reliability fallback. Both are switchable via one env var and genuinely functional, not unverified stretch goals. - Agent runtime: Amazon Bedrock AgentCore Runtime, the same Detector
logic deployed to a real managed runtime (
shipagentcore/), callable in-process for local development or remotely for the deployed path, toggled the same way - Retrieval: Amazon Bedrock Titan Embeddings, a small local numpy cosine-similarity store (the sourced regulation corpus is a few dozen chunks, so a hosted vector database would be pure overhead at this scale)
- Compute: two AWS Lambda functions, the webhook (fast-ack: verify, screen, dispatch) and an independently-scaling fragment processor (the actual model call, one fragment per invocation)
- Queueing: Amazon SQS, with a dead-letter queue for fragments that fail repeatedly and a concurrency cap on the processor so parallel reviews stay within the account's real request-rate limit
- State: Amazon DynamoDB, four tables.
ship-alerts(every write idempotent, so a webhook redelivery or a retried job can't create a duplicate and can't silently re-open a decision a human already made),ship-repos(which repositories SHIP watches. Connecting one is a point write from Gate's dashboard, not an environment-variable redeploy, seesrc/storage/repo_store.py),ship-status(a precomputed release summary, so Relay's voice path is oneGetItemregardless of how many findings exist, since Alexa+ allows Relay half a second to answer and aggregating on read would have blown that budget as findings accumulate), andship-device-connections(which device is reachable under which name, for the push path below) - Web: FastAPI (the webhook route and the Gate console, one deployable app, wrapped for Lambda via Mangum)
- Voice/MCP: MCP Python SDK
(Streamable HTTP, spec
2025-11-25) powers Relay (src/relay/), deployed as its own Lambda with SnapStart enabled. A fresh ASGI app is built per invocation, a genuine SDK/Lambda incompatibility rather than a style choice; see the docstring at the top ofsrc/relay/server.py. - Push: an Amazon API Gateway WebSocket API plus a small dedicated
Lambda (
device_gateway_handler.py) handling connect/disconnect/register, deliberately separate from Relay so Relay's own IAM role stays scoped to exactly what answering a question requires
- An AWS account with Bedrock model access enabled for at least one
Claude model (a brand-new account may need a one-time use-case
submission and/or an AWS Support request before this works. See
the troubleshooting note below if
bedrock:InvokeModelfails with a quota or subscription error). - The AWS CLI, configured (
aws configure) with a scoped IAM identity, not root credentials. - Python 3.12+,
pip, andnpm(for the AgentCore CLI). - A GitHub repo to protect, and a personal access token with read access to it.
git clone https://github.com/pandayv/ship-engine.git
cd ship-engine
python3 -m venv .venv && source .venv/bin/activate
pip install -r requirements.txtThe regulation text this project cites is already checked into
rag_corpus/ (GDPR Art. 32, EU AI Act Art. 10/14 + Annex
III §5(b), OWASP LLM06:2025), sourced verbatim, not paraphrased. Nothing
to do here unless you're extending the taxonomy with a new category; see
src/taxonomy.py for the registry a new category needs
to join.
export SHIP_MODEL_BACKEND=bedrock
export SHIP_DETECTOR_MODE=in_process
python3 -m src.agents.detectorThis runs a hardcoded violation through the real agent end to end and prints the structured verdict. It's the fastest way to find out if Bedrock access, model availability, or region config needs fixing before anything else. If it fails with an access or quota error, the model needs enabling in the Bedrock console first (Model access → request access), and a brand-new AWS account specifically may need its own quota raised via a support case. This is an account-activation gate, not a code problem.
python3 -m src.storage.alert_store
python3 -m src.storage.repo_storealert_store calls create_table_if_not_exists() and then a self-test
write/read/resolve cycle against the real ship-alerts table.
repo_store creates ship-repos, the table backing Gate's
connected-repos view. Needs dynamodb:CreateTable plus the basic
item-level actions on your IAM identity first.
npm install -g @aws/agentcore
cd shipagentcore
agentcore deploy --yesNote the deployed runtime ARN from the output. You'll need it in step 7.
If Detector's prompt, schema, or RAG corpus ever changes, both
src/agents/detector.py and shipagentcore/app/ship_diagnostician/main.py
need the same change and a fresh deploy. See
tests/test_taxonomy_consistency.py,
which exists specifically to catch the two copies drifting apart before
that reaches production silently. (The AgentCore app's own folder/runtime
name, ship_diagnostician, predates this project's Detector naming pass.
Left as-is deliberately, since renaming it mints a brand-new runtime ARN
for no visible benefit. It's internal deployment plumbing, not something
this README's naming otherwise touches.)
aws sqs create-queue --queue-name ship-fragment-queue-dlq \
--attributes MessageRetentionPeriod=1209600
DLQ_ARN=$(aws sqs get-queue-attributes \
--queue-url "$(aws sqs get-queue-url --queue-name ship-fragment-queue-dlq --query QueueUrl --output text)" \
--attribute-names QueueArn --query Attributes.QueueArn --output text)
aws sqs create-queue --queue-name ship-fragment-queue \
--attributes "{\"VisibilityTimeout\":\"960\",\"RedrivePolicy\":\"{\\\"deadLetterTargetArn\\\":\\\"$DLQ_ARN\\\",\\\"maxReceiveCount\\\":3}\"}"The fragment-processor Lambda needs its own execution role (permission to
consume the queue, call bedrock-agentcore:InvokeAgentRuntime against
the ARN from step 5, and write to the DynamoDB table from step 4).
Both Lambdas in this project deploy from the same zip. Build it once,
including real Linux dependency wheels: the --platform/--only-binary
flags matter even if you're building on macOS or Windows. Skip them and
the zip will contain the wrong platform's compiled packages, failing at
import time on Lambda rather than at build time.
mkdir -p build && cp -r src fragment_lambda_handler.py lambda_handler.py build/
pip install -r requirements-lambda.txt -t build/ \
--platform manylinux2014_aarch64 --only-binary=:all: --python-version 3.12
cd build && zip -r ../ship-webhook.zip . -x "*.dist-info/*" && cd ..(Building for arm64/manylinux2014_aarch64 above to match Lambda's
cheaper Graviton architecture. Switch both the --platform flag here and
--architectures below to x86_64 consistently if you'd rather not deal
with cross-compiling.)
aws lambda create-function --function-name ship-fragment-processor \
--runtime python3.12 --architectures arm64 \
--role <YOUR_FRAGMENT_PROCESSOR_ROLE_ARN> \
--handler fragment_lambda_handler.handler \
--timeout 900 --memory-size 512 \
--zip-file fileb://ship-webhook.zip \
--environment "Variables={SHIP_DETECTOR_MODE=agentcore,SHIP_MODEL_BACKEND=bedrock,SHIP_AGENTCORE_RUNTIME_ARN=<ARN_FROM_STEP_5>}"
aws lambda create-event-source-mapping --function-name ship-fragment-processor \
--event-source-arn <FRAGMENT_QUEUE_ARN> --batch-size 1 \
--scaling-config MaximumConcurrency=5MaximumConcurrency is the knob that keeps parallel fragment reviews
within your account's real Bedrock request-rate limit. Raise it once you
know what that limit actually is for your account.
Needs its own execution role: permission to send to the queue from step 6, and, if you also want the fast-ack path itself to fall back to processing inline with no queue configured, the same Bedrock/DynamoDB permissions as step 6's role.
aws lambda create-function --function-name ship-webhook \
--runtime python3.12 --architectures arm64 \
--role <YOUR_WEBHOOK_ROLE_ARN> \
--handler lambda_handler.handler \
--timeout 30 --memory-size 512 \
--zip-file fileb://ship-webhook.zip \
--environment "Variables={
SHIP_DETECTOR_MODE=agentcore,
SHIP_MODEL_BACKEND=bedrock,
SHIP_AGENTCORE_RUNTIME_ARN=<ARN_FROM_STEP_5>,
SHIP_FRAGMENT_QUEUE_URL=<QUEUE_URL_FROM_STEP_6>,
GITHUB_TOKEN=<a token with read access to that repo>,
GITHUB_WEBHOOK_SECRET=<a random secret you generate>,
SHIP_DASHBOARD_TOKEN=<a second random secret you generate>
}"
aws lambda create-function-url-config --function-name ship-webhook \
--auth-type NONEGITHUB_WEBHOOK_SECRET and SHIP_DASHBOARD_TOKEN must be two genuinely
different values. The webhook's signature check and the dashboard's auth
are deliberately independent, so compromising one can't silently disable
the other.
Open https://<your-function-url>/dashboard/repos?token=<SHIP_DASHBOARD_TOKEN>
and connect <owner>/<repo>. A payload naming any other repo gets
rejected before it can spend your GitHub token or Bedrock quota (see
src/storage/repo_store.py). That link only
needs to be visited once. The token in the URL establishes a session
cookie, and every page from there on (Active, History, Connected repos)
is just a normal link with no secret in it. Visiting /dashboard cold
prompts a login form instead. An optional
SHIP_ALLOWED_REPOS=<owner>/<repo> environment variable pre-seeds this
same allowlist without needing the table at all, useful for a first
bring-up before step 4's tables exist, or as a fallback if DynamoDB is
briefly unreachable.
Then, in the target repo's Settings → Webhooks: the Lambda Function URL
from step 7, content type application/json, secret matching
GITHUB_WEBHOOK_SECRET above, event: Pull requests.
curl https://<your-function-url>/health
# {"status":"ok"}Open a real PR against the target repo containing something Screener
would flag (raw PII reaching an external call is the easiest to trigger)
and confirm an alert appears at
https://<your-function-url>/dashboard?token=<SHIP_DASHBOARD_TOKEN>.
src/
agents/
screener.py # Fast regex/AST pre-filter + per-fragment isolation
detector.py # Strands Agent — RAG-grounded semantic judgment
triage.py # Risk-based routing, per-category thresholds
api/
main.py # Webhook route, fast-ack + fragment dispatch
dashboard.py # Gate — the human review console
github_client.py # Real PR-diff fetching via the GitHub API
github_writeback.py # Posts findings/decisions back to the PR (comment + commit status)
aws/
bedrock_session.py # Shared Bedrock session + adaptive-retry config
rag/
chunker.py # Regulation text -> citable paragraph chunks
vector_store.py # Local embedding store + retrieval
storage/
alert_store.py # DynamoDB — idempotent alert persistence
repo_store.py # DynamoDB — which repos SHIP watches
status_store.py # DynamoDB — precomputed release summary Relay's voice path reads
device_store.py # DynamoDB — which display device is reachable under which name
taxonomy.py # Single source of truth: detector <-> corpus <-> Screener trigger mapping
relay/
server.py # Relay — the MCP server; the voice/screen modality split lives here
modality.py # Enforces the spoken-response character cap and forbids line breaks
readmodel.py # Voice path (reads the precomputed summary) vs. screen path (reads alerts)
push.py # Delivers a payload to one named device over its open connection
rag_corpus/ # Sourced regulation/standard text, verbatim
shipagentcore/ # AgentCore Runtime deployment of Detector
docs/
alexa-simulator.html # The web-simulated Alexa+ experience — calls real, deployed Relay
ship-display.html # Any-device receiving surface — names itself, waits for a push
ship-render.js # Finding-rendering logic shared by both pages above
architecture.html # Full pipeline diagram
scripts/
healing_loop.py # Periodic corpus-grounding check, decoupled from the hot path
benchmark_models.py # The adversarial benchmark behind the Nova Lite model choice
precompute_embeddings.py # Regenerates rag_corpus/'s cached embeddings
lambda_handler.py # Webhook Lambda entrypoint
fragment_lambda_handler.py # Fragment-processor Lambda entrypoint
relay_lambda_handler.py # Relay's Lambda entrypoint — builds a fresh app per invocation, see server.py
device_gateway_handler.py # WebSocket connect/disconnect/register Lambda entrypoint
tests/ # 200 tests, no AWS credentials required to run
The full review pipeline, Screener through Gate, is live and deployed,
with the PR's own merge button actually gated by a real commit status.
Verified end to end against real Bedrock and a real GitHub webhook, not a
mocked demo. Freezing an alert sets ship/compliance to failing
(promotable to a required check in branch protection) and resolving one
recomputes it, so the loop from detection to a human decision to the
PR's own merge button actually closes.
On top of that, Relay (the MCP server), the Alexa+ simulated experience, and real cross-device push are also live. See Ask, don't read above for the full walkthrough. Every claim in that section is checked against the deployed endpoint, the same standard the rest of this README holds itself to.
Things known and deliberately not built yet, not overlooked:
- Approving an alert doesn't yet push the suggested patch back to the PR automatically. A human still applies it themselves once they've reviewed it in Gate. A real GitHub-API integration away, not an architecture change.
- Healing Loop (
scripts/healing_loop.py) closes part of this: run periodically (deliberately not on Detector's per-fragment hot path, see the script's own docstring for why), it re-fetches each sourced regulation page and flags any chunk that no longer appears verbatim, so a citation is never silently resting on text a regulator has since amended. Verified against the real live sources rather than fixtures alone; caught and fixed two genuine false-positive causes (HTML entity decoding, CSS-rendered clause numbering) this way. Not yet wired to an actual schedule (cron/ EventBridge) or to an alert channel beyond its own stdout report. That part is still manual. - The Alexa+ CLI/device path itself isn't connected. It requires an
AWS account already registered by an Amazon Solutions Architect, a live
account relationship rather than a self-service step (see
FRICTION_LOG.md). The simulated web experience calls the identical, real Relay endpoint a live connection would, so nothing about the review logic itself is untested. Only the transport Alexa+'s own infrastructure would use to reach it is missing. - New detectors, beyond what's listed above, for risk patterns that need more than a single PR diff to prove: infrastructure/deployment context, or behavior observed across multiple files or over time. See the architecture doc for where that boundary sits and why.
The review engine (Screener → Detector → Triage → Gate, everything
through the GitHub write-back) started as a submission to the Agents for
Humans Hackathon (Strands SDK). During this hackathon's own submission
window (opened August 31, 2026), real new work was added and verified,
not a relabeling of the old submission: Relay (the MCP server), the
modality contract that keeps voice brief and screens detailed, the real
Alexa+ web simulation, and genuine cross-device push over a WebSocket API
Gateway are all built inside this submission period. None of it existed
before this window opened. git log tells the same story directly,
commit by commit.
MIT — see LICENSE.