Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 6 additions & 0 deletions dev/.env.docker-compose
Original file line number Diff line number Diff line change
Expand Up @@ -20,4 +20,10 @@ DJANGO_BATAI_NABAT_OIDC_CLIENT_SECRET=batai-local-dev-secret
DJANGO_BATAI_NABAT_OIDC_ISSUER=http://localhost:8081/auth/realms/NABAT
DJANGO_BATAI_NABAT_OIDC_BASE_URL=http://keycloak:8080/auth/realms/NABAT

# Only reachable if docker-compose-nabat-mock.yml's nabat-mock service is also up; otherwise
# inert, the same way the OIDC vars above are unused unless Keycloak is also up.
# In production environments, ensure that this environment variable is unset. The base settings
# module handles the real NABat graphql endpoint.
DJANGO_BATAI_NABAT_API_URL=http://nabat-mock:8082/graphql

VITE_API_ROOT=http://localhost:8000
6 changes: 6 additions & 0 deletions dev/.env.docker-compose-native
Original file line number Diff line number Diff line change
Expand Up @@ -11,4 +11,10 @@ DJANGO_BATAI_NABAT_OIDC_CLIENT_SECRET=batai-local-dev-secret
DJANGO_BATAI_NABAT_OIDC_ISSUER=http://localhost:8081/auth/realms/NABAT
DJANGO_BATAI_NABAT_OIDC_BASE_URL=http://localhost:8081/auth/realms/NABAT

# Only reachable if docker-compose-nabat-mock.yml's nabat-mock service is also up; otherwise
# inert, the same way the OIDC vars above are unused unless Keycloak is also up.
# In production environments, ensure that this environment variable is unset. The base settings
# module handles the real NABat graphql endpoint.
DJANGO_BATAI_NABAT_API_URL=http://localhost:8082/graphql

Comment thread
naglepuff marked this conversation as resolved.
VITE_API_ROOT=http://localhost:8000
25 changes: 0 additions & 25 deletions dev/keycloak/README.md

This file was deleted.

8 changes: 8 additions & 0 deletions dev/nabat_mock/Dockerfile
Original file line number Diff line number Diff line change
@@ -0,0 +1,8 @@
FROM python:3-slim

RUN pip install --no-cache-dir minio

COPY mock_server.py /app/mock_server.py
WORKDIR /app

CMD ["python", "mock_server.py"]
10 changes: 10 additions & 0 deletions dev/keycloak/NABAT-realm.json → dev/nabat_mock/NABAT-realm.json
Original file line number Diff line number Diff line change
Expand Up @@ -20,6 +20,16 @@
"access.token.claim": "true",
"claim.name": "preferred_username"
}
},
{
"name": "sub",
"protocol": "openid-connect",
"protocolMapper": "oidc-sub-mapper",
"consentRequired": false,
"config": {
"id.token.claim": "true",
"access.token.claim": "true"
}
}
]
},
Expand Down
107 changes: 107 additions & 0 deletions dev/nabat_mock/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,107 @@
# Mocking the NABat Integration

The integration with the NABat platform requires some trickier auth and API calls to test
fully. This directory contains helpful tools for local development testing, covering both
halves of the integration:

- **Keycloak** stands in for NABat's OIDC login, so the "open in batai" redirect and code
exchange can be exercised for real.
- **nabat-mock** stands in for NABat's GraphQL API (`BATAI_NABAT_API_URL`), so recording
fetches can run end to end against real infrastructure (the existing `minio` service)
instead of hitting `sciencebase.gov`.

## Running the mock stack

In the top-level directory of the repository there is a docker compose file that can be used
in conjunction with whichever docker compose files you already use for development. This means
you only need to spin up these services when required. To chain docker compose files, simply
use the `-f` flag multiple times. For example, with local development:

```bash
docker compose -f docker-compose.yml -f docker-compose-nabat-mock.yml up
```

## Keycloak Configuration

Keycloak configuration is defined in [NABAT-realm.json](./NABAT-realm.json). It sets up 2
clients: a proxy "NABat" and a client for the locally running BatAI application. It also sets
up a test user (this would be the analog to your account with NABat). The username and
password for this user are both `testuser`.

This realm issues lightweight access tokens, so any claim batai needs has to be re-added by
an explicit protocol mapper - including `sub` (the user's Keycloak ID), which Keycloak
otherwise strips from lightweight tokens by default even though it's normally unconditional
per the OIDC spec. `create_recording_annotation` reads `sub` directly, so the `profile` scope
carries an `oidc-sub-mapper` for it. Keep that in mind if this realm export ever gets
regenerated from Keycloak's admin UI - it's easy to lose.

## nabat-mock

[mock_server.py](./mock_server.py) stands in for NABat's GraphQL API. BatAI never uses a real
GraphQL client - every query is a plain string with literal IDs, POSTed as `{"query": "..."}`
- so rather than running an actual GraphQL server, the mock just pattern-matches on substrings
in the query text and returns canned JSON shaped like NABat's real responses.

It's currently scoped to the single-recording fetch flow only (not NABat's "file list"
feature). Any `recording_id` resolves successfully except `0`, which is reserved to simulate
"not found / access denied" and exercise the existing 403 handling.

Every recording fetched this way also gets one already-vetted species seeded onto it, so
`create_nabat_recording_from_response()` picks it up and creates a matching
`NABatRecordingAnnotation` automatically on first fetch - there's something to see in the
annotations list immediately rather than an empty one. The seeded annotation's email matches
the Keycloak test user (`testuser@example.com`), so it's visible when browsing as them.
Pushing an annotation back to NABat (`.../push-to-nabat`) is mocked too and always succeeds.

This seeding requires a local `Species` row with a matching pk to already exist - species
sync itself isn't mocked. Defaults to pk `1`; override `SEED_ANNOTATION_SPECIES_ID` /
`SEED_ANNOTATION_EMAIL` if your database doesn't have that row or you want a different one.
Most dev databases already have Species data from prior real NABat usage, before this mock
existed.

Presigned URLs are generated for real against the `minio` service already in
`docker-compose.yml`, so the download and spectrogram-generation steps run against real
infrastructure. nabat-mock itself never connects to minio - presigning is pure local
signing - it only needs to know the *bucket name*, not to reach it. Until an object
actually exists at `recordings/example.wav`, the presigned URL it returns will 404 when
downloaded.

Seed that object (and create the bucket) with [upload_recording.py](./upload_recording.py),
once minio is up:

```bash
docker compose -f docker-compose.yml -f docker-compose-nabat-mock.yml up -d minio
uv run dev/nabat_mock/upload_recording.py
```

With no path given, it uploads the repo's own [assets/example.wav](../../assets/example.wav).
Pass a different path to use your own file instead - it must actually be a `.wav` (checked
both by extension and by parsing it with Python's `wave` module). The project's own
virtualenv already has the `minio` package (a transitive dependency via
django-minio-storage), so `uv run` needs no extra install.

### Running celery natively (not in Docker)

Whoever downloads the presigned URL (the celery worker) must be able to resolve the host
baked into it. nabat-mock signs against `MINIO_ENDPOINT`, defaulting to the compose-network
name `minio:9000` - fine if celery is also containerized. If you run celery natively (see
[native-development.md](../native-development.md)), it can't resolve `minio`; only the
published port on `localhost` is reachable from the host.

If you run celery locally for your development, make sure to set `MINIO_ENDPOINT=localhost:9000` in your environment.

```bash
MINIO_ENDPOINT=localhost:9000 docker compose -f docker-compose.yml -f docker-compose-nabat-mock.yml up
```

## Testing the full flow

Once the mock stack is running alongside BatAI, run [./print-nabat-auth-url.sh](./print-nabat-auth-url.sh)
to generate a URL.

Paste the URL into your browser, and if this is the first time you're going through the
workflow or KC doesn't have an active token for `testuser` you'll need to log in as
`testuser`. This redirects to your local BatAI application, exchanges the code with
Keycloak, and then kicks off a real fetch of the recording (backed by nabat-mock and minio)
and real spectrogram generation - the same pipeline production runs, driven entirely by
local infrastructure.
199 changes: 199 additions & 0 deletions dev/nabat_mock/mock_server.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,199 @@
#!/usr/bin/env python3
"""Minimal stand-in for NABat's GraphQL API, for local dev/testing.

BatAI never uses a real GraphQL client: every query is a plain string,
interpolated with literal IDs and POSTed as {"query": "..."}, with no
variables and no schema. So instead of running a GraphQL server, this
dispatches on substrings in the query text and returns canned JSON shaped
like NABat's real responses.

Scope (for now): only the single-recording fetch flow, i.e. the query shapes used by:
- bats_ai/core/tasks/nabat/nabat_data_retrieval.py (fetchAcousticAndSurveyEventInfo)
- bats_ai/core/views/nabat/nabat_recording.py (bare presignedUrlFromAcousticFile,
used as an access check by get_email_if_authorized / generate_nabat_recording;
and the updateAcousticFileVet mutation, used to push an annotation to NABat)
Anything else (species sync, NABat file lists) returns a GraphQL-shaped error
rather than crashing, so it's obvious a handler needs to be added rather than
failing confusingly downstream.

The combined query's response seeds one already-vetted species for every recording
fetched (see SEED_ANNOTATION_*), so create_nabat_recording_from_response() picks it
up and creates a matching NABatRecordingAnnotation automatically on first fetch -
giving you something to see immediately rather than an empty annotation list. This
requires a local Species row with a matching pk to already exist (the species-sync
query isn't mocked); most dev databases already have one from prior real usage.

Presigned URLs are generated for real against the `minio` service already in
docker-compose.yml, so the full download + spectrogram-generation pipeline
runs end-to-end against real infrastructure. Run upload_recording.py to seed
the object they point at (it also creates the bucket - this service never
needs a live connection to minio at all, since presigning is pure local
signing and does no network I/O).

Whoever downloads a presigned URL (the celery worker) needs to be able to
resolve the host baked into it, and that host must match what the URL was
signed with. Set MINIO_ENDPOINT accordingly: the compose-network name
`minio:9000` if celery also runs in Docker (the default), or `localhost:9000`
if celery runs natively on the host (see dev/.env.docker-compose-native),
since only the published port is reachable from there.
"""

from __future__ import annotations

from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
import json
import logging
import os
import re

from minio import Minio

logging.basicConfig(level=logging.INFO)
logger = logging.getLogger("nabat-mock")

MINIO_ENDPOINT = os.environ.get("MINIO_ENDPOINT", "minio:9000")
MINIO_ACCESS_KEY = os.environ.get("MINIO_ACCESS_KEY", "minioAccessKey")
MINIO_SECRET_KEY = os.environ.get("MINIO_SECRET_KEY", "minioSecretKey")
MINIO_BUCKET = os.environ.get("MINIO_BUCKET", "nabat-mock")
PORT = int(os.environ.get("PORT", "8082"))

# The single shared object every mocked recording_id resolves to. Per-ID fixture
# files aren't needed: any recording_id "just works" against this one object.
RECORDING_OBJECT_KEY = "recordings/example.wav"

# Reserved recording_id that always resolves to "not found", to exercise the
# existing 403 handling in get_email_if_authorized / generate_nabat_recording.
NOT_FOUND_RECORDING_ID = 0

# The "already vetted in NABat" species every newly-fetched recording seeds an
# annotation for. Must be a pk that actually exists in the local Species table
# (see the module docstring). Email matches dev/nabat_mock/NABAT-realm.json's test
# user, so the seeded annotation is visible when browsing as them.
SEED_ANNOTATION_SPECIES_ID = int(os.environ.get("SEED_ANNOTATION_SPECIES_ID", "1"))
SEED_ANNOTATION_EMAIL = os.environ.get("SEED_ANNOTATION_EMAIL", "testuser@example.com")

ACOUSTIC_FILE_ID_RE = re.compile(r'acousticFileId:\s*"?(\d+)"?')

# Presigning is pure local signing - no network I/O - as long as a region is given
# (otherwise minio-py falls back to a live GetBucketLocation request). So this never
# actually needs to connect to minio; it just needs MINIO_ENDPOINT to be whatever host
# the downloader (celery) can resolve. See the module docstring.
minio_client = Minio(
MINIO_ENDPOINT,
access_key=MINIO_ACCESS_KEY,
secret_key=MINIO_SECRET_KEY,
secure=False,
region="us-east-1",
)


def extract_id(pattern: re.Pattern, query: str) -> int | None:
match = pattern.search(query)
return int(match.group(1)) if match else None


def presigned_recording_url(recording_id: int) -> str | None:
if recording_id == NOT_FOUND_RECORDING_ID:
return None
return minio_client.presigned_get_object(MINIO_BUCKET, RECORDING_OBJECT_KEY)


def build_presigned_only_response(query: str) -> dict:
"""Mirrors nabat_recording.py's QUERY: presignedUrlFromAcousticFile only."""
recording_id = extract_id(ACOUSTIC_FILE_ID_RE, query)
url = presigned_recording_url(recording_id) if recording_id is not None else None
return {"data": {"presignedUrlFromAcousticFile": {"s3PresignedUrl": url} if url else None}}


def build_combined_response(query: str) -> dict:
"""Mirrors nabat_data_retrieval.py's fetchAcousticAndSurveyEventInfo query."""
recording_id = extract_id(ACOUSTIC_FILE_ID_RE, query)
url = presigned_recording_url(recording_id) if recording_id is not None else None

return {
"data": {
"presignedUrlFromAcousticFile": {"s3PresignedUrl": url} if url else None,
"surveyEventById": {
"createdBy": "mock@nabat.org",
"createdDate": "2024-01-01T00:00:00",
"eventGeometryByEventGeometryId": {
"description": "Mock survey event geometry",
"geom": {"geojson": None},
},
"acousticBatchesBySurveyEventId": {
"nodes": [
{
"id": "mock-batch",
"acousticFileBatchesByBatchId": {
"nodes": [
{
"autoId": SEED_ANNOTATION_SPECIES_ID,
"manualId": SEED_ANNOTATION_SPECIES_ID,
"vetter": SEED_ANNOTATION_EMAIL,
"speciesByManualId": None,
}
]
},
}
]
},
},
"acousticFileById": {
"fileName": f"mock_recording_{recording_id}.wav",
"recordingTime": "2024-01-01T00:00:00",
"s3Verified": True,
"sizeBytes": 0,
},
}
}


def build_update_vet_response(query: str) -> dict:
"""Mirrors nabat_recording.py's UPDATE_QUERY (updateAcousticFileVet mutation).

Always succeeds. update_nabat_species only checks for an "errors" key, never
reads the payload, so no fields here need to reflect what was actually sent.
"""
return {"data": {"updateAcousticFileVet": {"acousticFileBatchId": 1}}}


def build_response(query: str) -> dict:
if "updateAcousticFileVet" in query:
return build_update_vet_response(query)
if "fetchAcousticAndSurveyEventInfo" in query:
return build_combined_response(query)
if "presignedUrlFromAcousticFile" in query:
return build_presigned_only_response(query)
return {"errors": [{"message": "nabat-mock has no handler for this query yet"}]}


class Handler(BaseHTTPRequestHandler):
def log_message(self, format_, *args):
logger.info("%s - %s", self.address_string(), format_ % args)

def do_POST(self):
length = int(self.headers.get("Content-Length", 0))
body = self.rfile.read(length)
try:
payload = json.loads(body)
query = payload.get("query", "")
except json.JSONDecodeError:
query = ""

response_body = json.dumps(build_response(query)).encode("utf-8")

self.send_response(200)
self.send_header("Content-Type", "application/json")
self.send_header("Content-Length", str(len(response_body)))
self.end_headers()
self.wfile.write(response_body)


def main():
server = ThreadingHTTPServer(("0.0.0.0", PORT), Handler) # noqa: S104 (container-internal)
logger.info("nabat-mock listening on :%s", PORT)
server.serve_forever()


if __name__ == "__main__":
main()
Original file line number Diff line number Diff line change
Expand Up @@ -2,8 +2,9 @@
# Prints a Keycloak authorization URL you can paste into a browser to kick off the
# NABat -> batai redirect flow for real, the same way clicking "open in batai" in the
# NABat portal does (see scripts/USGS/sampleUrl.txt for the legacy apiToken-URL equivalent).
# Since batai doesn't handle the `code` param yet, the landing page will just show today's
# apiToken-less route - this is for exercising the Keycloak leg of the flow in a real browser.
# With docker-compose-nabat-mock.yml's nabat-mock service also up, this drives the full
# pipeline end to end: Keycloak login, code exchange, and then a real fetch + spectrogram
# generation against the mocked NABat GraphQL API and the existing minio service.
set -euo pipefail

KEYCLOAK_URL=${KEYCLOAK_URL:-http://localhost:8081/auth}
Expand Down
Loading