ProofScout is a local-first, trust-bound market intelligence scanner and proof-of-work generation pipeline. It surfaces remote, project-first opportunities by evaluating real-world signals, enforcing strict human review gates, and keeping raw evidence payloads immutable.
- Human Review Gate: ProofScout never automatically sends messages, submits applications, or triggers unreviewed side effects. All outreach drafts and proof kits require explicit operator approval.
- Immutable Evidence: Captured raw payloads and provenance records are stored immutably in SQLite. Recomputations and outcome updates never overwrite raw historical artifacts.
- Fail-Closed Security: Missing consents, expired approvals, unknown providers, or forbidden credential patterns abort execution immediately. No CAPTCHA bypasses or unattended credential automation.
- Transparent Scoring: Ranking adjustments, outcomes, and risk penalties are explicitly itemized in score explanations. No opaque black-box weights or hidden model biases.
- Completely Offline Verification: The test suite runs 100% offline without live network dependencies or external services.
- Node.js >= 22 (uses Node's native
node:sqlite) - npm
npm install
npm run buildnpm test
npm run typecheckThis walkthrough takes an empty checkout to a ranked shortlist backed by live crawl data, with every step's output visible. All fetching is approval-gated; nothing is sent anywhere.
npm install
npm run build
npm test # expect: 34 files / 302 tests passedlive-approval.json is a ScopedProviderApproval: it whitelists provider ids, limits lanes (structured, social), expires (expiresAt), and caps requests (requestBudget). No provider is fetched unless named here — this is the core invariant (ADR 0006/0013/0014/0015). Each custom board key (e.g. wise-sr) maps to a first-party ATS JSON endpoint in boardUrls.
node dist/src/cli.js scan --approval live-approval.json --days 30 --format json > reports/live-scan.jsonExpect ~43 providers: 39 structured ATS boards (Greenhouse, Ashby, Lever, Workable, SmartRecruiters) plus the social-text adapter (Reddit PullPush lead snippets from r/forhire and r/remotejobs). The JSON diagnostics array reports per-provider status and record counts; shortlist holds the top-ranked opportunities, each with a reachable application route. Use --format markdown for the human-readable version.
Two opt-in flags shape the shortlist (defaults unchanged — the raw top-10 by score):
# More than 10 cards
node dist/src/cli.js scan --approval live-approval.json --days 30 --format json --limit 25
# Only roles whose TITLE matches your focus keywords (AI/agents/data/automation), not the highest-paying generic roles
node dist/src/cli.js scan --approval live-approval.json --days 30 --format json --relevant-only --limit 25Relevance is scored against TargetProfile.focusKeywords (ADR 0018) with word-boundary matching; --relevant-only requires a title-level hit. The keyword list has two tiers — experience (data analyst, business analyst, business intelligence, analytics engineer, data engineer) and target (applied AI, agent, forward deployed, automation) — so the shortlist reflects what the CV already supports, not only the roles being aimed for.
node reports/generate-live-explorer.mjs # consumes the saved scan JSON
start reports/live-scan.html # standalone, read-onlyThe generator consumes the saved JSON by default; add --fresh to re-crawl inline instead.
Every candidate lands in one of three lanes — shortlist, needs_review, rejected — per the review model. Track your progression on opportunities without touching immutable evidence:
node dist/src/cli.js outcome-set <opportunity-id> researching
node dist/src/cli.js outcome-list --format markdownOutcomes persist in data/evidence.db (file-backed since commit 984b569), so watchlists, diffs, and alerts survive between runs.
node dist/src/cli.js schedule-run --days 30 --approval live-approval.json --format markdown
node dist/src/cli.js alerts-list --format markdownschedule-run diffs each run against the stored evidence and alerts only on new or materially changed shortlist-ready opportunities.
node dist/src/cli.js validate --fixture test/fixtures/validated-local-scan-v1.json
node dist/src/cli.js evaluate --fixture test/fixtures/validated-local-scan-v1.json
node dist/src/cli.js review-report --fixture test/fixtures/validated-local-scan-v1.json --policy shortlist-ready
node dist/src/cli.js holdout-gate --manifest test/fixtures/holdout/sanitized-holdout-v1.json \
--policy test/fixtures/holdout/holdout-policy-v1.jsonThe fixture evaluate baseline is deliberately frozen (documented 57.14% raw-fixture precision; the production shortlist-ready policy surfaces 100% precision on shortlist). See .scratch/precision-remediation/validation-report.md.
--eligibility classifies every ranked record against the profile's declared constraints (remote-only, hiring regions + residence country, sponsorship-for-relocation) into eligible / needs_confirmation / ineligible, without changing scores or the default shortlist. A value other than all also filters the main shortlist to that group:
node dist/src/cli.js scan --approval live-approval.json --days 30 --format markdown --eligibility allCountry fences are decided by the profile's residenceCountry (default exampleland), never by SEA/APAC proximity; an unknown arrangement or silent region stays needs_confirmation, never an eligibility guarantee.
source-report re-runs the approved scan and emits a per-source contribution report — unique qualified opportunities, shortlist-ready counts, stale records, same-source duplicates, cross-provider shared groups, raw diagnostic volumes, verbatim collection failures, and trailing weekly chosen-to-pursue yield from persisted outcomes. Missing capture data is reported as unavailable, never estimated:
node dist/src/cli.js source-report --approval live-approval.json --days 30 --format markdownThe same report is attached to any scan JSON with --source-metrics.
ProofScout provides a unified CLI entry point via proofscout (or node dist/src/cli.js):
proofscout <command> [options]Run a local scan across authorized structured and social providers:
# JSON output
proofscout scan --profile remote-project-first --days 30 --approval approval.json
# Human-readable Markdown report
proofscout scan --profile remote-project-first --days 30 --approval approval.json --format markdownProofScout enforces strict gates to prevent regressions before activating providers or deploying rules:
Validates schema conformance, field types, and integrity constraints against labeled corpus fixtures:
proofscout validate --fixture test/fixtures/validated-local-scan-v1.jsonMeasures classification precision, recall, and F1 metrics:
proofscout evaluate --fixture test/fixtures/validated-local-scan-v1.jsonGenerates an auditable decision report partitioning candidates into shortlist, review, and rejected lanes:
proofscout review-report --fixture test/fixtures/validated-local-scan-v1.json --policy shortlist-readyValidates live-extracted opportunities against sanitized holdout truth labels across structured, visual, and social lanes:
proofscout holdout-gate --manifest test/fixtures/holdout/sanitized-holdout-v1.json --policy test/fixtures/holdout/holdout-policy-v1.jsonProviders are strictly locked behind cryptographically auditable, time-bounded approval files (ScopedProviderApproval):
allowedProviderIds: whitelist of enabled providers.allowedLanes: structured, social, visual.expiresAt: strict ISO-8601 expiration.requestBudget: maximum number of requests allowed.
Instantly revoke approval or disable specific providers:
# Revoke an entire approval
proofscout provider-revoke --approval approval.json --reason "Compliance audit"
# Revoke a specific provider
proofscout provider-revoke ats-greenhouse --approval approval.json --reason "Rate limit policy"
# Temporarily disable a provider or shorten expiry
proofscout provider-disable ats-greenhouse --approval approval.json
proofscout provider-disable --approval approval.json --expires-at 2026-09-10T00:00:00.000ZAtomically purge all raw evidence, derived opportunities, and bundles for a provider:
proofscout source-delete ats-greenhouse --format markdownTwo optional social-lane providers built on ticket 02's authorized seams. They activate only when the scoped approval names reddit-pullpush and/or x-ingest and the approval's allowedLanes includes social — kill-switch and revoke apply by provider id as usual. (Live since commit 8006b24: main() wires the social adapter whenever the approval opens the lane.)
- Reddit via PullPush mirror (
reddit-pullpush): direct Reddit JSON is blocked (403/302 as probed 2026-09-07), so polling goes through the Arctic Shift / PullPush submission mirror. Polite personal-use rates only: at most one request per subreddit per 20 seconds (REDDIT_PULLPUSH_MIN_INTERVAL_MS) with exponential backoff on 429/5xx through the sharedRateLimiter. Every record enters as a lead snippet (isLeadSnippet: true, confidence capped,hiring-signalunless a concrete route is present). Removed/deleted bodies keep the title with a[post body removed or unavailable; title only]marker. - X via operator file ingest (
x-ingest): X search is login-walled with no public endpoint, so ProofScout never calls X itself. The operator runs the/last30daysskill separately and supplies its saved raw Markdown file;parseLast30DaysFile/loadLast30DaysFileturnAll Items by Sourceentries (X, Reddit, HN) into lead-snippet records with original URLs and attribution preserved. Malformed files yield zero records, never a scan failure.
Record candidate progression without mutating immutable evidence. Outcomes adjust ranking transparently (ignored: -20, rejected: -30, contacted: -10, interview: +15, hired: +20, researching: +5):
# Record an outcome
proofscout outcome-set opp-123 interview --note "Passed technical review"
# Record a typed reason; omitting --reason on a later update clears it
proofscout outcome-set opp-124 rejected --reason wrong-country --note "US-only fence"
# Generate a review-only proof-of-work kit for a stored opportunity
proofscout kit opp-123 --format markdown
# List recorded outcomes
proofscout outcome-list --format markdownRun scheduled scans with automatic opportunity diffing (new-opportunity, route-changed, compensation-changed, deadline-passed, risk-flag-added):
proofscout schedule-run --days 30 --approval approval.json --format markdownSee docs/scheduler.md for configuring crontab, systemd timers, or Windows Task Scheduler.
Alerts page only for new or materially changed shortlist-ready opportunities (status === "verified", actionable route, confidence >= 0.7, no risk flags). needs_review noise is suppressed:
proofscout alerts-list --format markdown
proofscout diffs-list --format markdownTrack strategic companies, founder handles, or high-conviction keywords. Watchlist matches lacking reachable application routes enter the review queue with evidence as hiring-signal, never as qualified opportunities:
proofscout watchlist-add company "Acme Corp"
proofscout watchlist-add author "@founder_alice"
proofscout watchlist-add keyword "founding engineer"
proofscout watchlist-list
proofscout watchlist-remove "Acme Corp"- ADR 0001: Personal Radar First
- ADR 0002: Evidence-First Human Gate
- ADR 0003: Hybrid Source and Visual Lane
- ADR 0004: Precision and Decision Coverage
- ADR 0005: Shortlist Readiness and Gate
- ADR 0006: Two-Gate Provider Activation
- ADR 0007: Human-Supplied Real-Source Holdout
- ADR 0008: Sanitized Holdout Storage
- ADR 0009: Configurable Holdout Gate
- ADR 0010: Two-Pass Holdout Label Review
- ADR 0011: Holdout Gate CLI Interface
- ADR 0012: Policy-Bound Holdout Manifest
- ADR 0013: Human Approval Before Provider Activation
- ADR 0014: Scoped Expiring Provider Approval
- ADR 0015: Provider Approval Lifecycle
- ADR 0016: Strict Compensation Policy for Non-USD and Non-Annual Ranges
- ADR 0017: Banding Bounded Annual USD Cash Ranges by Midpoint
- ADR 0018: Relevance Gate and Focus Keywords for the Live Shortlist
- Scheduling Guide