Open-source, 100% reproducible AI Agent Runtime Security Benchmark & Sandbox Environment (RFC-010 Draft Protocol).
-
Updated
Aug 6, 2026 - HTML
Open-source, 100% reproducible AI Agent Runtime Security Benchmark & Sandbox Environment (RFC-010 Draft Protocol).
Adversarial security benchmark for agent authorization: does a compromised agent's policy-violating proposal become an unauthorized external effect? 73 trials, nine families, an independent oracle, per-mechanism ablation, confidence intervals. 0 unauthorized effects in 61 attack trials (95% CI [0.0%, 5.9%]). Reproduction is partial.
Open deterministic security tests for unsafe multi-agent handoffs and authority escalation.
Open-source benchmark for adversarial evidence attacks on LLM-based cybersecurity auditors, targeting ACM AsiaCCS 2027.
Vendor-neutral benchmark measuring how MCP security proxies/gateways DEFEND against 22+ attack vectors — crosswalked to NIST AI RMF & OWASP LLM/Agentic Top 10. CI-gated, reproducible, DOI-cited. Submit your tool to the leaderboard.
Deterministic security benchmark for tool-using AI agents
Open AI-for-security validation benchmark: non-LLM scorer + a SOTA-validation loop. Labeled positive corpus withheld pending coordinated disclosure.
Cross-framework AI agent red-teaming benchmark. Tests LangChain, CrewAI, AutoGen, LlamaIndex & OpenAI Agents SDK for prompt injection, scope violations, and jailbreaks against a shared, OWASP ASI-aligned attack payload set.
Local-first workbench to run, inspect, compare, report, and gate OpenAI Codex Security scans.
GitHub action for Maester
ReplayBench-IoT: reproducible IoT replay-defense benchmark with Monte Carlo sweeps, CI, static demo, and hardware-validation artifacts.
FreightSkillBench is a reproducible benchmark for evaluating document-to-transaction integrity, prompt-injection risk, and security controls in AI-enabled shipping and logistics workflows.
Product-security LLM benchmark harness for realistic AppSec, supply-chain, and LLM application security evaluations.
Reproducible benchmark for smart-contract security tools, measuring precision, recall, and false positives against executable PoCs and versioned ground truth.
The core repository for the Maester module with helper cmdlets that will be called from the Pester tests.
Add a description, image, and links to the security-benchmark topic page so that developers can more easily learn about it.
To associate your repository with the security-benchmark topic, visit your repo's landing page and select "manage topics."