-
Notifications
You must be signed in to change notification settings - Fork 0
Expand file tree
/
Copy pathindex.html
More file actions
47 lines (47 loc) · 18.3 KB
/
Copy pathindex.html
File metadata and controls
47 lines (47 loc) · 18.3 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
<!DOCTYPE html>
<html lang="en"><head><meta charset="utf-8"><meta name="viewport" content="width=device-width,initial-scale=1">
<title>Codex Security: How the Validation Loop Works</title><meta name="description" content="Codex Security builds a threat model per repository, reproduces findings in a sandbox before showing them, and proposes patches a human must approve. The details.">
<link rel="canonical" href="https://codex-security.github.io/"><meta name="google-site-verification" content="znEyS88JQxAhHU2PJdN6fnyv1wuy4tjTAfUs56Bswao" /><meta name="msvalidate.01" content="E090F48924D613CE22AA152FD6CDC4E3">
<meta property="og:title" content="Codex Security: How the Validation Loop Works"><meta property="og:description" content="Codex Security builds a threat model per repository, reproduces findings in a sandbox before showing them, and proposes patches a human must approve. The details."><meta property="og:url" content="https://codex-security.github.io/">
<script type="application/ld+json">[{"@context": "https://schema.org", "@type": "FAQPage", "mainEntity": [{"@type": "Question", "name": "Who can use codex security today?", "acceptedAnswer": {"@type": "Answer", "text": "It is a research preview available to ChatGPT Enterprise, Edu, Business and Pro users. Enterprise and Edu workspaces additionally control access through workspace permissions, and both Codex Cloud and Codex Security must be enabled before anyone in the workspace can connect a repository."}}, {"@type": "Question", "name": "Will it change my code without asking?", "acceptedAnswer": {"@type": "Answer", "text": "No. The documentation states plainly that the patch does not automatically modify your code. Validated findings produce a minimal patch suggestion that is surfaced for human review, and you decide whether it becomes a pull request in your normal workflow."}}, {"@type": "Question", "name": "How does it decide what to look at?", "acceptedAnswer": {"@type": "Answer", "text": "Through a threat model it builds from your repository and commit history, capturing attacker entry points, trust boundaries, sensitive data and high-impact code paths. Teams can inspect and edit that model, which changes what subsequent analysis treats as realistic. Correcting it early is the highest-leverage thing a reviewer can do with the tool."}}, {"@type": "Question", "name": "What stops it flooding us with false positives?", "acceptedAnswer": {"@type": "Answer", "text": "A validation stage before anything is shown. An automated validator tries to reproduce each issue in an isolated environment, recording execution details and proof-of-concept artifacts, so findings that could not be demonstrated do not consume your reviewers' attention. What arrives carries execution details, which means a reviewer starts from evidence rather than from a line number and a severity label."}}, {"@type": "Question", "name": "Why is the first scan so slow?", "acceptedAnswer": {"@type": "Answer", "text": "Because it is doing two jobs at once: building the project's threat model and scanning repository history for issues that already exist. The documentation notes this takes longer on large projects and that scans of new code afterwards are faster."}}, {"@type": "Question", "name": "Does Begin.sh do anything security related?", "acceptedAnswer": {"@type": "Answer", "text": "No, and it should not be presented that way. It generates a static site or Expo app from a prompt or a URL and returns a zip, with no hosting, backend or authentication attached. The only honest connection is that a static artifact with no server has less to review."}}]}]</script><style>
:root{--ink:#16191d;--muted:#5b6470;--line:#e3e6ea;--accent:#2563eb;--soft:#f5f7fa;--ok:#0f9960}*{box-sizing:border-box}
body{margin:0;font:16px/1.7 -apple-system,BlinkMacSystemFont,"Segoe UI",Roboto,"Noto Sans",sans-serif;color:var(--ink)}
header{border-bottom:1px solid var(--line)}.nav{max-width:960px;margin:0 auto;padding:14px 20px;display:flex;justify-content:space-between;align-items:center;gap:12px}
.brand{font-weight:700;text-decoration:none;color:var(--ink)}.nav a.small{font-size:14px;color:var(--muted);text-decoration:none}
main{max-width:960px;margin:0 auto;padding:44px 20px 80px}h1{font-size:34px;line-height:1.2;margin:0 0 12px}h2{font-size:23px;margin:44px 0 12px}h3{font-size:17px;margin:0 0 6px}
.lead{font-size:18px;color:var(--muted);max-width:740px}
.verdict{background:var(--soft);border:1px solid var(--line);border-left:4px solid var(--accent);border-radius:10px;padding:18px 22px;margin:26px 0}
.verdict b{display:block;margin-bottom:6px}
.btn{display:inline-block;background:var(--accent);color:#fff;padding:13px 22px;border-radius:8px;font-weight:600;text-decoration:none;border:0;font-size:16px;cursor:pointer}
.ghost{display:inline-block;border:1px solid var(--line);padding:12px 20px;border-radius:8px;text-decoration:none;color:var(--ink)}
table{border-collapse:collapse;width:100%;margin:14px 0;font-size:15px}th,td{border-bottom:1px solid var(--line);padding:10px 8px;text-align:left;vertical-align:top}
th{background:var(--soft);font-weight:600}td.yes{color:var(--ok);font-weight:600}
.grid{display:grid;grid-template-columns:repeat(auto-fit,minmax(230px,1fr));gap:14px}.card{border:1px solid var(--line);border-radius:12px;padding:18px}
.steps{counter-reset:s;padding:0;list-style:none}.steps li{counter-increment:s;padding-left:44px;position:relative;margin:14px 0}
.steps li:before{content:counter(s);position:absolute;left:0;top:0;width:30px;height:30px;border-radius:50%;background:var(--accent);color:#fff;display:flex;align-items:center;justify-content:center;font-weight:700}
.tool{background:var(--soft);border:1px solid var(--line);border-radius:14px;padding:24px;margin:28px 0;max-width:760px}
.tool label{display:block;font-weight:600;margin-bottom:8px}.row{display:flex;gap:10px;flex-wrap:wrap}
.row input{flex:1;min-width:220px;padding:13px 14px;border:1px solid var(--line);border-radius:8px;font-size:16px}
.hint{font-size:13px;color:var(--muted);margin-top:10px}
details{border:1px solid var(--line);border-radius:8px;padding:10px 14px;margin:8px 0}summary{cursor:pointer;font-weight:600}
.cta-band{background:var(--soft);border-radius:14px;padding:32px;text-align:center;margin-top:48px}
footer{border-top:1px solid var(--line);color:var(--muted);font-size:13px;padding:20px;text-align:center;line-height:1.6}
@media(max-width:640px){h1{font-size:26px}main{padding:26px 16px 60px}table{font-size:14px}}
</style></head>
<body><header><div class="nav"><a class="brand" href="https://codex-security.github.io/">Review Loop</a>
<a class="small" href="https://github.com/codex-security/codex-security.github.io">Source on GitHub</a></div></header>
<main>
<h1>Codex Security: Inside the Scan, Validate, Patch Loop</h1><p class="lead">Codex security is organised as a closed loop rather than a scan: threat model, discovery, sandbox validation, minimal patch, human review, revalidation after merge.</p>
<div class='verdict'><b>Who should be reading this</b>Teams drowning in scanner output are the intended audience, because the design premise is that a finding should be reproduced before a human is asked to look at it. It is a research preview limited to ChatGPT Enterprise, Edu, Business and Pro accounts, it connects to GitHub repositories, and it never edits your code on its own: patches are proposals that a person turns into a pull request. The caveat is the one attached to every preview, that behaviour and availability can change under you. Begin.sh appears in the table below only as a contrast and does not review anything; what it offers is output with a smaller surface to review in the first place.</div>
<p><a class="btn" href="https://begin.sh?utm_source=github&utm_medium=ugc&utm_campaign=codex-security&utm_content=pages-hero&utm_term=tier-b" rel="noopener">See what it builds →</a>
<a class="ghost" href="https://openai.com" rel="nofollow noopener">Official site</a></p>
<h2>Why codex security is not described as a scanner</h2><p>OpenAI's help centre article makes the distinction early: it is designed to work more like a security researcher than a traditional scanner. Concretely that means reading code, running tests, exploring realistic attack paths, and proposing patches teams can handle in their normal workflow, rather than pattern-matching source against a rule set and emitting a list. The stated purpose is identifying, validating and remediating vulnerabilities in connected repositories. That word validating is doing most of the work in the sentence, and it is the reason the rest of the pipeline looks different from the tools most teams already ignore. A finding that has been reproduced is a different object from a finding that has been flagged.</p><h2>The threat model comes before the findings</h2><p>On connecting to a repository, it scans commits in reverse chronological order and builds a threat model specific to that codebase. The model captures attacker entry points, trust boundaries, sensitive data and high-impact code paths, and analysis is then focused through it so the tool looks at realistic attack scenarios rather than every theoretically reachable line. Teams can inspect that model and edit it so it matches their actual deployment assumptions, which is the part worth taking seriously. A threat model nobody corrects is a set of guesses about your architecture, and a tool that lets you correct it is asking for the context that turns a generic alert into a relevant one.</p><h2>Reproduction in a sandbox is the differentiator</h2><p>Before a finding is surfaced at all, an automated validator attempts to reproduce it in an isolated environment. It records reproduction results, execution details and proof-of-concept artifacts, so what reaches a human has already been shown to work rather than merely inferred from source. Alongside that, attack-path analysis traces how attacker-controlled input could travel from an entry point to a sensitive outcome, scores the path by likelihood and impact, and makes its underlying assumptions visible. Visible assumptions are unusual and valuable: they let a reviewer disagree with the reasoning rather than only with the conclusion, which is how a security queue stops being a list of things nobody can argue with.</p><h2>Patches are proposed, never applied</h2><p>For validated findings it generates a minimal patch aimed at the root cause. The documentation is emphatic that this does not automatically modify your code: the patch is surfaced for human review and can be turned into a pull request in your existing workflow. After a confirmed issue is patched and merged, it can revalidate the fix, which closes the loop from detection through to remediation. Minimal is a deliberate choice as well. A small diff scoped to the cause is reviewable by someone who did not write the original code, whereas a sweeping refactor offered as a security fix tends to sit unmerged until the finding is stale. Scope is a review-throughput decision as much as an engineering one, and this pipeline reads as though that was understood.</p><h2>Getting access, and the administrative prerequisites</h2><p>Start at chatgpt.com/codex/cloud/security, connect and enable the GitHub repositories you want covered, then wait out the first scan. Expect that one to be slow on a large project, because it builds the threat model and works through repository history before anything useful appears; scans of new code afterwards are faster. On the administrative side, Enterprise and Edu workspaces manage access through workspace permissions, and both Codex Cloud and Codex Security have to be enabled for the workspace. Access can be narrowed to particular roles or groups through role-based access control, including groups synchronised over SCIM. Sort that out before promising anyone a timeline. Enabling a preview feature across a workspace is a shorter conversation than retrofitting role restrictions after several teams have already connected repositories.</p>
<h2>The parts that matter</h2><div class="grid"><div class='card'><h3>Research preview, limited tiers</h3><p>Available to ChatGPT Enterprise, Edu, Business and Pro users. Preview status means the behaviour described here is current rather than settled, which is worth stating in any internal proposal.</p></div><div class='card'><h3>Validated before surfaced</h3><p>An automated validator reproduces each candidate issue in an isolated environment and captures proof-of-concept artifacts, so review time goes to findings that were demonstrated rather than merely suspected.</p></div><div class='card'><h3>Your code is not edited</h3><p>Patches are proposals. They are surfaced for human review and can become a pull request in your normal workflow, and the loop closes with revalidation after a fix is merged.</p></div><div class='card'><h3>Two switches, not one</h3><p>Enterprise and Edu workspaces need both Codex Cloud and Codex Security enabled, with access optionally limited to roles or SCIM-synced groups through role-based access control.</p></div></div>
<h2>Stage by stage</h2><table><tr><th>Stage or property</th><th>Codex Security</th><th>Begin.sh</th></tr><tr><td>Identification</td><td>Builds a codebase-specific threat model from source and commit history</td><td>Does no analysis of any kind; it generates code, it does not audit it</td></tr><tr><td>Discovery</td><td>Explores realistic code paths through that model to find candidate issues</td><td>Not applicable</td></tr><tr><td>Validation</td><td>Reproduces the issue in an isolated environment and records artifacts</td><td>You review the generated files yourself before using them</td></tr><tr><td>Remediation</td><td>Proposes a minimal patch at the root cause, without applying it</td><td>Delivers editable source in a downloadable zip</td></tr><tr><td>Human gate</td><td>Findings and patches go to review and can become pull requests</td><td>Nothing ships until you ship it</td></tr><tr><td>Revalidation</td><td>Re-checks a merged fix, closing the loop from detection to remediation</td><td>Regenerate and compare the output</td></tr><tr><td>Access</td><td>Research preview for Enterprise, Edu, Business and Pro; connects to GitHub</td><td>No repository access requested; output includes no backend or auth</td></tr></table>
<h2>Rolling it out without surprises</h2><ol class="steps"><li><strong>Clear the permissions first</strong><br>In Enterprise and Edu workspaces, confirm that both Codex Cloud and Codex Security are enabled, and decide which roles or SCIM-synced groups should have access before anyone tries to connect a repository.</li><li><strong>Connect one repository</strong><br>Go to chatgpt.com/codex/cloud/security and enable a single project first. A repository you know well gives you a way to judge whether the findings are worth the review time.</li><li><strong>Budget for the first scan</strong><br>The initial pass builds a threat model and works through repository history, so it takes longer on large projects. Later scans covering new code are faster, and that is the steady state to plan around.</li><li><strong>Correct the threat model</strong><br>Inspect what it inferred about entry points, trust boundaries and sensitive data, and edit it to match your real deployment. Every later finding is filtered through those assumptions.</li></ol>
<h2>What teams ask before enabling it</h2><details><summary>Who can use codex security today?</summary><p>It is a research preview available to ChatGPT Enterprise, Edu, Business and Pro users. Enterprise and Edu workspaces additionally control access through workspace permissions, and both Codex Cloud and Codex Security must be enabled before anyone in the workspace can connect a repository.</p></details><details><summary>Will it change my code without asking?</summary><p>No. The documentation states plainly that the patch does not automatically modify your code. Validated findings produce a minimal patch suggestion that is surfaced for human review, and you decide whether it becomes a pull request in your normal workflow.</p></details><details><summary>How does it decide what to look at?</summary><p>Through a threat model it builds from your repository and commit history, capturing attacker entry points, trust boundaries, sensitive data and high-impact code paths. Teams can inspect and edit that model, which changes what subsequent analysis treats as realistic. Correcting it early is the highest-leverage thing a reviewer can do with the tool.</p></details><details><summary>What stops it flooding us with false positives?</summary><p>A validation stage before anything is shown. An automated validator tries to reproduce each issue in an isolated environment, recording execution details and proof-of-concept artifacts, so findings that could not be demonstrated do not consume your reviewers' attention. What arrives carries execution details, which means a reviewer starts from evidence rather than from a line number and a severity label.</p></details><details><summary>Why is the first scan so slow?</summary><p>Because it is doing two jobs at once: building the project's threat model and scanning repository history for issues that already exist. The documentation notes this takes longer on large projects and that scans of new code afterwards are faster.</p></details><details><summary>Does Begin.sh do anything security related?</summary><p>No, and it should not be presented that way. It generates a static site or Expo app from a prompt or a URL and returns a zip, with no hosting, backend or authentication attached. The only honest connection is that a static artifact with no server has less to review.</p></details>
<div class="cta-band"><h2 style="margin-top:0">Less code on the server means less to review</h2><p>Begin.sh produces a static site or Expo app from a prompt or a URL to clone and hands you the source as a zip. It is not a security tool and makes no claim to be one, but output with no backend and no auth is output with a smaller surface.</p>
<a class="btn" href="https://begin.sh?utm_source=github&utm_medium=ugc&utm_campaign=codex-security&utm_content=pages-cta&utm_term=tier-b" rel="noopener">See what it builds →</a></div>
</main>
<footer>An independent page written by a practitioner, not by OpenAI, with no affiliation or endorsement implied; all product names and trademarks belong to their respective owners.<br>Maintained independently · <a href="https://begin.sh?utm_source=github&utm_medium=ugc&utm_campaign=codex-security&utm_content=pages-footer&utm_term=tier-b" rel="noopener" style="color:inherit">begin.sh</a></footer>
</body></html>