Skip to content

Harden json parsing - #41

Merged
ChrisJefferson merged 7 commits into
masterfrom
harden-json-parsing
Sep 26, 2026
Merged

ChrisJefferson merged 7 commits into
masterfrom
harden-json-parsing

Conversation

@ChrisJefferson

@ChrisJefferson ChrisJefferson commented Sep 26, 2026 •

Copy link
Copy Markdown
Member

This significantly overhauls our testing.

It adds bounds checking, deals with UTF-8 correctly.

Also we previously outputted 'inf' and 'nan' but wouldn't parse them. Now we follow Python's 'unofficial standard' on handling these numbers, and also parse them back in (and the old GAP ones just in case anyone made them and now has files they can't read).

Also fix that we outputted some floating point numbers in non-valid format.

Closes #4

@ChrisJefferson

Copy link
Copy Markdown
Member Author

Replaces and does significantly better than #40

@codecov

codecov Bot commented Sep 26, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 97.15808% with 16 lines in your changes missing coverage. Please review.
✅ Project coverage is 97.25%. Comparing base (5207534) to head (43e0863).

Files with missing lines Patch % Lines
gap/parse.gi 96.49% 14 Missing ⚠️
gap/json.gi 96.29% 1 Missing ⚠️
gap/utf8.gi 99.12% 1 Missing ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##           master      #41      +/-   ##
==========================================
+ Coverage   96.50%   97.25%   +0.74%     
==========================================
  Files           6       10       +4     
  Lines         143      801     +658     
==========================================
+ Hits          138      779     +641     
- Misses          5       22      +17     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

ChrisJefferson and others added 5 commits September 26, 2026 10:24
Limit nesting to 1024 containers, validate raw UTF-8 in parsed strings, and use the same strict UTF-8 byte rules when serialising GAP strings. Keep the existing permissive number grammar and Latin-1 fallback for malformed GAP byte strings.

Also make parse errors stable and fix trailing-data checks for embedded NUL bytes.

Co-Authored-By: Codex GPT-5 <noreply@anthropic.com>
Normalise GAP's finite float spellings to the JSON number grammar and validate every finite spelling before writing it.

Use Python's NaN and Infinity spellings for non-finite output, accept those spellings on input, and retain support for the lowercase spellings emitted by earlier json releases.

Co-Authored-By: Codex GPT-5 <noreply@anthropic.com>
Import the upstream parsing corpus and record the current parser's deliberately lenient cases. Adapt the test work from #39 so it can land independently of the pure GAP implementation and address #4 first.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Codex GPT-5 <noreply@anthropic.com>
Expect malformed UTF-8 and trailing NUL data to be rejected, while retaining the deliberately generous number grammar and Python-compatible non-finite extensions. Assert that all 316 corpus files are present.

Co-Authored-By: Codex GPT-5 <noreply@anthropic.com>
Retain only the constant-time normalisation required for GAP float spellings, and move the JSON number grammar checker into the tests. Reject non-real float objects explicitly.

On a 200,000-float list this reduces GapToJsonString time from 514-519 ms to 285-292 ms in interleaved runs, with identical output.

Co-Authored-By: Codex GPT-5 <noreply@anthropic.com>
Make the kernel extension optional by adding pure GAP parsing and string escaping backends. Keep the existing GAP serialisation methods as the shared output implementation, with the kernel continuing to accelerate common objects when available.

Match the hardened parser contract for UTF-8, nesting, generous numbers, non-finite values and errors. Compare both parsers across JSONTestSuite and run the full suite in CI without building the kernel extension.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Co-Authored-By: Codex GPT-5 <noreply@anthropic.com>
Include the two 100,000-level malformed inputs that previously overflowed the parser stack. The 1,024-level nesting limit now rejects both safely in the kernel and pure GAP implementations.\n\nRun all 318 upstream corpus files in the conformance test.\n\nCo-Authored-By: Codex GPT-5 <noreply@anthropic.com>
This was referenced Sep 26, 2026
@ChrisJefferson
ChrisJefferson merged commit 10d3cab into master Sep 26, 2026
7 checks passed
@ChrisJefferson
ChrisJefferson deleted the harden-json-parsing branch September 26, 2026 12:06
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Import or otherwise use testsuite from "Parsing JSON is a Minefield 💣"

2 participants