feat: JSON5 numbers, string escapes and property names - #89
Merged
Merged
Conversation
… NaN Numbers like `.5`, `5.` and `5.e3` now parse, and so do `Infinity`, `-Infinity`, `+Infinity`, `NaN` and `-NaN`. Two new `ParseOptions` switches gate them: `allow_bare_decimal_point_numbers` and `allow_non_finite_numbers`. Both default to `true`, like the other JSON5 switches. The `serde_json` conversion turns `.5` into `0.5` and `5.` into the float `5.0`. `serde_json` cannot hold `Infinity` or `NaN`, so the deserializer reports them as out of range, and the AST and CST conversions keep their fallback to a string. A property name like `Infinity_count` or `NaN-key` still scans as one word, and `[true-1]` still parses as `[true, -1]` when commas can be missing.
Strings now accept the JSON5 escapes: `\'` in a double-quoted string, `\"` in a single-quoted string, `\v`, `\0`, `\xHH`, a backslash before a line terminator (a line continuation), and a backslash before any other character except a digit, which gives that character. A new `ParseOptions` switch, `allow_extended_string_escapes`, gates them and defaults to `true`. With it off, the old errors for `\'` and `\"` stay. A line comment now also ends at a lone `\r`, which is a line terminator in JSON5 and in JavaScript, and error messages count it as a line break. U+2028 and U+2029 do not end a line comment, so this crate keeps agreeing with other JSONC parsers. The CST keeps a lone `\r` as whitespace, so the text round-trips unchanged.
JSON5 property names follow ECMAScript identifiers. So `$` can appear
anywhere, and the reserved words `true`, `false` and `null` are valid
names: `{ $id: 1 }` and `{ true: 1 }` now parse, in both the AST and
the serde parsers. `allow_loose_object_property_names` already gates
unquoted names, so this adds no new switch. A name that starts with a
keyword, like `true$` or `NaN$`, now scans as one word.
A lone `\r` now ends a line comment, but the CST stored it as whitespace, so edits like `append` and sorting treated the comment as still open and could place a new value inside the comment.
There was a problem hiding this comment.
Copilot review overview
🟢 Approval recommended
No unresolved review issues remain, and the review indicates approval readiness.
Review effort: Lite
Findings: None
What changed in this PR
Adds configurable JSON5 number formats, extended string escapes, loose property names, and CR line-comment handling across parsing and conversion layers.
Changes:
- Supports bare decimal and non-finite numbers.
- Adds extended string escapes and loose object keys.
- Updates scanner, parser, AST, serde, CST, and error handling.
| File | Description |
|---|---|
src/tokens.rs |
Recognizes loose property-name tokens. |
src/string.rs |
Decodes extended JSON5 escapes. |
src/serde.rs |
Supports serde conversions and tests. |
src/scanner.rs |
Scans new JSON5 syntax and CR comments. |
src/parser.rs |
Handles expanded object-key syntax. |
src/parse_to_value.rs |
Adds value parsing coverage. |
src/parse_to_ast.rs |
Propagates options and builds AST nodes. |
src/lib.rs |
Updates configuration documentation. |
src/errors.rs |
Adds error kinds and CR-aware positions. |
src/cst/mod.rs |
Preserves formatting and converts new literals. |
src/common.rs |
Normalizes bare decimal numbers. |
src/ast.rs |
Converts JSON5 numbers to serde values. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
3 tasks done
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
This adds JSON5 syntax that the parser rejects today. After it, every valid file in the json5/json5-tests corpus parses, except
todo/unicode-escaped-unquoted-key.json5, which the corpus itself files undertodo/.true.5,-.5,5.,5.e3allow_bare_decimal_point_numbers(new)Infinity,-Infinity,NaNallow_non_finite_numbers(new)"it\'s",'\x41','\v','\0',\then a line breakallow_extended_string_escapes(new)$and keywords in unquoted property names{ $id: 1 },{ true: 1 }allow_loose_object_property_names(existing)\rends a line comment// note\r1Each commit is one area and passes the tests alone.
Behavior notes
serde_jsoncannot holdInfinityorNaN. The serde deserializer reports them as "Number is out of range", the same as1e400. The AST and CST conversions use the existing fallback to a string.serde_jsonconversion,.5gives0.5, and5.gives the float5.0. The AST and the CST keep the raw text.\ris a line terminator in JSON5 and in JavaScript, and Microsoft jsonc-parser also ends a comment there. Error positions now count it as a line break. The CST stores it as whitespace, so the text round-trips.[true-1]still parses as[true, -1]when commas can be missing.Infinity_countandNaN-keystill scan as one word.true$now also scans as one word, where 0.33.2 gave an error.allow_extended_string_escapes: false, the old errors for\'in a double-quoted string and\"in a single-quoted string stay.Breaking change
ParseOptionsandScannerOptionsget three new pub fields, so an exhaustive struct literal no longer compiles.ParseErrorKindgetsBareDecimalPointNumbersNotAllowedandNonFiniteNumbersNotAllowed.ParseStringErrorKindgetsExpectedTwoHexDigits.This is the same kind of change as #61 (0.28.0) and #63 (0.29.0), so it needs a 0.x minor release. To check the semver breaks:
Downstream: dprint-plugin-json
dprint-plugin-json 0.24.0 uses
ParseOptions::default(), andgen_string_litrewrites string text without the new escapes. When it moves to this release, it would change string values:'say \"hi\"'gives output that does not parse. The dprint-plugin-json test suite still passes against this branch, so its tests do not catch this.There are two fixes:
gen_string_litthe new escapes. This does not need the new switch, so it can merge before this release.allow_extended_string_escapes: false. This needs this release.I can send a PR for option 1.
Testing
cargo testwith--features serde,--features preserve_order,--all-featuresand--release --all-features, on each commit.CstRootNode::parse. Each parsed file round-trips byte for byte.json5crate.false, these still apply:$and keyword property names\rafter a line comment