Skip to content

feat: JSON5 numbers, string escapes and property names - #89

Merged
dsherret merged 4 commits into
dprint:mainfrom
leemr:json5-conformance
Sep 27, 2026
Merged

dsherret merged 4 commits into
dprint:mainfrom
leemr:json5-conformance

Conversation

@leemr

@leemr leemr commented Sep 25, 2026

Copy link
Copy Markdown
Contributor

Summary

This adds JSON5 syntax that the parser rejects today. After it, every valid file in the json5/json5-tests corpus parses, except todo/unicode-escaped-unquoted-key.json5, which the corpus itself files under todo/.

JSON5 syntax Example Switch, default true
Leading or trailing decimal point .5, -.5, 5., 5.e3 allow_bare_decimal_point_numbers (new)
Infinity and NaN Infinity, -Infinity, NaN allow_non_finite_numbers (new)
String escapes "it\'s", '\x41', '\v', '\0', \ then a line break allow_extended_string_escapes (new)
$ and keywords in unquoted property names { $id: 1 }, { true: 1 } allow_loose_object_property_names (existing)
A lone \r ends a line comment // note\r1 none

Each commit is one area and passes the tests alone.

Behavior notes

  • serde_json cannot hold Infinity or NaN. The serde deserializer reports them as "Number is out of range", the same as 1e400. The AST and CST conversions use the existing fallback to a string.
  • In the serde_json conversion, .5 gives 0.5, and 5. gives the float 5.0. The AST and the CST keep the raw text.
  • A lone \r is a line terminator in JSON5 and in JavaScript, and Microsoft jsonc-parser also ends a comment there. Error positions now count it as a line break. The CST stores it as whitespace, so the text round-trips.
  • JSON5 also ends a line comment at U+2028 and U+2029, but Microsoft jsonc-parser does not. This keeps the current behavior there, so the two parsers agree on where such a comment ends.
  • [true-1] still parses as [true, -1] when commas can be missing. Infinity_count and NaN-key still scan as one word. true$ now also scans as one word, where 0.33.2 gave an error.
  • With allow_extended_string_escapes: false, the old errors for \' in a double-quoted string and \" in a single-quoted string stay.

Breaking change

  • ParseOptions and ScannerOptions get three new pub fields, so an exhaustive struct literal no longer compiles.
  • ParseErrorKind gets BareDecimalPointNumbersNotAllowed and NonFiniteNumbersNotAllowed.
  • ParseStringErrorKind gets ExpectedTwoHexDigits.

This is the same kind of change as #61 (0.28.0) and #63 (0.29.0), so it needs a 0.x minor release. To check the semver breaks:

cargo semver-checks check-release --baseline-version 0.33.2 --all-features

Downstream: dprint-plugin-json

dprint-plugin-json 0.24.0 uses ParseOptions::default(), and gen_string_lit rewrites string text without the new escapes. When it moves to this release, it would change string values:

input:  "line\<LF>cont"    value: linecont
output: "line\\ncont"      value: line\ncont

'say \"hi\"' gives output that does not parse. The dprint-plugin-json test suite still passes against this branch, so its tests do not catch this.

There are two fixes:

  1. Teach gen_string_lit the new escapes. This does not need the new switch, so it can merge before this release.
  2. Pass allow_extended_string_escapes: false. This needs this release.

I can send a PR for option 1.

Testing

  • cargo test with --features serde, --features preserve_order, --all-features and --release --all-features, on each commit.
  • Every file in json5/json5-tests (ceb24d4) through CstRootNode::parse. Each parsed file round-trips byte for byte.
  • The decoded values of the new escapes and of the finite new numbers match the json5 crate.
  • The three new switches do not cover all changes. With them set to false, these still apply:
    • $ and keyword property names
    • a lone \r after a line comment
    • the error kind or position for some invalid input

leemr and others added 4 commits September 25, 2026 15:55
… NaN

Numbers like `.5`, `5.` and `5.e3` now parse, and so do `Infinity`,
`-Infinity`, `+Infinity`, `NaN` and `-NaN`. Two new `ParseOptions`
switches gate them: `allow_bare_decimal_point_numbers` and
`allow_non_finite_numbers`. Both default to `true`, like the other
JSON5 switches.

The `serde_json` conversion turns `.5` into `0.5` and `5.` into the
float `5.0`. `serde_json` cannot hold `Infinity` or `NaN`, so the
deserializer reports them as out of range, and the AST and CST
conversions keep their fallback to a string.

A property name like `Infinity_count` or `NaN-key` still scans as one
word, and `[true-1]` still parses as `[true, -1]` when commas can be
missing.
Strings now accept the JSON5 escapes: `\'` in a double-quoted string,
`\"` in a single-quoted string, `\v`, `\0`, `\xHH`, a backslash before
a line terminator (a line continuation), and a backslash before any
other character except a digit, which gives that character. A new
`ParseOptions` switch, `allow_extended_string_escapes`, gates them and
defaults to `true`. With it off, the old errors for `\'` and `\"` stay.

A line comment now also ends at a lone `\r`, which is a line terminator
in JSON5 and in JavaScript, and error messages count it as a line
break. U+2028 and U+2029 do not end a line comment, so this crate keeps
agreeing with other JSONC parsers. The CST keeps a lone `\r` as
whitespace, so the text round-trips unchanged.
JSON5 property names follow ECMAScript identifiers. So `$` can appear
anywhere, and the reserved words `true`, `false` and `null` are valid
names: `{ $id: 1 }` and `{ true: 1 }` now parse, in both the AST and
the serde parsers. `allow_loose_object_property_names` already gates
unquoted names, so this adds no new switch. A name that starts with a
keyword, like `true$` or `NaN$`, now scans as one word.
A lone `\r` now ends a line comment, but the CST stored it as whitespace,
so edits like `append` and sorting treated the comment as still open and
could place a new value inside the comment.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟢 Approval recommended

No unresolved review issues remain, and the review indicates approval readiness.

Review effort: Lite
Findings: None

What changed in this PR

Adds configurable JSON5 number formats, extended string escapes, loose property names, and CR line-comment handling across parsing and conversion layers.

Changes:

  • Supports bare decimal and non-finite numbers.
  • Adds extended string escapes and loose object keys.
  • Updates scanner, parser, AST, serde, CST, and error handling.
File Description
src/​tokens.rs Recognizes loose property-name tokens.
src/​string.rs Decodes extended JSON5 escapes.
src/​serde.rs Supports serde conversions and tests.
src/​scanner.rs Scans new JSON5 syntax and CR comments.
src/​parser.rs Handles expanded object-key syntax.
src/​parse_to_value.rs Adds value parsing coverage.
src/​parse_to_ast.rs Propagates options and builds AST nodes.
src/​lib.rs Updates configuration documentation.
src/​errors.rs Adds error kinds and CR-aware positions.
src/​cst/​mod.rs Preserves formatting and converts new literals.
src/​common.rs Normalizes bare decimal numbers.
src/​ast.rs Converts JSON5 numbers to serde values.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

@dsherret dsherret left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks!

@dsherret
dsherret merged commit 4aff55d into dprint:main Sep 27, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants