Appearance
jsonfast — the JSON reader
Every dialect in aowlparser keeps every byte, because a rewriting front end must. That is the wrong shape for reading a large document, so src/jsonfast.nim is the opposite trade: throw the whitespace away and go fast.
Measured
Best of 30, DOM-building, parse-only, on real documents:
| reader | 9.9MB catalog | 1.5MB source index | 1.3MB protocol |
|---|---|---|---|
| jsonfast (borrowed, zero copy) | 2343 MB/s | 2872 MB/s | 1964 MB/s |
| jsonfast (reused parser) | 1784 MB/s | 2383 MB/s | 1710 MB/s |
| jsonfast (one-off) | 1698 MB/s | 2249 MB/s | 1654 MB/s |
V8 JSON.parse (node 25) | 621 MB/s | 775 MB/s | 573 MB/s |
CPython json (C accelerated) | 225 MB/s | 270 MB/s | 312 MB/s |
aowljson (ref tree) | 167 MB/s | 343 MB/s | 199 MB/s |
tests/json/bench.sh runs that table on your machine. That is 3.4–3.8× V8 and 6.3–10.6× CPython on the zero-copy path, and never below 1.6 GB/s on any shape.
Three rows, because they answer three different questions and quoting only the best one would be marketing. borrowed (parseBorrowed) copies nothing and is the row comparable to simdjson, which borrows the same way — and carries the same hazard: the document and its views die with the caller's buffer. reused is what a server does: newJsonDoc once, parseInto per document. one-off is what a script does.
It is still not a simdjson clone. simdjson builds a two-stage SIMD structural index over the entire document before parsing anything: a different architecture, not a tuning gap.
What that rewrite would actually buy — measured before writing it
Stage 1 of that design is a SIMD sweep that writes the offset of every structural byte into an index, and it is a hard ceiling on the whole approach, so it was measured first as a standalone C program:
| document | stage 1 alone | structural bytes |
|---|---|---|
| 9.9MB catalog | 2895 MB/s | 7.6% |
| 1.5MB source index | 2428 MB/s | 8.3% |
| 1.3MB protocol | 3071 MB/s | 8.3% |
Stage 2 then walks that index — roughly 8% of the bytes — and builds the tape. Combining the two serially puts the whole design in the neighbourhood of 1.6–1.7 GB/s on the catalog against 1430 today: about +20%, for a rewrite of the entire front half and a much subtler correctness story (quote state and escape runs have to be reconstructed from the index rather than seen directly).
That is a real option, not a closed door — but it is worth knowing the size of the prize before spending the week, which is why the number is here.
Where the speed comes from
- A flat tape, not a tree of refs. One 16-byte entry per value in one
seq: no allocation per value, no pointer chase. A container stores the index one past its last descendant, so skipping a 10MB sub-object isi = node.next. - Zero-copy strings, lazy numbers. A string is an offset and a length until someone asks for it; a number is its lexeme until someone wants its value. Most values in a large document are never read at all.
- No recursion. Depth is an explicit bounded stack, so
[[[[…]]]]is a named error rather than the stack overflow recursive-descent JSON parsers are famous for — and that one is chosen by attackers on purpose. - A pre-sized tape. Growing by doubling copies the whole tape each time; on 10MB that was 9% of total runtime spent in
memcpybefore the parse had learned anything. - A jump table for the parser state, worth another 2%: an if-chain makes every token pay, and whichever order you choose penalises one document shape.
- Not copying the document — worth 17–28%, the largest single win, and invisible until a profile put
memcpyandmemsetnext to each other. Input is taken bysinksoparseInto(p, readFile(path))moves, andparseBorrowedskips ownership entirely. - Reading through a byte pointer rather than a string — another +11%, with one code path for owned and borrowed bytes alike.
- A hand-grown tape instead of
seq.add— worth +43%, and the reason profiling beat intuition. nimony'sseq.addasks the allocator about the block on every append: 26% of all instructions were mimalloc bookkeeping, about 390 instructions per value for what should be a 16-byte store. - SIMD scanning (
src/jsonfast_simd.c) for the two loops that consume nearly every byte — whitespace, and the run to a string's next",\or control character — sixteen bytes at a time. +7% on the object-heavy catalog, +67% on the string-heavy index.
The C file, and why it is there
This is the one part of aowlparser that is not pure nimony. nimony rejects addr s[i], so neither SIMD nor the word-at-a-time trick that approximates it can be written in the language today — filed as an aowlsem requirement rather than accepted as a design choice.
The C functions are pure scanners: they find the next interesting byte and return its offset. Every grammar decision stays in the nimony source, so the C file cannot disagree with the parser about JSON — it can only be wrong about where the next quote is, and 494k prefix comparisons would say so. -d:jfPure compiles the scalar loops instead: slower, identical answers, and the gate is run both ways.
Two optimisations that did not pay, recorded because a plausible-sounding one that loses is worth more than a guess. A byte-classification TABLE for "is this whitespace / does this end a string", replacing the compare chains, cost 12% — the compares were already cheap and the table added a load. A parallel is the open container an object stack, meant to avoid re-reading the container's tape node, cost 6% (the node is written recently enough to still be in cache, so the miss it avoided was not happening). And var s = src — the obvious spelling — memcpy'd the whole document before parsing a byte of it.
Correctness is measured, not asserted
A reader has no round-trip to hide behind: it discards whitespace by design, so byte-exactness cannot check it at all. What can is an outside implementation. tests/json/tfast.nim holds jsonfast to CPython's json module on:
- every
.jsonfile on the dev machine — 10,029 of them, compared on value counts per kind, an FNV digest of every decoded string, and the exact sum of every integer; - 494,323 sampled PREFIXES of those files, compared on accept/reject.
602,466 checks, all agreeing.
The prefixes are the important half. A corpus of valid documents only proves a reader is permissive enough; a prefix is malformed in a different way each time, and accept/reject agreement over half a million of them is what catches the dangerous direction — accepting what is not JSON.
Strictness is RFC 8259, which is stricter than CPython: NaN and Infinity are not JSON, so the oracle is configured to reject them rather than letting the reader inherit CPython's extension. Two things the gate learned to say out loud: files that are not UTF-8 (neither side is truth there), and files that changed on disk mid-run — a real machine's .json is full of live state, and one such file rewrote itself during a sweep and made a correct parser look wrong by two integers.
Using it
nim
import jsonfast
let doc = parse(src)
if not ok(doc):
echo doc.err, " at offset ", doc.errPos # errors are values, never raised
let v = view(doc)
echo v{"user"}{"name"}.str("") # chain-safe, allocation-free
echo v{"tags"}.at(0).str("")
for tag in v{"tags"}.items: echo tag.str("")
for k, val in v.pairs: echo k.str(""), " = ", val.num(0)A missing key yields an invalid view, and every accessor returns its default for one — so a chain through absent data cannot fault and allocates nothing on the way.
On the command line
sh
aowlparser jsonlint config.json # RFC 8259 strict; exit 1 + line:col
aowlparser jsonq api.json user.name # extract one field
aowlparser jsonq api.json 'items[2].id' # dot-separated keys, [n] indicesjsonlint is deliberately stricter than most linters: NaN, Infinity, trailing commas and comments are not JSON, and a tool that quietly accepts them is how they reach a file someone else's parser has to read. The exit codes are the contract — 0 valid, 1 invalid, 2 misuse — and tests/robust.sh pins them, because a command no gate runs is a command that rots.
jsonq prints a scalar as itself (1.5e3 comes back 1.5e3, not 1500.0 — numbers keep their spelling) and a container as its size; dumping a subtree is render's job.
With aowljson
nim
import jsonfast_aowljson
var err = ""
let v = parseJsonFast(src, err) # drop-in for aowljson.parseJsonThe drop-in is not the fast path, and the docs should say so. Building a ref-per-value tree with a string copy per value costs roughly six times the parse itself, so parseJsonFast lands near aowljson.parseJson however fast the scanner is. If you want the throughput above, stay on the tape and use views.
Building the tree did surface a real defect in aowljson, since fixed there: []= rescans every existing key, making a parser that uses it quadratic in object width and silently collapsing the duplicate keys a faithful reader must keep. addPair is the parser's door.

