Skip to content
nihilai-collectivePublic

About

A few classes for extremely fast json parsing/serializing in modern C++. Possibly the fastest json parser in C++. Possibly the fastest json serializer in C++.

Topics

Resources

Stars

102 stars

Watchers

5 watching

Forks

Repository files navigation

Commit Activity License C++23

Jsonifier is fully RFC8259 compliant.

A high-performance C++ library for validating, serializing, parsing, prettifying, and minifying JSON data — very rapidly.

It achieves this through the usage of SIMD instructions as well as compile-time hash maps for efficient key lookups during parsing.


Compiler Support

Compiler Status
MSVC Visual Studio (Latest)
GCC GCC (Latest)
CLANG Clang (Latest)

Operating System Support

OS Status
Windows Windows (Latest)
Linux Ubuntu (Latest)
Mac macOS (Latest)

Android: JSONIFIER_PLATFORM_ANDROID is detected and handled distinctly from JSONIFIER_PLATFORM_LINUX (separate mmap/unistd.h wiring, its own entry in the supported-platform check), for NDK cross-compilation. It is not part of the CI matrix above, so treat it as best-effort rather than continuously verified.

CPU Architecture Support

Jsonifier automatically detects and optimizes for your CPU architecture:

  • x64 / AMD64 — 64-bit extension of x86 with enhanced memory addressing
  • AVX — 128-bit vector registers for SIMD operations
  • AVX2 — 256-bit vector registers with additional integer operations
  • AVX-512 — 512-bit vector registers for maximum parallelism (requires F + BW + VBMI2 support detected together)
  • PCLMULQDQ — carry-less multiplication support, detected and used where available
  • ARM-NEON — SIMD instructions for ARM processors
  • ARM-SVE2 — scalable vector extensions for ARM processors ⚠️ experimental — the SVE2 backend is new and still under active development; a handful of parsing cases are not yet handled correctly. NEON remains the recommended path for production ARM builds until SVE2 correctness is fully verified. Feedback and bug reports on SVE2-specific behavior are very welcome.

On x64, the configured tier is the highest one compiled in: the SIMD code is compiled once per tier (AVX, AVX2, AVX-512) into a single binary, and jsonifier_core<> picks the best tier the running CPU supports via cpuid/xgetbv on first use. A binary configured for AVX-512 also runs on AVX2 and AVX machines. See Runtime selection.

Manual configuration and cross-compilation are supported by pre-defining JSONIFIER_CPU_INSTRUCTIONS (plus JSONIFIER_SVE2_VECTOR_BITS for SVE2 targets) at CMake configure time to skip native feature detection.


Features

Compile-Time Reflection

Structures are registered via template specialization using member pointers as non-type template parameters. No macros, no code generation, no runtime type registry — the entire schema is known to the compiler, which means key lookups collapse into compile-time hash-map dispatch and dead branches get eliminated before they exist. Custom JSON key names (kebab-case, digit-prefixed, reserved words) are handled at compile time via makeJsonEntity<&member, "custom-name">() with zero runtime overhead.

SIMD-Accelerated Stage-1

Structural indexing runs through a batched-drain, fused-scan architecture: character classification, quote-scope tracking, and UTF-8 validation happen in the same pass over the input, with cross-SIMD-width validation state carried between blocks via a dedicated register-scoped validator. The structural index itself is stored as a compact array of offsets into the source buffer rather than raw pointers, keeping the index footprint small and cache-friendly. Everything the compiler can know statically about the target architecture — cache-line size, vector width, alignment — is baked into the binary as constexpr.

UTF-8 Validation, By Default

Full UTF-8 validation runs as part of every parse, always, built on a Lemire/simdjson-derived table-lookup validator with both a bulk-buffer path and a streaming, register-scoped path that carries state across chunk boundaries. Validation is fused directly into string scanning rather than requiring a separate pass. Correctness is backed by the full Markus Kuhn UTF-8 stress-test corpus, alignment-sweep fuzzing, and page-boundary overrun checks.

Purpose-Built Hash Maps

Key lookups during parsing use compile-time-generated hash maps specialized for object size (1, 2, 3+ fields), each picking a different strategy based on what's fastest for that cardinality. No runtime hashing, no bucket walks — the lookup is generated for the specific set of keys your struct declares.

Complete JSON Support

Full RFC8259 compliance. All types (objects, arrays, strings, numbers, booleans, null), full Unicode with proper surrogate-pair handling, all escape sequences, and jsonifier::raw_json_data for preserving arbitrary sub-trees verbatim.

Flexible Parsing Modes

  • Default parsing — single pass over the raw buffer, compile-time hash-map key dispatch, keys accepted in any order
  • Known-order parsing — tries the declaration-order key first, then a self-tuning per-position memo, then the hash map; adds a fused literal-match fast path for minified input
  • Partial reading — parse unordered or partial JSON structures
  • Arbitrary data — work with unknown JSON via raw_json_data
  • Generic (schema-free) parsing — lazy, On Demand-style jsonifier::generic::parser with re-readable values, out-of-order field access, JSON Pointers and iterateMany for NDJSON streams

Every parsing-family test in the suite is run across all eight combinations of partialRead, knownOrder, and nullTerminated.

Safety & Reliability

Comprehensive error reporting with source location tracking, backed by a compact, purpose-built status-code model — parsing, validation, minification, and prettification each propagate failure immediately on the first error rather than continuing to walk a broken buffer. Configurable maximum nesting depth (parse_options::maxDepth, default 1024) guards against malicious or runaway input.

Continuous integration runs AddressSanitizer and UndefinedBehaviorSanitizer on every push across every supported platform and compiler. The conformance test suite covers the full RFC8259 pass/fail battery — both the classic jsonchecker corpus and the full JSONTestSuite Y/N corpus — plus dedicated UTF-8 correctness tests including the full Markus Kuhn stress-test set, chunk/block-boundary crossings, unaligned-pointer sweeps, and an mmap page-boundary fault check.


CI/CD with unit-tests

Jsonifier uses GitHub Actions to continuously test across multiple platforms and compilers with sanitizers enabled. The test suite is built on rt-ut, fetched directly via CMake FetchContent, and runs on every push and pull request:

name: unit-tests
on:
  push:
    branches: [ "**" ]
  pull_request:
    branches: [ "**" ]
  workflow_dispatch:

jobs:
  build:
    strategy:
      fail-fast: false
      matrix:
        include:
          - os: ubuntu-latest
            compiler: clang
            name: "Ubuntu Clang"
          - os: ubuntu-latest
            compiler: gcc
            name: "Ubuntu GCC"
          - os: macos-latest
            compiler: clang
            name: "macOS Clang"
          - os: macos-latest
            compiler: gcc
            name: "macOS GCC"
          - os: windows-latest
            compiler: msvc
            name: "Windows MSVC"
    runs-on: ${{ matrix.os }}

This ensures memory safety and undefined behavior detection across all supported platforms. Note that sanitizers are automatically disabled for the GCC-on-macOS combination (unsupported), and UBSan has no effect under MSVC (no equivalent runtime) — Clang-cl or a Clang/GCC build is needed for UBSan coverage there.

The full suite passes on Ubuntu (Clang/GCC), macOS (Clang/GCC), and Windows (MSVC) on every push.


Comprehensive Test Suite

Jsonifier includes an extensive test suite that runs on every push across all supported platforms and compilers with ASAN and UBSAN enabled where supported.

Test Categories

Test Category Description
Conformance Tests Full RFC8259 compliance testing against two corpora, each wired up independently: conformance.hpp drives the classic jsonchecker set (77 of 79 fail*.json on disk + all 27 pass*.json, each checked against a specific expected parse_statuses error), and JSONTestSuite.hpp drives the JSONTestSuite Y/N corpus (all 188 n_*.json + all 95 y_*.json; the i_*.json implementation-defined cases are intentionally unused). 387 documents total, each run across all eight partialRead / knownOrder / nullTerminated combinations
Round-Trip Tests 27 serialize → parse → compare documents covering primitives, raw pointers, unique_ptr, and nested objects — 54 checks per configuration, run across the four partialRead / knownOrder combinations
Float Validation 64 edge cases including denormals, subnormal boundaries, round-half-to-even cases, infinities, and extreme exponents
Integer Validation Bounds testing for signed (24 pass / 11 fail) and unsigned (16 pass / 11 fail) integers, from zero through the full int64/uint64 range
String Validation 35 pass cases (Unicode, escape sequences, control characters, multi-byte emoji, ZWJ sequences, surrogate pairs) and 26 fail cases (malformed escapes, invalid \u sequences, unterminated strings)
UTF-8 Validation 202 tests: the full Markus Kuhn UTF-8 stress-test corpus, standalone 1–4 byte sequence and chunk/block-boundary tests, second-byte boundary tests per lead-byte class, fused string-parser validation (escape-aware, surrogate-pair-aware), an unaligned-pointer sweep, an unaligned invalid-sequence sweep, an mmap page-boundary fault check, and a width-transition sweep across body lengths 24–224
Bounds/Truncation Progressive truncation of full JSON payloads (Canada, CitmCatalog, Discord, Google Maps, Instruments, Marine IK, Mesh, Random, Twitter, Twitter Partial) — 10 corpus documents, minified and prettified (20 files), across all eight configurations (160 truncation runs), to validate graceful failure on malformed/truncated input
Parsing Tests Parse/serialize/minify/prettify/validate correctness across the full real-world payload suite, both minified and prettified, across all eight partialRead / knownOrder / nullTerminated combinations
Intrinsics Tests Direct correctness testing of the SIMD abstraction layer — comparison, bitmask, logical, saturating-subtract, shift, cross-register alignment, and load/store round-trips at every unaligned offset
Error & Core Tests 10 tests for construction, equality, line-number reporting, and control-character escaping of the error-reporting model, plus 11 jsonifier_core copy/move/self-assignment tests
Generic Parsing 75 tests for the schema-free On Demand-style parser: in-order, reverse-order, and missing-key access, JSON pointers, escaped and non-ASCII strings, invalid-UTF-8 rejection, and multi-document streams
raw_json_data 41 tests for type detection, int/uint/double round-trips, key and index access, contains, size, deep nesting, equality, and serialization (compact, prettified, and default-constructed)
Number Serialization 527 tests: 76 fastio digit-count/integer/float tests, 83 zmij double-serialization tests, and 368 digit-boundary tests for i_to_str and jsonifier::toString across all eight integer types
Type Coverage Primitives, containers (vector, array, map, unordered_map), tuples, optional, shared_ptr, enums, nested structs, renamed/escaped keys — 69 dedicated unit tests
Internal Containers & Utilities Direct correctness tests for the library's own building blocks — the fixed-size allocator, jsonifier::array, the tuple implementation and its iterator, the compile-time hash and hash-map generators, comparators and the string-literal comparator, enum-name and member-name reflection, the jsonifier::string class, stage-1 tape emission, and the minifier/prettifier/printer output paths — 624 dedicated unit tests

What Gets Tested

  • 69 dedicated unit tests covering reflection, renamed fields, optionals, enums, shared_ptr, nested structs, containers, tuples, maps, and escaped keys
  • 624 additional unit tests directly exercising the library's internal containers and utilities — allocator, jsonifier::array, tuple/iterator, compile-time hash and hash-map generation, comparators, enum/member-name reflection, jsonifier::string, stage-1 tape emission, and the minify/prettify/print output paths
  • 75 generic-parser tests, 41 raw_json_data tests, and 527 number-serialization tests
  • 7,030 assertions in total per platform (Windows MSVC, AVX2 backend), all passing
  • 265 conformance fail cases + 122 pass cases (77+27 from jsonchecker, 188+95 from JSONTestSuite), each asserting the exact expected parse_statuses value — 3,096 conformance assertions per platform across all eight configs
  • 27 round-trip documents (54 checks per configuration) including edge cases (null, empty, large numbers, raw pointers, unique_ptr, special floats)
  • 64 float edge cases (plus dedicated serialization-fidelity checks) and 24+16 int/uint pass cases with 11+11 matching fail cases, backed by separate digit-boundary/toString tests across all eight integer types
  • 35 string pass cases + 26 fail cases, including full Unicode/emoji/ZWJ/escape coverage
  • A full UTF-8 correctness gauntlet — basic sequence tests, the complete Markus Kuhn stress corpus, second-byte boundary tests, fused string-parser tests, unaligned-pointer and unaligned-invalid-sequence sweeps, an mmap page-boundary fault check, and a width-transition sweep
  • Direct SIMD intrinsics correctness tests, independent of parsing, covering every primitive operation across all supported backends
  • Memory safety — No leaks, double-frees, or use-after-free (ASAN)
  • Undefined behavior — No signed overflow, null pointer dereference, or invalid casts (UBSAN)
  • Bounds checking — Proper, graceful handling of truncated/malformed input across the full real-world payload suite
  • Unicode and emoji — Full UTF-8 support, including ZWJ sequences and surrogate pairs
  • Edge cases — Infinity, NaN, denormal numbers, integer overflow boundaries

Every parsing-family category above runs across all eight combinations of partialRead, knownOrder, and nullTerminated (round-trip runs the four partialRead × knownOrder combinations), and the full suite passes on Ubuntu (GCC/Clang), macOS (GCC/Clang), and Windows (MSVC).

Running Tests Locally

git clone https://github.com/nihilai-collective/Jsonifier.git
cd Jsonifier

cmake -B build -DJSONIFIER_UNIT_TESTS=ON

cmake -B build -DJSONIFIER_UNIT_TESTS=ON -DJSONIFIER_ASAN=ON -DJSONIFIER_UBSAN=ON

cmake --build build --target jsonifier-unit-tests
./build/unit-tests/jsonifier-unit-tests

Quick Example

#include <jsonifier>

struct event {
    int64_t id{};
    std::string name{};
    std::optional<std::string> logo{};
    std::vector<int64_t> topicIds{};
};

struct catalog {
    std::unordered_map<std::string, event> events{};
    std::string schema_version{};
};

template<> struct jsonifier::core<event> {
    using value_type = event;
    static constexpr auto parseValue = createValue<
        &value_type::id,
        &value_type::name,
        &value_type::logo,
        &value_type::topicIds>();
};

template<> struct jsonifier::core<catalog> {
    using value_type = catalog;
    static constexpr auto parseValue = createValue<
        &value_type::events,
        makeJsonEntity<&value_type::schema_version, "schema-version">()>();
};

int main() {
    jsonifier::jsonifier_core<> parser;

    catalog data;
    std::string json = R"({"events":{"42":{"id":42,"name":"Concert","logo":null,"topicIds":[1,2,3]}},"schema-version":"1.0"})";
    parser.parseJson(data, json);

    std::string output;
    parser.serializeJson(data, output);

    return 0;
}

Note the makeJsonEntity<&value_type::schema_version, "schema-version">() — Jsonifier maps the C++-legal schema_version member to the kebab-case "schema-version" key in JSON entirely at compile time, with zero runtime cost.

Warning: Include only . Direct inclusion of internal headers may cause unrelated code in the including translation unit to become uncompilable.

The jsonifier_core<> type is templated on an initial scratch-buffer size in bytes (default 1MB) — e.g. jsonifier::jsonifier_core<4 * 1024 * 1024> parser; for workloads that consistently deal with larger documents. It is an alias for jsonifier_core_collection. On x64 the collection holds one parser core per compiled SIMD tier (AVX, AVX2 and AVX-512, up to the configured one) and forwards each call to the best tier the running CPU supports. A binary configured for AVX-512 therefore also runs on AVX2 and AVX machines. The cores share one string buffer and error list. See Backends and the Core Collection.


Documentation

Getting Started

  • Installation — Install via vcpkg, CMake FetchContent, or source
  • Quick Start — Five-minute onramp with a working example

Core Usage

Optimization

Output Formatting

  • Prettifying — Pretty-print JSON with customizable indentation
  • Minifying — Minify JSON for compact output

Advanced Topics


Requirements

  • CMake 3.28 or later
  • C++23 compliant compiler (MSVC 2022 v19.40+, GCC 14+, Clang 18+)
  • Supported CPU (x64, ARM64 with NEON or SVE2*)

* SVE2 support is experimental — see the CPU Architecture Support section above.


License

This library is licensed under the MIT License. See the LICENSE file for details.


Acknowledgments

  • SIMD parsing techniques inspired by simdjson
  • Reflection interface inspired by Glaze
  • Dragonbox algorithm for float conversion
  • FastFloat for number parsing
  • Unit test harness: rt-ut
  • Raymond, because, Thanks.
  • Paul Mandarino. "Back yourself up, back your words up."

Contributing

Contributions are welcome! Please feel free to submit pull requests or open issues on GitHub.


Star this repository if you find it useful! ⭐

About

A few classes for extremely fast json parsing/serializing in modern C++. Possibly the fastest json parser in C++. Possibly the fastest json serializer in C++.

Topics

Resources

Stars

102 stars

Watchers

5 watching

Forks

Releases

Sponsor this project

Packages

Used by

Contributors

Languages