
Who Simdjson is for#
Database or analytics engine authors embedding a parser
ClickHouse, Apache Doris, StarRocks, and QuestDB embed simdjson for ingest. The On-Demand API fits column-oriented ingestion patterns where only specific fields need to be extracted from each row. NDJSON parsing exceeds 3 GB/s with the multithreaded API.
Skip if:
If your database is schema-first and JSON arrives as a single defined field type rather than flexible payloads, a simpler embedded parser may be sufficient.
C++ backend engineers with JSON-heavy hot paths
If profiling shows JSON parsing consuming a significant share of a service's CPU, swapping in simdjson can recover that headroom. The On-Demand API requires minimal changes to code that already pulls specific fields from JSON objects.
Skip if:
If JSON parsing is not a measured bottleneck, the migration overhead outweighs the gains. nlohmann::json's API is simpler and more widely documented for teams new to C++ JSON handling.
Python or Rust developers needing faster JSON parsing
pysimdjson and cysimdjson wrap the C++ core for Python, giving faster parse times than the standard json module for large documents. The simdjson-rs Rust port reimplements the algorithm natively for Rust projects.
Skip if:
For small JSON documents (under a few kilobytes), the overhead of calling into C++ from Python can exceed the parse time. The standard json module is fine below that threshold.
AI/ML engineers processing large JSONL datasets
Training data for large language models often arrives as JSONL (newline-delimited JSON). simdjson's multithreaded NDJSON parsing exceeds 3 GB/s, making it a practical choice for ingestion pipelines where reading and parsing raw data is the primary throughput constraint.
Skip if:
If your pipeline is already GPU-bound downstream and JSON parsing is not the bottleneck, the integration cost is not justified.
The problem it solves#
JSON parsing is a hidden bottleneck in high-throughput server applications. Every analytics event, API response, and database record arrives as JSON and must be parsed before any computation can begin. Conventional byte-by-byte parsers branch on every character, stalling the CPU pipeline and leaving SIMD execution units idle on every modern processor.
At scale, this adds up. A server processing 1 GB/s of inbound JSON spends more CPU time deserializing text than running the actual business logic. Teams either accept the overhead, write fragile hand-rolled parsers for hot paths, or pay for commercial services that abstract away the parsing cost while introducing their own pricing and lock-in.
How it solves it#
SIMD-accelerated parsing
Processes 32-64 bytes of JSON per CPU cycle using AVX2, SSE4.2, NEON, and AVX-512 instructions. On an Intel Xeon Gold processor, simdjson reaches 5.11 GB/s parsing twitter.json into native C++ structs, compared to 0.85 GB/s for RapidJSON and 0.15 GB/s for nlohmann::json.
On-Demand API
The On-Demand front-end parses only the JSON fields you actually access in code, skipping the rest of the document. For payloads where you need two or three fields from a 50-field object, this cuts parse time proportionally without requiring you to specify a schema upfront.
UTF-8 validation at 13 GB/s
Validates UTF-8 encoding at 13 GB/s as a separate pass. Every string returned by simdjson is guaranteed valid UTF-8, with no silent truncation or substitution for malformed input.
Single-header distribution
Ships as two files: simdjson.h and simdjson.cpp. Drop them into any C++ project and compile with g++ 7 or later or clang++ 6 or later. No build system changes, no CMake configuration, no package manager required.
Bindings for 20+ languages
Official or community-maintained bindings exist for Python (pysimdjson, cysimdjson), Rust (simdjson-rs), Go (simdjson-go), Java (simdjson-java), Node.js (simdjson_nodejs), Ruby, PHP, C#, R, Erlang, Lua, Haskell, Zig, Perl, Nim, and Dart.
C++26 static reflection
In C++26 mode, the library supports one-line serialization and deserialization to and from native C++ structs via compile-time reflection, without macros or manual field mapping. Serialization benchmarks at 11.2 GB/s on the twitter.json dataset.
Strengths and trade-offs#
Strengths
- Peer-reviewed performance claimsThe simdjson design is documented in two peer-reviewed papers: 'Parsing Gigabytes of JSON per Second' in VLDB Journal (2019) and 'On-Demand JSON: A Better Way to Parse Documents?' in Software: Practice and Experience (2024). Benchmark numbers are reproducible via the published experiment repository.
- Production adoption at scaleUsed in production by Node.js, ClickHouse, Meta Velox, Google Pax, Apache Doris, Milvus, StarRocks, and QuestDB. These high-volume systems confirm correctness and reliability across diverse architectures, not just benchmark conditions.
- Runtime CPU detectionSelects the optimal SIMD implementation at runtime with no configuration. An x86 machine with AVX-512 gets the fastest path; one without falls back to AVX2 or SSE4.2 automatically. ARM AArch64 and NEON are also supported, so the same binary works across cloud providers.
- Dual Apache-2.0 / MIT licenseUsers can choose between Apache-2.0 and MIT. Both are permissive and allow commercial use without copyleft restrictions. This is useful for teams whose legal policies prefer one form over the other.
Trade-offs
- -Non-C++ language bindings may lag the core libraryThe core library is C++. Non-C++ callers depend on foreign-function interfaces or community bindings, which add a maintenance surface and may lag the C++ library on new features. Python users choosing between pysimdjson and cysimdjson face different API shapes; neither is an official canonical binding.
- -No compatibility layer for existing JSON library codeLinking simdjson into projects that already use RapidJSON or nlohmann::json throughout requires replacing call sites. There is no drop-in compatibility shim. For large codebases, the migration effort may not be justified unless profiling confirms JSON parsing as a measurable bottleneck.
- -64-bit platforms onlyRequires a 64-bit system. 32-bit targets, including many embedded environments and legacy systems, are not supported.
Simdjson vs alternatives#
simdjson vs RapidJSON
Both libraries target high-performance JSON parsing in C++, but simdjson was built to exploit SIMD hardware that RapidJSON predates. RapidJSON has a larger body of tutorials, Stack Overflow answers, and existing production integrations.
| Feature | simdjson | RapidJSON |
|---|---|---|
| License | Apache-2.0 / MIT | MIT |
| Parse throughput | ~5 GB/s | ~0.85 GB/s |
| API style | DOM or On-Demand | SAX or DOM |
| C++ standard | C++17 required | C++03 compatible |
| Distribution | Single header | Single header |
simdjson is the better choice for new code where parse throughput is a priority and a C++17 compiler is available. RapidJSON is worth considering for projects that must support older compilers (pre-C++17) or that have significant existing code against the RapidJSON API where migration cost outweighs performance gains.
simdjson vs nlohmann::json
nlohmann::json is the most widely used C++ JSON library for developer ergonomics, with template-based APIs that map directly to C++ types. simdjson is roughly 33x faster on the twitter.json benchmark, but the two libraries serve different primary goals.
| Feature | simdjson | nlohmann::json |
|---|---|---|
| License | Apache-2.0 / MIT | MIT |
| Parse throughput | ~5 GB/s | ~0.15 GB/s |
| API ergonomics | Moderate | High |
| Error handling | Error codes or exceptions | Exceptions |
| DOM tree | Optional (On-Demand skips it) | Default |
nlohmann::json is the better starting point for most application code where JSON parsing is not a bottleneck, because the familiar C++ container syntax reduces implementation complexity. Switch to simdjson when profiling identifies JSON as a hotspot, or when building a new high-throughput service from scratch.
Quick start#
Clone the repository and build simdjson with CMake, or add the single-header files directly.
```bash
git clone https://github.com/simdjson/simdjson.git
```What it's built on#
- Languages
- C++
FAQ#
What makes simdjson faster than other JSON parsers?
simdjson uses SIMD (Single Instruction, Multiple Data) vector instructions to process 32-64 bytes of JSON per CPU cycle instead of one byte at a time. On modern x86 and ARM processors, this translates to over 5 GB/s parsing throughput, compared to 0.85 GB/s for RapidJSON and 0.15 GB/s for nlohmann::json on the same twitter.json benchmark.
Can I use simdjson in a language other than C++?
Yes. Community and official bindings exist for Python, Rust, Go, Java, Node.js, Ruby, PHP, C#, R, Erlang, Lua, Haskell, Zig, Perl, Nim, and Dart. The Node.js runtime itself uses simdjson internally. Binding quality varies; Python has multiple active options including pysimdjson, cysimdjson, and fastpysimdjson.
Is simdjson free for commercial use?
Yes. simdjson is dual-licensed under Apache-2.0 and MIT. Both licenses allow commercial use, modification, and redistribution without copyleft requirements. You pick whichever license fits your project's legal requirements.
How do I add simdjson to an existing C++ project?
Download simdjson.h and simdjson.cpp from the singleheader directory on GitHub and include them in your build. No package manager or CMake configuration is required, though packages are available in most Linux distro repositories. Compile with g++ 7 or later or clang++ 6 or later.
Does simdjson work on ARM and non-x86 platforms?
Yes. simdjson detects CPU capabilities at runtime and selects the best available SIMD implementation automatically. Supported architectures include x86-64 (AVX2, SSE4.2, AVX-512), ARM64 (NEON), LoongArch, and RISC-V, with no configuration changes needed.
Similar open-source tools#
Renart
SQL and Python pipelines, notebooks, and dashboards in Git
jsreport
JavaScript report server with PDF, Excel, DOCX, and HTML output
FckSignups
Open-source tools that work instantly, no signup required
Open Wearables
Open source health API for wearable device developers
Typesense
Fast typo-tolerant search engine with instant results
Omnara
Open-source agent deployment API. Self-host or use Omnara Cloud.

