Orisha vs Actix and ASP.NET on SQLite

· 8 min read

A recent YouTube benchmark — eight languages, one API, Postgres on a $12 server — hit the wall everyone hits: Postgres ate the CPU before the languages differentiated. So we ran the question the video actually wanted to ask: put SQLite in-process and let the language be the variable.

Three contestants: Orisha (our compile-time HTTP server), Actix (Rust, rusqlite), and ASP.NET Core (Microsoft.Data.Sqlite). Same database file — 10,000 rows, WAL mode — same endpoint shape, wrk -t4 -c256 -d10s, request IDs drawn uniformly at random so no hot-row cache does the work.

Three legs, because “a database endpoint” is not one thing:

  • cold: open → prepare → bind → step → close, per request. The naive driver path.
  • hot: one connection and prepared statement per worker thread. What a real service does.
  • pkg: the cold path written through koru-libs/sqlite3’s public API — what a Koru program actually compiles.

The results

LegOrishaActixASP.NET
plaintext153,246135,550151,194
cold14,0539,27410,335
hot155,744141,688154,768
pooled——153,290
pkg (cold through the library)14,304——

Requests per second from one cited pass — wrk -t4 -c256 -d10s, random IDs, WAL — on this machine, 2026-10-03. The finding sections below quote the specific runs that isolated each change; run-to-run variance moves the absolutes a few percent, not the ordering.

Orisha leads every leg we ran. The numbers aren’t the story, though — why they moved is.

the whole argument in twenty seconds: a near-tie on the steady-state path, then the cold leg collapses an order of magnitude onto sqlite's own per-open machinery
one wrk pass, two scales — the steady-state panel is the request path deciding; the cold panel, ~10x smaller, is sqlite's per-open machinery deciding

Finding 1: our own regression was a missing writev

Mid-benchmark Orisha’s plaintext path sagged to ~129k — below a historical 165k on this machine, below .NET. The cause: reply wrote the response head and body as two write() syscalls and re-scanned the raw request for Connection: close — a scan the pump had already done at request arrival. One writev and the arrival-time flag restored it: 129k → 159k, ahead of the previous record.

That’s ~23% of the request path in one syscall and one redundant scan. A benchmark’s first job is finding your own bugs; it did.

Finding 2: the cold leg is sqlite-bound, not driver-bound

Early on, Orisha’s cold leg “won” 3× — 33k vs ~10k. Suspicious. Phase isolation on the same sqlite build told the truth: open+close is ~36% of per-request cost, prepare ~36%, step+finalize ~28%. The win wasn’t our driver — it was macOS’s system libsqlite3, a thin build that registers less per open.

So we vendored the exact amalgamation rusqlite’s bundled feature compiles — same sqlite3.c, same flags — and the advantage evaporated: everyone lands at ~10k. Per-request cold open is bounded by sqlite’s own machinery (schema load, function registration, pager init), and once the internals are identical, the language stops mattering. We retired the 3× claim rather than let it stand on a library difference.

Finding 3: then the build flags became the variable

The amalgamation is the same source — what you compile into it is a choice. The bundled flag set enables FTS3/FTS5, RTREE, SOUNDEX, STAT4, load-extension, memory-management bookkeeping — all of which register work at every sqlite3_open, and a cold leg pays it ten thousand times a second.

A lean build — same source, none of those features — changed the cold leg to 16.7k. Two profiling finds compounded it on the hot side: WAL removes an fcntl shared-lock dance on every read transaction (a property of the database file, so every contestant inherits it identically), and SQLITE_OPEN_PRIVATECACHE (which only takes effect with SQLITE_OPEN_URI — we gave rusqlite the same flag) gives each worker connection its own page cache off the global pcache1 mutex. Hot leg: 131k → 161k, above our own plaintext ceiling in some runs.

This is the malleable-toolchain thesis one layer down: compile-time specialization applies to the dependency, not just the program.

Finding 4: the driver should make the cold pattern unwritable

sqlite3_prepare can’t happen at compile time — a prepared statement dies with its connection. But static SQL is compile-time knowledge, so koru-libs/sqlite3 now caches prepared statements per connection, keyed by the SQL text: the default transform emits get-or-prepare + reset/clear instead of prepare + finalize. ADO.NET gets this via its statement cache and pooled connections; the Koru library does it with a hashmap and no pool.

And the library exercise surfaced a real bug doing what these packages exist to do: the first-ever text binding through {{id}} interpolation hit a SQLITE_TRANSIENT translate-c defect — every prior binding had been an int. Fixed in the package.

What’s honest here

  • Cold leg: Orisha’s lead is a build choice, fully disclosed in build_sqlite.sh — FLAVOR=bundled reproduces rusqlite’s exact flags for verification.
  • Hot leg: identical internals, plus flags available to every driver (we set them symmetrically).
  • Pooled: only .NET ships the idiomatic pool; a pooled Orisha leg would land near its hot leg since there’s no statement cache tax left to remove.
  • Everything above is on this machine (Apple M2 Pro, macOS) — a $12 Linux box will move the absolute numbers; the mechanism ratios should hold.

The verdict isn’t “Koru wins a benchmark.” It’s that the boundary moved: once the library internals match, the language is off the hook for the database leg — and once the toolchain is malleable, the library internals are yours to choose.