LINQ vs Koru, Run Verbatim

· 5 min read

This is the LINQ code from the screenshot everyone has seen — a record, a threshold filter, an interpolated select, a materialize:

public enum ThreatLevel { Low, Elevated, High, Critical }

public record Session(string IpAddress, int ThreatScore, ThreatLevel ThreatLevel);

public class ThreatScanner
{
    public List<string> LatestResults(IEnumerable<Session> activeSessions, int threshold)
    {
        return activeSessions
            .Where(s => s.ThreatScore >= threshold)
            .Select(s => $"{s.IpAddress} [{s.ThreatLevel}]")
            .ToList();
    }
}

We compiled the same task with Koru and ran its emitted C# against this method, verbatim, on the same runtime — 10M sessions, managed arms min of 7, the native arm min of 3 whole-process, net8.0 arm64. The baseline is the code as posted: IEnumerable<Session> in, lazy producer behind it.

codemsvs LINQ as posted
LINQ — the method above, verbatim, lazy IEnumerable source — the code as posted2505—
Koru → C# — koruc --lang=cs14811.69× faster
handwritten C# — foreach15181.65× faster
handwritten C# — fused loop15221.65× faster
LINQ — the same method, fed a materialized Session[] — its specialized fast path15291.64× faster
Koru → Zig — the same task.k, koruc build --release=fast, whole process4445.6× faster

The emitted C# beats LINQ on every arm — 1.69× on the posted signature, 1.03× even after LINQ is handed a prebuilt array and its specialized iterators. It is indistinguishable from the handwritten loops. The same source compiled native runs 5.6× the posted code — no GC in the loop, allocation is memcpy instead of a managed heap. Three consecutive runs held the ordering; laptop numbers carry a ~±50ms band.

This is the Koru source and the C# the compiler emitted for its hot loop:

Verbatim from task.k (the @as casts are the real spelling). when on ! each is the Where; level-name is a cond cascade; fmt.blk is the Select; collect is the ToList — a disclosed ~proc shim (List.Add) standing in for custody.

Verbatim from output_emitted.cs, line-wrapped for print. The Where is an if inside the for; the Select is a C# interpolated string — the same DefaultInterpolatedStringHandler LINQ's lambda gets; the ToList is the same List.Add LINQ ends in. No enumerator, no delegates — the discharge comment is the compiler telling you the GC owns the string.

the same argument in thirty seconds: LINQ's pipeline mapped onto one task.k, then the two emits side by side — zig spells alloc→use→free, the GC owns it on .NET — and the six-arm race on one axis.

Every arm produces identical hit count and total characters — a stage that never ran shows up as a wrong checksum.

What the run caught

First measurement had the emitted C# at 2915ms — nearly 2× LINQ. Three emitter defects and one harness asymmetry, all fixed:

  • fmt:ln produced a { text = … } record — new { text = … }, a heap-allocated anonymous object per element. fmt.blk binds the string directly: 2915 → 2707ms.
  • Every {{ }} placeholder called __koru_str(dynamic v) — a signature ported from the JS emitter, where untyped parameters are free. On C#, a long crossing dynamic is a heap allocation per call — 6M boxes, ~750ms. Now __koru_str<T>(T v); the JIT specializes per operand. Ships in every emitted file’s preamble.
  • The placeholders emitted as + concat — a ToString substring per placeholder plus the concat result. Now fmt.blk|cs emits the interpolated-string form $"…" — DefaultInterpolatedStringHandler, one pooled buffer, one allocation. Numeric placeholders get a bare {expr} hole: digits are written into the buffer, zero allocs.
  • Our own checksum was inside the timed loop (hits/chars statics per element) while LINQ’s ran after the stopwatch — the generated arm now checksums post-timer like every other arm.

Where it breaks down

The walls, as they stand today:

  • capture/@typeInfo — KORU047. Custody-bearing closures refuse to emit on cs at all; collect is a disclosed ~proc shim (List.Add) standing in that exact spot. An absence, not a slowdown.
  • The GC floor. On zero-allocation work the emitted C# sits at the handwritten floor (0.3ns/call on effect dispatch). The moment the task allocates per element — this one — the native arm pulls 3.3–5.6× ahead and stays there. That’s the managed heap, not a defect.
  • dynamic at opaque boundaries. A |cs body authoring a shape the declaration can’t name — sub-two-field records, erased member names — detonates as RuntimeBinderException at runtime rather than refusing at compile time.
  • Concurrency — the procs aren’t ported. Koru is colorless by design: channels, spawn, ! wait exist, and async is a proc-internal HOW, never a signature. On cs the pump corpus refuses today — KORU047 NoCsProcBody (690_314: “tor ‘tick-wait’ has no |cs proc”) — the concurrency machinery has no |cs body yet. Whether a port lowers to Task-colored emit or stays colorless inside the proc boundary is the open question — not a language limit.

Sessions are derived from i, not read from a Session[]. These are laptop numbers with a ~±50ms band. tests/benchmarks/012_threat_scanner/run.sh reproduces the table.