LINQ vs Koru, Run Verbatim
This is the LINQ code from the screenshot everyone has seen — a record, a threshold filter, an interpolated select, a materialize:
public enum ThreatLevel { Low, Elevated, High, Critical }
public record Session(string IpAddress, int ThreatScore, ThreatLevel ThreatLevel);
public class ThreatScanner
{
public List<string> LatestResults(IEnumerable<Session> activeSessions, int threshold)
{
return activeSessions
.Where(s => s.ThreatScore >= threshold)
.Select(s => $"{s.IpAddress} [{s.ThreatLevel}]")
.ToList();
}
} We compiled the same task with Koru and ran its emitted C# against this method, verbatim, on the same runtime — 10M sessions, managed arms min of 7, the native arm min of 3 whole-process, net8.0 arm64. The baseline is the code as posted: IEnumerable<Session> in, lazy producer behind it.
| code | ms | vs LINQ as posted |
|---|---|---|
LINQ — the method above, verbatim, lazy IEnumerable source — the code as posted | 2505 | — |
Koru → C# — koruc --lang=cs | 1481 | 1.69× faster |
handwritten C# — foreach | 1518 | 1.65× faster |
| handwritten C# — fused loop | 1522 | 1.65× faster |
LINQ — the same method, fed a materialized Session[] — its specialized fast path | 1529 | 1.64× faster |
Koru → Zig — the same task.k, koruc build --release=fast, whole process | 444 | 5.6× faster |
The emitted C# beats LINQ on every arm — 1.69× on the posted signature, 1.03× even after LINQ is handed a prebuilt array and its specialized iterators. It is indistinguishable from the handwritten loops. The same source compiled native runs 5.6× the posted code — no GC in the loop, allocation is memcpy instead of a managed heap. Three consecutive runs held the ordering; laptop numbers carry a ~±50ms band.
This is the Koru source and the C# the compiler emitted for its hot loop:
run = for(0..(@as(usize, @intCast(n))))
! each i when (@as(i64, @intCast(i)) * 31) % 100 >= threshold
|> level-name(l: @as(i64, @intCast(i)) % 4): lvl
|> std/fmt:fmt.blk {10.0.0.{{ i % 256:d }} [{{ lvl:s }}]}: s
|> collect(s)Verbatim from task.k (the @as casts are the real spelling). when on ! each is the Where; level-name is a cond cascade; fmt.blk is the Select; collect is the ToList — a disclosed ~proc shim (List.Add) standing in for custody.
for (long __koru_item3 = 0; __koru_item3 < ((long)((n))); __koru_item3++) {
var i = __koru_item3;
if (((long)((i)) * 31) % 100 >= threshold) {
var lvl = main_module.level_name_event.handler(
new level_name_event.Input { l = (long)((long)((i)) % 4) });
var s = $"10.0.0.{(i % 256)} [{__koru_str(lvl)}]";
var __koru_p_s = (string)(s);
results.Add((string)__koru_p_s);
// The host owns memory — a Koru string is a managed C# string and the
// GC settles the allocated! debt. Empty by construction.
}
}Verbatim from output_emitted.cs, line-wrapped for print. The Where is an if inside the for; the Select is a C# interpolated string — the same DefaultInterpolatedStringHandler LINQ's lambda gets; the ToList is the same List.Add LINQ ends in. No enumerator, no delegates — the discharge comment is the compiler telling you the GC owns the string.
Every arm produces identical hit count and total characters — a stage that never ran shows up as a wrong checksum.
What the run caught
First measurement had the emitted C# at 2915ms — nearly 2× LINQ. Three emitter defects and one harness asymmetry, all fixed:
fmt:lnproduced a{ text = … }record —new { text = … }, a heap-allocated anonymous object per element.fmt.blkbinds the string directly: 2915 → 2707ms.- Every
{{ }}placeholder called__koru_str(dynamic v)— a signature ported from the JS emitter, where untyped parameters are free. On C#, alongcrossingdynamicis a heap allocation per call — 6M boxes, ~750ms. Now__koru_str<T>(T v); the JIT specializes per operand. Ships in every emitted file’s preamble. - The placeholders emitted as
+concat — aToStringsubstring per placeholder plus the concat result. Nowfmt.blk|csemits the interpolated-string form$"…"—DefaultInterpolatedStringHandler, one pooled buffer, one allocation. Numeric placeholders get a bare{expr}hole: digits are written into the buffer, zero allocs. - Our own checksum was inside the timed loop (
hits/charsstatics per element) while LINQ’s ran after the stopwatch — the generated arm now checksums post-timer like every other arm.
Where it breaks down
The walls, as they stand today:
capture/@typeInfo— KORU047. Custody-bearing closures refuse to emit on cs at all;collectis a disclosed~procshim (List.Add) standing in that exact spot. An absence, not a slowdown.- The GC floor. On zero-allocation work the emitted C# sits at the handwritten floor (0.3ns/call on effect dispatch). The moment the task allocates per element — this one — the native arm pulls 3.3–5.6× ahead and stays there. That’s the managed heap, not a defect.
dynamicat opaque boundaries. A|csbody authoring a shape the declaration can’t name — sub-two-field records, erased member names — detonates asRuntimeBinderExceptionat runtime rather than refusing at compile time.- Concurrency — the procs aren’t ported. Koru is colorless by design: channels,
spawn,! waitexist, and async is a proc-internal HOW, never a signature. On cs the pump corpus refuses today —KORU047 NoCsProcBody(690_314: “tor ‘tick-wait’ has no |cs proc”) — the concurrency machinery has no|csbody yet. Whether a port lowers toTask-colored emit or stays colorless inside the proc boundary is the open question — not a language limit.
Sessions are derived from i, not read from a Session[]. These are laptop numbers with a ~±50ms band. tests/benchmarks/012_threat_scanner/run.sh reproduces the table.