Koragraph Benchmarks
How we measure against everyone else.
Every head to head we run, published with the corpus, the questions, the gold sets and the scorer, so the other system’s column can be reproduced by anyone who doubts it.
Runs
2 benchmarksKoragraph vs CodeGraph, GitNexus, Graphify
September 2026Code-graph extraction accuracy, declarations and edges, across twelve languages scored by independent compiler front ends
Koragraph wins 34 of 39 language and plane cells on recall and leads precision on all 39, against CodeGraph, GitNexus and Graphify, every number produced by an independent compiler front end rather than the tool under test.
34 of 39
zero cells lost outright
recall cells won, language by plane
39 of 39
every language, every plane
precision cells led
87.6%
best rival 45.1%
JavaScript intra-repo call recall
12 · 36
CPython ast, go/parser, Roslyn, tsc
languages and repositories, judge-free oracles
See the numbers →
Koragraph vs Graphify
August 2026Declaration extraction, indexing speed and retrieval, across ten languages and 102 repositories
Koragraph indexes 99.70% of the declarations in 102 repositories against Graphify’s 82.07%, does it with 3.78× less CPU, and beats Graphify’s best retrieval result at any depth while spending 45% fewer tokens.
99.70%
Graphify 82.07%
declaration recall across 10 languages
3.78×
ahead on 17 of 17 repositories
less CPU to index
0.824
Graphify’s best 0.769, on 45% fewer tokens
retrieval recall against their ceiling
4 of 6
2 ties, none lost
context budgets won on retrieval
See the numbers →
What is next
Larger repositories, more systems, and the same rule every time: an independent referee rather than either side’s own definition of a correct answer, the corpus pinned to commit SHAs, and every input published so the comparison can be run against us. A rerun on a new corpus is published as a new benchmark rather than as an edit to an old one, so a number you saw once does not change underneath you.
