Tools measure. AI explains.
Every score traces back to evidence you can click.
Submit a repo. EVALON clones it, runs real static analysis, then grounds three sequential AI agents in what those tools actually found — so no score on your scorecard is a raw model opinion.
Runs on local models. Your code is never sent to a third-party API.
- •0 functions exceed the complexity threshold
- •86% docstring coverage across core modules
The pipeline
Three tools run before a single AI token does.
Measure
radon, semgrep, and doc-coverage tools run against the cloned repo first — complexity, security findings, test presence — before a single AI token is generated.
Ground
Three agents run one at a time, never in parallel: Repository Understanding, then Code Quality, then Innovation. Each reads the tool output directly — none of them guess.
Explain
Every criterion ships with the specific evidence that produced it. Click any score on your scorecard to see exactly what the tools found.
Watch a hackathon judge itself
Submission counts, score distributions, and tech-stack trends update live over SSE as evaluations complete — no refresh, ever.
A mentor that's read the evidence
Ask why a score landed where it did. Answers cite the same static analysis and agent findings your scorecard is built on, never a generic tip.
Never a blank screen
If a model is busy or a tool times out, EVALON falls back to static analysis and says so plainly — one failure never crashes an evaluation.
See your own repo's evidence.
Submit your hackathon repo and get a scorecard where every number opens into the findings behind it.
Join as a participant