EVALONAdmin sign in

Tools measure. AI explains.

Every score traces back to evidence you can click.

Submit a repo. EVALON clones it, runs real static analysis, then grounds three sequential AI agents in what those tools actually found — so no score on your scorecard is a raw model opinion.

Runs on local models. Your code is never sent to a third-party API.

Sample evaluationn = 47 submissions
Your scoreHackathon average
Code Quality82
  • •0 functions exceed the complexity threshold
  • •86% docstring coverage across core modules

The pipeline

Three tools run before a single AI token does.

01

Measure

radon, semgrep, and doc-coverage tools run against the cloned repo first — complexity, security findings, test presence — before a single AI token is generated.

02

Ground

Three agents run one at a time, never in parallel: Repository Understanding, then Code Quality, then Innovation. Each reads the tool output directly — none of them guess.

03

Explain

Every criterion ships with the specific evidence that produced it. Click any score on your scorecard to see exactly what the tools found.

Watch a hackathon judge itself

Submission counts, score distributions, and tech-stack trends update live over SSE as evaluations complete — no refresh, ever.

A mentor that's read the evidence

Ask why a score landed where it did. Answers cite the same static analysis and agent findings your scorecard is built on, never a generic tip.

Never a blank screen

If a model is busy or a tool times out, EVALON falls back to static analysis and says so plainly — one failure never crashes an evaluation.

See your own repo's evidence.

Submit your hackathon repo and get a scorecard where every number opens into the findings behind it.

Join as a participant