skip to content
// evaluation

Evaluation methodology_

Methodology only — no public scores yet.

Limited beta · read-only GitHub access

// datasetspublic, with ground truth we did not write
// isolationthe same sandbox a customer gets
// repetitionsseveral runs per cell, all published
// logsno score ships without its logs
// methodology

How scoring will work.

Which datasets.

A benchmark will qualify if it is public, has ground truth we did not write, has a licence that allows us to run it, and covers a stack the registry serves. Every dataset considered and rejected will be listed, with the reason.

Isolation.

Each benchmark task will run in the same single-use sandbox as a customer run, with the same network rules. Workflows will not see other tasks, other runs, or the scoring code.

Models.

Model versions will be pinned by identifier and date, and retired versions will stay in the table marked retired.

Repetitions.

Every workflow × benchmark × profile cell will be run several times; the table will show the mean and the download will hold every run.

What score means.

For datasets with a list of known vulnerabilities, score is recall: found ÷ known, where “found” requires the report to name the file and the root cause. Datasets with a native metric will use it.

Contamination.

A workflow may not embed dataset answers, task names or file paths from a benchmark. Workflow versions will be diffed against the datasets before scoring.

Our own workflows.

Workflows authored by Midkernel staff will be marked in the table and scored by the same harness, with no other treatment.

Logs.

Results will be published only with their retained logs and method. A score without its logs is a claim, not a result.

// status
No public scores exist yet, and the evaluation harness repository is not open. Internal evaluation of the live workflows comes first; results will be published in research write-ups, each with its retained logs and method.
// faq
Are there public scores today?

No. Nothing is scored publicly yet. When results are published, each will come with its logs, methodology, model versions and dates — this page is the commitment to how that will work.

Why should I trust the eventual benchmark?

You shouldn't have to. The datasets will be public, the methodology is on this page, and the logs behind every cell are meant to be downloadable so you can check any cell yourself.

Do you score closed scanners?

No. The bench will score workflows we can run and publish. A closed product can be compared on your own code by running a workflow next to it.

Can I suggest a benchmark?

Write to hello@midkernel.com. It needs public ground truth and a licence that lets us run it.

Read the workflows while you wait.

Limited beta · read-only GitHub access