The harness matters more than the model
Updated 2026-09-16: this post described the bench and the registry as they are designed, in the present tense. There are no public scores yet and the evaluation harness repository is not open — see Evaluation for the method and the current status. The live workflows are single Markdown files with a small manifest at midkernel/playbooks.
Two teams can point the same model at the same code and get different findings at a different cost. Ask a vendor why their scanner found a bug and the answer is usually the model. Ask why it missed one and the answer is usually also the model. Neither answer is complete. The incomplete part is the interesting one.
same repository, same model
- harness A
- harness B
The model did not change. The instructions did.
What a model is given
A model does not audit a repository. A harness does, using the model. The harness decides which files the model may read and in what order, which tools it can call, how many times it may try, what counts as a finding, what counts as proof, and when to stop.
Every one of those decisions is written down somewhere (in a system prompt, a tool manifest, a loop with a budget). Every one of them changes the outcome.
A harness that reads only changed files has a different scope from one that traces routes across the repository. Requiring supporting evidence before reporting a finding changes the review method too. These choices can affect coverage, cost and error rates; their effects need to be measured on real runs. The model can stay the same while the instructions change.
- 01which files
- 02which tools
- 03finding, proof, and stop
This is not a new observation. Aikido Security made the same argument earlier this year in a post titled "How Aikido finds more vulnerabilities than Mythos at half the cost". Their conclusion was that the harness still makes the difference. We agree with the argument and draw a different conclusion from it.
If the harness is what matters, it should be open
A harness is text. A manifest and a set of instructions. Text is the thing software teams already know how to handle: it goes in a repository, it has a commit hash, it gets diffed and reviewed, it gets versioned, it can be forked when it is wrong.
Closed scanners treat the harness as the product and hide it. That is a reasonable business decision and a poor engineering one, because it leaves the customer with the one question that cannot be answered from the outside: what did it actually do? A finding without its trace is an opinion. A missed bug without the log is a mystery.
Midkernel's position is that the harness is the thing to publish. Every workflow in the registry is public text: a small manifest that names the workflow and the pipeline that executes it, and Markdown instructions that say how to audit and what counts as proof. If the workflow made a bad call, you can read the call. Recording the exact workflow revision behind each run is being finished across the platform.
Open here means you can read the text that ran. It does not mean the customer's repository is public.
If the harness is what matters, it should be scored
The second half of the argument is that a claim about a harness should be checkable. Vendor benchmarks are run by vendors, on tasks chosen by vendors, with results summarised by vendors. That is not an accusation. It is just where the incentives sit.
The bench exists to move the incentives. The design: every workflow in the registry runs on public datasets (CyberGym and EVMbench among them) in the same sandbox a customer gets, with the methodology on a public page and the logs behind every cell downloadable — and results publish only with those logs. The methodology is published before the results; no scores exist yet.
The point is not to declare a winner. It is to make "which one should I run on this?" a question with an answer you did not have to take on faith.
A score without a log is another opinion. The cell is useful because you can open it.
What this means for you
If you run one AI scanner today, you are running one harness, and you cannot see it. The practical version of this post's argument is small: read the harness before you trust its output, and compare two on your own code before you buy either. Both of those are the point of Midkernel once hosted runs ship.
The model will keep getting better. So will the harness, faster, because more people can work on it.