Measurements
Since August 2026, the CG Arena has run controlled comparisons of how a model performs on the same task with and without Code Guardian. This page summarizes the method and the current record.
Method
- The same model runs the same task with and without Code Guardian, 5 runs per arm.
- An additional arm with Code Guardian serves as a noise floor.
- Scoring is mechanical, against a table fixed in advance, checked with an exact permutation test.
Record
With and without Code Guardian, same model, same task:
- 10 runs, all with a verdict.
- 7 times Code Guardian performs measurably better.
- 3 times the result is not decidable.
- The most recent run is not decidable (p = 0.19).
- No run shows Code Guardian performing worse.
A second scenario compares two ways of working that BOTH run with Code Guardian (orchestrated against classic). It says nothing about with or without and is listed separately: 3 runs, 1 of them with no difference shown, 2 without a verdict.
Limits of the measurement
- The sample is small, and the tasks come from the manufacturer's own scenarios, not from your project.
Limits of the product
- Code Guardian protects against a careless agent, not against one that deliberately tries to bypass it.
- Under PowerShell, only part of the command guard scripts see the commands, Git Bash is the supported path.
- Full tool depth, meaning static analysis and mutation testing, exists only for PHP/Laravel and for JS/TS/Vue. Other languages get gates, reflexes and the slop scanner.
- No legal advice.
PRUEFNACHWEIS.md
The package includes PRUEFNACHWEIS.md. It is generated when the package is built and describes the package's own test layers. It proves nothing about your project.
In one sentence
The CG Arena shows a repeated but numerically small advantage on the manufacturer's own scenarios, never a guarantee for any particular project.