What we measure
Task correctness, security findings, regressions, time to verified completion, human corrections and token use.
THE EVIDENCE PROGRAM
The strongest product story is one that can be inspected. Our blind evaluation is in progress. This is where methods and validated results will be published.

Task correctness, security findings, regressions, time to verified completion, human corrections and token use.
Comparable tasks, disclosed agent and model versions, fixed conditions and blind assessment where appropriate. Reports will include sample sizes, dates and uncertainty.
CodePulse product capabilities are documented. Comparative performance findings remain unpublished while evaluation is incomplete. No model ranking or improvement percentage is claimed here.
CodePulse
Bring your agent. Keep control of your application. Put evidence behind the release.