Code written by a model and reviewed by the same model has been reviewed by nobody.
CounterProof is an independent adversarial review practice for modern codebases — machine-written code above all. We deliver the one thing you cannot produce in-house: an independent, signed, evidence-graded security assessment, built for the scrutiny of regulators, acquirers, insurers, and
enterprise customers.
WHY YOU’RE HERE
Nobody buys a code review. They buy what it unlocks.
You’re not reading this because you woke up wanting a security assessment. Something is asking you for
evidence.
- A regulator. The EU Cyber Resilience Act requires you to “apply effective and regular tests and reviews of the security” of your product (Annex I Part II(3)), and your technical documentation must contain “reports of the tests carried out to verify conformity” (Annex VII(6)) — kept for at least ten years after placing on the market, or the support period if longer (Article 13(13)). Article 14 notification of actively exploited vulnerabilities and severe incidents starts 11 September 2026; the Regulation applies in full from 11 December 2027. Most products may self-assess — which is not
relief, it is exposure: the burden of producing records that stand up falls entirely on you. And the option narrows with classification. For important Class I, internal control is available only where harmonised standards, common specifications, or a qualifying certification scheme have been applied
in full — and several do not yet exist. For important Class II it is not available; critical
products go to certification under Article 8(1) first. Annex III products that qualify as free and
open-source software have their own route (Article 32(5)) if their technical documentation is public. - An acquirer or investor. Technical diligence is where AI-assisted codebases now get discounted.
An independent assessment on the table changes that conversation before it starts. - A cyber-insurer. Underwriters increasingly price the gap between “we test internally” and “an
independent party tested and signed.” - An enterprise customer. Their security questionnaire has no checkbox for “our model reviewed its
own output.”
Three of these four never accept your own word about your own code — independence is the property they
are buying, and it is the one property no team can supply for its own work. The fourth, the regulator,
will often take your self-assessment — and then hold you to every record behind it. Either way the
evidence has to hold. That is what we produce.
WHAT YOU GET
The deliverable is the point.
An engagement produces a CounterProof Assessment under a persistent engagement identifier
(CPR-YYYY-NNN — what your maintainers, your acquirer, or your fix commits cite). Every
finding is either confirmed against your source — exact file and line, reproduction path, impact
classification — or explicitly graded as plausible-only. Nothing padded, nothing scanner-generated,
nothing you can’t act on.
Every finding names the evidence rung it actually reached — source trace, compile-proof, test, or
live reproduction — and never implies a higher one. Every path:line citation is machine-resolved
against the exact revision we reviewed before the report leaves our hands. You can check our work.
That is deliberate.
It is built to be handed over:
- To an authority — structured to be included in your technical documentation as an Annex VII(6)
test report, evidencing the Part II(3) testing duty. It supports the file you assemble; it is not the
file, and it does not verify conformity across Annex I. - To a diligence team — evidence-graded, reproducible, scoped, with our name on the risk.
- To your own engineers — short enough to read, precise enough to fix. We can stay through the
patch and independently verify each fix against the finding it targets, recording what we verified
and at which revision — so you end with a record of what was checked, not a list of notes.
WHY THIS CAN’T BE DONE IN-HOUSE
You could run the same models. You can’t be independent.
- Running multiple AI models over your own code replicates our tooling and loses the property that
matters. Your team chooses what the reviewers see, frames the questions, and judges the answers —
and every one of those choices carries your assumptions straight back into the review. That is not
a discipline failure; it is structural. The author of a system cannot be its own adjudicator. - So even a flawless internal review is still your own word about your own code. That is exactly what
an acquirer, an insurer, an enterprise security team — and, for the product classes the CRA routes to
third-party conformity assessment, the process itself — will not take on trust. What they are buying is a third party willing to sign. We are
not a notified body and we do not certify conformity; we produce the independent evidence that
supports your assessment and is structured to be re-examined by anyone else’s. - The method doesn’t care who — or what — wrote your code. Independence is missing from human-written
code just as often. Machine-written code only makes the gap impossible to ignore.
WHERE THE METHOD COMES FROM
We didn’t design this method. We earned it.
CounterProof wasn’t built as a product. It accreted, rule by rule, while we built and secured our own
threshold-cryptography payment infrastructure — where a defect that slips through doesn’t cost a
customer relationship; it costs the operator its own funds. We are that operator. That is why the
method assumes every reviewer, including itself, is wrong until proven otherwise.
Every rule exists because its absence burned us on our own code first. Citations are machine-resolved
because unresolved ones drifted. Findings face independent refutation because confident
single-reviewer conclusions were wrong exactly where confidence ran highest. We name the evidence rung
reached because we once claimed a higher one than we stood on. The method is scar tissue, organized.
And it has never stopped running. Our own codebase — much of it machine-written — goes through the
same adversarial review we sell, continuously, under a standing internal register. That register is
not a claim; it is a paper trail — findings, refutations, withdrawals, and verified fixes on our own
code — and within an engagement we can open it to you. This is the workshop: where we hold ourselves
to the method before we hold your code to it. We don’t call reviewing our own code independent — we
are the authors, and saying so plainly is the discipline working, not failing. Independence is the
boundary we hold for you. We were our own first customer, and we remain our hardest one.
The rules that came out of it, and that every engagement runs under:
- A standing withdrawal contract. If a finding is shown to be wrong, we retract it in writing —
the report’s footer says so, in advance, on every report. A review brand that never withdraws is one
that never admits error, and nobody should believe it. - Nothing ships unrefuted. A finding reaches you only after independent model families — not a
second pass of the same one — have tried to refute it and failed, with disagreements settled by
reading your source, not by vote. - No unmeasured number, ever. We state operation counts; we never quote a wall-clock, throughput,
or rate we did not measure. That rule applies to our marketing too — which is why this page contains
no detection percentage. No honest one exists. - We keep our own error log. Within an engagement, under NDA, we walk you through our methodology —
including a standing record of review errors we have made, the ones outside reviewers caught that we
had missed, and the gaps we know remain. Ask another provider for theirs. - Specialists, not generalists. Rust, protocol implementations, cryptographic code, payments, and
anything that touches other people’s money or data.
THE DISCLOSURE THAT COMES WITH THAT
You would find this anyway. Better you hear it from us.
The origin above is also a conflict, and we would rather name it than have your diligence team name
it for us.
CounterProof is part of a group that builds payment and digital-asset custody infrastructure. That is where this method was forged — and it means that if you build in those markets, our affiliate may be adjacent to you, or competing with you.
So: before any engagement, we tell you exactly what the group builds and where it operates, in writing. You decide whether that is acceptable, and you decide before you have paid us anything. If it is not acceptable, that is a legitimate answer and we would rather hear it at the start.
What our independence claim does and does not cover, precisely: we never sign an assessment on code from inside our own walls. That is structural and absolute. It is a statement about whose code we review, not a claim to have no commercial interests anywhere near your sector — no review firm with
real domain expertise can honestly claim the latter, and we are not going to pretend otherwise.
METHOD
How it works.
Every assessment is produced under the CounterProof protocol — a multi-lineage adversarial review
in which independent model families each attack the findings. Every finding must survive adversarial
refutation before it reaches you, and every finding we label confirmed has been verified against
your source; anything that survived refutation but not confirmation is labelled plausible-only and
says so. Full methodology documentation, including its limits and our own recorded errors, is
available under NDA within an engagement.
ENGAGEMENTS
Three ways in.
- Assessment. Fixed-scope adversarial review of a repository or release candidate. The signed
report, evidence-graded, under itsCPRidentifier. - Assessment + Remediation Verification. We stay through your fixes and independently verify each
patch against the finding it targets. The report records what was verified, at which revision, and
what remains open. - Standing Review. Recurring review of named releases, delivered as Annex VII(6)-shaped test
reports you retain in your own records. Note the word regular in Part II(3): a one-off assessment
does not earn it. This produces evidence relevant to that duty — it does not discharge it.
Already running your own model reviews? Good — you’ve met the noise. Bring us the findings you can’t adjudicate.
THE CLOCK IS NOT OURS. IT’S BRUSSELS’.
- 11 September 2026 — reporting obligations for actively exploited vulnerabilities begin.
- 11 December 2027 — the Regulation applies in full to products with digital elements sold in the EU.
- Today — the testing and vulnerability-handling records you’ll need then are the ones you start
producing now.
[ Talk to us before the deadline does → ]
Why we exist: this practice funds and sharpens infrastructure we are building to hold other people’s money — which is why the method had to be real for us long before it was ever for sale.
CounterProof is a practice of Clavestra Capital Limited (Malta, C 113987). We review code; we do not perform statutory audits, and we are not a notified body or conformity-assessment body under the CRA. Independent by structure: we never sign an assessment on code from inside our own walls — see the disclosure above for what that does and does not cover. Proof attests; CounterProof refutes — CounterProof is the adversarial sibling of Clavestra Proof: the discipline of falsification, applied to code.
Sample finding (sanitized)

