Evidence, not promises

What it actually did, misses included.

Most AI companies show you a demo. This page is the other thing: what our system did at real companies, on private data, scored against bars we published before we had results. Where a number is not measured yet, it says so rather than showing a figure we invented.

The rule

This page only moves forward. Entries are dated when written, never backdated. Nothing here gets quietly deleted when it stops flattering us.

01

The pilot scorecard

Anonymized running totals from live deployments: how many cases the system processed, how many were accepted with only minor edits, how many the reviewer sent back, and how long review took against the manual baseline. Updated as the pilot runs, not at the end.

Cases processed Measurement in progress
Accepted with minor edits Measurement in progress
Sent back by the reviewer Measurement in progress
Reviewer minutes vs manual baseline Measurement in progress

Not yet measured

We would rather show you four empty rows than four numbers we cannot produce on request. These fill in as the scored runs complete.

02

The deployment runbook

The actual step list we run to stand this up inside a customer's walls, with the sensitive specifics blacked out. It exists so you can see the shape of the work before you commit: what is configuration, what is engineering, what we need from your team, and where the week actually goes.

Setup takes about a week. Your team's involvement is a few hours across it, not a second job.

Redacted runbook Available on the call
Published version In preparation

Redaction in progress

Redaction is the slow part. A runbook with customer specifics left in it would tell you we are careless with exactly the thing you are worried about.

03

The accuracy benchmark

We wrote down what would count as success before running the scored benchmark, so the bar cannot move to meet the result. These three thresholds were recorded in our technical feasibility log during the cycle of 22 May to 11 June 2026. They are reproduced here unchanged.

At least 70% of outputs usable with review

A qualified reviewer can work from the draft rather than starting over.

At least 80% of citations correct

The link points at the passage that actually supports the claim.

Zero fabricated claims surviving sign-off

The one bar with no tolerance. A fabrication that reaches a signature is a failure of the whole design, not a bad score.

Scored result against these bars Measurement in progress

Pre-registered

The result posts here the week it exists, hit or miss. If we miss a bar, the miss goes on this page next to the bar we set.

04

Lab notebook

Dated excerpts from the build log, including the weeks that went badly. The entry that changed the product most is the one where the system was confidently wrong.

Prep work is a bottleneck, not an annoyance.

Embedded sessions with an engineering firm and a law firm. Quote turnaround ran about a week; the first accurate responder usually won the job; years-old pricing was being reused, producing at least one zero-margin job. At the law firm, intake and conflict checks were consuming the most senior people. We recorded the win-rate claim as single-source at the time and flagged it for broader validation rather than treating it as settled.

It invented specifications, confidently, and experts caught it.

On clean inputs the prototype produced coherent, cited, on-domain drafts. On messy real documents it fabricated specifics wherever grounding was weak — and did it fluently enough to look right. Domain experts caught the fabrications in drafts they had otherwise rated usable. That located the gating factor precisely: grounding, correct citations, and a human checkpoint. A larger model would produce more convincing fabrications, not fewer. The review step exists because of this week, and the three accuracy bars above were written down in the same cycle, before any scored result.

The privacy objection became the reason to buy.

Privacy and data residency came up repeatedly as blockers to generic cloud AI. We demonstrated the agent running on-prem, on the customer's own confidential files, with no data leaving their boundary. Posture shifted from cautious to willing to proceed. The abstract promise had not moved them; watching it run on their own machines did.

Money followed the demonstration, not the pitch.

A bounded pilot offer with an explicit ask converted to a paid engagement for an intake and conflict-check workflow, immediately after the private on-prem demonstration. We also recorded what did not happen: our original plan of ten cold conversations was the wrong instrument, and depth with a working demonstration beat breadth of outreach.

Dated when written

Entries are excerpted from our running experiment logs. Dates are the cycles the work happened in.

05

Limitations

What this system does not do. Published here rather than discovered by you in month two.

  • It does not run unsupervised on messy documents. That finding is why the review step exists, and we have no plans to remove it.
  • It does not replace your expert's judgment. It removes the digging that precedes the judgment.
  • It does not make claims it cannot cite. Where evidence is insufficient it flags the gap and asks, which means you will sometimes get a question instead of an answer.
  • It is not certified. We hold no compliance certifications and do not claim any. We verify the data boundary with a review and publish what we find.
  • The accuracy benchmark is not finished. The bars are published above; the scored result is not in yet.

Current

If a limitation here stops being true, it moves to the notebook with a date rather than disappearing.

The standing offer

If a claim here is not backed, tell us.

We will fix it or cut it, and the change goes in the notebook with a date. That is the deal this page is making with you.

Book a call