Skip to content

Contract AI engineering, Available now, Same-business-day reply

A model that guesses costs you more than one that says nothing.

I build document automation for compliance, security and accessibility work, retrieval systems that cite every claim to a named source file and refuse to answer when the evidence isn't there.

Every number on this site comes from a benchmark harness in a public repository. Run it yourself.

No discovery-call funnel. Describe the workflow, get a written read on whether it's worth automating.

Drafted from the customer's own policy base 1 of 20

Click the flagged row. That blank is the product. A fabricated encryption claim in front of a prospective customer is a problem no amount of throughput makes up for.

What a confident wrong answer actually costs

Throughput is easy to sell and easy to build. These three are the reasons your team hasn't shipped the automation yet.

Regulatory exposure

A wrong answer becomes an audit finding

A questionnaire response is a written representation. When the drafting tool invents one, the failure surfaces in an audit or a deal review, and by then the paper trail is yours.

Engineering time

Output nobody trusts gets checked by hand

If any answer might be invented, a human re-reads all of them. The automation saved nothing. Grounding with enforced citations is what makes review a spot-check instead of a rewrite.

Reputation

Your customers only see the wrong one

Ninety accurate answers do not offset one false claim about encryption or pen-test cadence. The tools I build leave a gap and flag it rather than fill it.

What I build

Python, FastAPI, n8n. Retrieval, evaluation harnesses and the review step designed in from the start rather than bolted on afterwards.

Services in detail
  • Retrieval Document-grounded RAG where every claim traces to a source, and unsupported claims don't get made.
  • Compliance Questionnaire and RFP drafting from a company's own policy base, with the human review step designed in.
  • Media Transcript and caption pipelines for accessibility obligations, WCAG, the European Accessibility Act.
  • Pipelines Ingest, dedup, scoring and summarisation workflows on n8n and FastAPI.

Work, with the numbers attached

vendor-security-questionnaire-rag

Case study

Measured on 20 CAIQ v4.0.2 questions: 13 usable as drafted, 7 correctly abstained, 0 fabricated.

Turns a vendor security questionnaire into a reviewable first draft from a company's own policy documents. Citation grounding is enforced twice, independently, the second check verifies that every sentence the model cites actually appears in the retrieved evidence, and forces the answer back to "not found" when it doesn't.

anchor-align

Boundary error on untouched words: 4.0 ms, against 370.9 ms for the difflib baseline. Synthetic corpus; real-transcript validation still open.

Recovers caption timing after a human has edited the transcript, surviving both small corrections and whole sentences moved elsewhere, then segments the result into WebVTT cues that obey the standard constraints and never overlap.

Source, benchmarks and known gaps

eu-ai-newsletter

Running weekly in production.

Weekly EU AI Act tracking for compliance and legal teams: RSS ingest, deduplication against the five outlets covering the same event, relevance scoring, summarisation, and HTML issue assembly on n8n with a FastAPI sidecar.

Source

How I work

I benchmark before I claim. Every number on this site comes from a harness in a public repository that you can run yourself, and the caveats sit next to the results, including the ones that don't flatter the work.

When a measurement says a feature makes things worse, it ships disabled. Anchor-align's phonetic matching is off by default for exactly that reason, and the reason is written down in the repo.

Known gaps are tracked as open issues rather than left out of the README. You can read what doesn't work yet before you hire me.

Describe the workflow. I'll tell you if it's worth automating.

A written read on scope, risk and where a human still has to sit, replied to the same business day. If the honest answer is "don't build this", you get that instead.

Ready to test it on your own documents? A three-day proof run is €400, credited in full against a pilot.

Send a brief