Most work in the demo and break under real traffic. I design the architecture, evaluations, and infrastructure that hold — then hand your team the keys.
Shipping LLM systems since GPT-3 · 300+ engineers trained
Shipped with evidence — not vibes.
Example trace — a production RAG eval, anonymized. Hover a span.
Trusted by teams who ship
From Fortune 500 to Series A — and the AI-infra companies themselves. The common thread: they needed it to work, not just demo.
The record
Building production LLM systems before “AI engineer” was a job title.
The capability stays after the engagement ends. You keep the muscle.
Enterprise, fintech, funds, and the eval companies other labs rely on.
From the lab · private beta 2026
The eval layer I kept rebuilding for clients, now a product. One line of code turns on OpenTelemetry-native observability, a swapped-position judge panel, and human-in-the-loop quality gates.
Explore evalOSWhat I do
RAG, agents, and eval loops that hold under real traffic — not benchmark traffic.
For funds and acquirers: what's actually under the hood, and what breaks at scale.
Senior technical direction through high-growth phases, without the $500K hire.
Nothing below $10K. Priced on outcomes, not hours.
How we engage
I integrate with your team — standups, PRs, Slack — and stay until it ships. Not a deck and a disappearance.
A specific problem, scoped and solved. Fixed scope, fixed timeline, fixed price. You get the deliverable, not a report about it.
The bench
Not contractors — collaborators. Each could run their own shop.
I take on limited engagements. If it's not a fit, I'll tell you. If it is, we move fast.
Response within 48 hours · no sales process · direct conversation