The expensive failures were never crashes. They were dashboards that kept looking right after the input went wrong.
I design AI-augmented systems, direct the AI that builds them, and verify them end-to-end. The part I am actually good at is the part most people skip: writing down what wrong looks like before building the thing that could be wrong.
What I take on
AI validation architecture
You have an agent or a pipeline that produces confident output, and nobody can tell you when it is wrong. I design the states where it has to say I don't know, and the structural checks that make a missing answer impossible to mistake for a clean one.
Pricing & quoting intelligence
Quotes go out at the wrong margin for a year and nobody notices, because every individual quote looks reasonable. I build the layer that puts the declared rule and what actually shipped side by side, so the gap is a number instead of a feeling.
Adaptive learning systems
Assessment that changes what it asks next and can show why it changed. Running in production against a real multi-year academic record, with the scoring rules versioned so an old result still means what it meant when it was issued.
Three hackathon submissions in one month.
Not portfolio screenshots — public repos under an OSI license, built inside a deadline against systems I had never touched. Each one shipped with the bugs I found in the host system on the way, filed upstream where anyone can read them. Judging outcomes are listed as they stand today, won or not.
unmeasured
Apache-2.0Build with DataHub: The Agent Hackathon
Submitted Aug 7, 2026 · Judging Aug 17 – Aug 31
A catalog agent that reports FRESH, STALE, or UNMEASURED — and refuses the fourth option, which is a confident number derived from a clock that measures the wrong thing. Datasets nobody declared an SLA for become visibly unmeasured instead of silently absent from the report.
Filed upstream while building
- datahub#18753 — Platform bug: three storage layers returned three different answers about the same write. Closed as completed by a maintainer, who traced the cause to a v1.6.x fix.
- static-assets#218 — The sample dataset's own setup script silently overwrote the metadata the exercise depends on.
- static-assets#219 — The documented planted fault did not match the committed data.
- static-assets PR#220 — The fixes, each tested against a live instance. Open.
unclosed
Apache-2.0OpenSearch Agent Skills Hackathon
Submitted Aug 3, 2026 · Submission window closes Aug 17
Log root-cause analysis that audits its own premise. Before naming a cause it checks whether the window it looked at can support the claim — because the same alert window, sampled at two different moments, produced two different verdicts.
Filed upstream while building
- Hackathon issue #92
- PR #101 — Contributed to the upstream skills repo. Open.
Qwen Cloud Hackathon — Autopilot track
Submitted Jul 7, 2026 · In judging
A quoting autopilot for facilities maintenance: two-anchor pricing with a human approval gate. Built from the same domain I run in production, rebuilt from scratch inside the submission window.
Got a system producing a number nobody can verify?
Tell me what it outputs and how you would know if it were wrong. If I can help I will say so, and if the honest answer is that you do not need me, I will say that instead. Scope and rates come after that conversation, not before it.