DAM Butler MCP
Natural-language retrieval across 235,000+ global brand assets, with an architecture adopted and shipped into daily enterprise workflows in Breville Group global GTM teams.
Most AI demos look impressive but rarely survive contact with a real workplace: messy data, unclear ownership, and people who have better things to do than babysit an agent. This is my sweet spot: scaling demos into product that survives the messiness of enterprise environments.
I work between product leadership and applied AI engineering, turning ambiguous problems into AI systems people can actually use. That includes agent workflows, harness engineering, evaluation harnesses and local-first products. One end of that range is a production system helping global GTM teams search 235,000+ brand assets in natural language. The other end is smaller experiments that never leave my own machine.
Before building AI-native products became my full-time job, I built products across games, SaaS and hardware, including two startups that were acquired. Different industries, same instinct: find the part of the system that doesn't work for humans, then make it less stupid.
Natural-language retrieval across 235,000+ global brand assets, with an architecture adopted and shipped into daily enterprise workflows in Breville Group global GTM teams.
A local-first, provenance-aware AI architecture for fragmented, high-stakes information.
A mobile-first financial literacy app for Indonesian Gen Z, built with loop and graph engineering. Micro-lessons plus a paper-trading sandbox with virtual Rupiah, live in beta.
A local-first document-cleaning tool for neuro-spicy folks, built around privacy, accessibility and user control.
Smaller builds that fed the flagship work. Same disciplines, lower stakes.
Local-first MCP with strict offline boundaries: where the local-agent patterns behind Vivid Clean and VAI Santé started.
Execution guardrails between agent recommendations and real orders: the same safety-envelope discipline the flagship work depends on.
End-to-end product shipping with real users and real payments: proof the product judgement in the flagship work is not theoretical.
Data rigour from before LLMs existed: regression, sentiment analysis and a live app, still running.
More on GitHub →
Panel with Tom Glover (MYOB) and Kristina Ryan (Go1), moderated by JJ Lecocq (Vercel). The interesting part was the cross-industry contrast: legacy cleanup, shadow-AI sprawl, and a modern stack running inside a 94-year-old hardware company, all wrestling with the same tension between moving fast and building foundations that hold.
Read the recap →
Panel discussion on enterprise AI program evaluation, hosted by Clutch Events in Sydney.
Event details →I always start with the why when bringing AI into a problem, because the how changes over time and should stay modular and flexible. Then the what: the work people are already trying to do, the decisions, handoffs, bottlenecks and failures. The models come last, and they depend on what we're trying to achieve and what success looks like.
I use prototypes to test value and architecture together. A convincing demo is useful, but it isn't evidence that the system will survive contact with real users.
Before scaling, I turn expectations into representative tasks, evaluation rubrics, failure modes and review gates. Same discipline behind chain-of-custody memory and execution guardrails on live trading.
Model capability matters. So do state, context, tools, permissions, memory, observability and recovery. Most of the product lives in that surrounding system. This is the actual work behind CoworkOS (Anthropic) and WorkspaceOS (OpenAI), deployed for Finance, Legal, IP, NPD, Test Kitchens and GTM teams.
Production behaviour is the final argument. I look for adoption, failure patterns and the moments where people stop trusting the system.
Contact
I'm interested in frontier and applied AI roles where product judgement and hands-on building belong in the same job. If you're working on agent systems, evaluations, enterprise AI or trustworthy deployment, I'd like to hear what you're trying to make work.