I build AI systems that take
safety as seriously as capability.

Most AI demos look impressive but rarely survive contact with a real workplace: messy data, unclear ownership, and people who have better things to do than babysit an agent. This is my sweet spot: scaling demos into product that survives the messiness of enterprise environments.

I work between product leadership and applied AI engineering, turning ambiguous problems into AI systems people can actually use. That includes agent workflows, harness engineering, evaluation harnesses and local-first products. One end of that range is a production system helping global GTM teams search 235,000+ brand assets in natural language. The other end is smaller experiments that never leave my own machine.

Before building AI-native products became my full-time job, I built products across games, SaaS and hardware, including two startups that were acquired. Different industries, same instinct: find the part of the system that doesn't work for humans, then make it less stupid.

Flagship systems

Four systems
  1. DAM Butler MCP

    Natural-language retrieval across 235,000+ global brand assets, with an architecture adopted and shipped into daily enterprise workflows in Breville Group global GTM teams.

    productionenterprise
    MCP · Vercel · Brandfolder API · ChatGPT Enterprise
  2. VAI Santé

    A local-first, provenance-aware AI architecture for fragmented, high-stakes information.

    researchactive
    Python · Mermaid · evaluation harness
  3. Koinaku

    A mobile-first financial literacy app for Indonesian Gen Z, built with loop and graph engineering. Micro-lessons plus a paper-trading sandbox with virtual Rupiah, live in beta.

    betaproduction
    Next.js 16 · TypeScript · Supabase · Capacitor
  4. Vivid Clean

    A local-first document-cleaning tool for neuro-spicy folks, built around privacy, accessibility and user control.

    shippedaccessibility
    Bash · Python · pandoc

Experiments

Smaller builds that fed the flagship work. Same disciplines, lower stakes.

Four experiments
  1. Espresso Horoscope MCP

    Local-first MCP with strict offline boundaries: where the local-agent patterns behind Vivid Clean and VAI Santé started.

    Python · Next.js 15 · LM Studio · MCP
  2. Vivid Alpaca

    Execution guardrails between agent recommendations and real orders: the same safety-envelope discipline the flagship work depends on.

    Python · Dash · Alpaca API · multi-agent
  3. Almost

    End-to-end product shipping with real users and real payments: proof the product judgement in the flagship work is not theoretical.

    Next.js 14 · Anthropic API · Fraunces
  4. Sourdough Intelligence

    Data rigour from before LLMs existed: regression, sentiment analysis and a live app, still running.

    R · IBM Watson NLP · regression

Speaking

Two engagements
  1. Panel on stage at Vercel Ship Sydney 2026: The AI org, people, structure, and the decisions that matter
    Vercel Ship Sydney · 2026

    The AI org: people, structure, and the decisions that matter

    Panel with Tom Glover (MYOB) and Kristina Ryan (Go1), moderated by JJ Lecocq (Vercel). The interesting part was the cross-industry contrast: legacy cleanup, shadow-AI sprawl, and a modern stack running inside a 94-year-old hardware company, all wrestling with the same tension between moving fast and building foundations that hold.

    Read the recap →
  2. Speaker card for Vivid Savitri at the Sydney Enterprise AI and Automation Summit 2026
    Sydney Enterprise AI and Automation Summit · 2026

    Is your AI program hitting the mark?

    Panel discussion on enterprise AI program evaluation, hosted by Clutch Events in Sydney.

    Event details →

Writing

One piece
  1. LinkedIn · 21 August 2026

    AI watermarking is the new scarlet letter

    Every major lab is shipping watermarks as a transparency win. What actually happens: a mark that reads as cheating no matter why the AI was used, landing hardest on people with dyslexia, or other disabilities, and neurodivergent writers who rely on it to write at all. Spell-check and academic fraud getting flattened into the same signal is the actual problem.

    Read the article →

How I build

Five principles

Start with the problems and the workflow, not the tech stack

I always start with the why when bringing AI into a problem, because the how changes over time and should stay modular and flexible. Then the what: the work people are already trying to do, the decisions, handoffs, bottlenecks and failures. The models come last, and they depend on what we're trying to achieve and what success looks like.

Build the smallest system that can answer the real question

I use prototypes to test value and architecture together. A convincing demo is useful, but it isn't evidence that the system will survive contact with real users.

Define what "good" means

Before scaling, I turn expectations into representative tasks, evaluation rubrics, failure modes and review gates. Same discipline behind chain-of-custody memory and execution guardrails on live trading.

Engineer the harness

Model capability matters. So do state, context, tools, permissions, memory, observability and recovery. Most of the product lives in that surrounding system. This is the actual work behind CoworkOS (Anthropic) and WorkspaceOS (OpenAI), deployed for Finance, Legal, IP, NPD, Test Kitchens and GTM teams.

Ship, observe and revise

Production behaviour is the final argument. I look for adoption, failure patterns and the moments where people stop trusting the system.

Contact

Get in touch

I'm interested in frontier and applied AI roles where product judgement and hands-on building belong in the same job. If you're working on agent systems, evaluations, enterprise AI or trustworthy deployment, I'd like to hear what you're trying to make work.