Back to DimensionsReliability-First
Dimension: BUILD
Reliability-First
AI Architectures.
I don't build models; I build the systems that prove models work. I built my own multi-agent harness — Exponential OS — with a constitution enforcing engineering invariants, an agentic memory layer for context management, model routing, and composable skills. Then I used it to ship real product.
2026 - Present
Exponential OS — Multi-Agent Harness (exponentialos.io)
Architect and builder
- Built his own multi-agent harness: a generative-principle constitution enforcing engineering invariants as structural and semantic gates before any output ships.
- Agentic memory subsystem for context management — hydrates relevant context on pre-prompt hooks and distills lessons to a long-term index at session end, so agents do not lose knowledge across sessions.
- Control plane for task routing, model selection and cross-family verification cascades; composable skills and MCP server integrations; lifecycle hooks across the full turn.
- exponential-developer plugin: cross-LLM agentic SDLC workflow, nine stages and five hard gates. Acceptance criteria and evals fixed before code; a change ships only if it beats baseline. Parallel agent teams in isolated git worktrees, model right-sizing by task class, cross-LLM jury with cascading escalation, vision-model review of rendered UI, GitHub Actions CI, SonarQube and security scanning.
- Co-Dialectic plugin (open source): prompt improvement, context management, token efficiency, hallucination reduction. github.com/thewhyman
- jury skill: cross-family review panel — cheap models first, escalating to premium only on conflict; the independent-verification gate before anything ships.
- In daily production use, and used to deliver the AI Fund work — the shipped products are the proof it holds up in production, not a demo.
May 2026 - Jul 2026
AI Fund — Engineer in Residence (Andrew Ng's Venture Studio)
Engineer in Residence
- Owned the full loop across two major pivots: customer discovery, ICP selection, contextual inquiry and Mom Test interviews, with demand validated before scaling the build.
- Validated demand against ~150 companies and 36 grounded outreaches; 19 of 22 confirmed the technical gap but not commercial urgency, so he killed the wedge on evidence with minimal sunk cost.
- Built brand-voice personas on embeddings, style-RAG and prompt tuning; designed a four-stage LLM workflow with layered auditors — deterministic checks for what is mechanically verifiable, semantic LLM judges for what is not — plus eval suites gating each stage, targeting AI slop reduction and humanized output.
- Model evaluation and selection: stood up the measurement spine first (two-judge ensemble scoring three axes — voice fidelity, coherence, audience fit), then ran controlled fine-tuning experiments on Oumi against Qwen base models. Full fine-tuning on a small, structurally-uniform corpus regressed coherence and fluency (catastrophic forgetting); he diagnosed the cause, moved to parameter-efficient LoRA, and set the decision ladder: prompt, then RAG, then fine-tune last.
- PostHog KPI instrumentation to verify shipped features moved behavior.
- The EIR concluded July 2026 having completed the exploration on social-media post adaptation; the technical build shipped but the commercial signal was not strong enough to advance to fund.
2019 - 2025
Google Platform Engineering (~$40B Portfolio)
Lead Architect & Senior SWE Manager
- Delivered $500M+ ROI by engineering the global platform handling Google's $40B annual data center and office construction portfolio.
- Architected critical system migrations including DC Security VM integration and AODocs platform migrations.
- Alleviated 'Code Purple' delays through a 400% performance improvement on backend scheduling architectures.
- Navigated complex Google Security protocols to deploy Vizzy, the first-ever 1P Android app launched to the public Play Store.
2009 - 2013
Trellis FISMA/NIST Private Cloud Migration
Manager, Solutions Architecture (NIST-Certified Architect)
- Led enterprise-wide regulatory compliance (FISMA, NIST, PCI) for the nation's third-largest student loan guarantor.
- Architected and deployed a FISMA-certified private cloud PaaS from scratch, eliminating $1M in software licensing costs.
- Engineered rigorous static/dynamic code testing paradigms complying with OWASP security standards to secure federal contracts.
2015 - 2018
Charles Schwab Cloud-Native Transformation
Technical Director
- Led the enterprise-wide transition to Pivotal Cloud Foundry (PaaS) to support millions of daily transactions across a strictly regulated financial infrastructure.
- Authored the '12 Cloud Native Principles', creating the reference architecture for 7 trading and compliance platforms.
- Compressing multi-quarter infrastructure delivery cycles from months to weeks by standardizing deployment pathways.
2026
Medicaid FQHC Copilot
Independent Applied AI Engineer
- Engineered a HIPAA-compliant ReAct-style agent to automate Medicaid eligibility calculations for caseworkers.
- Designed a 5-layer defense architecture, replacing naive LLM math with a deterministic Python ground-truth engine.
- Enforced strict data governance using per-patient memory scoping to adhere to PII and HIPAA regulations.
Published Thought Leadership
The Applied AI Philosophy.
© 2026 The Why Man Hub
What brings you here?
Pick the closest fit and I will point you at the right conversation.
Hiring
Recruiter or hiring manager
Senior engineering leadership, applied AI architecture, forward-deployed roles
Continue
Building
Collaborator or founder
Agentic systems, harness architecture, 0→1 product work
Continue
Advising
Consulting or advisory
AI strategy, eval architecture, engineering governance
Continue
Prefer to ask first? Interview my AI concierge — bottom right of any page.