I turn operational problems into tested agent systems, decision tools, and reusable workflows.
This is my public build record: working artifacts, adversarial evaluations, documented failure cases, and explicit boundaries.
One reference path across three libraries:
Decide → Act → Prove
testbench Consequence Rail MandateBound
- Decide: Constitutional Agent Testbench evaluates structured agent JSON against a declared policy.
- On pass → Act: Consequence Rail reserves recourse, executes, and settles or compensates.
- On dispute → Prove: MandateBound runs evidence-readiness simulation for review.
npm run bootstrap
npm run demo # decide → act (settled)
npm run demo:dispute # decide → act → prove (compensated)-
Constitutional Agent Testbench: Deterministic Python policy evaluation and PrecedenceTrace for structured agent responses.
-
Consequence Rail: Recourse-gated execution, recovery preflight, and signed settlement receipts.
-
MandateBound: Evidence readiness and deterministic dispute replay for UCP/AP2 agentic commerce.
-
TraceCanary: Desktop GUI and CLI for detecting privacy regressions in OpenTelemetry GenAI exports with synthetic canaries.
-
Agent Team: Bounded specialist-agent workflows with independent auditing.
- Corridor Lab: Desktop GUI and CLI for comparing fictional cross-border payment routes across cost, speed, liquidity, and failure assumptions.
- Unconventional Moves: Practical, non-obvious approaches with reversible 48-hour tests.
- Hermes Parallel Follow-ups: Drop-in patches and regression tests that preserve message boundaries and route independent follow-ups while Hermes is busy.
Business problem → specification → implementation → adversarial evaluation → acceptance
I direct problem selection, product direction, requirements, business judgment, evaluation, rights review, and final acceptance for the projects published here.