▸ ls ~/projects

PROJECTS

Things I've built where data & agents meet.

abacus.exe

Abacus

A data analyst you can interrogate

Data · AI agents · 2026

A data analyst you can interrogate. It does not just answer; it investigates. Ask "why did revenue jump?" and it breaks the change across every dimension into a waterfall that reconciles to the dollar. Ask for cohort retention or an anomaly scan and it computes those live too, against a 125,000-row warehouse running on SQLite-in-WebAssembly in your browser. A pixel crew of four talks through each step. A public 35-question eval shows the browser engine matches its Python twin to the cent. Paste your API key and a live LLM takes the planner's seat, with guardrails that keep it from ever writing raw SQL.

STACK

    ↳ How it fitsOne semantic layer defines every metric exactly once for both engines. The deterministic parser and the LLM planner emit the same plan schema, checked one key at a time before the guarded compiler writes any SQL. Your key stays in your tab. Calls go straight from the browser to the provider.

    oracle.exe

    Oracle

    Multi-agent reasoning engine

    Agentic AI · 2026

    Put your data in front of a skeptic, an optimist, a fact-checker, and a wildcard. Each agent runs on a different frontier model. They debate until they reach consensus, then hand back an answer with a confidence score pulled straight from their disagreements.

    STACK

      ↳ How it fitsEach agent lives in its own LangGraph node and talks to Claude, GPT, Gemini, or Grok. They really do disagree. A moderator weighs the arguments and turns the spread of opinion into the confidence score. The disagreement is the metric.

      armada.fleet

      Armada

      A control plane for browser fleets

      Distributed Systems · Agentic Automation · 2026

      Armada runs fleets of real browsers as a managed service. Operators build multi-tenant campaigns in a React console, then a scheduler spreads the work across the day. Each run opens a live Chromium session, hands the page's accessibility tree to a frontier model, and accepts one schema-checked action in return.

      STACK

        ↳ How it fitsThe less glamorous platform work keeps the fleet alive. The scheduler handles cron, jitter, and ramp-up. Warmed persistent identities come from a pool, health-scored proxies disable themselves, failures are classified with alerting, and costs are metered against a hard ceiling that stops a runaway run. Every action must match a strict JSON schema, so the model can only click, fill, scroll, or stop. I can VNC into any live worker and watch it think. Postgres and Redis hold it together, unattended, on a single Mac Mini in the corner.

        Live Demo ↗ Production system is private; this demo replays recorded fixtures.
        operator.call

        Operator

        AI receptionist that answers the phone

        Voice AI · Twilio · 2026

        Operator answers real phone calls, talks naturally, figures out what the caller needs, and books the appointment on a live calendar. No hold music. No phone tag. It uses a real Twilio number, so callers just hear someone pick up the phone.

        STACK

          ↳ How it fitsTwilio streams in the call audio and Whisper turns it into text. An agent works out what to say, checks the calendar for an open slot, speaks the answer, and writes the booking. It all moves fast enough to feel like a normal conversation.

          Live Demo ↗ Production system is private; this demo replays recorded fixtures.
          churn.ipynb

          Churn Radar

          Telecom customer churn prediction

          Machine learning · 2026

          Churn Radar flags the telecom customers most likely to cancel before they do. It was trained on 7,043 real customer records from the public IBM Telco dataset, then exported to score customers live in your browser and show the exact reasons behind every prediction.

          STACK

            ↳ How it fitsI cleaned the data in pandas, built an explicit one-hot feature matrix, and tested logistic regression against gradient boosting. The readable model wins, with an AUC of 0.845 versus 0.830. Each prediction is the sum of named coefficients, so the live demo shows exactly why a customer got that score. It is the model's own arithmetic, not a post-hoc explainer.

            helmsman.yaml

            Helmsman

            Agentic DevOps copilot

            DevOps · 2026

            A side project for making DevOps on-call a little less painful. It watches the cluster. When something breaks, it reads the logs, explains the likely cause in plain English, and drafts the kubectl fix with the rollback staged up front. A human has to approve it before anything touches production.

            STACK

              ↳ How it fitsIt calls kubectl for pods, logs, and events, matches the evidence to vetted failure signatures, then drafts a fix from templates with a risk grade and staged rollback. The approval gate lives in code, not convention. The subprocess layer refuses mutating commands until a human hits approve.

              cue.live

              Cue

              Live-call copilot for new hires

              Voice AI · Onboarding · 2025

              Cue is a real-time assist panel for newly onboarded team members. It follows the live call and reads the screen, surfacing the right answer, account detail, or next thing to say while the rep is still talking. New hires can handle calls with confidence while they ramp up, without stalling or putting anyone on hold. It follows the conversation, not just the transcript.

              STACK

                ↳ How it fitsIt captures the call audio and the rep's screen, then transcribes locally with almost no lag. That context streams to a model, which drafts the next answer or step for a glanceable side panel. New team members get live guidance without breaking the flow of the call.

                Live Demo ↗ Production system is private; this demo replays recorded fixtures.
                forge.dockerfile

                Forge

                Isolated runtime for agents

                Infra · 2024

                Forge gives every agent a disposable Linux sandbox for running code, browsing, and calling tools safely. Each task gets a fresh container, scheduled across the cluster and torn down as soon as the agent finishes.

                STACK

                  ↳ How it fitsA Go orchestrator schedules one Docker sandbox per agent task on Kubernetes and streams stdout and results back over NATS. Redis tracks state, so nothing leaks from one run to the next.

                  Live Demo ↗ Production system is private; this demo replays recorded fixtures.
                  atlas.db

                  Atlas

                  Knowledge agent for ops

                  RAG · 2026

                  Ask Atlas, “how do we rotate the API keys?” It answers from your runbooks, past incidents, and live system state, with citations, so nobody has to dig through a wiki at 2 a.m.

                  STACK

                    ↳ How it fitsLangChain splitters chunk the runbooks, incidents, PDFs, CSVs, and diagram notes, then they are embedded into Postgres with pgvector. A hand-rolled hybrid ranker combines cosine similarity and keywords behind a relevance floor. Groq drafts the cited answer quickly, and Redis caches hot queries for instant repeats. The live demo runs the same pipeline entirely in your browser with real embeddings and no server.

                    conductor.db

                    Conductor

                    Data-pipeline reliability agent

                    Data engineering · 2026

                    When the nightly pipeline derails, Conductor reads the run ledger and probes the warehouse with real SQL. Maybe a vendor renamed a column, join keys changed shape, or a file loaded twice. It explains the break in plain English with quoted evidence, then drafts a repair with the rollback staged. A human clears every fix before a single row gets rewritten.

                    STACK

                      ↳ How it fitsThe staging-and-facts pipeline really runs, with SQL and data-quality checks after every task. Diagnoses use deterministic signature matches over probed evidence, and fixes come from vetted templates. Destructive repairs snapshot the table first. The demo replays a genuine end-to-end run: the data actually broke, and it actually got repaired.

                      pulse.wire

                      Pulse

                      Live macro briefing desk

                      Business intelligence · 2026

                      Pulse writes a morning brief from live data. Your browser pulls numbers straight from public feeds: ECB exchange rates, the US Treasury's Debt to the Penny ledger, and World Bank indicators. It charts them and writes the day's plain-English brief on the spot. There is no server or key, and no stale screenshot. The numbers are whatever the world says when you open it.

                      STACK

                        ↳ How it fitsThe whole brief runs client-side. The browser builds the dollar scoreboard and calculates the debt's pace in dollars per second, clearly labeled as an extrapolation between daily Treasury postings. It also handles the math for each resident. Every claim carries its source and fetch time. If a feed is down, Pulse says so. The same engine ships as a Python CLI with offline fixture tests.

                        coming-soon

                        More on the way

                        Projects in progress

                        Coming soon

                        I'm always building something. A few more projects are in progress and will show up right here.