Brandon Robinson United States, open to relocation. Contact

Systems I built, and what they proved.

The systems on this page run on n8n, Python and TypeScript with Postgres behind them: a 92-node production pipeline delivered to a client, an outbound program measured until its own numbers ended it, and the tools my own work runs on. I specify them, build the workflow graphs myself, direct the code, run them in production, and shut them down when the evidence says so.

I am looking for AI operations, implementation and program leadership roles: the person who puts a system into production and makes a team's use of it real.

Applicant pipeline schematic Three scheduled chains: intake, outreach and replies. A person approves every name before outreach. Every chain writes to one shared log. every hour on the hour every 15 min intake dedupe score, route a person approves approved only daily cap, pacing send, keep thread new replies match to person classify, update one log of everything
Applicant pipeline. Three chains, one gate, one log.92 nodes, April to July 2026

Written up in full

Full write-ups. Each one ends with what happened after delivery.

  • Three scheduled chains and an error subworkflow, delivered to a regional distributor and run by their team.
    Client production, April to July 2026n8n, Google Sheets, Gmail, Postgres
  • Built with the compliance and audit chain in from the start, then shut down on that chain's evidence instead of run on sunk cost.
    My practice, April to August 2026n8n, Apify, NeverBounce, Google Sheets
  • A deterministic path and two local language models scored on the same public cases, with an attestation gate before any number ships.
    My practice, June 2026Python, standard library only, local modelscode

Systems my own work runs on

One line each. The code is the write-up.

  • A read-only MCP server that turns two operational Notion databases, an audit trail and an action queue, into four native query tools for AI coding clients, with enum values validated against the live schema. Specified by me, code written under my direction, run by me daily.
    Live since May 2026; hosted May to September, local sinceNode.js, TypeScriptcode
  • Six workflows that read receipt email and bank statements into a budget, idempotent by design, with a four-state review that routes only the records needing a decision to me. Specified and directed by me.
    Running since May 2026; 93 transactions into 19 categories, May to August 2026n8n, Claude, Notion
  • A work queue for AI coding agents
    Sessions claim work from a board, stop for a person's approval before executing, and write blockers to the record the moment they hit one. Seven skills and a template, in daily use.
    Since July 2026Claude Code, Notion

Applicant intake and re-engagement pipeline

Client production, April to July 2026. Built by me in n8n; the client's team operated it.

A regional distributor was losing hires between "applied" and "first contact". The gap showed up in their own sheet: rows arriving faster than the contacted column filled in, and follow-up depending on who had time that day. I built the system that closes that gap: three independently scheduled chains that take raw applicant rows, deduplicate and score them, queue the qualified ones for a person to approve, send the re-engagement email, and match every reply back to its candidate.

Applicants processed10,580
Classification coverage100%, with 71% filtered out before a person saw them
Nodes in the final graph92, across three chains and an error subworkflow
Where it rann8n on a hardened Linux server, Postgres behind it, Google Sheets as the operator's desk
What decidesKeyword scores and hard pre-gates. No model makes the call.

How it works

Every routing decision, every duplicate, every send and every halt lands in one log sheet, so the question "what did the system do with this applicant" always has an answer. Sends are capped per day and paced three minutes apart. A classified failure halts the chain and writes a halt entry instead of cascading.

The full node map of the production workflow: 92 nodes in three horizontal chains with an error subworkflow, drawn from the workflow export with node names only.
Every node in the graph as it ran, drawn from the workflow export. Names only; no parameters, data or credentials. Open it to read the nodes.

Decisions I'd defend

  • Keyword scores decide who gets contacted, not a language model.Every decision writes its reason to the log and gives the same answer next month. An operator can read the rule and move a threshold. A model would have added a token cost and a second thing to audit on a chain that emails real people. Two months later my evaluation harness measured a related question on public data: the deterministic path within 0.8 points of the best local model at zero token cost.
  • The operator's desk is a spreadsheet.The client's team already worked in Sheets, so approval is flipping one cell and nobody learns a new tool. The cost is weaker validation and no locking, so the pipeline re-reads the sheet on every run and never trusts a value it wrote last time.
  • A classified failure halts the chain instead of retrying.An email cannot be unsent, and a retry loop is the one way to contact someone twice. A halt writes a log entry and waits for a person. The daily cap is read from the log itself, so a restart cannot reset it.

What it proved

The client's team ran the approval gate themselves and released 29 names for outreach, the way it was designed to be run. They used it once, on a stale list, and concluded email did not work for them. It did not become part of how they hire.

Both halves are true, and the second one changed how I scope and hand off every build since.

The number that matters here is not throughput. A three-day silent failure was found by reading the destination sheet, not the run status. The run status said "success".

What is shown here, and what is not

  • Node map, drawn from the export, names onlyshown
  • Log sheet schema and one redacted incident timelineshown
  • Client name, applicant data, sheet, mailbox, servernever

An outbound email program, measured until its own numbers ended it

My practice, April to August 2026. Full stop on 9 August 2026.

The practice ran on outbound: find small firms losing revenue to a broken intake, reach them with a credible offer. I built the sending system with the compliance and audit chain in from the first send, because the one thing an outbound program cannot recover from is its own domain reputation. Then the chain did its job in the other direction.

Real sends, counted at the mailbox240
Genuine replies0
Sent bodies carrying a postal address and an opt-out238 of 238 verified
Authentication in a real DMARC aggregateDKIM 29 of 29, SPF 29 of 29
Sourcing pipeline behind itFive workflows, 87 nodes, verified against the running instance

How it works

Every send-authorizing parameter lives in one configuration sheet: the kill switch, the send window, the per-run caps, the bounce and complaint thresholds. A suppression ledger is checked before every send, and opt-outs are captured from replies twice a day. Every send writes a fourteen-column audit row, and if that write fails the batch halts and the row goes to an inbox instead of being lost.

Decisions I'd defend

  • The kill switch is a cell in a sheet, not a line in a workflow.Send policy has to be changeable by the person accountable for it without an edit to running code. A false value in one cell stops every run. It was flipped to false when the evidence came in.
  • Compliance is built in, not checked after.The footer address and opt-out are resolved at one chokepoint the send cannot bypass, which is why 238 of 238 verified bodies carried them. It also meant I could rule out compliance as the cause of silence in one read.
  • Stop on the evidence.240 sends and no genuine reply is a signal, and the options were warm more domains or find out why. Authentication, blacklisting, domain age and volume were each eliminated from the system's own records. What remained was content filtering, confirmed by the provider's own notice. The program was ended on that finding rather than kept alive on what it had cost.

What it proved

Commercially zero, and provably so. The sourcing pipeline did its job; the outbound program it fed did not, and its own instrumentation said so clearly enough to act on.

Ending a program you built, on its own data, is the part of this I would want a hiring manager to notice.

What is shown here, and what is not

  • Gate design, audit row schema, the shutdown reasoningshown
  • Recipients, sender domains, the suppression ledger, any message bodynever

AI output evaluation harness

My practice, June 2026. Scoped and directed by me; the code was written under my direction.

A tool for a question most evaluation tooling assumes away: is the language model the right tool for this classification task, or does a cheaper, auditable deterministic method do the job? The harness runs both over the same labelled cases, scores each against the gold labels, routes low-confidence cases to human review, and emits a comparison report a reader can check.

AttestationPassed on 11 known-answer checks before any metric was allowed out
Test set120 public AG News cases; no client data
Deterministic pathEmbedding nearest-neighbour, 80.0% at zero token cost
Language models, localmistral 80.8%, llama3 79.2%
DependenciesNone beyond the Python standard library and a local model server

Decisions I'd defend

  • No metric ships until the scorer proves itself on hand-computed answers.The only real risk in an offline evaluation tool is a confident wrong report. The attestation suite is that risk's gate, and it is the reason the tool is Tier 1 rather than something heavier.
  • Zero dependencies, local models.Anyone can run it on a laptop with nothing to install, and every report carries the dataset hash, the model digests and the seed, so a number can be reproduced instead of trusted.
  • Public data only.The pattern it generalizes came from a client pipeline; the cases it runs on come from a public benchmark. The client's data never enters the repository.

What it proved

The deterministic path came within 0.8 points of the best local model at zero token cost. That answers whether a model was warranted for the task rather than assuming it was.

It was built to portfolio grade and never shipped to a user. That is the honest state of it.

What is shown here, and what is not

  • The repository, design notes and a sample reportshown
  • The client pipeline's data, referenced as anonymized narrative onlynever

Personal financial operations pipeline

Personal, running since May 2026. Specified and directed by me.

Six workflows on self-hosted n8n: receipt email is read by a language model into structured rows, bank statements are read by a vision model, and a daily job links every row to its budget line. Idempotent by design, so re-running a month writes nothing twice. A four-state review routes only the records that need a decision to me; I do not look at every record, and the design does not ask me to. Between May and August 2026 it wrote 93 transactions into a 19-category database. Nothing from it is shown here beyond the design; it is my money.