Brandon RobinsonUnited States, open to relocation. Contact
Systems I built, and what they proved.
The systems on this page run on n8n, Python and TypeScript with Postgres behind them: a 92-node production pipeline delivered to a client, an outbound program measured until its own numbers ended it, and the tools my own work runs on. I specify them, build the workflow graphs myself, direct the code, run them in production, and shut them down when the evidence says so.
I am looking for AI operations, implementation and program leadership roles: the person who puts a system into production and makes a team's use of it real.
Applicant pipeline. Three chains, one gate, one log.92 nodes, April to July 2026
Written up in full
Full write-ups. Each one ends with what happened after delivery.
A read-only MCP server that turns two operational Notion databases, an audit trail and an action queue, into four native query tools for AI coding clients, with enum values validated against the live schema. Specified by me, code written under my direction, run by me daily.
Live since May 2026; hosted May to September, local sinceNode.js, TypeScriptcode
Six workflows that read receipt email and bank statements into a budget, idempotent by design, with a four-state review that routes only the records needing a decision to me. Specified and directed by me.
Running since May 2026; 93 transactions into 19 categories, May to August 2026n8n, Claude, Notion
A work queue for AI coding agents
Sessions claim work from a board, stop for a person's approval before executing, and write blockers to the record the moment they hit one. Seven skills and a template, in daily use.
Since July 2026Claude Code, Notion
Applicant intake and re-engagement pipeline
Client production, April to July 2026. Built by me in n8n; the client's team operated it.
A regional distributor was losing hires between "applied" and "first contact". The gap showed up in their own sheet: rows arriving faster than the contacted column filled in, and follow-up depending on who had time that day. I built the system that closes that gap: three independently scheduled chains that take raw applicant rows, deduplicate and score them, queue the qualified ones for a person to approve, send the re-engagement email, and match every reply back to its candidate.
Applicants processed
10,580
Classification coverage
100%, with 71% filtered out before a person saw them
Nodes in the final graph
92, across three chains and an error subworkflow
Where it ran
n8n on a hardened Linux server, Postgres behind it, Google Sheets as the operator's desk
What decides
Keyword scores and hard pre-gates. No model makes the call.
How it works
Every routing decision, every duplicate, every send and every halt lands in one log sheet, so the question "what did the system do with this applicant" always has an answer. Sends are capped per day and paced three minutes apart. A classified failure halts the chain and writes a halt entry instead of cascading.
Every node in the graph as it ran, drawn from the workflow export. Names only; no parameters, data or credentials. Open it to read the nodes.
Decisions I'd defend
Keyword scores decide who gets contacted, not a language model.Every decision writes its reason to the log and gives the same answer next month. An operator can read the rule and move a threshold. A model would have added a token cost and a second thing to audit on a chain that emails real people. Two months later my evaluation harness measured a related question on public data: the deterministic path within 0.8 points of the best local model at zero token cost.
The operator's desk is a spreadsheet.The client's team already worked in Sheets, so approval is flipping one cell and nobody learns a new tool. The cost is weaker validation and no locking, so the pipeline re-reads the sheet on every run and never trusts a value it wrote last time.
A classified failure halts the chain instead of retrying.An email cannot be unsent, and a retry loop is the one way to contact someone twice. A halt writes a log entry and waits for a person. The daily cap is read from the log itself, so a restart cannot reset it.
What it proved
The client's team ran the approval gate themselves and released 29 names for outreach, the way it was designed to be run. They used it once, on a stale list, and concluded email did not work for them. It did not become part of how they hire.
Both halves are true, and the second one changed how I scope and hand off every build since.
The number that matters here is not throughput. A three-day silent failure was found by reading the destination sheet, not the run status. The run status said "success".
An outbound email program, measured until its own numbers ended it
My practice, April to August 2026. Full stop on 9 August 2026.
The practice ran on outbound: find small firms losing revenue to a broken intake, reach them with a credible offer. I built the sending system with the compliance and audit chain in from the first send, because the one thing an outbound program cannot recover from is its own domain reputation. Then the chain did its job in the other direction.
Real sends, counted at the mailbox
240
Genuine replies
0
Sent bodies carrying a postal address and an opt-out
238 of 238 verified
Authentication in a real DMARC aggregate
DKIM 29 of 29, SPF 29 of 29
Sourcing pipeline behind it
Five workflows, 87 nodes, verified against the running instance
How it works
Every send-authorizing parameter lives in one configuration sheet: the kill switch, the send window, the per-run caps, the bounce and complaint thresholds. A suppression ledger is checked before every send, and opt-outs are captured from replies twice a day. Every send writes a fourteen-column audit row, and if that write fails the batch halts and the row goes to an inbox instead of being lost.
Decisions I'd defend
The kill switch is a cell in a sheet, not a line in a workflow.Send policy has to be changeable by the person accountable for it without an edit to running code. A false value in one cell stops every run. It was flipped to false when the evidence came in.
Compliance is built in, not checked after.The footer address and opt-out are resolved at one chokepoint the send cannot bypass, which is why 238 of 238 verified bodies carried them. It also meant I could rule out compliance as the cause of silence in one read.
Stop on the evidence.240 sends and no genuine reply is a signal, and the options were warm more domains or find out why. Authentication, blacklisting, domain age and volume were each eliminated from the system's own records. What remained was content filtering, confirmed by the provider's own notice. The program was ended on that finding rather than kept alive on what it had cost.
What it proved
Commercially zero, and provably so. The sourcing pipeline did its job; the outbound program it fed did not, and its own instrumentation said so clearly enough to act on.
Ending a program you built, on its own data, is the part of this I would want a hiring manager to notice.
What is shown here, and what is not
Gate design, audit row schema, the shutdown reasoningshown
Recipients, sender domains, the suppression ledger, any message bodynever
AI output evaluation harness
My practice, June 2026. Scoped and directed by me; the code was written under my direction.
A tool for a question most evaluation tooling assumes away: is the language model the right tool for this classification task, or does a cheaper, auditable deterministic method do the job? The harness runs both over the same labelled cases, scores each against the gold labels, routes low-confidence cases to human review, and emits a comparison report a reader can check.
Attestation
Passed on 11 known-answer checks before any metric was allowed out
Test set
120 public AG News cases; no client data
Deterministic path
Embedding nearest-neighbour, 80.0% at zero token cost
Language models, local
mistral 80.8%, llama3 79.2%
Dependencies
None beyond the Python standard library and a local model server
Decisions I'd defend
No metric ships until the scorer proves itself on hand-computed answers.The only real risk in an offline evaluation tool is a confident wrong report. The attestation suite is that risk's gate, and it is the reason the tool is Tier 1 rather than something heavier.
Zero dependencies, local models.Anyone can run it on a laptop with nothing to install, and every report carries the dataset hash, the model digests and the seed, so a number can be reproduced instead of trusted.
Public data only.The pattern it generalizes came from a client pipeline; the cases it runs on come from a public benchmark. The client's data never enters the repository.
What it proved
The deterministic path came within 0.8 points of the best local model at zero token cost. That answers whether a model was warranted for the task rather than assuming it was.
It was built to portfolio grade and never shipped to a user. That is the honest state of it.
The client pipeline's data, referenced as anonymized narrative onlynever
Personal financial operations pipeline
Personal, running since May 2026. Specified and directed by me.
Six workflows on self-hosted n8n: receipt email is read by a language model into structured rows, bank statements are read by a vision model, and a daily job links every row to its budget line. Idempotent by design, so re-running a month writes nothing twice. A four-state review routes only the records that need a decision to me; I do not look at every record, and the design does not ask me to. Between May and August 2026 it wrote 93 transactions into a 19-category database. Nothing from it is shown here beyond the design; it is my money.