Wind-whipped Convict Lake under freshly snow-covered peaks in the Eastern Sierra

AI coding agents, made trustworthy · any provider

AI made code cheap to write. I make it trustworthy.

I help engineering teams put their AI coding agents on a hardened process: gates that check outcomes, sub-agents with small contexts, an outer loop that reviews every PR, and evals built from real feedback. Short sprints, in your repo, on your tools.

One engagement at a time. 1 and 2 week engagements, remote or onsite.

Jerry Fernholz
Jerry FernholzVP of Software Engineering at Broadvoice, 2018 to 2024. Founder of Rentables. Building Based Budget with AI agents.
Convict Lake, Eastern Sierra
One ticket, one deploymentRelease process I built as VP of Software Engineering at Broadvoice (2018 to 2024). Operational efficiency up 30%
100% uptime SLAs, made profitableThrough architecture and observability work
40+ engineers, 6 time zoneseNPS of 90+, and a team scaled 650%
Founder, Rentables (2012 to 2018)From rails new to hundreds of paying customers

The problem

Clean diffs. Bad premises.

Agents are very good at producing code that looks right. The tests pass, the diff is tidy, the PR description is confident. What goes wrong is upstream: the agent misread the ticket, tested that a method was called instead of what it did, or shipped a feature nothing downstream can use.

Review every line

Safe, but you gave back most of the speed you bought.

Trust the output

Fast, until a quiet failure reaches a customer.

There is a third option: make the process catch it. Rules the agent actually follows, gates that check outcomes, and a loop that tightens itself every time something slips. A few rules from my own repo:

  • Tests assert outcomes, not that a method was called.
  • No silent failures.
  • A green ticket that cannot feed its downstream consumer is not done.
src/billing/invoice.test.ts
+ it("sends invoice on close", async () => {
+   await closeMonth(account);
+   expect(mailer.send).toHaveBeenCalled();
+ });
- // TODO: verify totals match ledger
✓ 48 passing
Gate: test asserts a call, not an outcome. Which invoice? What total? Blocked until it checks the result.

Offers

Short engagements with a clear finish line.

Fixed scope, fixed price. No long transformation programs, no retainer required.

Agent Readiness Review

1 week · remote

A straight answer on where your AI coding workflow will hurt you, and a ranked list of fixes.

  • Review of agent configs, repo rules, CI gates, and recent AI-assisted PRs
  • Written findings with severity and effort
  • 60-minute readout with your leads
  • Fee credited toward a Sprint booked within 30 days
$5,000fixed
Founding clients. I'm taking up to 3 teams at 35% off list in exchange for a written case study and a reference call. Ask about it on the call.

Not sure which fits? The 30-minute call is for exactly that. If none of them fit, I will say so.

How it works

Two loops. One builds, one watches.

This is the setup I run on my own product every day. In a sprint I adapt it to your repo, your tools, and your team's habits. It's plain files in your repo, so it works with whichever coding agent you use.

Inner and outer loop An inner loop of five steps (capture, plan, kickoff, open PR, ship) with gates and evals at the center, surrounded by an outer loop where a Chief of Staff agent and a Product agent watch every PR and feed corrections back to the builder. OUTER LOOP: WATCHES EVERY PR FEEDS FIXES BACK TO THE BUILDER Chief of Staff Product bot Capture Plan Kickoff Open PR Ship Gates + evals small-context sub-agents self-improving commands INNER LOOP
  1. CaptureA request becomes a crisp ticket with acceptance criteria, before any code exists.
  2. PlanThe work is split into small PRs, each with a named gate: how we'll know it actually works. Bad premises die here, cheaply.
  3. KickoffA test-author sub-agent writes failing behavior tests first, then an implementer builds against them. Each keeps a small context, so each follows its rules.
  4. Open PRGates run. Tests must assert outcomes, and stale docs block the PR. An architecture-reviewer sub-agent reviews in rounds until it's clean.
  5. ShipA person merges. The gate is checked against reality, not just a green build, and the agent fixes its own commands so the next run is better.
The outer loop sits above the builder. A chief-of-staff agent turns decisions into briefs and checks each PR against its brief; a product agent checks user-facing changes. Neither changes code or merges. Eval sets curated from real feedback keep score over time.

Proof

I run this every day before I sell it.

Live case study

Based Budget

A personal-finance app with an AI assistant named Finn, at getbased.app. I build it with the harness above, a few hours a day, as a product owner more than a programmer.

With this setup, I don't even look at code any more.
VP of Software Engineering, Broadvoice (2018 to 2024)Created a one-ticket, one-deployment process and incremental agile releases that improved operational efficiency by 30%. Made 100% uptime SLAs profitable to offer through architecture and observability. Led 40+ engineers across 6 time zones with an eNPS of 90+, and scaled the team 650%. Also led a zero-downtime migration to AWS.
A single pane of glassBuilt a near-real-time data pipeline that gave customers one view of everything. It's why every Sprint starts with a baseline and ends with before and after numbers.
Founder, Rentables (2012 to 2018)Accounting and operations SaaS for property managers, built from rails new to hundreds of paying customers: a double-entry ledger with 1M+ entries and $20M+ in ACH volume. I know the owner's side of shipping software.
Harnesses since GPT-3Several generations of AI dev tooling, from early prompt scripts to today's multi-agent pipeline with an outer review loop.
Outer loop, in daily useOn Based Budget, an AI chief of staff and a product reviewer read every PR against the brief before I merge. Bots never change code; they write the brief and check the result.
LinkedIn · Sep 2026

"AI made code cheap to write. It didn't make being wrong any cheaper."

Read the post
LinkedIn · Mar 2026 · 96k views

"To everyone letting Claude / Grok / Copilot write their unit tests: please pause for a second."

Read the post and the comments
LinkedIn

More on agents, gates, and evals

Follow on LinkedIn
Jerry relaxing in a red Adirondack chair on a deck, his dog lying beside him
Off the clock in Mammoth, with the dog.

About

Hi, I'm Jerry.

I started as an electrical engineer (BS from UCLA, MS from CSUN), first as a satellite engineer at General Dynamics, then as a principal engineer and integrated product team lead at LinQuest, building satellite network simulation for US Space Command.

In 2012 I founded Rentables and took it from rails new to hundreds of paying customers. From 2018 to 2024 I was VP of Software Engineering at Broadvoice, where the work I'm proudest of was process: one ticket, one deployment, and a team of 40+ engineers across six time zones that kept an eNPS of 90+ while it grew 650%.

These days I live in the LA area, own rental properties, and get up to Mammoth whenever I can. Most days I'm building Based Budget with AI agents. On the side, my wife Christy and I make Pets Forever Storybooks.

Consulting is part-time on purpose. I take one engagement at a time, keep it short, and give it my full attention. If your team wants AI agents it can trust, I'd like to hear what you're working on.

FAQ

Questions teams ask

How available are you?

Part-time by design. I run one engagement at a time, roughly 10 to 15 hours a week, scheduled as 1 or 2 week sprints. Next openings start in January 2027, and I’m booking them now.

Which tools do you work with?

Whatever your team already uses: Claude Code, Cursor, Codex, Copilot, Grok, or a mix. My daily driver right now is the Grok Build CLI, and I used Claude Code before that. The harness is plain files in your repo (commands, agent definitions, rules, gates, and evals), so it keeps working if you switch providers.

Remote or onsite?

Remote by default. For the 2-week Sprint I can also work onsite, in LA or elsewhere, with travel and lodging covered by you. Onsite works best for the kickoff and the two working sessions.

Will you sign an NDA?

Yes. A mutual NDA before I see any code is standard. I work inside your environment with the access you grant, and nothing leaves it.

We tried rules files. Agents ignore them. Why would this work?

Rules alone often do get ignored, usually because one agent is holding too much context. Small sub-agents, gates that block on outcomes, and an outer loop that checks the work are what make rules stick. If you're skeptical, start with the Readiness Review and judge the findings.

Do you write the code?

The harness writes the code. My job is to set it up so your engineers trust what it produces, and to leave them able to run and improve it without me.

How will leadership know it's working?

Every Sprint starts with a baseline and ends with before and after numbers: how much of the work goes through the agents, PR throughput and review time, and what the gates caught. If you want those numbers on one page your leadership checks weekly, I can add a small dashboard. Ask on the call.

What happens after the sprint?

A 30-day check-in is included. After that you can book another sprint if you want one. There is no retainer.

Do you do cloud migrations or MVP builds?

Not as offers. I've led a zero-downtime migration to AWS, but projects like that take months and I keep engagements short.

Contact

Book a 30-minute call.

Tell me what your team is building and how agents fit in today. You'll leave with at least one thing to try, whether or not we work together.