Agents on a leash

2026 · Tooling and design

I let AI agents run on their own, on every push, all day. Here's what it took to keep them useful, honest and cheap.

Claude Code (headless, read-only), Swift, SwiftUI, AppKit, zsh, launchd, GitLab API, Jira API

QAnelita’s checklists land in sticky notes that float over my screen, each check written in plain language a non-engineer could act on. Since September 14 a headless, read-only Claude Code run has added checks to them after each push I make at my day job, reading the diff, the ticket and the screenshot diffs. When a merge touches my branch’s screens, or a new push can reach every screen, QAnelita unticks the sticky’s checks and writes under each one why it’s back.

back: new push <sha8> — its changes can reach every screen
back: main moved to <sha8> — its changes touch this branch's screens

I started QAnelita, named after my chihuahua Canelita, on August 17 as a night shift in the cloud. Each morning its claims landed for me to judge. The cloud never saw a push in time, so the push became the clock and the run moved to my machine.

Cap the runs at 25 a day

On September 17, one failed fetch of my open merge requests wiped QAnelita’s record of what it had checked, so the next run checked them all again. Unchanged pushes were also checked again whenever the base branch moved. That made 19 Opus runs for 2 tickets between 18:24 and 20:34, until one sticky held 69 checks and the account hit its usage limit. The next day the record started loading from disk, and a gate went in. In work hours it reads Taquito, my menu-bar usage meter named after my other chihuahua, and refuses to run at 60% of the 5-hour window, so the rest stays mine.

Another loop ran unseen for four days. Two tickets each had two open MRs sharing one entry in that record, so each MR saw the other’s push as new and ran Opus again, 27 to 107 times a day in all. Since September 21 one sticky follows one MR, and a hard cap allows 25 runs a day at any hour. Each refusal is logged.

budget: the account is 100% through its 5-hour window (work-hours cap 60%), skipping — off-hours runs are uncapped
budget: 25 claude runs today, the daily cap is 25 — something may be looping; QA_BRAIN_MAX=… overrides by hand

Hitting the cap also writes a line on the HEALTH sticky. I added it on August 27, after || true had hidden a three-day outage.

A JSON endpoint instead of scraping

Taquito used to scrape the usage page and kept asking me to sign in when I already was. A HAR capture turned up the JSON endpoint behind that page, which gives the exact percent the gate now reads.

Where no machine was running Taquito yet, its history is estimated from local cost data, calibrated against its overlap with real readings. That makes part of the history a guess. Estimates are drawn at reduced opacity, with “(estimated)” on hover, and a real reading always replaces one.

The first eval leaked its answer

On QAnelita’s first day I held out a past MR as a blind eval case, and 12 min later declared attempt 1 invalid. Its rubric cites the cases behind each rule, which makes each rule auditable and the rubric an answer key. Two models had read the defect straight out of it. That afternoon a sanitizer and a leak check went in to keep that evidence out of an eval, a check on the judge that Parity Bench needed too.

Read-only up to the save button

The model can only Read, Grep and Glob, and an alarm kills any run still going after 15 min. I paste the QA notes QAnelita drafts into the ticket myself, and a sticky’s form-filling links stop before the save button, so the save stays the tester’s.

Checking each diff against its ticket

My lead said I sneak unrequested changes into tickets. Guilty: animated icons. So every run also lists, in plain language, the changes it judges the ticket didn’t ask for. Before it went live, it found 64 of them in 92 of my merged MRs.

QAnelita still runs on every push, and I mark its checks good or bad in one click. A type of check with three bad marks and no good ones is never suggested again. Without that rule, QAnelita’s README says, “the report becomes noise and stops being read, which is the only way this system actually fails.”