Clear the backlog overnight
Queue well-scoped tickets before you log off and leave the app open, with the Mac awake. They run in parallel, highest priority first, and the morning starts with draft PRs instead of a to-do list.
Haiyo takes each ticket through plan, failing tests, implementation, review and verification — in parallel, in isolated git worktrees, while you’re away. What comes back is a draft PR with the evidence attached.
| Ticket | Task | Project | Stage | Status | Draft PR |
|---|
storefront workspace
You don’t babysit a run. You queue work, leave, and come back to a short list of pull requests worth your attention. Three panels from the app, replayed.
Eligible tickets launch on their own up to your parallel limit — highest priority first, ties to the oldest. A ticket that depends on another waits, then starts with that plan and PR as context.
Finished runs land in one queue across projects, with the ones that need you most at the top. Each row carries blast radius, criteria score and confidence — approve in one click, or send plain-language changes back to the same PR.
Set token and dollar caps per task, a daily cap for the whole fleet, and a stuck timeout that stops runs that go silent. Failures are classified and routed: retry, pause, or hold for you.
| Flaky | 0 | retry |
| Missing credentials | 0 | held |
| Rate limited | 0 | resumes |
Eight jobs, each built on features that ship in the app today.
Queue well-scoped tickets before you log off and leave the app open, with the Mac awake. They run in parallel, highest priority first, and the morning starts with draft PRs instead of a to-do list.
Turn on issue intake, then tag a GitHub issue autopilot and it becomes a task by itself. The draft PR says Closes #N, and its link is posted back on the issue. Linear and Jira import too.
Split a feature into tasks that depend on each other. Each waits for its upstream, then builds on that plan and PR instead of starting cold.
Paste a requirements document. It’s broken into tasks with acceptance criteria and dependencies, then run as one batch in dependency order.
Point it at a test that fails one run in six. It reproduces, fixes the cause, and ships only when the suite and browser E2E pass — screenshots attached.
Give a ticket a schedule — every few hours or daily at a set time — and it re-queues itself while the app is open: dependency bumps, lint debt, test sweeps.
Not quite right? Write what to change. The agent re-implements, re-runs the checks and updates the same draft PR — no new branch, no lost context.
Start a new project from a plain-language idea on one opinionated stack — Vite and React, Hono on Cloudflare Workers — then deploy after merge and check the live URL.
Four example runs: three draft PRs in the app’s real format, and one it refused to call done.
Example data. The structure — status line, risk reasons, test plan, criteria with evidence, review with the confidence verdict — is what the app writes into every draft PR.
Hands-off is the point. Safety comes from hard technical limits — not from asking you to click “approve” on every step.
A separate review pass marks each criterion met or unmet. One it can’t verify counts as unmet — and is listed as a blocker in the PR.
Any unmet criterion or failed review rates Low. A high-risk diff or missing criteria caps the rating at Medium.
Runs never touch your working copy or each other. Rejecting a run rolls its branch back.
Set per-task token and dollar caps and the agent is stopped mid-run when it crosses them. A daily cap holds the queue, and a stuck timeout kills silent runs and records why.
An infrastructure change that destroys anything waits for approval at every autonomy level, and no setting turns that off. Rollbacks are never automated.
Work happens in git worktrees inside your own repositories, with your existing GitHub login. Agent environments are stripped of credentials they don’t need.
One setting moves merge, deploy and infrastructure behaviour together. The individual switches stay available underneath.
A universal macOS build for Apple silicon and Intel. The first-run setup checks your machine for you.
Pick a folder, clone a GitHub repo or start from an idea. Import an issue or write a brief, then press Start TDD run.
Finished runs end in a draft pull request with the evidence attached. What happens next is your autonomy level. Fully gated never merges without you; Supervised only logs what auto-merge would have done. Autonomous merges only when the run is rated high confidence with every criterion met, the diff is low-risk and within size limits, CI ran and passed, and every changed file is in a category you’ve allowed. No categories are allowed out of the box, so nothing merges until you choose them.
Your repository stays on your Mac — runs happen in git worktrees inside it. The AI model sees the code it needs for the ticket it’s working on, and GitHub receives the branch and the draft PR.
The test stage still writes a real test file into your project before any implementation starts. The app never shows a passing test stage it didn’t actually run.
Set a stuck timeout and silent runs are stopped by a watchdog; set caps and runaway spend is stopped too. Every failure is classified: flaky tests retry a bounded number of times, missing credentials or an unclear ticket hold the task and tell you exactly what’s needed, and rate limits pause and resume on their own. If the reviewer can’t confirm every acceptance criterion, the run stops before the PR and says which one.
Yes. Two optional gates pause the run after planning and before opening the PR. Rejecting at either gate rolls the branch back. Once a PR exists, request changes in plain words and the same PR is updated.