Hand it the ticket.
Wake up to the pull request.

Haiyo takes each ticket through plan, failing tests, implementation, review and verification — in parallel, in isolated git worktrees, while you’re away. What comes back is a draft PR with the evidence attached.

Illustration with example data — not an app screen Scroll to follow one ticket ↓
01 — One run

One ticket, start to finish.

00

Ticket

Runs

storefront workspace

⑂ main

Add rate limiting to the login endpoint

⑂ agent/add-rate-limiting-to-the-login-endpoint
Queued
0%
Queued
Plan
Write tests
Implement
Verify
Ship
Waiting for a ticket…
Acceptance criteria
6th login attempt within 60s returns 429
Counter resets after the window
Other routes are unaffected
systemIdle — waiting for a run
3 changed files
rateLimit.tssrc/server/middleware · +14
auth.tssrc/server/routes · +2 −1
rate-limit.test.tstests/api · +42
src/server/middleware/rateLimit.ts
Built from the app’s real screens · example data · timing compressed Playing Scroll to advance ↓
02 — The fleet

Built for a backlog, not a chat.

You don’t babysit a run. You queue work, leave, and come back to a short list of pull requests worth your attention. Three panels from the app, replayed.

Eligible tickets launch on their own up to your parallel limit — highest priority first, ties to the oldest. A ticket that depends on another waits, then starts with that plan and PR as context.

  • Every run gets its own git worktree and branch, so parallel runs never collide.
  • Run all walks the whole dependency graph; cycles are refused when you link tasks.
  • Quit mid-run and it resumes from the last finished stage.

Finished runs land in one queue across projects, with the ones that need you most at the top. Each row carries blast radius, criteria score and confidence — approve in one click, or send plain-language changes back to the same PR.

  • Risk comes from the real diff: file count, churn, and auth, payment, migration or CI paths.
  • Requested changes are re-implemented, re-verified and pushed to the same draft PR.

Set token and dollar caps per task, a daily cap for the whole fleet, and a stuck timeout that stops runs that go silent. Failures are classified and routed: retry, pause, or hold for you.

  • Reaching the daily cap holds the next run and says why — it never silently overspends.
  • One blocked ticket doesn’t stop unrelated ones.
storefront boardParallel limit 3 · P0 first
0 runningRun all
Backlog 0
In progress 0
Final review 0
Done 0
InboxFinished runs waiting for a human
0 to review
Control towerEstimated spend, caps and fleet failures
0 runs today
Spend today
$0.00 of $12.00 cap
Failures by class
Flaky0retry
Missing credentials0held
Rate limited0resumes
Stuck · paused · held
Next run held — projected spend reaches today’s cap. “Daily dependency update” starts when the cap resets or you raise it.
03 — Use cases

What people hand it.

Eight jobs, each built on features that ship in the app today.

01

Clear the backlog overnight

Queue well-scoped tickets before you log off and leave the app open, with the Mac awake. They run in parallel, highest priority first, and the morning starts with draft PRs instead of a to-do list.

10 tickets → 3 at a time → inbox at 9:00
02

Label an issue, get a pull request

Turn on issue intake, then tag a GitHub issue autopilot and it becomes a task by itself. The draft PR says Closes #N, and its link is posted back on the issue. Linear and Jira import too.

issue labelled → task created → PR linked on the issue
03

Ship a feature in dependent steps

Split a feature into tasks that depend on each other. Each waits for its upstream, then builds on that plan and PR instead of starting cold.

API → UI → docs · Run all
04

Turn a spec into a task graph

Paste a requirements document. It’s broken into tasks with acceptance criteria and dependencies, then run as one batch in dependency order.

spec → tasks + criteria → one batch, in order
05

Fix the flaky test everyone ignores

Point it at a test that fails one run in six. It reproduces, fixes the cause, and ships only when the suite and browser E2E pass — screenshots attached.

red → root cause → green ×30 → evidence
06

Keep maintenance on a schedule

Give a ticket a schedule — every few hours or daily at a set time — and it re-queues itself while the app is open: dependency bumps, lint debt, test sweeps.

“Daily dependency update” · 06:00
07

Iterate on a PR in plain words

Not quite right? Write what to change. The agent re-implements, re-runs the checks and updates the same draft PR — no new branch, no lost context.

“also filter by status” → re-run → same PR updated
08

Go from idea to a live web app

Start a new project from a plain-language idea on one opinionated stack — Vite and React, Hono on Cloudflare Workers — then deploy after merge and check the live URL.

idea → scaffold → PRs → deploy → live check
04 — Example runs

What comes back.

Four example runs: three draft PRs in the app’s real format, and one it refused to call done.

Example data. The structure — status line, risk reasons, test plan, criteria with evidence, review with the confidence verdict — is what the app writes into every draft PR.

05 — Safety

Autonomous, with floors.

Hands-off is the point. Safety comes from hard technical limits — not from asking you to click “approve” on every step.

Acceptance criteria

Graded one by one, with evidence.

A separate review pass marks each criterion met or unmet. One it can’t verify counts as unmet — and is listed as a blocker in the PR.

Confidence

High has to be earned.

Any unmet criterion or failed review rates Low. A high-risk diff or missing criteria caps the rating at Medium.

Isolation

Every run in its own worktree.

Runs never touch your working copy or each other. Rejecting a run rolls its branch back.

Budgets

Caps that stop the run, not just warn.

Set per-task token and dollar caps and the agent is stopped mid-run when it crosses them. A daily cap holds the queue, and a stuck timeout kills silent runs and records why.

Destructive operations

Always stop for a human.

An infrastructure change that destroys anything waits for approval at every autonomy level, and no setting turns that off. Rollbacks are never automated.

Local

Runs on your Mac.

Work happens in git worktrees inside your own repositories, with your existing GitHub login. Agent environments are stripped of credentials they don’t need.

Autonomy level

Pick how far it goes on its own.

One setting moves merge, deploy and infrastructure behaviour together. The individual switches stay available underneath.

FLOORAt every level, a destructive infrastructure change stops for your approval.
06 — Get started

Give it the ticket you’d give a teammate.

Step 1

Install the app

A universal macOS build for Apple silicon and Intel. The first-run setup checks your machine for you.

Step 2

Pass the checks

  • Gitrequired
  • GitHub CLI, signed inrequired
  • An AI coding agent, signed inrequired
  • Your project’s toolchainrequired
  • Docker, Playwrightoptional
Step 3

Add a project, write a ticket

Pick a folder, clone a GitHub repo or start from an idea. Import an issue or write a brief, then press Start TDD run.

Read the FAQ Early access · macOS
07 — FAQ

Asked first.

Does it merge my code?

Finished runs end in a draft pull request with the evidence attached. What happens next is your autonomy level. Fully gated never merges without you; Supervised only logs what auto-merge would have done. Autonomous merges only when the run is rated high confidence with every criterion met, the diff is low-risk and within size limits, CI ran and passed, and every changed file is in a category you’ve allowed. No categories are allowed out of the box, so nothing merges until you choose them.

Where does my code go?

Your repository stays on your Mac — runs happen in git worktrees inside it. The AI model sees the code it needs for the ticket it’s working on, and GitHub receives the branch and the draft PR.

What if my project has no tests?

The test stage still writes a real test file into your project before any implementation starts. The app never shows a passing test stage it didn’t actually run.

What happens when a run gets stuck or fails?

Set a stuck timeout and silent runs are stopped by a watchdog; set caps and runaway spend is stopped too. Every failure is classified: flaky tests retry a bounded number of times, missing credentials or an unclear ticket hold the task and tell you exactly what’s needed, and rate limits pause and resume on their own. If the reviewer can’t confirm every acceptance criterion, the run stops before the PR and says which one.

Can I approve the plan before it writes code?

Yes. Two optional gates pause the run after planning and before opening the PR. Rejecting at either gate rolls the branch back. Once a PR exists, request changes in plain words and the same PR is updated.