pstack Lets Agents Merge Their Own PRs Overnight, but Only After Fresh Verifiers Sign Off

Crafting seamless user experiences with a passion for headless CMS, Vercel deployments, and Cloudflare optimization. I'm a Full Stack Developer with expertise in building modern web applications that are blazing fast, secure, and scalable. Let's connect and discuss how I can help you elevate your next project!
pstack is the open-source Cursor plugin Lauren Tan credits for landing about 2,500 pull requests in August 2026 without reviewing each one before merge.
Tan builds Grok Bot at SpaceXAI. She walked through the setup with Matt Pocock in a live interview on October 2, 2026. The short version she gives is that agents merge while she sleeps and she samples the commit history in the morning. The long version, which lives in the pstack repository, is stricter than that sentence. The agent that wrote a PR never merges on its own verdict. A separate swarm of verifiers has to return a clean result on the exact patch, and CI has to pass on the rebased head.
Three things replaced line-by-line review in her setup. A scripted verification skill exists for every app. Lint rules make the common mistakes impossible to write. Verifier agents open the app and try to break the change. Whether a team can copy that depends on cost and on whether its changes can be reverted.

The 2,500 figure is self-reported, and most of it is maintenance
Tan posted the talk behind the number on September 21, 2026, under the title "how i shipped 2,500 PRs last month to production", so the month in question is August. The Neuron, in a September 10, 2026 explainer, cites her own guide for a more exact 2,462 PRs in production for August. Tan's own guide (Part 1, August 31, 2026) put it at roughly 2,000 PRs a month, according to The Neuron, and several summaries repeat that figure. In the interview she is looser still, saying "2,000 or however many." Nobody outside her team has audited any of these counts.
She also deflates it herself. "They're not obviously like 2,500 features," she says of those PRs. A large share is what she calls gardening: restructuring code and deleting patterns agents would otherwise copy. Reading the number as 2,500 shipped features overstates it badly.
| Figure | Value | Source and date |
|---|---|---|
| Guide, Part 1 | Roughly 2,000 PRs a month | Lauren Tan via The Neuron, August 31, 2026 |
| Talk title | 2,500 PRs in a month | Lauren Tan, September 21, 2026 |
| August close | 2,462 PRs | The Neuron, September 10, 2026 |
| Interview wording | "2,000 or however many" | Matt Pocock interview, October 2, 2026 |
| pstack version | 0.15.9 | cursor/plugins repository, read October 5, 2026 |
| pstack contents | 26 workflow skills, 24 principle skills, 23 playbooks, 2 subagents | same |
Some background on the person. Tan worked on the React team at Meta; the pstack README says she is still on the React core team working on React Compiler. She joined Cursor in March 2026 and now works on Grok Bot, SpaceXAI's always-on agent product, which The Next Web reported entered beta on August 11, 2026. The same report notes that SpaceX agreed in June 2026 to buy Cursor's parent Anysphere in a $60 billion all-stock deal.
A verification skill gives the agent hands and eyes before anything else
Tan's first recommendation does not depend on pstack at all. The single most important skill in any toolkit, she says, is verification. The agent can launch the app, click through it the way a user would, and pull traces or heap snapshots when it needs to debug.
The lesson came from pain. In early April 2026 she was fixing lag in Cursor's agents window by hand, reading flame graphs and heap snapshots herself and relaying what she saw to the agent. She calls that being "the meat proxy between my agent and Chrome DevTools." The first skill she wrote at Cursor was verification. Once the agent could see the result of its own change, the loop closed and she could step back. Every app at Cursor and SpaceXAI now has a verification skill, and those skills are maintained automatically.
pstack turns that into a generator. The create-verification-skill instructions have the agent interview the repository, then write a project-local skill with six fixed sections:
- Launch: the exact command that starts the app and how to tell it is ready.
- Doctor: one read-only check that says whether this instance is worth driving.
- Drive: real selectors and commands from this repo, no placeholders.
- Evidence: what to capture, where it goes, and a rule that the proof must come from the real user path.
- Cleanup: kill only what you started, and never delete the evidence.
- Helpers: any shipped script must be executable and its invocation shown.
The file ends with a line worth pinning above a desk: "A generated skill that was never executed is a draft, not a deliverable."
One design choice inside the skill matters more than it looks. Tan splits every task into the parts that need judgment and the parts that are mechanical, and she moves the mechanical parts into a CLI the skill ships. Before that, each agent rebuilt its own throwaway test harness, some of which worked and some of which did not, and all of which were discarded. The CLI itself is glue around Playwright and the Chrome DevTools Protocol. Its value is that every agent calls the same thing.
When agents repeat a mistake, she changes the repo, not the prompt
The earliest versions of Grok Bot were, in Tan's words, "eight god files" of at least 10,000 lines each. Agents kept appending. Her fix was structural. Inside an internal framework she calls Dune, every feature lives in its own directory, there is one supported way to do each thing, and lint rules reject the rest. Dune is not open source; her description is "an internal Next.js for our Electron apps."
Her standing habit is to watch where agents fail and ask one question each time: can this become a lint rule, so the mistake cannot be written again?
pstack's correct skill, added to the repository on October 3, 2026, the day after the interview, encodes the order of operations. A mistake class counts once it has happened twice. Then you fix it at the highest level that works:
- Architecture: one owner per piece of state, one supported way per task, delete the old way an agent would copy.
- Types, then lint or CI: make the bad state unrepresentable; if it still compiles, add a check whose error message names the fix.
- Tests of behavior: rewrite or delete any test that would still pass if every function returned nothing.
- Docs and agent rules last, and only for judgment calls, because nothing fails when an agent skips them.
Every new check has to be proven against a real past mistake. The order inverts what most teams do. The reflex is to add a sentence to AGENTS.md; this skill puts that step at the bottom.
Nothing merges overnight without a clean swarm verdict and green CI
Secondhand accounts of Tan's setup compress it to "the AI checks itself and merges." The autopilot-full playbook splits the roles further than that. One owner agent carries each PR from build to merge, but permission to merge comes from a root agent that aggregates independent verifiers. The owner's own opinion of its work does not count.
| Gate | What the playbook requires |
|---|---|
| Who merges | One owner agent per PR; items the operator names stay with the operator, and no agent merges them |
| When verification runs | At the owner's code-ready head SHA, and again on every later push that changes the patch |
| What verifiers do | Parallel independent lanes: rerun the gates, prove the behavior live on the real surface, audit the diff while distrusting the PR description |
| Merge condition | The verdict matches the patch that merges, and CI passes on a head freshly rebased onto trunk |
| Audit and stop | The root agent audits every owner hourly; an operator stop reaches every owner as a zero-writes order |

In the interview Tan describes the same machinery from the outside. Full autopilot spawns a batch of verifier agents per PR. They open the application, click around like a human, look for regressions, and fix what they find until the PR is in a mergeable state. In the morning she reads the commit history, reverts what looks wrong, and adds a new lint rule.
The pstack guide's chapter on running work overnight frames the handoff as a contract with four parts: a goal, a finish condition, permissions, and an escape hatch. Its example is short enough to quote in full:
/poteto-mode im going to bed. migrate every caller to the new parser in a fresh worktree off <base>.
done means zero old callers, all parser fixtures pass, old api deleted.
keep a decision log. don't ask me before committing.
/loop until done. if you're truly stuck after a few hours, stop and write up why.
The guide's warning is blunt: a duration is not a finish condition. Ask for four hours of work and you get four hours of motion.
One more line from pstack's principles matters here. "Never block on the human" tells the agent to proceed and let the human correct afterward, but it reserves confirmation for irreversible actions. That exception is where the whole approach has its limit.
The work arrives because she stopped being the proxy for bug reports
The 2,500 PRs did not come from 2,500 chats. Bug reports land in Slack, Linear, and social channels. Tan used to read them and relay them to an agent. Now a few Grok Bots subscribe to those channels and forward new reports into Cursor Projects.
Projects shipped on September 10, 2026, and Cursor's changelog describes it as a coordinator that delegates tasks to thousands of subagents and performs recurring work without being prompted. Tan calls each coordinator a chief of staff: it writes no code, only splits tasks, delegates, and tracks progress. She says she runs more than ten of them at once, one on Grok Bot desktop performance, one on user-reported bugs, and so on.
Why not one agent per bug? Because a burst of reports often shares a root cause. Fixing them separately duplicates work and hides the actual problem, which may sit a layer above where any single report points.
I like one smaller detail more than the headline numbers. She runs an agent that scans for React foot-guns, but it is not allowed to fix them. It appends findings to a document. Every few days she reads the document and usually finds that a dozen entries are one problem. A queue forces the big-picture look that pure execution skips.
The bill comes in tokens, setup time, and a hard boundary
Tokens first. Tan says full autopilot is "quite token intensive," and that it can be tuned down from ten verifier agents to one. The Neuron reports a side-by-side test by Rob O'Shaughnessy: the same project took about 30 minutes without pstack and about one hour with it. The pstack run also caught three false claims the agent had hallucinated during development. Doubling the wall-clock time to catch three errors before merge is a good trade only if those errors would have been expensive in production.
Setup time second. She is explicit that this is "very hard to get to this point" and that she does not want to sell it as something you can do by installing pstack. It took months of watching agents fail and adding one guardrail at a time.
The boundary third. Pocock asked about domains where most PRs are one-way doors: data loss, medical software, finance. Tan's answer was that it depends on whether the work can be verified programmatically. Software mostly can. Domains that cannot be verified that way will struggle to reach this point, and she called it a good question she does not have the answer to.
Put together, skipping per-PR review does not remove quality control. It moves the cost from human reading time to tokens, lint rules, and verification scripts. That move only works when changes are reversible and behavior is machine-checkable.
Where a team without a verification skill should start
A team without a verification skill should adopt pstack's approach in four steps, starting with one verification skill and ending with autopilot-stack, not autopilot-full.
- Build a verification skill for one app and run it once end to end. Without this, the rest of the stack has nothing to stand on.
- Mine your own transcripts for the places you keep correcting the agent. Tan's recall skill came from re-explaining the previous chat's context every time she opened a new one, and pstack's automate-me skill drafts a personal mode skill from your history.
- Take every mistake that has happened twice through the correct skill's order, with docs last.
- For the first overnight runs, use autopilot-stack rather than autopilot-full. It runs the same owner loop but ships nothing, leaving one linear stack with a verifier's verdict on every link for you to review and land in the morning.
Both people in the interview agree that skills are just process written down. Pocock's own skills repository on GitHub had 276,225 GitHub stars on October 5, 2026, and Tan still argues everyone should end up with their own set. She also expects skills to shrink: last year's versions spelled out exact commands, and with current models those lines can be deleted, leaving only the workflow.
Frequently asked questions
Does pstack let AI agents merge pull requests without any human review?
pstack's autopilot-full playbook does let an owner agent carry a PR through to merge. The merge is allowed only after a separate swarm of verifier agents returns a clean verdict on the exact patch and CI passes on the rebased head. Lauren Tan said in the October 2, 2026 interview that installing pstack alone does not get a team to her level; the verification skills and lint rules have to exist first.
Does pstack work outside Cursor?
pstack documents only the Cursor install path, /add-plugin pstack, and relies on Cursor's built-in /loop command and cloud agents for overnight runs. pstack's skills are Markdown files under an MIT license and the README invites forks, but the documentation does not say whether the playbooks work in other agent tools.
Is pstack's overnight merging safe for finance or healthcare systems?
Lauren Tan did not claim pstack's overnight merging is safe for finance or healthcare systems; in the October 2, 2026 interview she said it depends on whether the work can be verified programmatically. For systems where most changes cannot be reverted, she said she does not have an answer. pstack's own principles reserve human confirmation for irreversible actions.
How much does pstack's verification cost in tokens?
pstack publishes no token figures. Lauren Tan described full autopilot as quite token intensive and said the verifier count can be reduced from ten to one. The Neuron reported on September 10, 2026 that one comparison run took about one hour with pstack against about 30 minutes without it.
Sources
- Matt Pocock, LIVE: Poteto (creator of pstack) on shipping 1,000's of PR's a month at SpaceX
- Lauren Tan, how i shipped 2,500 PRs last month to production (talk recording)
- Lauren Tan, The Complete Guide to pstack Pt. 1
- cursor/plugins, pstack README
- cursor/plugins, pstack autopilot-full playbook
- cursor/plugins, pstack create-verification-skill
- cursor/plugins, pstack correct skill
- cursor/plugins, pstack guide: Run work while you sleep
- Cursor Changelog, Cursor Projects
- The Neuron, pstack explained: Lauren Tan's system for trustworthy AI agents
- The Next Web, SpaceXAI launches Grok Bot as the agent race moves to office work
- GitHub, mattpocock/skills
Author Insight
Most teams should copy autopilot-stack and leave autopilot-full for later. Tan's own conditions explain why. She lets agents merge while she sleeps because every app has a verification skill and a bad merge can be reverted. She also accepts the token bill of a verifier swarm, which she calls "quite token intensive" herself. Remove any one of those, say a change set dominated by database migrations, or a repo whose verification skill has never been run. Then the right move is to let agents take work to a merge-ready stack while a person keeps the merge button. Hand it over once several weeks of morning audits turn up nothing worth reverting.





