turbodiff
why I built this

Every software company will (eventually) run a software factory

No, software factories will not replace coding agents or deployment pipelines. I imagine a future where they co-exist; connected in a loop that turns signals into verified changes. That's what I'm building towards with Turbodiff.

The loop I kept seeing

Looking at how production bugs get fixed today, I see the same sequence every time. An alert fires. Someone reads it in Slack, decides it is real, and opens a ticket. Another person reproduces the problem, writes a fix (most probably with the help of AI), waits for the checks, code review, merges, deploys, and watches the dashboard. And this cycle repeats weekly.

To me, that is already a loop with clear inputs and outputs. The problem is that people are responsible for moving information between these tools: an alert becomes a Slack message, then a ticket, then a branch, then a PR. And while there is some value in having human judgement in the process, I don't think copying context from one tool to the next is a good use of anyone's time.

My bet is that software teams will eventually run this loop as a system instead of relying on habit. That's where software factories come in. From my perspective, it does not replace coding agents, CI, deployment tools or human involvement. It augments the human loop and automates the repetitive tasks.

The loop I envision

The same six steps keep happening: observe, diagnose, fix, test, deploy, and learn. In practice, I imagine one pass looking like this:

Observe01

Most signals that already exist: an error rate, a repeated support message, a page losing traffic, or some custom event.

Diagnose02

From there, I have an agent inspect the signal against the real code and decide whether it is a bug, a regression, a gap, or just noise.

Fix03

Before any code changes, I want a short plan with the files, acceptance criteria, and risks. Once approved, the agent builds the change in a sandbox and opens a PR.

Test04

For testing, I run the repository checks and ask a separate agent to verify each criterion, including launching the app and taking screenshots.

Deploy05

Deployment stays in the existing pipeline. The factory opens a normal PR and only merges when I allow it (merging can also be automated, but it's risky).

Learn06

After deployment, I check whether the original signal went away. That result helps me handle the next similar signal.

Whenever a new signal arrives, I want the loop to run again instead of waiting for the next planning cycle. Turbodiff currently handles the three lit stages in the middle. The remaining ones are still on my list (to complete the entire factory loop).

This is bigger than bugs

Error rates are a natural starting point because the manual loop is easy to see there. But the same approach will work for other things that can be measured and changed in code:

  • Analytics. When a funnel step drops after a release, I want to find the change that caused it, propose a fix, and check whether conversion recovers.
  • Customer support. Twenty messages about the same confusing setting look like a product signal to me. The fix might be copy, a better default, or a small feature.
  • SEO monitoring. When a page loses rank or a crawl report finds broken structured data, I want to trace the problem back to a concrete code change.
  • Custom events. Any event I would normally ask a person to investigate is one the factory can look at first.

Either way, I still want the work to end in a normal PR. The goal is simply to automate more of the path from the first signal to that PR.

Keeping the tools that already work

This is not about replacing the tools that already work. Coding agents can write the code, CI can run the checks, and the existing pipeline can deploy it. What is missing is the connection between the different systems/tools: getting a signal, making a plan, checking the result, and leaving a useful record of what happened.

That is the idea shaping Turbodiff. It checks out the real branch in a sandbox, runs the repository's own check command, and opens a normal PR (if needed). The reviewers, merge policy, and deployment all remain under my control.

Proof is the hard part

As soon as you remove people from the handoffs, you implicitly remove a lot of judgment from the process. So you need another way to trust the result. Reading every line of every generated change does not scale. A PR description written by the coding model does not count as evidence to me; it is still a claim from the same model.

Anyone can generate code. We ship proof.

the line on the landing page, and the reason the product exists

That is why Turbodiff preserves the approved plan, normalized change, checks, and review artifacts. The rest of the loop is not useful unless its output can be inspected.

What Turbodiff can do today

Currently, Turbodiff runs the middle of the loop: plan, build, verify, and ship. You describe a feature, an agent plans it against the repository, and you approve the plan. A sandbox builds it and opens a PR. An (automated) reviewer fixes what it finds, and a separate verifier adds screenshots before the merge. This is also how I build Turbodiff itself. Agents wrote roughly two thirds of the commits currently on its main branch.

What I believe

Today I still see people connecting alerts, Slack messages, tickets, emails, and PRs by hand. My bet is that much of this work will be automated in the next few years. Better models alone will not get us there. But with good integrations and verification, automations get easier and more reliable.

My goal is to build that system in the open, let teams run it on their own infrastructure, and leave enough evidence for them to judge the result.

That is why I'm building Turbodiff. If it sounds like a loop you are running by hand, star the repository, or point it at a repo and let it build something for you.

taggedsoftware factoryagentscode review