Building an AI software factory means assembling five layers: signal intake that decides what to build, spec generation that removes ambiguity, an agent substrate that writes code, deterministic validation that checks it, and approval gates where humans decide. Most teams buy an agent first and wonder why nothing improved — it's the easiest layer to acquire and the least valuable on its own. The agent was never the bottleneck. Everything upstream and downstream of it is.

This piece covers what mechanically happens inside each pipeline stage, or what a spec should contain. For that, see Automating Feature Development and Spec-Driven Development. This piece is about a different question: for each layer, do you build it, buy it, or assemble it from parts — and who inside your org is on the hook when it breaks?

What layers does an AI software factory actually need?

Five, and skipping any of them just relocates the failure to a human who now has to catch it manually.

1. Signal intake — buy it

Ingesting support tickets, sales call notes, NPS verbatims and app reviews, then clustering them into themes, is a solved problem with mature vendors. Building this in-house is a multi-quarter distraction with no defensible IP at the end of it. Owner: product. Failure mode: the data exists, gets clustered beautifully, and then sits in a dashboard nobody routes anywhere — signal with no path to the backlog is just a report.

2. Spec generation — decide deliberately

This is the layer most teams under-invest in because it looks like "just writing docs." It's actually the layer that determines everything downstream, because agents don't fail on hard problems, they fail on ambiguous ones. Buying a spec-generation tool gets you speed and a forcing function for structure; building it gets you specs that encode your own domain vocabulary and invariants. Either is defensible. What's not defensible is skipping it and feeding agents raw tickets. Owner: product and engineering jointly — this is the one layer with two names on the door, because a spec written by product alone omits technical constraints, and one written by engineering alone omits intent. Failure mode: specs that read as complete but leave the exact decision an agent needs to make correctly, unstated.

Should you build or buy the agent substrate?

Buy, and buy in a way that keeps you swappable. The coding-agent market is moving faster than any internal roadmap can track — the model that's best in class today is a commodity in eighteen months. If your orchestration logic, your prompts, and your review workflow are all wired directly into one vendor's agent API, you've architected yourself into a migration project. Treat the agent as a replaceable component behind an interface you control, not the foundation you build on. Owner: platform/engineering. Failure mode: deep vendor coupling that turns "switch agents" into a quarter-long re-platforming effort instead of a config change.

What should never be bought?

Deterministic validation. Nobody sells your invariants, because nobody else has them. Your schema constraints, your business rules, your "this field can never be negative" logic — that's yours, and it has to live in code, not in a prompt. The operating principle is simple: anything with a verifiable answer gets checked deterministically, never asked of a model. If a test can confirm it, write the test. Don't ask an LLM to confirm it and call that validation. Owner: platform/engineering. Failure mode: teams that let the agent "self-check" its own output and treat a plausible-sounding confirmation as proof.

How should approval gates be structured?

Build them around the review culture you already have — don't import a generic gate and bolt it on. The design decision that matters is routing by confidence and blast radius, not applying one gate to everything. A one-line copy change and a database migration should never clear the same bar. Owner: engineering leadership, because gate design is a risk-tolerance decision, not a tooling decision. Failure mode: a uniform gate applied to every PR decays into a rubber stamp within a month, because reviewers can't sustain scrutiny on low-risk changes and that fatigue bleeds into the high-risk ones.

In what order should you turn these layers on?

Crawl, walk, run — and the reason isn't caution for its own sake, it's diagnosability.

Teams that switch everything on simultaneously can't tell which layer failed when a bad PR lands. Was the signal wrong, the spec ambiguous, the agent off, or the gate too loose? With four variables moving at once, you're debugging the whole factory to find one broken part. Sequenced rollout turns that into a single-variable problem every time.

What should you automate first, and what never?

First: well-tested, reversible modules — internal tooling, CRUD endpoints, UI components with existing test coverage, anything where a bad outcome is cheap to detect and cheap to undo. Never, indefinitely: schema migrations, billing logic, and auth. These share one property — the cost of an undetected error is asymmetric and often irreversible. That's a human-authorship line, not a maturity milestone you graduate past once your agents get better. AI in the SDLC covers which stages AI is actually competent at; this is the separate question of which stages you let it touch regardless of competence.

Who should own the factory as a product?

A named person or a platform team — not a committee, not "whoever has time." Unowned internal tooling rots; this is not a maybe. Without a single owner accountable for the pipeline's health, gate decay, cost creep, and validation drift all happen silently because no one's job depends on catching them. If you can't name the owner today, you don't have a factory, you have a demo.

How do you govern cost before it governs you?

Track token spend per feature from day one, not after the bill surprises someone. Set per-run ceilings (stop a single agent run from spending $40 on a one-line fix) and separate period caps (a monthly budget per team or project). You need a usage ledger before you scale, not after — retrofitting cost visibility onto a system already running hundreds of agent runs a day is far harder than building it in at ten runs a day. Full mechanics in Cost Per Feature with AI.

How do you measure whether the factory is actually working?

Two metrics, tracked separately, borrowed and adapted from DORA: lead time for changes, and change failure rate split specifically by agent-originated PRs versus human-originated ones. That separation is the only way to learn whether your gates are doing their job. If agent-originated change failure rate is climbing while human rate stays flat, your gates are too loose for the risk you're actually taking on — and you'd never see that signal in a blended number.

Should you build the whole stack yourself or buy it assembled?

Both are legitimate, and the honest answer depends on what you're optimizing for. Assembling it yourself — buying signal intake and the agent substrate, building spec generation and validation in-house, wiring your own gates — gives you maximum control and forces your team to understand every seam, which pays off when something breaks at 2am. It costs more in integration time and ongoing maintenance, and the ThoughtWorks Radar is a reasonable place to track which component vendors are worth trusting with each layer as the market shifts.

Buying it assembled trades some of that control for speed to a working loop and someone else owning the integration seams. VocxAI turns customer feedback into shipped code — it ingests signals from your support and feedback tools, prioritises what to build, and runs an AI agent pipeline from PRD to pull request with human approval at every gate. That's the assembled option, not the only option. If you have the headcount and the appetite to own five vendor relationships and the glue between them, building it yourself is a completely defensible call — see Martin Fowler's writing on build-vs-buy tradeoffs for the general case this is a specific instance of. What's not defensible is buying an agent, skipping the other four layers, and calling it a factory.

See how VocxAI builds this for you

VocxAI connects your customer signals to your revenue data and surfaces a ranked, revenue-weighted product backlog - automatically, every week.

Sign up for free