Baseline and target
Every goal starts from a measured baseline, not from a feeling that things are bad. Without a baseline there is nothing to compare the result against, so the Factory refuses to start.
Stigmera Factory
Stigmera Factory takes a goal — not a prompt — and runs the whole loop: hypotheses → tickets → code → gates → deploy → measure → next iteration. This page is the detail behind that sentence.
01 · Goals
You do not describe a feature. You pick a metric from a fixed catalogue, state where it stands today, where it should get to, and what must not get worse on the way.
Every goal starts from a measured baseline, not from a feeling that things are bad. Without a baseline there is nothing to compare the result against, so the Factory refuses to start.
Each metric is paired with one that would suffer if it were gamed. Sign-ups up while retention collapses is a failed hypothesis, not a win. This is what keeps Goodhart’s law out of the loop.
Before a goal runs, the measurement is checked against itself: split identical traffic in two and confirm the instrument reports no difference. An instrument that finds an effect where there is none cannot referee anything.
02 · Hypotheses
A hypothesis moves through proposed → running → verdict. The last step is the one that matters, and it is deliberately taken away from the agent that authored the work.
03 · Delivery
The interesting part of autonomy is not writing code — it is everything that has to be true before code is allowed to reach production. Tickets, branches, agent sessions, gates, deploy windows, incident handling and a retry policy that depends on why something failed. As of today that pipeline runs 76 steps across eight kinds.
Engineering standards are enforced by tests in the pipeline, not by a document nobody reads. Work that violates one comes back to the ticket.
The platform ships its own updates through its own pipeline, blue-green, in cluster. Everything on this site went through that same line.
A ticket carries its own history: who took it, which gate failed, how long each stage took, what the deploy reported. Nothing about a run is folded away.
04 · Economics
Autonomy that costs more than it delivers is a demo. The Factory prices its own work: what a delivered ticket cost, which runs burned budget without shipping anything, and where the money went.
05 · Memory
Agents coordinate through traces they leave in a shared environment — tickets, branches, standards, measurements, notes. That environment is also where the system keeps what it has learned.
What worked, what failed and why is written down where the next session will read it — instead of being re-derived from scratch every run.
Notes are periodically merged and pruned. Memory that only grows becomes noise, and noise is worse than nothing.
Sessions are scored and the harness is benchmarked against thresholds, so a change that makes the factory worse is visible as a regression.
06 · Controlled autonomy
Autonomy is granted in steps and can be taken back. Today the delivery loop runs autonomously and deploys wait for a human. Closing the loop end-to-end on business goals is in progress — and we ship it the same way we ship everything else: in public.
FAQ
They sell an engineer: a task goes in, a pull request comes out, and a human decides what happens next. We sell the loop: a goal goes in, and hypotheses, code, deploys and measurements keep coming until the metric moves. Different unit of work — theirs is a task, ours is a metric.
App builders hand you a project and stop. The Factory keeps the product and keeps working on it — measuring, iterating, shipping the next version.
factory.ai builds an autonomy stack for enterprise engineering teams. We build a product factory for solo builders and small teams. Different user, different unit of work.
AutoGPT had no measurement and no foundation under it. Here the only source of truth is the metric — each one paired with an antagonist metric so the loop cannot game itself — verdicts are issued by the platform from the numbers, and underneath sits a full SDLC with enforced gates.
It is the right question, and the answer is mechanisms rather than reassurance. Deploys stay on a manual trigger. Standards are enforced by tests. Every metric has an antagonist metric. An agent can propose a verdict but never confirm one — the platform does that from the numbers. Goals carry budget caps. Every step is a ticket you can read.
Waitlist
The Factory is in private access while the goal loop closes. Join the waitlist and we will write when it opens up.
We keep your email, plus the country and network address the request came from, so we know where interest comes from. Nothing else, nobody else — see Privacy.