Agentic Delivery Lab

Your agents work. Your delivery process doesn't.

Most teams I talk to are stuck between "a few developers use Copilot" and "we don't dare go further". The lab gets one of your teams past that point in six weeks, in the tools you already use, with people still making the decisions.

1 of 3 pilot spots left for Q1 2027 · German or English · on site in Switzerland or remote

Run record · on the work item and the pull request

Run #142 · Story 5.006
Trigger    Status → Ready · PO
Agent      implementation · model pinned
Skills     backend@1.4
Context    ADR-007 · arc42 §5
Decision   New dependency?
           → Awaiting decision
           → PO approved
Result     PR #318 · 14 tests green
Cost       38 min · CHF 2.40

Every change answers three questions: who asked, who decided, why.

Works in your toolchain Azure DevOps Jira & Confluence GitHub GitLab

What this looks like in practice

Three teams I've built agentic delivery with, anonymized.

SME · 300 employees
2.5 FTE
now deliver the output of a product team that used to have 16 engineers: 1.5 FTE in development and 1 FTE in product ownership and requirements, working with the agentic system.
SME · 30 employees
10×
more features shipped per period since the team started working with agents.
Startup · 2 people
Autonomous
development with agents. A software and tech lead steps in only where it matters.

Why most agent pilots stall

The pilot usually works. The demo is impressive. Then it quietly stops.

In almost every case I've seen, the agent wasn't the problem. Nobody had decided who reviews its work, when it should stop and ask, or what an item is allowed to cost. And the architecture it was supposed to respect only existed in the heads of three senior people.

Agents don't fail at the model. They fail at missing process, and at architecture they can't read.

That's why I treat agentic delivery as a change in how a team works, not as a tooling rollout. More on that in The Agentic Operating Model – Eleven Realities.

How the lab works

Six weeks, one team, and 5–10 items from your real backlog. No sandbox and no toy project.

Week 0
Preparation
We look at your toolchain, repositories, documentation and security rules, pick the team and the items, and measure where you start.
Week 1 · 3 days, on site or remote
Set it up together
The agent flow goes into your toolchain, with a decision policy, skills for your stack and agent-readable arc42 and ADRs. The first items go through while we're in the room.
Weeks 2–5
Real work
Your team runs real items through the flow. We meet once a week to tune, unblock and decide. Evals and cost tracking go live.
Week 6
Review
Before and after in numbers, a plan for the next teams, and the start of three months of Agent Care.

What your team has after six weeks

  • An agent flow running in your toolchain, from Ready to pull request
  • A team that actually works with it, not just one enthusiast
  • Measured results: lead time, cost per item, and how often a person had to step in
  • A decision policy and a run record on every item
  • A plan for rolling it out to the next teams

No black-box agents

Your platform already logs what happened. The lab adds who decided and why, the same way in Azure DevOps, Jira, GitHub and GitLab. It's built in from day one.

A decision policy, as code
A versioned file in your repository defines what agents may do on their own. New dependencies, schema changes, public APIs and security-relevant code go to a person. Changing the policy is itself a pull request.
A run record on every item
Who triggered the run, which agent, which model and skill versions, which ADRs it read, what it cost and who approved. Versions are pinned, so any run can be traced and repeated.
Architecture as the guardrail
Agents read your ADRs and arc42 docs before they build. If they want to deviate, they draft an ADR and a person decides.
Control in numbers
How many items go through on their own, how often people step in, cost per item and quality over time.

Agents run in your environment, on your runners. I don't host your code.

Who you'd work with

Patrick Roos

I'm Patrick Roos, a software architect and the author of workingsoftware.dev.

I'm not advising on this from the outside. I build agentic delivery systems with teams, from a two-person startup to an SME with a 16-person product team: agents that take work items through refinement, implementation, review and testing, inside real toolchains with all their access rules, plus architecture documentation that agents can actually use. The lab packages what I've learned there, so your team doesn't have to learn it the slow way.

The pilot spots are for teams that want to help shape the lab. You get a reduced price, I get a reference and an anonymized case study.

Further reading

What you can book

Every team starts from a different place, so I quote after a first call. Prices are fixed, not hourly, and model and token costs are billed directly by your provider.

6 weeks · 1 team
Agentic Delivery Lab
From setup to a team that ships with agents, including three months of Agent Care.
Monthly · per team
Agent Care
Evals, model updates, cost reports, new skills, a monthly report and a quarterly review. Cancel anytime.

All prices on request. Book a 30-minute call and you'll get a fixed offer for your team.

Questions

Do we need new tools?

No. The lab works in Azure DevOps, Jira and Confluence, GitHub or GitLab, whichever you already use.

Does our code leave the company?

Not to me. Agents run in your environment. Which model provider you use, and where, including data protection under the nDSG, is settled in week 0.

Isn't this just Copilot with extra steps?

Copilot helps one developer write code faster. The lab changes how a team delivers: who decides, who reviews and what gets measured.

What happens after the first three months of Agent Care?

It continues monthly per team, and you can cancel at any time.

On site or remote?

Both. The three days in week 1 can happen at your office or remotely, and the weekly check-ins are usually remote.

German or English?

Both. Workshops, documents and calls are in whatever language your team works in.

Let's talk for 30 minutes

Tell me where agents are stuck in your team. I'll tell you honestly whether the lab is a good fit, or whether something smaller makes more sense.

Book a 30-minute call