How we organise AI agents for software delivery with evidence and control
A practical guide to using AI agents in software delivery with verifiable requirements, tests, traceability, and human review.
Why this matters
Software operations improve when decisions are tied to real delivery constraints
These resources help technical leaders make clearer decisions about software modernization, delivery constraints, continuity, and operating risk.
7
Sections
6
Minutes
In this guide6 min read
Article
A guide to move from the decision to the operating criteria and next step without losing the thread.
A team asks to automate invoice intake. The request sounds small: read an email, extract a PDF, and create a record. The real workflow raises harder questions. Which mailbox is authorised? How do we detect a duplicate? What if the supplier does not match? Who approves an uncertain amount? Which evidence must remain available to explain the result?
An agent can write code before those questions have answers. That speed then turns into rework. At Eximus, we are organising agent-assisted delivery around a clearer path: evidence and requirements first, then a reviewable specification, a bounded task, tests, a pull request, and feedback. Each change should be understandable and resumable even when a different person or model takes over.
A specification is a working contract
We call this specification-guided development, or Spec-Driven Development. It does not require a long document for every small change. It means clarifying observable behaviour and its supporting evidence before implementation.
For invoice intake, a useful specification explains which messages enter the flow, how an existing invoice is recognised, which fields are extracted, when supplier details are checked, and when a person must decide. Acceptance criteria also cover difficult cases: an unreadable PDF, conflicting amounts, a repeated webhook, and an external system that is temporarily unavailable.
The specification states what the change will not do. Extracting data does not authorise an agent to approve an invoice, reconcile a bank transaction, or send an external message. That boundary makes it possible to delegate implementation while keeping business decisions with the right people.
The delivery path
Conceptual path: approved evidence → requirements and specification → human review → bounded work item with an owner → code and tests → technical gates → PR and human review. Open decisions return to the evidence; a failed gate returns to the work item; PR feedback can update the requirements. The exact controls depend on the risk of the change.
- Evidence. Meetings, documents, existing behaviour, and constraints keep a visible source. An agent's statement is not the source itself.
- Requirements. Confirmed facts, proposed interpretations, and open decisions are kept separate. Scope and responsibility questions are settled before building.
- Specification. The team defines the expected result, error cases, dependencies, acceptance criteria, and exclusions.
- Build readiness. The change has an identifiable technical surface, a test approach, the right permissions, and a recovery path. If uncertainty remains high, the next work item may be a focused investigation.
- Executable work item. Work is linked to a Story or Bug with a visible owner and state. Create a child Task when it represents distinct work under that item. An unowned chat does not authorise a repository change.
- Build and gates. The agent implements the slice, runs relevant tests, and records results. A green pipeline proves that specific checks passed; it does not prove business acceptance.
- PR and feedback. Review connects the code to its evidence. Feedback returns to the requirement or implementation that needs attention without losing the original decision.
Claude, Codex, and the source of truth
Claude and Codex can help analyse sources, draft a specification, implement a Task, or review a change. Roles depend on the Task; we do not claim that one model always owns a particular stage. The relevant Azure DevOps Story, Bug, or Task, repository, approved documents, and test results are the artefacts the team can inspect. An agent chat is working context, not the only project record.
Parallel work needs a simple rule: each unit has an active owner and a declared change surface. If two Tasks must modify the same module or data contract, the work is sequenced or integrated deliberately. Otherwise, the apparent time saving can become merge conflicts and inconsistent decisions.
A checkpoint makes work resumable
A useful work item checkpoint includes the requirement identifier, sources used, change surface, commit or PR, test commands and results, open decisions, and next step. When a test fails, the record explains how to reproduce it. When work pauses, another participant can resume without asking the agent to retell the entire chat.
Working material also needs boundaries. Approved sources, secrets, personal data, and permissions must remain within the project's rules. A meeting transcript may help identify a requirement, but it should not be copied in full into a PR when reviewers only need the decision and its reference. Traceability means keeping useful links and decisions, not duplicating every piece of source data.
Return to the invoice example. The agent implements document reading and a test for a valid attachment. A gate then finds that processing the same message twice creates two records. The work item receives a reproducible failure. The specification is clarified to require a deduplication key. The revised change passes the test, and the PR shows both the normal and duplicate cases. The system did not “learn by itself”: the team found a gap, made the behaviour testable, and kept the evidence.
Controls that match the risk
A small interface change may need a focused test and visual review. A change to permissions, financial calculations, or customer data needs stronger evidence: sources, rules, negative tests, an approval owner, and a rollback path.
The agent can prepare a change. Tests and pipelines can verify technical conditions. An authorised person must accept the behaviour and decide whether it can advance or be published. These are different decisions. Keeping them separate prevents automation from gaining authority that the project never granted.
A gate should also be understandable. For each type of change, the team should know which check blocks progress and who can resolve it: automated tests for known regressions, visual review for interface changes, security review for new permissions or sensitive data, and functional validation for business rules. If the test environment is unavailable, the result is “not verified,” not “approved by default.”
What improves, and what it costs
This flow asks for more work early: clarifying requirements, writing scenarios, and keeping evidence. It pays off when several systems, frequent changes, critical operations, or handovers are involved. For a small isolated fix, the specification may fit inside the same work item. An integration with exceptions and business decisions deserves more detail.
We would measure time from a clear request to a reviewable delivery, rework caused by ambiguous requirements, Tasks with tests and evidence, defects after release, and the time another person needs to resume a change. We do not present public performance figures for this workflow yet.
Start with one bounded flow
Choose a workflow with a visible problem and clear limits. Document a few real scenarios, identify the approval owner, define a small executable work item, and record the questions that appeared before and after the PR. If the first cycle leaves a specification people can discuss, a test that demonstrates behaviour, and a history that explains the decision, there is a sound basis for expanding agent use.
Can your team trace a change from the original request to its tests and release decision? Discuss a first traceable agent workflow with Eximus.
Related topics
Explore more on this topic
This article connects with other resources that explain the operating, commercial, and technical context behind the flow.
Recommended path
How to turn this topic into an execution path
More content
Keep exploring
ai agent operations
AI Agent Operations
How to introduce AI agents into real enterprise workflows with clear system boundaries, human oversight, and operational traceability.
ai agent operations
AI agents for real operations in the Netherlands
The first useful AI-agent deployment inside an enterprise is rarely the biggest one. It is usually the workflow with a clear boundary, real supervision, traceability, and an operating model the team can trust.
ai agent operations
Automated technical onboarding for Eximus collaborators
Eximus is automating technical onboarding from the signed contract: corporate identity, access, VPN, credentials, operating tools, and traceability in a governed flow from EximusHub.
ai agent operations
How Eximus started adopting OpenClaw
We started adopting OpenClaw early because Eximus already had a serious base in Azure, IaC, security, and enterprise integration. That let us test agents with control instead of treating them as an isolated demo.