The Scope of Work AI Builds Need Before Day One
TL;DR
If every decision in your business still runs through you, the scope of work is where that starts to change. A scope of work for an AI build has to do four things. Name every agent and its job. Set a pass or fail test for each one. Fix how much freedom each agent gets. Commit to a working prototype on one workflow inside 30 to 90 days. Get the document right and your team can run the system without you checking every step.
Why scoping an AI build is different
A normal app does what it is told. An agent chooses how to do the task.
That one difference changes the document. The scope has to cover judgment calls, not just features. Who decides when the agent is unsure? What can it send without asking? What happens when it gets it wrong?
Without those answers, projects drift. Agents get built with no clear owner. Nobody agrees on what "done" means. You end up with a pile of automation nobody trusts enough to switch on.
The scope is the cheapest step in the whole project. DevWiz covers the same stage for software builds in its guide to the AI software discovery phase.
The scope of work AI agents need, section by section
Six sections do the work.
| Section | What it says | The fight it stops |
|---|---|---|
| Agent list | Each agent's job in one line, what goes in, what comes out | Nobody knows who owns a broken step |
| Pass or fail tests | A measurable test for every deliverable | Arguments about what "done" means |
| Tools and data | Which systems each agent can touch | An agent with access it never needed |
| Risky actions | Sending money, emailing clients, deleting files | Finding out the hard way |
| Human owners | A named person for every escalation | "I thought someone else was watching" |
| Go-live checks | The tests the system must pass first | Shipping on hope |
UC Berkeley's Agentic AI Risk-Management Standards Profile backs the same ground. It calls for agent-specific policies that cover delegated decision-making, tool access and an agent's ability to set its own sub-goals. It also says to enforce least privilege on tool access.
One more line belongs in every scope: what does not need AI. Sometimes the right tool is a plain automation with no AI in it. Forcing AI into a process that does not need it adds risk for no gain.
How much freedom each agent gets
A new hire does not get the company card on day one. Your agents should not either.
Write three tiers into the scope:
- Tier 1. The agent suggests. A human approves every action.
- Tier 2. The agent acts on routine work. A human reviews on a set schedule.
- Tier 3. The agent acts inside a fixed boundary. It raises an alert on exceptions.
Start any agent that touches a client at Tier 1. Move it up only after it has held its accuracy over a set number of reviews.
Oversight has levels too, and they run the other way. The higher the level, the more senior the human. The Berkeley profile points to a three-level model. At level one, automated systems watch most actions. At level two, anomalies and high-stakes decisions go to human reviewers with the right expertise. At level three, the most serious issues go to a senior oversight group.
It also names a risk that is easy to miss: speed. An agent can act faster than anyone can watch. So the scope should say who can stop an agent, and how.
Two questions to settle in writing. Can an agent hand work to another agent? If so, how many layers deep before a human signs off?
Start with one agent
One agent that checks work and hands out tasks beats five agents tripping over each other.
Think of a kitchen. One head chef runs the pass. A pastry chef joins only when there is a job nobody else can do. Add cooks too early and you crowd the kitchen.
The vendors say the same. OpenAI's orchestration guide puts it plainly. Begin with a single agent. Bring in specialists only when they clearly improve how capabilities or policies are kept apart, how clear the prompts are, or how easy a run is to trace. Redis, in its guide to splitting work across sub-agents, suggests trying a single agent first, then better prompts, then tools, before you add agents. It also warns that the extra turns agents spend sharing context can make a multi-agent system cost more than a single one.
Your scope should ask for:
- One lead agent first. Specialists are added only where they clearly help.
- A cheap first check. A small model handles the easy tasks. A stronger one takes the rest.
- Simple paths. Straight lines where you can. Loops only where a task needs repeat checks.
- Handoff rules. A set way to pass context between agents so nothing is lost.
That last one matters most. The OrchBench research on multi-agent orchestration found that keeping task-critical information intact matters more than adding agents. It also found the gains from running agents in parallel shrink as coordination failures pile up. Our post on the six-field agent handoff contract shows what to write down.
What the first 90 days should deliver
A good scope does not promise a finished system on day one. It promises a working prototype fast, then builds from there.
- Foundation. Capture the founder's IP. List the tasks. Set the tiers. Map the risks.
- Prototype. Build one small, live workflow. Run the basic tests.
- Validation. Test real cases. Replay past decisions through the system. Try to break it on purpose.
- Handover. Write the runbooks. Train the team. Agree who does what when something fails.
We map those phases to the four functions in NIST's AI Risk Management Framework. Foundation covers govern and map. Prototype and validation are measure. Handover is manage. The Generative AI Profile is NIST's companion guide for applying them to generative AI.
Never sign a scope that skips the prototype. One workflow is enough to prove the point. In one build we describe in our Slack AI workflows playbook, a weekly report that took up to four hours by hand now runs in 22 minutes. Our post on the 90-day human-in-the-loop prototype walks through the structure.
What to measure once it is live
Bind the scope to four things:
- What gets tracked. How fast agents act, what they tried to access, and how far work was handed down.
- Who reads it. A named person, on a weekly or monthly rhythm.
- What happens on a breach. Clear steps, including how to shut an agent down.
- Ongoing tests. Regular replays of past decisions, so you catch drift before a client does.
Our guide to AI agent testing covers the checks in detail.
How to draft it
Start with the founder's head, not a template. Your process is the product. Capture it in plain steps before anyone writes code.
- List every task the founder does that an agent could take on.
- Turn each task into a role with clear inputs and outputs.
- Write the pass or fail test for each role.
- Set the autonomy tier.
- Write the risk line. What is the worst that happens if this agent is wrong?
Keep the language plain. If a developer or a lawyer has to ask what a line means, rewrite it.
Swap vague deliverables for ones you can measure. Not "improve efficiency". A response time. An error rate. A number both sides agreed.
Then review the draft with whoever will run the system day to day. They spot gaps a founder will not see.
Who needs to be in the room
| Role | What they own |
|---|---|
| Founder | The knowledge being encoded. Signs off the tiers |
| Operations lead | Where agents touch client work, and where mistakes hurt most |
| Technical lead | Turns the scope into a build. Flags what is not realistic |
| Named reviewer | Reads the results on the agreed schedule |
| The team | Confirms the workflow fits how they work |
Do not skip the last row. The tech work is the easy 20%. Resistance is the other 80%. A build the team did not help shape sits on a shelf.
Three example scopes
| Agent | Input | Output | Tier | Escalates when |
|---|---|---|---|---|
| Content | Founder's notes or a recorded call | A draft in the founder's voice | 1, then 2 once accuracy holds | Anything is due to be published |
| Client delivery | Intake form and past session notes | A session plan | 1, then 2 with weekly review | Pricing or contract terms change |
| Support triage | An inbound client message | A sorted ticket and a draft reply | 1 for refunds and complaints. 3 for simple questions, once tested | A refund or a complaint is mentioned |
Every row has the same shape: role, input, output, tier, escalation. That is what makes the document usable by more than one team.
Where most scopes fail
- Vague deliverables. "Make it smarter" is not a test.
- No named owner. When it breaks, everyone assumes someone else is watching.
- No risk map. High-stakes actions get the same oversight as trivial ones.
- Too many agents too soon. More cost, more places to fail.
- No prototype clause. A full build signed off before a small one is tested.
- Lost context. Handoffs that drop information.
Prompts are the easy part. Ownership is the part people skip.
How we scope a build
AI is an enthusiastic intern, not a magic button. So we run 10/80/10. That is 10% planning, 80% the AI doing the work, and 10% a human making it right. The scope of work is the first 10%.
Our method is Explore, Map, Transform. Explore pulls out the calls only you make. Map turns them into roles, tiers and tests. Transform builds them into AI employees with Claude Code, coordinated as one AI Operating System, so the output scales without you in every loop.
That is the shape of the 90-day program. If you want to see the build side, read how we build custom AI delivery systems with Claude Code.
The first page of any scope is the list of tasks only you can do today.
James Killick
Find the first workflow to scope
The best first scope covers the work that leans on you hardest. If you are not sure which workflow that is, measure it.
Take the Founder Bottleneck Assessment. It scores five dimensions in six minutes, then names your next move.
Sources
- Agentic AI Risk-Management Standards Profile, UC Berkeley Center for Long-Term Cybersecurity
- Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST
- OrchBench: Evaluating Multi-Agent Orchestration Plans in Isolation via Deterministic Simulation, arXiv
- Orchestration and handoffs, OpenAI developer guide
- Sub-agents and splitting context across agents, Redis
Frequently Asked Questions
James Killick
Founder
The AI Orchestrator. 10+ years building digital products and 200+ apps shipped, now helping $1M+ educators and consultants turn their IP into AI-powered delivery systems.
James Killick founded and runs The AI Orchestrators.
More from James Killick