The most useful AI project for your business probably does not look like an autonomous company. It looks like one tedious, recurring job completed faster, with fewer omissions and a clear person responsible for the result.
That may sound less exciting than hiring a team of “digital employees”. It is also much more likely to work.
AI assistants can now search documents, analyse information, draft materials and use business software. Some can carry a task through several steps rather than merely answering a prompt. But capability is not the same as reliability. Performance tends to fall when work becomes longer, less clearly specified, politically sensitive or difficult to check.
For a small or medium business, the sensible starting point is therefore one bounded, measurable assistant workflow. Give it a defined outcome, limited access and an acceptance test. Keep people in charge of money, commitments and exceptions. Expand only when repeated results justify it.
Assistant or agent? The plain-English version
An AI assistant helps a person do a job. You ask it to summarise a meeting, draft a proposal or compare suppliers; a person remains closely involved.
An AI agent has more freedom over how it completes a job. It can choose several steps, use approved tools, inspect the result and try again within set limits. For example, it might retrieve figures from a CRM, compare them with last month, flag anomalies and prepare a weekly report.
This is not a sharp technical divide, and you do not need to become preoccupied with the terminology. The practical question is: what can the system read, decide and change without asking someone?
Anthropic offers a useful distinction in its guide to building effective agents: a workflow follows predefined paths, while an agent dynamically chooses its steps and tools. Dependable business systems usually combine both. The workflow supplies the rails; AI handles a bounded area of ambiguity.
Start with an assistant. Introduce agent-like freedom only where choosing between steps creates genuine value.
Choose work with a clear finish line
A good first workflow has five qualities:
- it happens regularly;
- it consumes meaningful staff time;
- the necessary information is available and reasonably clean;
- the output can be checked against evidence; and
- a failure is reversible and limited in impact.
Useful starting points include:
A weekly management briefing
Let the assistant collect read-only data from sales, finance, marketing and support systems, then prepare a concise operating pack. Every figure should link back to its source, while recommendations remain proposals for a manager to approve.
Customer feedback analysis
An assistant can group reviews, support tickets and survey comments into themes, identify recurring complaints and draft a summary. A person should inspect the underlying examples before changing policy or making public claims.
Lead or supplier research
The system can gather agreed facts, record sources and prepare a shortlist against explicit criteria. It should not make unsupported judgements or contact people under your name without approval.
Proposal and campaign preparation
An assistant can assemble a first draft from approved service descriptions, case studies and brand guidance. Staff still check pricing, promises, customer details and final publication.
Internal knowledge maintenance
It can flag stale documents, turn resolved questions into draft guidance and identify contradictory instructions. A named owner approves changes to the official record.
These jobs are useful because they reduce preparation work without transferring consequential authority.
What not to automate first
Do not begin with the process that has the biggest theoretical saving. Begin with the one where you can learn safely.
Keep these areas outside an initial pilot:
- unrestricted access to bank accounts, cards or purchasing;
- hiring, dismissal, pay or employee discipline;
- legal commitments or regulated advice;
- unsupervised refunds and policy exceptions;
- public publishing where an error could damage the brand;
- security changes or production deployment without rollback; and
- any objective whose success you cannot measure.
The issue is not that AI can never contribute to these activities. It can prepare information or drafts. The issue is that errors involve rights, money, trust or irreversible consequences. Accountability remains with your business, not the software supplier and certainly not the assistant.
Klarna illustrates the need for balance. The company initially reported large gains from its customer-service assistant, including faster resolution and substantial volume handled by AI. It later reinvested in human customer-service capacity and emphasised customer choice. Automation can improve efficiency while still weakening service if containment and cost become the only measures.
Write a one-page operating contract
A vague instruction such as “improve our sales operation” is not delegation. It is an invitation to improvise.
Before connecting an assistant to business systems, write a one-page contract covering:
- Objective: What business outcome should improve?
- Baseline: What happens today, how long does it take and how often is it wrong?
- Deliverable: What exact report, draft or system update must be produced?
- Acceptance test: What must be true before the work is accepted?
- Allowed information and tools: What can it read or use?
- Prohibited actions: What must it never do?
- Budget: Set limits for cash, usage, time and retries.
- Approval gates: Which actions require a named person?
- Evidence: What sources, calculations or records must accompany the answer?
- Stop rules: When should it fail safely and escalate rather than guess?
For a weekly operating pack, an acceptance test might require every figure to trace to a source system, every recommendation to have an owner and next step, and missing data to be labelled rather than estimated.
This document is the real foundation of the project. Tool selection comes afterwards.
Put safeguards around the work
An assistant that can act in software is a new non-human user. Treat it accordingly.
Give it the least access needed. Read-only permissions are ideal for a first pilot. Use a separate identity rather than sharing an employee’s account, and never place broad credentials in prompts or documents.
Keep a run log showing what the assistant read, which tools it called, what it produced, how much it cost and who accepted it. Conversational memory is not an audit trail. Important state should live in a task record or existing business system.
Require approval before the assistant spends money, sends external communications, changes records, publishes material or creates a commitment. Add limits on retries and elapsed time, plus a straightforward kill switch.
Finally, verify the real result. A polished explanation that a task was completed is not proof that the CRM changed correctly or that a source supports the claim. Check records, links, calculations and downstream effects.
Understand the full cost
Software pricing can make these projects look exceptionally cheap. The model itself may cost pennies for a modest text job, but that is only one line in the budget.
The total includes the platform subscription, model usage, search or browser fees, integrations, storage, monitoring, human review, retries and rework. Browser-heavy and research-heavy tasks usually cost more than tidy text workflows.
The largest hidden cost is often supervision. At a hypothetical loaded staff cost of £30 an hour, five minutes of review across 1,000 jobs costs £2,500. Fifteen minutes costs £7,500. A cheap system that creates review debt may be worse value than a more capable model that succeeds more consistently.
Measure cost per accepted outcome, not cost per token. Also track time saved, correction rate, completion rate and whether the output was actually useful.
As an indicative range from current product and usage pricing, a solo experiment may cost roughly $30–$150 a month. A production small-business workflow may land around $150–$1,500 before human review, depending on volume and integrations. Treat these as planning bands, not quotations: prices and usage patterns change quickly.
A practical 30-day pilot
A month is long enough to test repeated work, but short enough to prevent an experiment becoming permanent by accident.
Days 1–5: select and measure
Choose one read-only or reversible workflow. Observe how it works today. Record staff time, turnaround, error rate and typical volume. Name one process owner and one person authorised to accept results.
Write the operating contract. If you cannot define a pass or fail test, choose a different workflow.
Days 6–10: build the smallest useful version
Connect only the minimum data and tools. Use fixed workflow steps wherever the process is predictable, with AI reserved for tasks such as classification, summarisation or drafting.
Prepare ten to twenty representative examples, including awkward cases and missing data. Remove or protect sensitive information that the pilot does not need.
Days 11–20: run in parallel
Run the assistant alongside the existing process rather than replacing it. Do the same job at least three times under comparable conditions; one successful demonstration proves very little.
For every run, record completion, factual corrections, review minutes, elapsed time, cost and usefulness. Capture failures rather than quietly fixing them and moving on. They reveal unclear instructions, poor data and undocumented exceptions.
Days 21–25: tighten the controls
Review where the system guessed, duplicated work or used the wrong source. Improve acceptance tests, permissions and stop rules before improving the prose.
If a person repeatedly corrects the same issue, turn that correction into an explicit rule or test. Do not add more agents merely to make the design look sophisticated.
Days 26–30: decide with evidence
Compare the pilot with the baseline. Continue only if it produces an accepted outcome more quickly, cheaply or consistently without creating unacceptable risk.
Your decision should be one of four things: stop; revise and test again; keep the workflow at its current authority; or expand one carefully defined permission. Do not jump from a useful report generator to unsupervised customer contact or spending.
Scale the workflow, not the mythology
Large organisations are experimenting with teams of specialised agents, but the relevant lesson for a smaller business is not the number of digital workers. It is the operating discipline around them.
BNY’s reported use of “digital employees”, for example, is notable because the bank is building identity, management and control structures around agents inside a regulated institution—not handing the bank to a model. PwC has reported targeted client gains from agent workflows, but those results are supplier-reported examples rather than a universal benchmark. Meanwhile, Forrester describes a wide gap between interest in agentic AI and meaningful production systems.
A second specialist assistant may help when the first workflow has a proven bottleneck—for example, separating data retrieval from checking. More agents can also multiply context loss, duplicated effort, cost and confident mistakes. Add complexity only when the measurements support it.
The businesses most likely to benefit from AI assistants will not be those that surrender the most control. They will be those that describe work clearly, provide reliable context, restrict authority and learn from every run.
Start with one Monday morning report, one research pack or one set of draft proposals. Make the finish line visible. Measure the result. Then earn the right to automate the next step.
