At a glance
- Most leaders approve AI one request at a time, usually on the strength of the best demo.
- Few can say which tasks sit inside AI’s frontier and which sit outside it, where quality falls.
- In the best-known field experiment, consultants using AI were 19 points less likely to be right outside it.
- Map the work first, test each task, and mark what stays human before anything is built.
The hardest AI decision in a service business is which work should never be handed to a model, and which work doesn’t need one at all.
We learned that on our own firm. Mapping Ludia’s four service lines and 34 functions produced 150 opportunities. Some of the rows that mattered most say stays human.
The frontier is jagged, and it doesn’t look jagged
AI’s strengths and weaknesses don’t follow how hard a task looks. In 2023, researchers from Harvard Business School and BCG gave 758 BCG consultants GPT-4 and a set of realistic assignments.
On tasks inside AI’s frontier, consultants using it completed 12.2 percent more tasks, worked 25.1 percent faster, and produced work rated more than 40 percent higher in quality.
On a task chosen to sit just outside that frontier, the same tool hurt. Consultants using AI were 19 percentage points less likely to reach the right answer than those working without it.
The researchers called the boundary a jagged frontier. Two tasks that look equally hard to a manager can sit on opposite sides of it, and the output sounds equally confident on both.
For a COO, the risk is a pilot that works on one side and gets scaled to the task next to it. Nobody notices until a customer does.
Every task gets one of three answers, and AI is only one
A map of the work should end in a decision for each task. In our Helios method, every row gets one of three answers.
Process: people own it. Safety calls, regulatory sign-off, pricing decisions, and second-person checks.
Workflow: the same inputs always produce the same output, so a rule does the work and no AI model is needed. Agent: judgment, language, or synthesis is central, so an agent drafts and a named person reviews and owns the result.
Across real inventories the mix runs roughly six in ten agents, three in ten workflows, and one in ten processes. About four in ten items need no model at all. In the first two business lines we mapped for a national power generation services operator, 54 rows split into 33 agents, 16 workflows, and 5 processes kept human.
A map that is all agents is a sales document.
Four questions decide which side of the line a task sits on
We sort each task with four questions, asked in this order. They are our method, drawn from mapping work, not a published study.
1. Does a wrong answer hurt someone or break a commitment? If a mistake touches safety, a price, a compliance obligation, or a person’s care, the row is a process. People own it, whatever the tool can do.
2. Do the same inputs always lead to the same output? Then it’s a workflow. Paying for a model on deterministic work buys cost and risk you don’t need.
3. Can a person check the draft faster than doing the work? If yes, the task is an agent candidate with a named reviewer. If checking takes as long as doing, the agent saves nothing.
4. Would your reviewer notice when it’s wrong? Outside the frontier, errors look like right answers. If the person reviewing can’t tell the difference today, keep the task human and test it again later.
The rows marked “stays human” are what earn trust
The people doing the work ask for human-owned rows, and the firms seeing results draw them.
Stanford researchers asked 1,500 workers about 844 tasks across 104 occupations. Workers usually wanted more human involvement than AI experts judged necessary. An equal partnership between person and agent was the top preference in 47 of the 104 occupations.
McKinsey’s 2026 State of AI survey of 1,719 respondents points the same way. The small group of AI high performers was far likelier to have defined when a person validates AI output: 65 percent, against 23 percent of the rest.
The survey can’t prove the validation causes the results. It does show that the firms with results decided who checks the work.
A technician who knows the agent will never make the safety call trusts it with the paperwork. That line is why the rest of the map gets used.
Where this map needs redrawing
The frontier moves, so treat the map as a living workbook. The BCG experiment used GPT-4 in 2023; tasks that sat outside the frontier then may sit inside it now, and the reverse happens after a process change. Revisit the rows that were close calls.
The evidence also has a gap. There are no randomized studies yet of field technicians, plant operators, or maintenance crews. For frontline work, test in your own operation against a measured baseline before you scale.
Some rows never move. Where a regulation or a safety rule assigns a sign-off to a person, such as a lockout/tagout release, the row stays a process however good the tools become.
What leaders should do before approving the next AI request
Map the whole business line before approving use cases. One ranked view of the work shows which requests matter and which capability serves five functions at once.
Ask the four questions of every task. The answer decides between a person, a rule, and an agent, before anyone picks a product.
Mark what stays human first. Those rows set the limits every agent works inside, and they make the rest believable.
Name an owner and a reviewer for every agent. A team name in the owner column means nobody.
Revisit the close calls as the tools change. The frontier moves; your map should move with it.
This is the work Helios 2 does for a business line, and the idea behind our view of AI. The next two pieces in the series cover what the map means for new hires and for the operating model around them.
How we know this
This point of view comes from Ludia’s Helios mapping work: our own firm (Helios 1) and client maps, described without client names.
- Jagged frontier experiment: Dell’Acqua, Mollick et al., “Navigating the Jagged Technological Frontier”, Harvard Business School working paper (Sept 2023); published in Organization Science (2026). 758 BCG consultants using GPT-4.
- Worker preferences: Shao et al., “Future of Work with AI Agents”, Stanford (June 2025). WORKBank survey of 1,500 workers, 844 tasks, 104 occupations.
- High performers and human validation: McKinsey, The State of AI (2026). 1,719 respondents surveyed May to June 2026; self-reported.
- Mapping figures: Ludia’s Helios 1 map of its own firm, and Helios 2 inventories, including the first two business lines of a national power generation services operator.



