According to RAND Corporation, which interviewed 65 experienced data scientists and engineers for a 2024 report on why artificial intelligence projects fail, more than 80% of those projects don't deliver what was expected — twice the failure rate of ordinary technology projects that don't involve AI.
The most-cited cause wasn't the model's computing power, and it wasn't budget. It was picking the wrong problem: teams chasing a technology that sounds impressive instead of solving a business task that was already quantified, with an owner and a clear measure of success.
That matters because the temptation inside a company is always the same: grab the most visible task — the one the director will remember at the quarterly meeting — and try to automate it first. That's exactly the task that tends to fail. The right question isn't "what impresses the most", it's "what kind of work does today's AI already handle well". And recent research answers that with unusual precision.
Where AI Is Already Good
The Organisation for Economic Co-operation and Development published, in May 2026, an index that measures the distance between what nearly 900 occupations require — across nine cognitive, social, and physical capabilities — and what today's AI actually delivers. The smaller the distance, the higher the exposure: the closer that kind of work sits to what AI already does well.
The result is sharp. In domains like creativity, the measured distance was just 0.1; in language, 0.4; in knowledge and memory, 0.5. In domains like social interaction, problem-solving, and metacognition — the ability to evaluate one's own reasoning — the distance climbs to 1.1 each. Fine manipulation of physical objects sits at 0.7.
At the occupation level, the pattern repeats: billing, typing, bookkeeping, and data-entry roles show up with distance near zero — most of what they involve is exactly the kind of work today's AI already handles well. Community and social-service occupations sit at the other extreme, with a distance of 6.4; legal and teaching occupations, at 5.8.
The OECD is explicit about what this index does not say: it is not a forecast of how many jobs will disappear. It is a map of where today's technical capability already reaches into work — the rest depends on adoption, cost, regulation, and each company's own choices. For anyone deciding what to automate first, that map is worth more than any success story: it shows where the odds of success are already high before the project even starts.
Why the Flashy Task Fails

According to the World Economic Forum's 2025 report on the future of jobs — a survey of more than a thousand companies across 22 industries and 55 economies — 40% of employers already plan to shrink teams in roles AI can automate. That number pushes the decision in the wrong direction: when the stated goal is "cut headcount", the task chosen tends to be the most visible one, not the most suitable one.
The trouble is that the visible task — answering any customer question, writing the entire sales proposal from scratch, deciding alone whether to approve a loan — almost always demands exactly the capabilities where the OECD's measured distance is largest: judgment about context, accountability for the outcome, understanding of a one-off situation. An agent — an assistant that carries out a task in steps, not just one that answers a question — can help a great deal in those situations, but it can rarely own them end to end on its own.
The side effect is the worst possible one: the ambitious attempt fails visibly, the team loses confidence in AI as a whole, and the company stops trying even on the repetitive, well-defined tasks where success was nearly certain. A failed automation charges twice: the time lost on the project, and the skepticism it leaves behind for the next attempt.
What Has to Be in Place
A good task to automate carries four checkable marks — and they apply just as much to a simple script as to a more sophisticated agent.
It happens often. If the task shows up once a quarter, the time spent designing, testing, and reviewing the process costs more than the task itself. Volume is what pays for the investment.
The standard for correct is clear. There's a right answer, an expected format, an objective benchmark to compare the result against. When "correct" depends on personal opinion or someone's mood that day, AI gets it wrong without anyone being able to point to exactly why.
The needed information already exists somewhere. A document, a spreadsheet, a system, a conversation history — it doesn't depend on guessing at context that only lives in one person's head.
The error is cheap to fix. When it goes wrong, someone notices quickly and the cost of correcting it is small. A task whose error only shows up months later, or whose error is expensive to undo, calls for human approval before any action — not direct automation.
When a task fails on any of these four marks, the right move isn't to force full automation — it's to give AI a supporting role, with human approval at the risk point, visible cost per use, and an audit trail of what was done. That's how Skyller was designed — approval scaled to risk and a record of every action as the default, not an extra setting — precisely for the tasks that sit in the middle ground between "automate outright" and "not even worth trying".
The Payoff of Choosing Well

The effect of choosing the first task well isn't just "it works". It's what comes after. A successful automation on the right task builds measurable trust: the team sees the time saved, the manager sees the error that didn't happen, and the next candidate task arrives with less resistance.
It also changes the kind of work left for people to do. When repetitive volume leaves someone's desk — reconciling a spreadsheet, answering the same question for the hundredth time, filling out the same form — the freed-up time doesn't turn into idleness: it turns into room for the judgment only a person can exercise, exactly the tasks the OECD measures as furthest from AI's current capability. A well-chosen automation doesn't compete with the team; it frees the team to do the part AI still can't.
For companies in Brazil and Latin America, this distinction matters even more: teams tend to run lean, each person carries several roles, and a visibly failed automation attempt costs trust exactly where there's the least time to rebuild it.
A Roadmap to Get Started
Before picking the next automation, it's worth running this list with the team that does the work day to day.
- List the tasks anyone on the team repeats, without exception, every month. If no one recalls a similar task from last week, it doesn't have enough volume to justify the investment.
- Ask whether there's a right way and a wrong way to do it. If the answer is "it depends on who's doing it", the standard for correct isn't clear enough yet.
- Confirm where the needed information lives. If the answer is "only in one person's head", the task isn't ready — document it first, then automate.
- Calculate how much an error costs and how long it takes someone to notice. An error that's expensive or slow to catch calls for human approval midway through, not end-to-end automation.
- Start with the most boring task on the list, not the flashiest one. It's counterintuitive, but it's exactly where the odds of success are highest — and where the team will notice the gain first.






