In 2024, Klarna announced that an OpenAI-powered chatbot was doing the work of 700 customer service agents. CEO Sebastian Siemiatkowski framed it as the future of work. The company paused hiring, reduced headcount by roughly 22%, and the AI was reportedly handling two-thirds of all customer chats within a month of launch. The numbers looked clean.

A year later, Klarna reversed course. Customer satisfaction had dropped on complex interactions. The projected savings hadn’t fully materialized. As Siemiatkowski put it to Bloomberg, “as cost unfortunately seems to have been a too predominant evaluation factor when organizing this, what you end up having is lower quality.” The company began recruiting human agents again, this time as an “Uber-style” remote workforce to handle the cases the AI couldn’t. A Klarna spokesperson told Fortune the firm is “very much still AI-first.” They aren’t abandoning AI; they’re walking back the replacement framing.

Klarna isn’t an outlier. Orgvue’s 2025 survey of 1,163 senior leaders found that 39% had made employees redundant because of AI deployments, and 55% of those leaders said they got the redundancy decisions wrong. McDonald’s wound down its AI drive-through pilot after a string of order errors. Duolingo walked back its “AI-first” hiring approach amid customer backlash. The pattern is consistent: aggressive automation, quiet retreat.

These weren’t bad bets on AI. They were bets on the wrong unit of analysis. Each company looked at customer service or content moderation or order-taking as a single workflow and asked whether AI could replace it. The right question isn’t whether AI can do “customer service.” It’s which parts of the customer service workflow AI can do well, and which parts it can’t.

Decompose Before You Score

“Customer service” isn’t a task. It’s a portfolio of dozens of distinct workflows: refund inquiries, password resets, billing questions, complex disputes, complaints from grieving family members trying to close an account. Some of those are nearly perfect AI candidates — high-volume, repetitive, verifiable. Others are pure human territory. Treating them as one thing is what created Klarna’s reversal. The AI wasn’t bad at customer service; it was forced to handle pieces of customer service it had no business handling.

The same is true of nearly every workflow worth automating. Invoice processing isn’t one task. It’s intake, three-way matching, exception handling, vendor communication, and approval routing. Sales operations breaks down the same way: research, outreach drafting, meeting prep, contract review, and relationship management. Each component has a different AI fit. Some are ideal for AI automation; others require a human touch. Most workflows are a mix.

This is the move most companies skip. They look at a workflow like “customer service,” conclude it’s a mixed bag for AI, and either go all-in (Klarna’s path) or walk away. The companies that ship do something different. They decompose first — breaking the workflow into its component tasks — and evaluate each piece on its own. Just as the discovery work in the previous post of this series surfaces individual use cases rather than entire functions, the right unit of analysis is the component, not the workflow.

Once you’ve decomposed, the five questions tell you which pieces AI handles well.

The Five Questions

For each component of a workflow, work through these in order. The more yeses, the stronger the AI candidate.

Is it repetitive? Does the task follow predictable patterns the same way each time? Invoice intake fits. One-off strategic decisions don’t. Repetitiveness is the single strongest predictor; patterns are exactly what AI is built to recognize.

Is it data-intensive? Does the work involve processing volume (many records, many documents, many inputs) where the bottleneck is throughput rather than insight? Reviewing 10,000 vendor invoices is data-intensive. Reviewing one strategic acquisition isn’t. Volume is where AI’s speed advantage compounds.

Is it pattern-recognition or content-generation? Is the work fundamentally about identifying trends in existing data, or producing standardized content from a template? Variance commentary, refund responses, and demand forecasting all sit here. Crafting a board narrative for a strategic pivot does not.

Is it labor-intensive? Are people currently spending real hours on it? A task that takes one person 15 minutes a week isn’t worth the integration cost. A task that consumes 200 hours a month across the team is the engagement that pays for itself in a quarter.

Is it verifiable? Can you tell quickly whether the output is correct? AI makes mistakes; you need a fast, cheap way to catch them. Numerical reconciliations are easy to verify. Open-ended creative judgments are not. Verifiability is what separates a task you can deploy in weeks from one that becomes a permanent QA project.

A component that scores well across all five (repetitive, high-volume, pattern-driven, labor-intensive, verifiable) is a strong AI candidate.

The Override

Even when a component scores well across the five questions, there are four categories where humans should stay in the lead regardless of what the rubric says.

Taste. Creative direction, brand voice, design judgment, executive communications. There’s no ground truth for the AI to check itself against. Tasks where “right” is a matter of judgment, not correctness, don’t have a verification path even if everything else looks like a fit.

Policy and compliance. Anything where being wrong creates regulatory exposure: HIPAA in healthcare, fair lending in financial services, COPPA if your customers include minors, employment law in HR. The Air Canada chatbot we discussed in the first post of this series is the cautionary example: a customer service component that scored well on the rubric, but where the cost of one wrong answer was a tribunal ruling.

Architecture and irreversible decisions. Strategic pivots, M&A choices, restructuring, executive hires. Even when these involve heavy data analysis, the cost of being wrong compounds over years rather than minutes. Use AI as input. Keep humans on the decision.

Emotional and relational stakes. Performance reviews, terminations, key client relationships, crisis communications. In these, trust is the deliverable. This is where the Klarna re-diagnosis lands: the high-volume “where’s my refund” component is a clean AI candidate. The “I’m calling because there’s a charge from my deceased husband’s account” component hits the emotional override. Both inside the same customer service workflow. The fix isn’t to kill the AI; it’s to route the override cases to humans and let AI handle the rest.

When a component lands in any of these four categories, the score doesn’t matter. The override is binary. But — critically — the override applies to components, not whole workflows. Hitting the override on one piece of a workflow doesn’t disqualify the rest.

Putting It to Work

In our work at Areté, the engagements that ship as working systems in weeks are the ones that decompose before they score. The ones that get abandoned in month six tried to automate a whole workflow without separating the AI-suitable components from the override territory.

When multiple components score well, prioritize where time saved compounds: high frequency, many people doing it, low effort to deploy. A 30-minute component done daily by 50 people produces thousands of hours a year. A four-hour component done quarterly by one person doesn’t, no matter how clean the score.

What to Do This Week

Schedule a 60-minute session with your leadership team. The agenda is simple. Spend 10 minutes reviewing the five questions and the override. Then pick one concrete workflow (something like “month-end close” or “new-customer onboarding,” not a function like “finance”) and spend 30 minutes decomposing it together. List every component. Score each one. Use the last 20 minutes to discuss the disagreements and pick one component to pilot.

The disagreements are the value. When your CFO scores “verifiable” a 5 and your COO scores it a 2 on the same component, you’ve surfaced a real question about whether you actually have a way to check the output. That conversation is more useful than any score. By the end of the hour, you should have one component you can run a 30-day pilot on: chosen by consensus, scored honestly, and clear of the override.

The Bottom Line

The companies that win with AI don’t aim at 100%, and they don’t aim at 80% either. They decompose. They send AI at the repetitive, high-volume, verifiable components and keep humans on the taste, the policy, the irreversible decisions, and the relationships. The best deployments treat AI as a thought partner — but humans remain the thought leaders. AI handles the parts of work that don’t require judgment, freeing people to focus on the parts that do.

This is the third in a series on the AI Value Map. Next: even with the right component chosen and scored correctly, AI deployments still fail, and they fail for three predictable reasons that have nothing to do with the model.