Quick Answer: A systematic review of 106 studies and 370 effect sizes from MIT Sloan found that human-AI combinations performed worse, on average, than the better of the two working alone. So the useful question for mid-market companies is not whether to add AI to an operation. It is how much work can be moved into a governed human-machine system where AI does what it is measurably good at, humans retain judgment and accountability, and someone is responsible for supervising the boundary between the two.

Key Takeaways

  • A 106-study, 370-effect-size review found hybrid human-AI teams underperform the best standalone option on average, so “hybrid always wins” is not a defensible claim (MIT Sloan).
  • In a 2025 field study of 5,172 customer-support agents, AI assistance raised issues resolved per hour by about 15%, with gains concentrated among newer and lower-skill agents (Quarterly Journal of Economics).
  • A preregistered study of 758 BCG consultants using GPT-4 found 25.1% faster work and over 40% higher quality inside the model’s competence, but 19 percentage points lower accuracy on a task placed deliberately outside it (Harvard Business School).
  • 84.7% of consumers still prefer a human agent to an AI one, and that preference barely changes even when the AI option guarantees resolution (Metrigy data via No Jitter).
  • AI adoption climbs sharply with company size: 11.9% among 10-49 employee firms versus 40% among firms with 250 or more, a gap the OECD attributes to fixed adoption costs most mid-market operators have not yet absorbed (OECD).
  • A lean four-person internal AI operations team runs well over $825,000 in fully loaded annual cost, against a US private-sector average of $46.60 an hour in total compensation, roughly a third of it benefits (US Bureau of Labor Statistics).
  • Gartner forecasts that more than 40% of agentic AI projects will be cancelled by the end of 2027, a forward-looking estimate of project abandonment, not a measured current failure rate (Gartner).

The Wrong Question Is “How Much Can AI Remove”

Most conversations about AI in customer operations, finance and back-office work start with the wrong unit of analysis. They ask how many people a given job requires once AI is introduced, as if a job were one indivisible task. Operations do not work that way. A customer-support role, an accounts-payable function or an IT help desk is a bundle of tasks with wildly different automation potential, reversibility and consequence, retrieval and drafting sit at one end, ambiguous exceptions and consequential judgment calls sit at the other. Treating the whole job as a single automation decision is what produces both the overclaiming and the eventual retreat that has defined the last two years of AI-in-operations coverage.

The more useful question, and the one this article answers with the evidence available in 2025 and 2026, is a second-order one: how much work can be delegated to a team that combines AI and human delivery, and still get results that hold up against the alternative. That question has a real, if conditional, answer.

The Evidence Against “Hybrid Always Wins”

It would be convenient to claim that adding AI to a human team automatically produces a better outcome than either working alone. The evidence does not support that. The MIT Sloan-led systematic review of 106 studies and 370 effect sizes found that, on average, human-AI combinations performed worse than the better standalone performer. Hybrid teams did not carry a built-in advantage. They carried an advantage under specific conditions, chiefly when work was decomposed so the machine handled search, extraction and drafting while people owned ambiguity, authorization and recovery, and when humans retained the ability to override the AI’s output rather than defer to it by default.

The clearest illustration of where that boundary sits comes from a preregistered study of 758 Boston Consulting Group consultants using GPT-4. On tasks inside the model’s competence, AI-assisted consultants completed 12.2% more tasks, worked 25.1% faster, and produced output rated over 40% higher in quality than the control group. On a task placed deliberately outside that competence boundary, the same AI-assisted consultants were 19 percentage points less accurate than consultants working without AI at all. The researchers named the pattern the jagged technological frontier: AI capability does not degrade gradually as tasks get harder, it falls off an uneven, hard-to-predict edge, and the same tool that makes someone faster and better on one task can make that person confidently wrong on the next. That combination of findings is why the honest framing for an AI-augmented operating model is conditional, not absolute, and why any partner claiming an unconditional hybrid advantage should be treated with the same skepticism as a vendor claiming 100% chatbot containment.

Where AI-Augmented Human Teams Actually Outperform

None of this means AI augmentation is weak. It means the credible case for it is narrower and more specific than most vendor claims suggest, which is also what makes it defensible. The strongest field evidence for AI-augmented human delivery in a live operation comes from a 2025 study published in the Quarterly Journal of Economics, covering 5,172 customer-support agents. AI assistance increased issues resolved per hour by roughly 15%, driven by shorter handling time and a modest rise in successful resolution. Crucially, the gains concentrated among lower-skill and less experienced agents, and the research also found improved customer sentiment, fewer requests to escalate to a manager, and evidence of better retention among newer staff. The mechanism matters more than the headline number: AI captured behaviours associated with the strongest-performing agents and disseminated them across a broader pool, effectively compressing part of the learning curve rather than simply replacing labour.

That mechanism is directly relevant to offshore delivery. A distributed team spread across customer experience, back-office, financial services and IT support functions benefits disproportionately when AI assistance helps newer or less specialized agents perform closer to the level of the team’s strongest performers, rather than treating AI purely as a headcount-reduction lever. Back-office processing shows a related pattern from a different angle: accounts-payable benchmarking commonly places invoice-processing cost near $10.89 per invoice for bottom performers versus roughly $1.77 for top performers, a sixfold spread associated with automation, standardization and process maturity rather than headcount alone (ProcureDesk, citing APQC benchmarking). The pattern across both examples is the same: the winning unit is the workflow, not either participant in isolation.

Why Customers Still Want a Human Escalation Path

Any AI-delivery strategy that ignores customer preference is optimizing for the wrong metric. The consumer research here is unusually consistent in direction even where exact figures vary by sample. Analyst firm Metrigy’s 2025-2026 study found 84.7% of consumers preferred a human agent to an AI one, and that preference barely moved even when the AI option guaranteed resolution, 80.1% still chose the human. The practical implication is that perceived control, accountability and an accessible escalation path materially affect whether customers accept an AI-heavy service model at all, independent of whether the AI is technically accurate.

Two documented cases make the stakes concrete. A Canadian tribunal held Air Canada responsible after its chatbot gave a passenger incorrect bereavement-fare guidance, rejecting the argument that the chatbot was a separate legal entity. The award itself was small, but the precedent is operationally significant: an AI interface remains the company’s representation, and automation does not transfer accountability away from the business that deployed it. Klarna offers the more strategically relevant example. After heavily promoting AI-led customer service and reducing human staffing, the company restored a stronger human-service option, with management acknowledging that an excessive focus on automation-driven cost reduction had affected service quality even while AI continued handling a majority of routine inquiries. Neither case proves AI failed at every interaction. Both prove that high automation volume and an acceptable customer experience are different objectives, and a governed model has to manage both.

How Much Does It Cost to Build This In-House?

Here is the part most AI commentary skips: none of the above happens automatically. Someone has to design the workflow decomposition, own the knowledge base the AI draws from, run quality assurance against a live scorecard, manage prompt and model versioning, watch for drift, and decide exactly which actions an AI system is and is not allowed to take on its own. That is an operating discipline, not a software subscription, and it requires a specific, scarce mix of specialists: AI workflow designers, knowledge engineers, QA and model-evaluation specialists, and people who understand both the operational process and the AI tooling well enough to keep the two aligned.

This is exactly where firm size becomes the deciding factor. Across OECD member countries, 2025 estimates place AI adoption at roughly 11.9% among firms with 10-49 employees, 20.4% among firms with 50-249 employees, and 40% among firms with 250 or more, a gap the OECD attributes largely to fixed adoption costs and complementary assets such as ICT skills, data infrastructure and prior digital investment that smaller and mid-market firms have simply not accumulated yet. The cost gap is concrete, not directional. Market estimates for a lean, four-person internal AI engineering group, covering senior AI/ML, data engineering, MLOps and backend integration, run from roughly $825,000 to over $1.1 million in fully loaded annual cost before recruiting, infrastructure and ramp time are even counted, against a US private-industry average of $46.60 an hour in total compensation, of which 30.1% is benefits load alone.

Cost componentUS in-house, no AIUS in-house, AI-augmentedOutsourced AI-augmented delivery
Human capacity for a fixed unit of workBaseline headcountPotentially lower headcount if productivity gains transfer to the specific workflowVendor-staffed; a credible partner discloses productive headcount and AI allocation
Recruiting and attrition riskBuyer bears in fullBuyer bears, including scarce AI-operations skills specificallyProvider bears contractually, priced into the contract
AI platform, integration and governance costNot applicableAdded on top of labour cost: platform, evaluation, monitoringIncluded in the vendor’s operating cost, must be stated explicitly, not hidden inside a seat rate
Specialist coverage (AI workflow design, QA, model evaluation)Not typically justified below significant scaleRequires hiring or contracting scarce specialists individuallyShared across the provider’s client base, the core economic argument for outsourcing this layer

No neutral, publicly available dataset compares identically scoped US in-house teams, US in-house AI-augmented teams and offshore AI-augmented outsourced teams under the same service levels and volume, so this table is a decision framework, not a market-average claim. Any specific savings figure a provider quotes should come from an auditable proposal for the buyer’s actual scope, not a percentage claimed in the abstract. Buyers evaluating outsourcing data security and compliance alongside this decision should ask for the same specificity on the AI-governance side that they already expect on the information-security side.

What a Governed AI-Augmented Operating Model Actually Requires

Forrester expects roughly 30% of enterprises to build service functions that mirror human management structures for AI specifically, onboarding, coaching, monitoring and unblocking AI systems the way a manager would develop a team of people. That prediction lines up with what the evidence above actually supports: operational autonomy is far less mature than headline AI-adoption numbers suggest, and most credible deployments preserve human responsibility for exceptions, approval and customer recovery rather than removing it.

A workable operating model therefore looks less like “add a chatbot” and more like a layered system. AI screens every interaction rather than a small manually sampled fraction. Humans calibrate the scorecard the AI is measured against. QA specialists investigate the highest-risk and lowest-confidence cases the system flags. Supervisors coach agents using what the AI surfaces. Operations teams update prompts, knowledge content and routing rules on an ongoing basis. A controlled sample is double-scored by humans to catch drift before it compounds. This is the same operational discipline that separates a serious BPO partner from a call-volume vendor, extended to cover the AI layer instead of stopping at the human one, and it is the layer an IT outsourcing partner’s AI Studio capability is specifically built to provide alongside core delivery.

Governance extends to where the data actually goes. A South African delivery operation processing client data typically acts as an operator under the Protection of Personal Information Act, with the buyer remaining the responsible party, which means a written processing agreement, documented instructions, incident-notification commitments and control over subprocessors including any AI model or cloud provider in the chain. South Africa is not on the European Commission’s adequacy list, so EU personal data flowing to a South African delivery team generally requires Standard Contractual Clauses or an equivalent Article 46 safeguard, not an assumption of automatic adequacy. None of this is exotic for a compliance-certified provider. It is exactly the diligence a buyer should already be applying to call center and customer support outsourcing generally, extended to cover model use, prompt handling and subprocessor flow-down specifically.

What Mid-Market Buyers Should Ask For

The commercial model has to change to match the operating model. A provider that charges purely by seat while keeping every AI-driven efficiency gain for itself is traditional labour arbitrage with an AI label attached. A provider that charges purely on automation “deflection,” without separately reporting resolution quality and repeat contact, has a direct incentive to hide service failure behind a favourable-looking automation number. Neither is the model a mid-market buyer should sign.

A credible AI-augmented outsourcing contract should separate and report, as a standing operating rhythm rather than a one-off audit: AI-only resolutions, human-assisted resolutions and human-only resolutions as distinct categories; escalations and failed automations; repeat contacts and reopened cases; corrections made after an AI action; cost per successfully completed outcome rather than per interaction or per seat; and human override rates. That level of transparency is what distinguishes measurable BPO ROI from a vendor’s own favourable metric, and it belongs in any BPO contract or SLA that includes an AI-augmented delivery component. Gartner’s own forecast that more than 40% of agentic AI projects will be cancelled by the end of 2027, a projection of future cancellations, not a measured failure rate today, is itself a reason to insist on outcome transparency from day one rather than discovering the gap two years into a contract.

The strategic question mid-market leaders should be asking is not how many people AI can remove from the payroll. It is how much of the work can be moved into a properly governed human-machine operating system, one that combines AI’s speed on well-defined tasks with an offshore team’s judgment, accountability, and ability to recover when a real interaction departs from the script, which it constantly does.

Frequently Asked Questions

Does adding AI to an offshore team automatically improve results?

No. A 106-study meta-analysis led by MIT Sloan found that human-AI combinations performed worse, on average, than the better standalone performer. Hybrid delivery outperforms only when work is deliberately decomposed so AI handles retrieval, drafting and pattern-matching while humans retain judgment, exceptions and override authority.

What does the evidence say about AI-assisted human agents specifically?

A 2025 Quarterly Journal of Economics study of 5,172 customer-support agents found AI assistance raised issues resolved per hour by roughly 15%, with the largest gains among newer and lower-skill agents, alongside improved customer sentiment and fewer manager escalations.

Do customers actually want AI-only service?

No. Metrigy’s 2025-2026 research found 84.7% of consumers preferred a human agent over an AI one, a preference that barely changed even when the AI option guaranteed resolution. Multiple other studies from the same period land in a similar range.

Why can’t a mid-market company just build this AI-operations capability in-house?

A lean four-person internal AI engineering team costs well over $825,000 in fully loaded annual expense before recruiting and ramp time, and OECD data shows AI adoption climbing sharply with firm size, from roughly 12% among small firms to 40% among large ones, reflecting the fixed costs and specialist capabilities smaller firms have not yet built.

What should a buyer ask an AI-augmented outsourcing partner to report?

Ask for the split between AI-only, human-assisted and human-only resolutions, escalation and failed-automation rates, repeat contacts, corrections made after an AI action, cost per successfully completed outcome, and human override rates, reported as a standing rhythm, not a one-time audit.

Is South African data governed differently when AI is part of the delivery model?

The underlying law does not change, but the scope of diligence does. A South African delivery operation typically acts as an operator under POPIA on the client’s instructions, and EU data specifically requires Standard Contractual Clauses since South Africa is not on the European Commission’s adequacy list. AI-specific diligence adds a model and subprocessor register and a written policy on whether client data trains shared models.

Does AI-augmented delivery mean fewer people are needed offshore?

Not in a simple sense. The evidence shows AI compresses part of the learning curve and raises the ceiling on what a distributed team can handle, but it also creates new work: prompt management, QA calibration, drift monitoring and escalation design. The realistic case is a shift in what offshore teams do, not a straightforward headcount reduction.

Afrishore’s IT Outsourcing division delivers the AI-augmented layer described here, agent assist, automated QA and virtual-agent tooling, alongside core business process outsourcing delivery in customer experience, back-office, financial services and technical support. To discuss where AI-augmented delivery fits a specific operation, contact Afrishore BPO for a no-obligation assessment.

Related reading: IT Outsourcing · AI-Powered Contact Centers · What Is CCaaS? · Business Process Outsourcing · The Onshore-Offshore Hybrid Team Model · Measuring BPO ROI · BPO Contracts and SLAs