Why the support-agent study matters

Across 5,179 customer-support agents, AI assistance increased issues resolved per hour by 14 percent on average and 34 percent for novice and lower-skilled workers, with minimal effects on the most experienced. The system appeared to diffuse patterns associated with stronger workers. That is a richer mechanism than “the model writes replies faster.”

The assistant sat at a point where a human had a customer, a live problem, organizational knowledge, and authority to act. Its answer could immediately change what happened next. Many enterprise copilots lack that clarity. They produce plausible text in a side panel, leaving the employee to discover whether, where, and how to use it.

A workflow has social and temporal structure

Microsoft’s randomized field experiment with 6,000 workers found that GenAI changed activities individuals could alter independently, such as email time, more readily than coordinated work such as meetings. This is a useful limit. An answer can accelerate one person while the organization’s approvals, incentives, handoffs, and shared practices remain fixed.

The human-centered unit is therefore not “user plus chatbot.” It is the work system. Who owns the decision? What evidence travels across the handoff? What must be recorded? Which expert is interrupted by exceptions? Where does a customer wait? Answer design includes the artifacts and coordination needed after generation.

Do not automate away the learning loop

Large gains for novice workers sound unequivocally positive. Yet if the assistant supplies the answer without exposing diagnosis, evidence, or feedback, the novice may become productive without becoming expert. The critical-thinking study suggests that GenAI shifts work toward verification, integration, and stewardship. Those skills need support inside the product.

Show why a recommendation applies, ask for a small judgment where it matters, and return outcome feedback. Let agents see when experts override the model and why. A system can distribute expert patterns while still making expertise legible. The alternative is an organization that becomes dependent on advice nobody can audit or improve.

Bounded writing offers a second clue

In a preregistered experiment with 453 professionals, ChatGPT reduced time on bounded writing tasks by 40 percent while raising evaluator-rated quality by 18 percent. The result is not a forecast for customer support, but it isolates another mechanism: the assistant absorbed low-value drafting effort while leaving a person to integrate and own the output.

The transferable design question is where expert judgment should enter. A system can suggest language, surface precedent, flag an exception, or recommend an action. These are different interventions with different risks. The commercial answer is not a longer response; it is the right judgment made available at the point where a person can use and contest it.

Look for learning, not only output

Choose one support intent with meaningful volume and variation. Randomize eligible cases between the current knowledge workflow and an assistant that presents a recommended response, supporting evidence, an uncertainty or exception check, and a structured escalation. Measure resolution per hour, repeat contact, customer outcome, unsafe acceptance, escalation quality, and agent learning on later unaided cases.

Segment by experience. If novices gain speed but not later unaided competence, add learning scaffolds. If experts slow down, let them collapse explanations and contribute corrections. If throughput rises while repeat contacts rise, the system may be producing fast closure rather than resolution. Judgment is valuable only when it survives contact with the next customer.