The easiest answer can make the next decision harder

A CHI study collected 936 real examples from 319 knowledge workers. Higher confidence in GenAI predicted less reported critical-thinking effort; higher confidence in one’s own ability predicted more. The hidden product variable is therefore not simply “effort saved,” but which effort disappears and what capacity disappears with it.

The CHI study of 319 knowledge workers found that confidence in GenAI predicted lower reported critical-thinking effort, while task self-confidence predicted more. The authors also observed a shift toward verification, response integration, and stewardship. Product design should support these new responsibilities rather than assuming the human remains meaningfully “in the loop” because a submit button exists.

Not all friction is failure

A confusing prompt, repeated data entry, and a wall of caveats are bad friction. Asking a clinician to confirm a contraindication, an analyst to compare two explanations, or a buyer to state a priority can be productive friction. It preserves agency and creates information the model cannot safely infer.

Tools-for-thought research asks how GenAI can augment cognition rather than replace it. That does not require making every interaction Socratic and slow. A few well-placed cognitive actions can change the quality of a decision: predict before reveal, choose criteria before ranking, inspect disagreement before accepting, and explain an override when accountability matters.

Workflow placement changes cognition

The Productive vs. Reflective CHI study emphasizes that different ways of integrating AI affect cognition and motivation. The same model can behave as a substitution engine, a critique partner, a source of alternatives, or a reflective mirror depending on when it enters and what controls surround it.

This timing matters commercially. Early generation can anchor a team on familiar ideas. Late critique may improve quality after the customer has developed a position. In routine work, early completion may be ideal. A product needs a portfolio of interaction patterns matched to task stakes, expertise, and learning goals rather than one promise of effortless assistance.

Measure capability after assistance

Assisted speed and quality are necessary metrics. Add transfer. Can the person handle a related case when the assistant is unavailable? Do they recognize a model error after repeated use? Do novices acquire expert patterns or simply route around understanding? Does a team retain shared knowledge, or does it depend on a private history inside the system?

These effects may take longer than a product A/B test. Use staged experiments: immediate task performance, a delayed near-transfer task, and later behavior in real work. Combine logs with interviews that reveal how people changed their process. The social-science measurement framework helps keep “critical thinking” from becoming another vague score.

Measure what remains after the assistant leaves

Choose a task where learning or judgment has continuing value. Compare direct answer delivery with a reflective design that first elicits a prediction or criteria, then shows the model’s answer and disagreement. Match time budgets fairly. Measure immediate quality, effort, confidence calibration, delayed transfer, and willingness to use the product again.

If reflection improves transfer but harms routine adoption, route it to novices, rare cases, or high-stakes decisions. If it adds effort without benefit, remove it. If direct answers improve both performance and learning, the task may not contain the productive struggle you assumed. Human-centered design does not worship effort. It protects the effort through which humans remain capable.