Users completing everyday assistant tasks across private, public, noisy, hands-busy, and visually complex contexts.
When should a multimodal LLM stay on-device or escalate to the cloud?
Test whether a context- and privacy-aware policy can choose text, voice, vision, on-device inference, or cloud escalation more effectively than a fixed interaction policy.
A context-aware modality and edge/cloud policy will improve first-attempt task completion by at least 10% while reducing unnecessary cloud escalation by at least 20%.
First-attempt task completion without switching modality or manually restarting the task.
+10% first-attempt completion and −20% unnecessary cloud escalation.
Evidence for when on-device and multimodal LLM features improve usefulness enough to justify device and cloud resources.
DECISION RULES
Proceed to a device-level pilot.
Retune the edge/cloud decision policy.
Do not add adaptive modality complexity.
Business outcomes are research targets, not guarantees. A null or negative result may still create substantial value by preventing investment in an ineffective product, feature, or campaign.
A policy that selects interaction modality and edge/cloud execution using task, environment, privacy, and device-state signals.
A fixed default modality and cloud policy for the same tasks.
Cloud escalation rate · Interaction latency · User correction rate · Perceived privacy and control