All projects
seeking partner

When does multimodal output justify its compute cost?

Measure whether image, voice, or animation improves successful outcomes enough to pay for additional latency and compute.

ProductValue & PricingAdoption & Retention
Chips, Cloud & AI Infrastructure
WHAT YOU CAN EXPECT

Evidence protects you in both directions.

01
Illustrative result · threshold met

−31% · cost per successful task

Product managerIllustrative result · threshold met
−31%

cost per successful task

Using generated visuals only for high-complexity steps improved correct completion enough to offset the added inference cost.

What your study would pin down

You get a modality policy by task: where image, voice, or animation changes completion enough to pay for itself, and where text should remain the default.

Decision

Use multimodality selectively instead of paying for it on every response.

Public precedent · not your resultLysol

GenAI creative cut cost per asset 80% and reached 10× production speed with near-parity in sales lift.

Lysol reports its best AI-generated asset delivered short-term sales-lift effectiveness nearly identical to its top traditional asset while cutting production cost per asset by 80%.

Source: Think with Google · 2026 ↗The precedent shows the effect can exist elsewhere. Your study determines whether, where, and how strongly it holds for your customers, workflows, and economics.
02
Decision

Evidence protects you in both directions.

Product managerThreshold not met · do not scale
What the data may show

No material benefit

If the intervention does not clear the predefined threshold, that is evidence against spending more to build, launch, or scale it in this context.

Decision

Avoid the compute investment.

Product managerThreshold met · value left unused
The other expensive error

A real opportunity can still be left on the table.

If the intervention clears the threshold but the business keeps the current approach, measurable savings, revenue, adoption, or risk reduction may remain unrealized.

Decision

Act only when the measured opportunity is large enough to justify the change.

03
Source

Public precedent · not your result

Comparable public caseMcDonald’s

AI voice ordering reached 100+ restaurants — then the IBM trial was ended.

McDonald’s ended the automated drive-through ordering test after mixed success and repeated complaints about order accuracy and interpretation. The company continued evaluating other voice-AI approaches.

Source: Associated Press ↗Comparable public case — not a claim that this study would have prevented the event.
Related published evidenceFactory workers + radiologists

Visual explanations improved real-world task performance by +7.7 pp in manufacturing and +4.7 pp in medicine.

Two preregistered experiments with domain experts found explainable visual heatmaps improved performance over black-box AI in both settings. Skipping the right modality can also leave measurable performance gains unused.

Source: Scientific Reports · 2024 ↗External research for context — not a promise that the same effect size will reproduce in your customers.
HYPOTHESIS

A targeted visual explanation will reduce time-to-correct-action by at least 20% while increasing inference cost by less than 12%.

PRIMARY METRIC

Compute cost per correctly completed task.

MEANINGFUL THRESHOLD

20% faster correct action at less than 12% additional inference cost.

BUSINESS TARGET

A positive unit-economic case for a premium multimodal feature.

DECISION RULES

Unit economics positive

Launch a limited premium pilot.

Comprehension rises but economics fail

Reduce modality frequency or complexity.

No material benefit

Avoid the compute investment.

Business outcomes are research targets, not guarantees. A null or negative result may still create substantial value by preventing investment in an ineffective product, feature, or campaign.

Population

Customers completing explanation-heavy tasks on desktop and mobile.

Intervention

Text plus a generated visual explanation selected only for high-complexity steps.

Comparator

Text-only response from the same model.

Secondary metrics

Latency tolerance · Comprehension · Preference · Repeat use

MAKE IT YOUR DECISION

Turn this research question into a decision for your business.

We adapt the population, intervention, thresholds, and economics to your customers. The result may tell you to scale, to stop spending, or to act on an opportunity you are currently leaving unused. Each of those is a useful business decision when the evidence is strong enough.

The goal is not a positive result. The goal is evidence strong enough to change a real decision.

Design this study for your business

What customer or commercial decision should the evidence strengthen?

You know the opportunity. Tell us the decision and choose the outcomes that would make the result useful.

Your role
What should the study help you measure?
Draft the conversation

Your email app opens with the draft; this site stores nothing.