The same sentence does not perform the same social action

A study comparing model responses with World Values Survey data across Danish, Dutch, English, and Portuguese found no consistent relationship between language capability and cultural alignment. Gemma models from 2B to 27B parameters showed one pattern; OpenAI’s turbo-series did not. Fluency was not a reliable proxy for cultural fit.

This is especially visible in a city such as Haifa, where Hebrew, Arabic, Russian, English, professional jargon, religious traditions, and different institutional experiences can meet inside one service. “Localize the copy” is too small a brief. The product must learn which facts, roles, examples, cautions, and forms of respect make an answer usable in a particular context.

Multilingual capability can conceal cultural bias

Research comparing multilingual competence with cultural alignment finds that models can write fluently while remaining more aligned with US value distributions than with local ones. This is not solved by switching the response language. The model may import assumptions about autonomy, disclosure, risk, or customer service that do not travel intact.

The commercial consequence is not only offense. A recommendation may target the wrong decision-maker, omit family or community considerations, misread trust in institutions, or use examples that fail to signal relevance. Customers then disengage from an answer that appears technically tailored while remaining culturally generic.

Culture is hard to turn into a score

Work on the unreliability of cultural-alignment evaluation warns against reducing complex values to decontextualized binary choices. A benchmark can create the appearance of precision while changing the construct it claims to measure. Culture is internally diverse, situational, contested, and dynamic. Country and language are poor substitutes for lived context.

A separate benchmark makes the product risk visible: across 1,284 multimodal queries from 16 countries and 14 languages, the best tested model reached 61.79 percent cultural awareness but only 37.73 percent compliance. Anthropology, sociology, history, and literary interpretation can reveal why an answer fails, for whom, and under which relationship; quantitative testing can then measure a specified intervention rather than a vague idea of cultural correctness.

From translated output to culturally grounded answer

Start with the decision journey in each target community. Who asks the question? Who is affected but absent? What constitutes credible evidence? Which examples are familiar? What must be said directly, and what requires contextual explanation? Which terms carry institutional or political history? The answers may justify different structure and escalation, not just different vocabulary.

Avoid personas built from stereotypes. Ground variants in observed needs and let users correct assumptions. Make provenance and uncertainty visible when local evidence is thin. Cultural relevance should increase a customer’s ability to understand and act without narrowing them to an imagined representative of a group.

Let the market change the answer, not only the language

Take one consequential answer in two or three language communities. Create a direct translation and a locally co-designed version based on interviews with customers and frontline staff. Keep policy and factual content equivalent. Measure task comprehension, perceived respect, willingness to continue, chosen action, and ability to identify exceptions or seek help.

Analyze variation within each language group rather than reporting a single national effect. If local adaptation helps one segment and harms another, expose a user choice or use task context instead of demographic inference. The business win is not a “culturally aligned” badge. It is fewer customers forced to translate the product back into their own world.