Frontier LLMs pick Japan 6 of 8 times for culture answers after instruction tuning
A new study finds instruction tuning concentrates cultural outputs on the US and Japan, shrinking global viewpoints.

Researchers from the University of the Basque Country and Cardiff University tested eight frontier LLMs on 31,680 culturally grounded, open-ended questions. The results show LLMs, after instruction tuning, disproportionately reference Japan and the United States, raising risks for cross-cultural deployments.
An April 2026 study says frontier LLMs display a “disproportionate prominence of Japan” when asked culture questions, and it gets specific: on average, six out of the eight evaluated models preferred referencing Japan during exogenous responses. In practice, that means when a model is asked a question in one language without direct country cues, it often “helpfully” picks Japan anyway.
That preference shows up in the exact experiment design, not vibes. The researchers built a multilingual set of 31,680 “open-ended yet culturally grounded questions,” covering 66 cultural subtopics grouped into 11 higher-level domains. Prompts were written in 24 languages, using standardized templates with no explicit regional hints, so the model’s internal priors should drive what countries appear in the answer rather than prompt engineering. Across models, when they answered in the same country implied by the prompt language, that part was predictable: French-language prompts tend to yield France. The more interesting part is the “exogenous” cases, where the model references a different country than the language suggests. There, Japan becomes the repeat winner, with the United States in second, then India, China, and France.
Why executives should care is pretty straightforward: if your product is global, your model is quietly deciding which cultures deserve the spotlight. The study frames this as an uneven regional representation in frontier model outputs, concentrated in a small set of dominant regions. And it does not just observe the end result. It tries to figure out when the bias emerges during training.
To do that, the researchers compared English responses from multiple model families both before and after instruction tuning, including Meta Llama, Qwen, and Gemma, and Mistral models. Before instruction tuning, base models displayed a broader set of cultural associations. Even though “the United States remains prominent,” the researchers note substantial references were also made to Japan, India, China, and several European countries.
Then instruction tuning enters the chat. After instruction tuning, the study reports a sharper preference: across all examined model families, instruction tuning “sharply increases alignment with the United States and Japan while reducing references to most other countries.” The researchers call out an important nuance for decision-makers: this convergence toward culturally dominant regions occurs even in models developed outside Western contexts. In other words, the change is not just “where the model came from.” The post-training step itself is homogenizing cultural perspectives.
The strongest compression effect shows up with supervised fine-tuning (SFT), described as an additional layer where a pre-trained LLM is trained on examples of correct responses, often using human-generated or human-curated datasets. The researchers found that “SFT sharply increases concentration on a small number of dominant regions (most notably the United States and Japan),” and that the effect is “only marginally” mitigated by additional instruction alignment. That detail matters because it suggests the “alignment” process is not merely improving helpfulness. It is injecting cultural bias into the response criteria by selecting which answers count as correct.
From a deployment standpoint, the second-order implication is chillingly simple: instruction-tuned systems may be less culturally diverse at the moment users ask cross-cultural questions, exactly when enterprises and game companies want broad, respectful coverage. The study explicitly warns of “important implications for the deployment of instruction-tuned models in cross-cultural or global applications, where preserving diverse cultural viewpoints may be critical.” If your model is steering conversations, customer support, educational experiences, localization, or even game dialogue toward two dominant cultures, the gap between “AI general intelligence” and “globally balanced representation” starts to look like a product risk, not a research curiosity.
There is also a governance angle. Boards and product leadership will recognize the dynamic: instruction tuning is often treated as a quality upgrade and risk reduction lever. But this work suggests it can also reduce cultural diversity, steering outputs toward limited perspectives. For leaders evaluating model vendors or internal pipelines, the takeaway is that “post-training quality” can come with a representational cost. In a market where frontier models are being rolled into global workflows quickly, the paper is essentially asking: what happens when your training process makes only a narrow slice of the world sound “right”?
This story's Key Insights and Take-aways are locked.
Create a free account to unlock Executive Actions for one credit.
Register to UnlockAlways free for Executives Club members. Join the Club
More in Technology

OpenAI says a rogue AI agent hacked Hugging Face during testing
The ChatGPT maker calls it an “unprecedented incident” after an autonomous agent accessed the open web and attacked Hugging Face.

Kratsios alleges Moonshot distilled Anthropic’s Fable for Kimi K3 development
A White House science official claims covert large-scale distillation, plus access to Nvidia GB300 hardware.

Lego’s $200 Donkey Kong arcade set lets Carl Merriam satisfy Miyamoto, reportedly
A $200 Lego arcade machine delivers a playable mini game and nudges even Mario’s creator toward approval.

