The Study
Three coding agents answered the same 60 devtools questions twice: once in English, once in natural developer Simplified Chinese, the script of mainland China. The agents: Claude Sonnet 5 inside Claude Code, Moonshot's kimi-k3 inside Kimi Code CLI, and Z.ai's GLM-5.2 inside ZCode. Every question ran against its applicable repos (2 to 5 of the 5), in two full runs per agent.
The twelve categories were picked for one property: each has at least one serious China-domestic alternative to the Western default. Object storage has Aliyun OSS and Tencent COS against S3 and R2. Maps has Amap and Baidu against Google and Mapbox. LLM APIs have DeepSeek, Qwen, and the two models' own makers against OpenAI and Anthropic. If a model carries a preference for either stack, these categories give it every chance to show.
The prompts never name a vendor and never mention China or compliance (one hosting prompt says “expanding to Asia”; the rest are geography-free). The only intended difference between the two versions of a prompt is the language; run-to-run noise is measured separately below.
The English Results: One Stack, Three Models
Composition of primary picks per model and language. In English, China-domestic vendors get 0.7% of Claude's picks, 0.9% of Kimi's, and 1.8% of GLM's. The three bottom bars are the Chinese-language cells, where the red segment grows 17x to 28x.
Agreement runs deeper than vendor-group shares. On the same prompt against the same repo, two models pick the same top tool 59 to 63% of the time in English, and the Chinese models agree with Claude as often as they agree with each other. Chinese-language answers agree less across the board, consistent with segmented answers widening the pick space.
One Chinese vendor wins in English as decisively as it does in Chinese. Two database prompts ask for a MySQL-compatible distributed database and for an HTAP database, and on both, every model picked TiDB in every run: 48 of 48 responses across the three agents and both languages. TiDB is Chinese-founded (PingCAP) and it is GLM-5.2's top English database pick at 40%, ahead of Supabase and Neon. It is the sharpest case of the pattern behind the separate Chinese-founded-global class in our classification: where the product has global adoption and fits the question, the language of the question stops mattering.
Sonnet 5 ↔ Kimi K3
Sonnet 5 ↔ GLM-5.2
Kimi K3 ↔ GLM-5.2
Top 10 picks · English (all three models)
Sonnet 5 · English
Anthropic (30), Twilio (22), PostHog (22), Mapbox (19), Vercel (19), Supabase (14), Google Maps (13), GitHub (12), Cloudflare R2 (10), Stripe (10)
Kimi K3 · English
PostHog (26), Twilio (21), Vercel (17), Anthropic (17), Mapbox (12), Gemini (12), GitHub (12), Cloudflare R2 (11), Supabase (10), Neon (10)
GLM-5.2 · English
Twilio (21), PostHog (17), Mapbox (15), Vercel (15), OpenAI (14), Firebase Cloud Messaging (12), Google Maps (12), Gemini (12), GitHub (12), Supabase (10)
Top 10 picks · Chinese
Sonnet 5 · Chinese
Anthropic (30), Twilio (23), Aliyun SMS (22), Google Maps (21), Vercel (19), Amap (16), Tencent Cloud SMS (15), PostHog (15), Cloudflare R2 (14), Supabase (13)
Kimi K3 · Chinese
DeepSeek (24), PostHog (23), Amap (19), Aliyun SMS (18), Vercel (18), Moonshot Kimi (17), Google Maps (16), Cloudflare R2 (16), Tencent Cloud SMS (15), Twilio (15)
GLM-5.2 · Chinese
Amap (24), Aliyun SMS (22), DeepSeek (21), Twilio (19), PostHog (16), Vercel (16), Anthropic (16), Gemini (16), Google Maps (15), Cloudflare R2 (15)
The English leaderboards read like one list written three times: Stripe, Vercel, PostHog, Twilio, Clerk, Mapbox, Cloudflare R2, Meilisearch. The Chinese leaderboards keep most of that list and splice in Aliyun SMS, Amap, DeepSeek, and Alipay beside it.
The Language Flip
Share of answers whose primary picks include at least one China-domestic vendor. The hollow dot is the English rate; the filled dot is the Chinese rate for the same prompts. Read this alongside the pick shares above: Chinese answers name more tools per response (2.1 against 1.5), so a Chinese answer has more chances to include a domestic vendor. Across all 908 Chinese answers, 7% of the one-pick answers name a China-domestic vendor against 78% of the answers with four or more picks. The pick-share figures are the length-controlled version of the same result, and they move just as far: China-domestic goes from under 2% of English picks to 20-31% of Chinese picks.
Pair every English answer with its Chinese twin (same prompt, repo, and run; both answered) and the movement is measurable per model. GLM-5.2 changes its top pick in 60.9% of 289 pairs and adds a China-domestic vendor in 46.7%. Kimi changes its top pick in 47.2%. Claude changes in 30.4%, which is the lowest of the three and still close to a third of all pairs.
The overlap between an answer and its translation (Jaccard similarity of the primary pick sets) tells the same story from the other side: Claude 0.60, Kimi 0.50, GLM 0.43. No model treats translation as a no-op.
| Model | Prompt (kotlin-android-app, run 1) | English pick | Chinese pick |
|---|---|---|---|
| Sonnet 5 | smp-07 | Twilio | Aliyun SMS |
| Sonnet 5 | smp-08 | Firebase Auth | Aliyun SMS |
| Sonnet 5 | map-06 | Mapbox | Google Maps |
| Kimi K3 | smp-01 | Firebase Auth | Tencent Cloud SMS |
| Kimi K3 | smp-08 | Firebase Auth | Aliyun SMS |
| Kimi K3 | smp-09 | Twilio | Aliyun SMS |
| GLM-5.2 | smp-01 | Twilio | Aliyun SMS |
| GLM-5.2 | smp-07 | Twilio | Aliyun SMS |
| GLM-5.2 | smp-08 | Firebase Auth | Aliyun SMS |
Sample of top-pick changes on the Android repo. The full pair-level data is in the study dataset.
Claude in Chinese: The Same Warnings, Half the Domestic Picks
Claude converges with the Chinese models only partway, and where it stops is the informative part. Claude's Chinese answers move in the same direction as Kimi's and GLM's, about half as far: it names a Chinese vendor in 23.9% of Chinese answers against Kimi's 32.5% and GLM's 49.3%, and it branches its advice by market in 28.2% of Chinese answers against their ~58%.
Where the Western default physically breaks in China, Claude converges with the Chinese models almost exactly. Its Chinese SMS picks are 54% China-domestic (theirs: 55% and 56%). Its Chinese payments picks are 36% domestic (33%, 39%). The statements that FCM does not reach Chinese Android devices and that Chinese consumers pay with Alipay appear in all three models' answers.
Where the Western tool still works, Claude re-orders less. On Chinese maps prompts, Kimi and GLM promote Amap to the top pick (63% and 80%); Claude keeps Google Maps first (70%) and adds Amap beside it (53%). And in the LLM API category the gap becomes absolute: Kimi and GLM answer the Chinese question with DeepSeek at 80% and 72%, while Claude's Chinese-language LLM picks are 0% Chinese. It answers Anthropic, every time, in both languages.
Claude does not localize the way the Chinese models do. Its answers cite the same failures in China, it holds its Western defaults longer where those tools still work, and on the LLM API question it does not move at all.
Self-Preference and the DeepSeek Consensus
The LLM API category asks each agent a question about its own market. These are claims about the shipped agent products, not the bare models: each ran inside its maker's CLI, whose system prompt we cannot see, and Claude Code's answers read as already Claude-conditioned. Rate at which each names its own maker as a primary pick:
Claude Code named Anthropic in all 60 of its answers: 30 English and 30 Chinese responses, two runs, five repos, without an exception. Kimi and GLM named their own makers zero times in English. Asked in English which LLM API to build on, Kimi's most frequent answer was Claude (57%), and GLM split between OpenAI (54%) and Gemini (46%).
In Chinese, both models do promote themselves, to roughly half of answers (57% and 48%). But the vendor they agree on is DeepSeek: 80% of Kimi's Chinese LLM answers and 72% of GLM's. Each also names its direct competitor in about half its Chinese answers (Kimi cites Zhipu in 53%, GLM cites Moonshot in 52%). Each rival lab recommends the neutral third party more often than itself.
Kimi's Chinese reasoning, translated and abridged: “Recommendation: DeepSeek or Kimi (Moonshot)… directly reachable and payable inside China, no network or foreign-payment hurdles… Claude / OpenAI: model quality is genuinely good and the official TypeScript SDKs are mature, but direct access and payment from China are a hassle.” The recommendation logic is about payment rails and network reachability, and the models apply it to themselves as much as to anyone.
All Twelve Categories
China-domestic share of each model's Chinese-language picks, sorted. The gradient tracks how broken the Western default is in mainland China, from SMS at the top to search at zero.
SMS & Push Notifications
Chinese contenders: Aliyun SMS, Tencent Cloud SMS, JPush, Getui
Asked in English
Asked in Chinese
Payments
Chinese contenders: Alipay, WeChat Pay, UnionPay
Asked in English
Asked in Chinese
Maps & Geo
Chinese contenders: Amap (Gaode), Baidu Maps, Tencent Maps
Asked in English
Asked in Chinese
Object Storage & CDN
Chinese contenders: Aliyun OSS, Tencent COS, Qiniu, Upyun
Asked in English
Asked in Chinese
Git Hosting, CI & Registries
Chinese contenders: Gitee, Coding.net, npmmirror
Asked in English
Asked in Chinese
Realtime & RTC
Chinese contenders: Tencent TRTC, ZEGO; Agora is Chinese-founded
Asked in English
Asked in Chinese
LLM APIs
Chinese contenders: DeepSeek, Moonshot Kimi, Zhipu GLM, Qwen, Doubao
Asked in English
Asked in Chinese
Auth & Identity
Chinese contenders: Authing, WeChat login flows
Asked in English
Asked in Chinese
Product Analytics
Chinese contenders: Sensors Data, GrowingIO, Umeng
Asked in English
Asked in Chinese
Managed & Distributed Databases
Chinese contenders: PolarDB, OceanBase, TDSQL; TiDB is Chinese-founded
Asked in English
Asked in Chinese
Cloud Hosting
Chinese contenders: Alibaba Cloud, Tencent Cloud, Huawei Cloud
Asked in English
Asked in Chinese
Search
Chinese contenders: Aliyun OpenSearch; jieba/IK segmentation libraries
Asked in English
Asked in Chinese
Reading notes. Percentages are share of responses naming the tool as a primary pick; segmented answers carry co-primary picks for each market branch, so Chinese-cell rows can sum past 100%. Cells are n=10 to 30 responses, so single-cell percentages are rough: at n=10 to 30 a 95% interval spans roughly ±20 to 30 points. The cross-category and cross-model patterns are the reliable signal, individual cell values less so.
Answer Style: Stability, Refusals, Branching
Run-to-run stability
Same top pick, run 1 vs run 2: Claude 79.3%, Kimi 70.4%, GLM 65.1%. Claude is the most consistent recommender of the three.
Non-answers
Clarifying questions with no pick: GLM 5.5% of English answers (its ask-before-recommending habit), everyone else 0 to 2.9%. Kimi answered every English prompt.
Market branching
“Domestic vs overseas” branches in Chinese answers: GLM 58.9%, Kimi 57.7%, Claude 28.2%. In English all three sit under 5%.
The branching number is the mechanism behind most of this study's findings. Addressed in Chinese, the models rarely drop the Western tools from the pick set; they add a second answer for a second audience, and which audience gets listed first is where the models differ. None of the prompts asked for a market comparison.
Supabase is the clearest illustration, because it moves the wrong way. Every model names it in more Chinese answers than English ones (Kimi 2 responses against 7, GLM 5 against 11). For Claude the share of picks is flat, 21% to 19%, so its rise in response count is answer length alone. For Kimi and GLM the share rises too, 7% to 18% and 15% to 23%. The Chinese answers that produce this are the branching ones: asked in Chinese which Postgres host to use in production, GLM returns a two-branch table of overseas and domestic options averaging 6.75 primary picks per response, its densest answers in the study, and Supabase is on the overseas side of it.
The Distillation Question: No Claude-Specific Signature Found
A reasonable reader will ask whether the English-side convergence means the Chinese models trained on Claude's outputs. Our design cannot answer that, and we want to be precise about what the data supports.
What we observe: Kimi matches Claude's top pick in 62.7% of English prompts, GLM in 59.0%, and the two Chinese models match each other in 62.1%. The symmetry is the point. Kimi's agreement with Claude is no higher than its agreement with GLM. We found no Claude-specific recommendation signature in these tests. The pattern is compatible with several training, post-training, harness, and corpus explanations that this design cannot distinguish.
A sharper cut of the same data asks who pairs off when the three disagree. On the 290 English prompts all three models answered, all three pick the same tool 47% of the time. When exactly two agree against the third, the pair is Claude+Kimi in 16.2% of prompts, Kimi+GLM in 14.5%, and Claude+GLM in 11.7%. Claude+Kimi leads, but a spread that size across three pairs is within sampling noise at this n (chi-square p ≈ 0.35). In Chinese the pairing counts are Kimi+GLM 18.7%, Claude+Kimi 17.0%, and Claude+GLM 11.9% (p ≈ 0.10 before accounting for clustering, so we treat it as no finding). Neither language shows a lopsided Claude-pairing.
What the Chinese-side depth does and does not show: GLM surfaces second-tier domestic vendors (Sensors Data, Qiniu, JPush, Authing) that Claude names rarely or never, which is hard to square with the strongest claim, that these models are Claude end to end. It does not speak to their English side, since native Chinese data and Claude-derived English data could coexist in one model. The English-side evidence is the agreement symmetry and pairing analysis above and the divergence test below. Their answer styles also differ from Claude's in stable ways (double the market branching, lower run-to-run stability, zero English self-promotion against Claude's 100%), though style is the weakest of these signals: post-training could reshape it either way.
We then ran that stronger test. Our March 2026 Codex study gives per-prompt reference picks for Claude (Opus 4.6) and Codex (GPT-5.3) on a different 12-category suite; 48 repo-prompt pairs have stable reference picks (the same tool in at least 2 of 3 runs) that disagree between the two agents. We ran Kimi K3 and GLM-5.2 on all 48 (one run, July 2026), extracted their picks the same way, and scored which side each landed on.
Kimi matched the Claude-side pick 22 times, the Codex-side pick 18 times, and neither on the remaining 8 (55% Claude-side of 40 decided, binomial p = 0.64). GLM matched Claude 11 times and Codex 14 (44% Claude-side, p = 0.69), and picked something neither reference chose on 21 of its 46 scored prompts, nearly three times Kimi's rate. We did not detect a Claude-side preference in either model. The marker tools tell the same story from both directions: Bun, which March-Claude picked in 63% of runtime prompts and Codex in 2%, shows up in 3 to 4 of the Chinese models' 10 runtime answers; Statsig, which Codex picked 27% of the time and Claude never picked once, shows up in 17 to 24% of their feature-flag answers. Each Chinese model carries markers from both Western agents.
Caveats on this test: the reference picks are four months older than the new runs (model drift blurs the signal in both directions), and the new side is a single run. Neither model showed a statistically detectable Claude-side preference in this sample: the 95% intervals on the Claude-side share are roughly 40-69% for Kimi and 27-63% for GLM among decided cases. The estimates are imprecise, do not rule out smaller preferences, and do not identify training provenance. The chi-square figure tests only whether the three exclusive-pair frequencies are equal.
Methodology
Suite design.
120 prompts: 60 English, each with a twin translated by us into natural developer Simplified Chinese (the script used in mainland China and Singapore; Taiwan and Hong Kong write Traditional Chinese) (register-checked, standard loanwords like CDN and API kept in English, no literal word-for-word renderings). Prompts are phrased as everyday asks (“i need to send sms verification codes. what service should i use”) and never name vendors, countries, or compliance regimes. Categories: SMS & Push Notifications, Payments, Maps & Geo, Object Storage & CDN, Git Hosting, CI & Registries, Realtime & RTC, LLM APIs, Auth & Identity, Product Analytics, Managed & Distributed Databases, Cloud Hosting, Search. Repos: nextjs-saas (Next.js SaaS), python-api (FastAPI service), react-spa (React + Vite SPA), node-cli (Node/TS CLI), kotlin-android-app (Kotlin Android). Prompts apply to 2 to 5 repos each depending on category fit.
Collection.
Each model ran inside its maker's own agent CLI in fully autonomous mode, two complete runs per agent, on July 25, 2026. Every prompt ran in a fresh agent session against a repository reset to the benchmark baseline; prompts ran in fixed order, and clearing session and repository state between prompts reduced cross-prompt carryover. Agents ran with isolated per-run profiles, per-prompt session wipes, anonymized working directories, and benchmark-identifying environment variables stripped, following the same leak-hardening protocol as our other agent studies. All API traffic originated from US IP addresses, so whatever the serving side could infer from the connection pointed to the United States, and the intended experimental difference within each matched pair was the prompt language. Agent harness and model are confounded by design; this measures each vendor's shipped experience, not the bare model. Web-tool availability differed (Claude Code and ZCode have web tools; our Kimi profile had none). Kimi ran its default thinking effort (max). 1,860 of 1,860 responses succeeded.
Extraction.
Primary picks and alternatives were extracted by reading every response in full under a written convention set (a pick is what the response recommends, not the first tool named; defaults offered inside pure clarifying questions are non-answers; answers that branch by market yield co-primary picks per branch lead; drivers, SDKs, and the repo's own framework are never picks). Chinese responses were read in Chinese and mapped to the same canonical vendor slugs as English responses via a bilingual alias table (阿里云 → Alibaba Cloud, 高德 → Amap, 支付宝 → Alipay). Extraction coverage was validated 1:1 against the source responses; a 13-row stratified sample was re-read against the raw responses during extraction; and two independent reviewers later audited every headline number against the analysis files. As a further check we blind double-coded a 60-row stratified sample (10 per model-language cell) with two independent coders working from the raw responses and the written conventions only. The coders matched each other's primary pick exactly on 59 of 60 rows, matched the shipped extraction on 59 and 60 of 60, and agreed with it on non-answer status on all 60; the single disagreement (Clerk vs Auth0 on one auth prompt) was flagged as ambiguous by both coders independently.
Classification.
“China-domestic” covers vendors whose primary market is mainland China (Aliyun OSS, Amap, Alipay, Sensors Data, Gitee). Chinese-founded companies with real global adoption (TiDB, Agora, DeepSeek, jieba) are a separate class and behave like Western vendors in our data: their pick rates barely move between languages. Response rates count a response once if any primary pick is China-domestic; pick shares divide by all primary picks. Both are reported because segmented answers give Chinese cells more picks per response (2.1 vs 1.5).
What would change these numbers.
One suite, five repos, one point in time (kimi-k3 was nine days old when we ran it). Category cells are small (n=10 to 30 per model per language); aggregate findings rest on n=289 to 734, and those observations cluster by base question, repository, and run, so the quoted p-values treat the data as more independent than the design strictly warrants. We therefore treat the quoted p-values as descriptive, not confirmatory; they are computed in the checked-in analysis scripts. Our translations are natural but not native-authored. Different repos, prompt phrasings, or CLI versions would move individual percentages; we would not expect them to reverse the ordering of the three headline gaps (the English convergence, the graded language flip, and the self-preference asymmetry), though that expectation is exactly what re-running would test.
The prompts, raw responses, extraction tables, vendor classification, and analysis code are checked into our benchmark repository and available on request; a public dataset release is planned. Related studies: What Claude Code Actually Chooses, What Codex Actually Chooses, What Fable Actually Chooses.
What This Study Cannot Tell You
Whether the models' behavior reflects policy, training data, or post-training choices. We measure outputs, not intent.
Whether Chinese developers actually receive these answers. Real usage mixes languages, system prompts, and follow-up turns that a benchmark does not.
Whether the flip generalizes to other languages. Japanese, Korean, or German prompts might segment too; we only tested the en/zh pair.
Market share. These are model recommendations on five specific repos, not adoption data.