Amplifying/agent-intelligence

Research · Full Report

What Kimi K3 & GLM-5.2 Actually ChooseChina's AI models vs Claude Sonnet 5

Edwin Ong & Alex Vikati · July 2026 · Overview · 简体中文版

This is the full data companion to the study: all 12 categories with per-cell pick tables, the cross-model and cross-language comparisons, the answer-style numbers, and the complete methodology. 1,860 responses, collected July 25, 2026, with zero runtime failures (42 responses gave no recommendation and are counted as non-answers).

The Study

Three coding agents answered the same 60 devtools questions twice: once in English, once in natural developer Simplified Chinese, the script of mainland China. The agents: Claude Sonnet 5 inside Claude Code, Moonshot's kimi-k3 inside Kimi Code CLI, and Z.ai's GLM-5.2 inside ZCode. Every question ran against its applicable repos (2 to 5 of the 5), in two full runs per agent.

The twelve categories were picked for one property: each has at least one serious China-domestic alternative to the Western default. Object storage has Aliyun OSS and Tencent COS against S3 and R2. Maps has Amap and Baidu against Google and Mapbox. LLM APIs have DeepSeek, Qwen, and the two models' own makers against OpenAI and Anthropic. If a model carries a preference for either stack, these categories give it every chance to show.

The prompts never name a vendor and never mention China or compliance (one hosting prompt says “expanding to Asia”; the rest are geography-free). The only intended difference between the two versions of a prompt is the language; run-to-run noise is measured separately below.

The English Results: One Stack, Three Models

Composition of primary picks per model and language. In English, China-domestic vendors get 0.7% of Claude's picks, 0.9% of Kimi's, and 1.8% of GLM's. The three bottom bars are the Chinese-language cells, where the red segment grows 17x to 28x.

Sonnet 5 · English
0.7%
Kimi K3 · English
0.9%
GLM-5.2 · English
1.8%
Sonnet 5 · Chinese
19.8%
Kimi K3 · Chinese
25.8%
GLM-5.2 · Chinese
30.5%
Western + DIY Chinese-founded global (TiDB, Agora, DeepSeek) China-domesticright column: China-domestic share

Agreement runs deeper than vendor-group shares. On the same prompt against the same repo, two models pick the same top tool 59 to 63% of the time in English, and the Chinese models agree with Claude as often as they agree with each other. Chinese-language answers agree less across the board, consistent with segmented answers widening the pick space.

One Chinese vendor wins in English as decisively as it does in Chinese. Two database prompts ask for a MySQL-compatible distributed database and for an HTAP database, and on both, every model picked TiDB in every run: 48 of 48 responses across the three agents and both languages. TiDB is Chinese-founded (PingCAP) and it is GLM-5.2's top English database pick at 40%, ahead of Supabase and Neon. It is the sharpest case of the pattern behind the separate Chinese-founded-global class in our classification: where the product has global adoption and fits the question, the language of the question stops mattering.

Sonnet 5 ↔ Kimi K3

English
62.7%
Chinese
48.8%

Sonnet 5 ↔ GLM-5.2

English
59%
Chinese
44.6%

Kimi K3 ↔ GLM-5.2

English
62.1%
Chinese
51.9%

Top 10 picks · English (all three models)

Sonnet 5 · English

Anthropic (30), Twilio (22), PostHog (22), Mapbox (19), Vercel (19), Supabase (14), Google Maps (13), GitHub (12), Cloudflare R2 (10), Stripe (10)

Kimi K3 · English

PostHog (26), Twilio (21), Vercel (17), Anthropic (17), Mapbox (12), Gemini (12), GitHub (12), Cloudflare R2 (11), Supabase (10), Neon (10)

GLM-5.2 · English

Twilio (21), PostHog (17), Mapbox (15), Vercel (15), OpenAI (14), Firebase Cloud Messaging (12), Google Maps (12), Gemini (12), GitHub (12), Supabase (10)

Top 10 picks · Chinese

Sonnet 5 · Chinese

Anthropic (30), Twilio (23), Aliyun SMS (22), Google Maps (21), Vercel (19), Amap (16), Tencent Cloud SMS (15), PostHog (15), Cloudflare R2 (14), Supabase (13)

Kimi K3 · Chinese

DeepSeek (24), PostHog (23), Amap (19), Aliyun SMS (18), Vercel (18), Moonshot Kimi (17), Google Maps (16), Cloudflare R2 (16), Tencent Cloud SMS (15), Twilio (15)

GLM-5.2 · Chinese

Amap (24), Aliyun SMS (22), DeepSeek (21), Twilio (19), PostHog (16), Vercel (16), Anthropic (16), Gemini (16), Google Maps (15), Cloudflare R2 (15)

The English leaderboards read like one list written three times: Stripe, Vercel, PostHog, Twilio, Clerk, Mapbox, Cloudflare R2, Meilisearch. The Chinese leaderboards keep most of that list and splice in Aliyun SMS, Amap, DeepSeek, and Alipay beside it.

The Language Flip

Share of answers whose primary picks include at least one China-domestic vendor. The hollow dot is the English rate; the filled dot is the Chinese rate for the same prompts. Read this alongside the pick shares above: Chinese answers name more tools per response (2.1 against 1.5), so a Chinese answer has more chances to include a domestic vendor. Across all 908 Chinese answers, 7% of the one-pick answers name a China-domestic vendor against 78% of the answers with four or more picks. The pick-share figures are the length-controlled version of the same result, and they move just as far: China-domestic goes from under 2% of English picks to 20-31% of Chinese picks.

0%20%40%60%Sonnet 51%23.9%Kimi K31%32.5%GLM-5.22.7%49.3%
asked in Englishasked in Chinese

Pair every English answer with its Chinese twin (same prompt, repo, and run; both answered) and the movement is measurable per model. GLM-5.2 changes its top pick in 60.9% of 289 pairs and adds a China-domestic vendor in 46.7%. Kimi changes its top pick in 47.2%. Claude changes in 30.4%, which is the lowest of the three and still close to a third of all pairs.

The overlap between an answer and its translation (Jaccard similarity of the primary pick sets) tells the same story from the other side: Claude 0.60, Kimi 0.50, GLM 0.43. No model treats translation as a no-op.

ModelPrompt (kotlin-android-app, run 1)English pickChinese pick
Sonnet 5smp-07TwilioAliyun SMS
Sonnet 5smp-08Firebase AuthAliyun SMS
Sonnet 5map-06MapboxGoogle Maps
Kimi K3smp-01Firebase AuthTencent Cloud SMS
Kimi K3smp-08Firebase AuthAliyun SMS
Kimi K3smp-09TwilioAliyun SMS
GLM-5.2smp-01TwilioAliyun SMS
GLM-5.2smp-07TwilioAliyun SMS
GLM-5.2smp-08Firebase AuthAliyun SMS

Sample of top-pick changes on the Android repo. The full pair-level data is in the study dataset.

Claude in Chinese: The Same Warnings, Half the Domestic Picks

Claude converges with the Chinese models only partway, and where it stops is the informative part. Claude's Chinese answers move in the same direction as Kimi's and GLM's, about half as far: it names a Chinese vendor in 23.9% of Chinese answers against Kimi's 32.5% and GLM's 49.3%, and it branches its advice by market in 28.2% of Chinese answers against their ~58%.

Where the Western default physically breaks in China, Claude converges with the Chinese models almost exactly. Its Chinese SMS picks are 54% China-domestic (theirs: 55% and 56%). Its Chinese payments picks are 36% domestic (33%, 39%). The statements that FCM does not reach Chinese Android devices and that Chinese consumers pay with Alipay appear in all three models' answers.

Where the Western tool still works, Claude re-orders less. On Chinese maps prompts, Kimi and GLM promote Amap to the top pick (63% and 80%); Claude keeps Google Maps first (70%) and adds Amap beside it (53%). And in the LLM API category the gap becomes absolute: Kimi and GLM answer the Chinese question with DeepSeek at 80% and 72%, while Claude's Chinese-language LLM picks are 0% Chinese. It answers Anthropic, every time, in both languages.

Claude does not localize the way the Chinese models do. Its answers cite the same failures in China, it holds its Western defaults longer where those tools still work, and on the LLM API question it does not move at all.

Self-Preference and the DeepSeek Consensus

The LLM API category asks each agent a question about its own market. These are claims about the shipped agent products, not the bare models: each ran inside its maker's CLI, whose system prompt we cannot see, and Claude Code's answers read as already Claude-conditioned. Rate at which each names its own maker as a primary pick:

100%100%Sonnet 50%56.7%Kimi K30%48.3%GLM-5.2
asked in English asked in Chinese

Claude Code named Anthropic in all 60 of its answers: 30 English and 30 Chinese responses, two runs, five repos, without an exception. Kimi and GLM named their own makers zero times in English. Asked in English which LLM API to build on, Kimi's most frequent answer was Claude (57%), and GLM split between OpenAI (54%) and Gemini (46%).

In Chinese, both models do promote themselves, to roughly half of answers (57% and 48%). But the vendor they agree on is DeepSeek: 80% of Kimi's Chinese LLM answers and 72% of GLM's. Each also names its direct competitor in about half its Chinese answers (Kimi cites Zhipu in 53%, GLM cites Moonshot in 52%). Each rival lab recommends the neutral third party more often than itself.

Kimi's Chinese reasoning, translated and abridged: “Recommendation: DeepSeek or Kimi (Moonshot)… directly reachable and payable inside China, no network or foreign-payment hurdles… Claude / OpenAI: model quality is genuinely good and the official TypeScript SDKs are mature, but direct access and payment from China are a hassle.” The recommendation logic is about payment rails and network reachability, and the models apply it to themselves as much as to anyone.

All Twelve Categories

China-domestic share of each model's Chinese-language picks, sorted. The gradient tracks how broken the Western default is in mainland China, from SMS at the top to search at zero.

0%15%30%45%60%SMS & Push Notifications54%55%56%Payments36%33%39%Maps & Geo30%29%42%Object Storage & CDN16%30%31%Git Hosting, CI & Registries23%24%26%Realtime & RTC18%24%29%LLM APIs0%36%35%Auth & Identity11%25%22%Product Analytics8%9%28%Managed & Distributed Databases5%10%19%Cloud Hosting12%7%14%Search0%0%0%
Sonnet 5Kimi K3GLM-5.2

SMS & Push Notifications

Chinese contenders: Aliyun SMS, Tencent Cloud SMS, JPush, Getui

Asked in English

Sonnet 5n=30
Twilio
73%
Firebase Cloud Messaging
20%
Firebase Auth
7%
Plivo
7%
Vonage
3%
Kimi K3n=30
Twilio
67%
Firebase Cloud Messaging
20%
Firebase Auth
13%
OneSignal
3%
Supabase Auth
3%
GLM-5.2n=29
Twilio
72%
Firebase Cloud Messaging
21%
Firebase Auth
10%
Clerk
3%

Asked in Chinese

Sonnet 5n=29 · CN 54%
Twilio
79%
Aliyun SMS
76%
Tencent Cloud SMS
48%
Firebase Cloud Messaging
21%
Firebase Auth
3%
Kimi K3n=30 · CN 55%
Aliyun SMS
53%
Tencent Cloud SMS
50%
Twilio
47%
Firebase Cloud Messaging
20%
Firebase Auth
13%
GLM-5.2n=29 · CN 56%
Aliyun SMS
76%
Twilio
66%
Firebase Cloud Messaging
21%
Tencent Cloud SMS
21%
JPush
21%

Payments

Chinese contenders: Alipay, WeChat Pay, UnionPay

Asked in English

Sonnet 5n=10
Stripe
100%
Apple Pay
20%
Google Pay
20%
Kimi K3n=10
Stripe
100%
Apple Pay
20%
Google Pay
20%
Stripe Link
10%
GLM-5.2n=10
Stripe
100%
Apple Pay
40%
Google Pay
40%
Paddle
10%
Lemon Squeezy
10%

Asked in Chinese

Sonnet 5n=10 · CN 36%
Stripe
100%
WeChat Pay
40%
Alipay
40%
Apple Pay
20%
Google Pay
20%
Kimi K3n=10 · CN 33%
Stripe
100%
Apple Pay
20%
Google Pay
20%
Baiwang Cloud
20%
Nuonuo
20%
GLM-5.2n=10 · CN 39%
Stripe
90%
Alipay
60%
WeChat Pay
50%
Apple Pay
20%
Google Pay
20%

Maps & Geo

Chinese contenders: Amap (Gaode), Baidu Maps, Tencent Maps

Asked in English

Sonnet 5n=30
Mapbox
63%
Google Maps
43%
Nominatim
17%
Pelias
13%
reverse_geocoder
10%
Kimi K3n=30
Mapbox
40%
Google Maps
30%
OpenStreetMap
17%
MapLibre
17%
Nominatim
13%
GLM-5.2n=30
Mapbox
50%
Google Maps
40%
MapLibre
23%
Pelias
17%
Apple MapKit
10%

Asked in Chinese

Sonnet 5n=30 · CN 30%
Google Maps
70%
Amap
53%
Mapbox
40%
Tencent Maps
23%
reverse_geocoder
20%
Kimi K3n=30 · CN 29%
Amap
63%
Google Maps
53%
Mapbox
33%
Nominatim
20%
PostGIS
13%
GLM-5.2n=30 · CN 42%
Amap
80%
Google Maps
50%
Mapbox
40%
PostGIS
13%
Pelias
13%

Object Storage & CDN

Chinese contenders: Aliyun OSS, Tencent COS, Qiniu, Upyun

Asked in English

Sonnet 5n=30
Cloudflare R2
33%
AWS S3
30%
Backblaze B2
23%
Supabase Storage
17%
Vercel
13%
Kimi K3n=30
Cloudflare R2
37%
Backblaze B2
23%
AWS S3
17%
Supabase Storage
13%
Cloudflare
10%
GLM-5.2n=30
AWS S3
30%
Cloudflare R2
27%
Backblaze B2
20%
Cloudflare
13%
Cloudflare Stream
7%

Asked in Chinese

Sonnet 5n=30 · CN 16%
Cloudflare R2
47%
AWS S3
33%
Backblaze B2
20%
Aliyun OSS
20%
Cloudflare
17%
Kimi K3n=30 · CN 30%
Cloudflare R2
53%
Aliyun OSS
33%
Tencent COS
30%
AWS S3
23%
Backblaze B2
20%
GLM-5.2n=28 · CN 31%
Cloudflare R2
54%
Aliyun OSS
46%
AWS S3
36%
Tencent COS
25%
Backblaze B2
18%

Git Hosting, CI & Registries

Chinese contenders: Gitee, Coding.net, npmmirror

Asked in English

Sonnet 5n=30 · CN 9%
GitHub
40%
GitHub Actions
20%
GitHub Packages
20%
Custom/DIY
10%
npmmirror
10%
Kimi K3n=30 · CN 11%
GitHub
40%
GitHub Actions
20%
GitHub Packages
13%
AWS CodeArtifact
7%
npm Official Registry
7%
GLM-5.2n=30 · CN 15%
GitHub
40%
GitHub Actions
20%
GitLab
13%
GitHub Packages
13%
npmmirror
13%

Asked in Chinese

Sonnet 5n=24 · CN 23%
GitHub Actions
25%
GitHub Packages
25%
GitHub
25%
npmmirror
17%
Tsinghua Mirror
8%
Kimi K3n=25 · CN 24%
GitHub
28%
GitHub Actions
24%
npmmirror
16%
Verdaccio
8%
GitHub Packages
8%
GLM-5.2n=30 · CN 26%
GitHub
40%
GitHub Actions
20%
Gitee
13%
npmmirror
13%
Verdaccio
10%

Realtime & RTC

Chinese contenders: Tencent TRTC, ZEGO; Agora is Chinese-founded

Asked in English

Sonnet 5n=30
AWS IVS
17%
LiveKit
17%
Yjs
17%
Daily.co
17%
Mux
13%
Kimi K3n=30
LiveKit
20%
Supabase Realtime
20%
Mux
10%
Stream Video
7%
Stream Chat
7%
GLM-5.2n=27
LiveKit
26%
Supabase Realtime
26%
Mux
22%
Daily
11%
Agora
7%

Asked in Chinese

Sonnet 5n=30 · CN 18%
Agora
37%
AWS IVS
20%
Yjs
20%
Cloudflare Stream
17%
Tencent Cloud Live
17%
Kimi K3n=30 · CN 24%
LiveKit
23%
Agora
20%
Tencent Cloud Live
17%
Aliyun Live
17%
AWS IVS
17%
GLM-5.2n=30 · CN 29%
Agora
40%
LiveKit
37%
Tencent TRTC
30%
Tencent Cloud Live
20%
Yjs
20%

LLM APIs

Chinese contenders: DeepSeek, Moonshot Kimi, Zhipu GLM, Qwen, Doubao

Asked in English

Sonnet 5n=30
Anthropic
100%
Claude Agent SDK
3%
Kimi K3n=30
Anthropic
57%
Gemini
40%
OpenAI
27%
Vercel AI SDK
17%
Groq
7%
GLM-5.2n=26
OpenAI
54%
Gemini
46%
Anthropic
19%
Vercel AI SDK
15%
LiteLLM
4%

Asked in Chinese

Sonnet 5n=30
Anthropic
100%
Kimi K3n=30 · CN 36%
DeepSeek
80%
Moonshot Kimi
57%
Anthropic
37%
Gemini
33%
OpenAI
23%
GLM-5.2n=29 · CN 35%
DeepSeek
72%
Anthropic
55%
Gemini
55%
Zhipu GLM
48%
Qwen
41%

Auth & Identity

Chinese contenders: Authing, WeChat login flows

Asked in English

Sonnet 5n=20
Clerk
50%
Auth0
20%
WorkOS
20%
Supabase Auth
15%
NextAuth.js
5%
Kimi K3n=20
Clerk
45%
Supabase Auth
30%
WorkOS
20%
Firebase Auth
10%
NextAuth.js
5%
GLM-5.2n=19
Clerk
47%
Supabase Auth
26%
Auth0
21%
NextAuth.js
16%
WorkOS
16%

Asked in Chinese

Sonnet 5n=20 · CN 11%
Clerk
40%
Supabase Auth
25%
WorkOS
20%
NextAuth.js
10%
Auth0
10%
Kimi K3n=20 · CN 25%
Clerk
55%
Supabase Auth
30%
WorkOS
20%
Firebase Auth
15%
Aliyun PNVS
10%
GLM-5.2n=20 · CN 22%
Clerk
25%
Authing
25%
Supabase Auth
25%
Auth0
20%
WorkOS
20%

Product Analytics

Chinese contenders: Sensors Data, GrowingIO, Umeng

Asked in English

Sonnet 5n=30
PostHog
73%
Plausible
17%
Firebase Analytics
13%
Amplitude
7%
Umami
7%
Kimi K3n=30
PostHog
87%
Plausible
17%
Umami
10%
Sentry
7%
Firebase Crashlytics
7%
GLM-5.2n=30
PostHog
57%
Firebase Cloud Messaging
17%
Plausible
17%
Mixpanel
13%
TelemetryDeck
3%

Asked in Chinese

Sonnet 5n=30 · CN 8%
PostHog
50%
Firebase Analytics
27%
Mixpanel
23%
Amplitude
20%
Umami
13%
Kimi K3n=30 · CN 9%
PostHog
77%
Amplitude
20%
Umami
20%
Sensors Data
17%
Firebase Crashlytics
10%
GLM-5.2n=30 · CN 28%
PostHog
53%
Sensors Data
30%
Amplitude
30%
Umeng
20%
Firebase Cloud Messaging
17%

Managed & Distributed Databases

Chinese contenders: PolarDB, OceanBase, TDSQL; TiDB is Chinese-founded

Asked in English

Sonnet 5n=20
Neon
45%
TiDB
40%
Supabase
35%
Vitess
15%
SingleStore
10%
Kimi K3n=20
Neon
50%
TiDB
40%
Vitess
10%
PlanetScale
10%
Supabase
10%
GLM-5.2n=20 · CN 6%
TiDB
40%
Supabase
25%
Neon
25%
SingleStore
20%
Crunchy Bridge
15%

Asked in Chinese

Sonnet 5n=20 · CN 5%
Neon
55%
TiDB
40%
Supabase
40%
RDS
20%
OceanBase
10%
Kimi K3n=20 · CN 10%
TiDB
40%
Supabase
35%
Neon
35%
RDS
25%
Google Cloud SQL
15%
GLM-5.2n=20 · CN 19%
Supabase
55%
TiDB
40%
Neon
35%
Crunchy Bridge
20%
PolarDB
15%

Cloud Hosting

Chinese contenders: Alibaba Cloud, Tencent Cloud, Huawei Cloud

Asked in English

Sonnet 5n=27
Vercel
56%
AWS
26%
Google Cloud
22%
Supabase
22%
Render
22%
Kimi K3n=30
Vercel
50%
Cloudflare Pages
27%
Supabase
27%
AWS
17%
Google Cloud
13%
GLM-5.2n=25
Vercel
56%
Render
20%
Fly.io
20%
Supabase
20%
AWS
16%

Asked in Chinese

Sonnet 5n=28 · CN 12%
Vercel
57%
AWS
21%
Render
21%
Supabase
18%
Alibaba Cloud
11%
Kimi K3n=30 · CN 7%
Vercel
53%
Google Cloud
17%
Cloudflare Pages
17%
Supabase
17%
AWS
10%
GLM-5.2n=26 · CN 14%
Vercel
54%
Google Cloud
19%
Render
15%
Cloudflare Pages
15%
Alibaba Cloud
12%

Reading notes. Percentages are share of responses naming the tool as a primary pick; segmented answers carry co-primary picks for each market branch, so Chinese-cell rows can sum past 100%. Cells are n=10 to 30 responses, so single-cell percentages are rough: at n=10 to 30 a 95% interval spans roughly ±20 to 30 points. The cross-category and cross-model patterns are the reliable signal, individual cell values less so.

Answer Style: Stability, Refusals, Branching

Run-to-run stability

Same top pick, run 1 vs run 2: Claude 79.3%, Kimi 70.4%, GLM 65.1%. Claude is the most consistent recommender of the three.

Non-answers

Clarifying questions with no pick: GLM 5.5% of English answers (its ask-before-recommending habit), everyone else 0 to 2.9%. Kimi answered every English prompt.

Market branching

“Domestic vs overseas” branches in Chinese answers: GLM 58.9%, Kimi 57.7%, Claude 28.2%. In English all three sit under 5%.

The branching number is the mechanism behind most of this study's findings. Addressed in Chinese, the models rarely drop the Western tools from the pick set; they add a second answer for a second audience, and which audience gets listed first is where the models differ. None of the prompts asked for a market comparison.

Supabase is the clearest illustration, because it moves the wrong way. Every model names it in more Chinese answers than English ones (Kimi 2 responses against 7, GLM 5 against 11). For Claude the share of picks is flat, 21% to 19%, so its rise in response count is answer length alone. For Kimi and GLM the share rises too, 7% to 18% and 15% to 23%. The Chinese answers that produce this are the branching ones: asked in Chinese which Postgres host to use in production, GLM returns a two-branch table of overseas and domestic options averaging 6.75 primary picks per response, its densest answers in the study, and Supabase is on the overseas side of it.

The Distillation Question: No Claude-Specific Signature Found

A reasonable reader will ask whether the English-side convergence means the Chinese models trained on Claude's outputs. Our design cannot answer that, and we want to be precise about what the data supports.

What we observe: Kimi matches Claude's top pick in 62.7% of English prompts, GLM in 59.0%, and the two Chinese models match each other in 62.1%. The symmetry is the point. Kimi's agreement with Claude is no higher than its agreement with GLM. We found no Claude-specific recommendation signature in these tests. The pattern is compatible with several training, post-training, harness, and corpus explanations that this design cannot distinguish.

A sharper cut of the same data asks who pairs off when the three disagree. On the 290 English prompts all three models answered, all three pick the same tool 47% of the time. When exactly two agree against the third, the pair is Claude+Kimi in 16.2% of prompts, Kimi+GLM in 14.5%, and Claude+GLM in 11.7%. Claude+Kimi leads, but a spread that size across three pairs is within sampling noise at this n (chi-square p ≈ 0.35). In Chinese the pairing counts are Kimi+GLM 18.7%, Claude+Kimi 17.0%, and Claude+GLM 11.9% (p ≈ 0.10 before accounting for clustering, so we treat it as no finding). Neither language shows a lopsided Claude-pairing.

What the Chinese-side depth does and does not show: GLM surfaces second-tier domestic vendors (Sensors Data, Qiniu, JPush, Authing) that Claude names rarely or never, which is hard to square with the strongest claim, that these models are Claude end to end. It does not speak to their English side, since native Chinese data and Claude-derived English data could coexist in one model. The English-side evidence is the agreement symmetry and pairing analysis above and the divergence test below. Their answer styles also differ from Claude's in stable ways (double the market branching, lower run-to-run stability, zero English self-promotion against Claude's 100%), though style is the weakest of these signals: post-training could reshape it either way.

We then ran that stronger test. Our March 2026 Codex study gives per-prompt reference picks for Claude (Opus 4.6) and Codex (GPT-5.3) on a different 12-category suite; 48 repo-prompt pairs have stable reference picks (the same tool in at least 2 of 3 runs) that disagree between the two agents. We ran Kimi K3 and GLM-5.2 on all 48 (one run, July 2026), extracted their picks the same way, and scored which side each landed on.

Kimi matched the Claude-side pick 22 times, the Codex-side pick 18 times, and neither on the remaining 8 (55% Claude-side of 40 decided, binomial p = 0.64). GLM matched Claude 11 times and Codex 14 (44% Claude-side, p = 0.69), and picked something neither reference chose on 21 of its 46 scored prompts, nearly three times Kimi's rate. We did not detect a Claude-side preference in either model. The marker tools tell the same story from both directions: Bun, which March-Claude picked in 63% of runtime prompts and Codex in 2%, shows up in 3 to 4 of the Chinese models' 10 runtime answers; Statsig, which Codex picked 27% of the time and Claude never picked once, shows up in 17 to 24% of their feature-flag answers. Each Chinese model carries markers from both Western agents.

Caveats on this test: the reference picks are four months older than the new runs (model drift blurs the signal in both directions), and the new side is a single run. Neither model showed a statistically detectable Claude-side preference in this sample: the 95% intervals on the Claude-side share are roughly 40-69% for Kimi and 27-63% for GLM among decided cases. The estimates are imprecise, do not rule out smaller preferences, and do not identify training provenance. The chi-square figure tests only whether the three exclusive-pair frequencies are equal.

Methodology

Suite design.

120 prompts: 60 English, each with a twin translated by us into natural developer Simplified Chinese (the script used in mainland China and Singapore; Taiwan and Hong Kong write Traditional Chinese) (register-checked, standard loanwords like CDN and API kept in English, no literal word-for-word renderings). Prompts are phrased as everyday asks (“i need to send sms verification codes. what service should i use”) and never name vendors, countries, or compliance regimes. Categories: SMS & Push Notifications, Payments, Maps & Geo, Object Storage & CDN, Git Hosting, CI & Registries, Realtime & RTC, LLM APIs, Auth & Identity, Product Analytics, Managed & Distributed Databases, Cloud Hosting, Search. Repos: nextjs-saas (Next.js SaaS), python-api (FastAPI service), react-spa (React + Vite SPA), node-cli (Node/TS CLI), kotlin-android-app (Kotlin Android). Prompts apply to 2 to 5 repos each depending on category fit.

Collection.

Each model ran inside its maker's own agent CLI in fully autonomous mode, two complete runs per agent, on July 25, 2026. Every prompt ran in a fresh agent session against a repository reset to the benchmark baseline; prompts ran in fixed order, and clearing session and repository state between prompts reduced cross-prompt carryover. Agents ran with isolated per-run profiles, per-prompt session wipes, anonymized working directories, and benchmark-identifying environment variables stripped, following the same leak-hardening protocol as our other agent studies. All API traffic originated from US IP addresses, so whatever the serving side could infer from the connection pointed to the United States, and the intended experimental difference within each matched pair was the prompt language. Agent harness and model are confounded by design; this measures each vendor's shipped experience, not the bare model. Web-tool availability differed (Claude Code and ZCode have web tools; our Kimi profile had none). Kimi ran its default thinking effort (max). 1,860 of 1,860 responses succeeded.

Extraction.

Primary picks and alternatives were extracted by reading every response in full under a written convention set (a pick is what the response recommends, not the first tool named; defaults offered inside pure clarifying questions are non-answers; answers that branch by market yield co-primary picks per branch lead; drivers, SDKs, and the repo's own framework are never picks). Chinese responses were read in Chinese and mapped to the same canonical vendor slugs as English responses via a bilingual alias table (阿里云 → Alibaba Cloud, 高德 → Amap, 支付宝 → Alipay). Extraction coverage was validated 1:1 against the source responses; a 13-row stratified sample was re-read against the raw responses during extraction; and two independent reviewers later audited every headline number against the analysis files. As a further check we blind double-coded a 60-row stratified sample (10 per model-language cell) with two independent coders working from the raw responses and the written conventions only. The coders matched each other's primary pick exactly on 59 of 60 rows, matched the shipped extraction on 59 and 60 of 60, and agreed with it on non-answer status on all 60; the single disagreement (Clerk vs Auth0 on one auth prompt) was flagged as ambiguous by both coders independently.

Classification.

“China-domestic” covers vendors whose primary market is mainland China (Aliyun OSS, Amap, Alipay, Sensors Data, Gitee). Chinese-founded companies with real global adoption (TiDB, Agora, DeepSeek, jieba) are a separate class and behave like Western vendors in our data: their pick rates barely move between languages. Response rates count a response once if any primary pick is China-domestic; pick shares divide by all primary picks. Both are reported because segmented answers give Chinese cells more picks per response (2.1 vs 1.5).

What would change these numbers.

One suite, five repos, one point in time (kimi-k3 was nine days old when we ran it). Category cells are small (n=10 to 30 per model per language); aggregate findings rest on n=289 to 734, and those observations cluster by base question, repository, and run, so the quoted p-values treat the data as more independent than the design strictly warrants. We therefore treat the quoted p-values as descriptive, not confirmatory; they are computed in the checked-in analysis scripts. Our translations are natural but not native-authored. Different repos, prompt phrasings, or CLI versions would move individual percentages; we would not expect them to reverse the ordering of the three headline gaps (the English convergence, the graded language flip, and the self-preference asymmetry), though that expectation is exactly what re-running would test.

The prompts, raw responses, extraction tables, vendor classification, and analysis code are checked into our benchmark repository and available on request; a public dataset release is planned. Related studies: What Claude Code Actually Chooses, What Codex Actually Chooses, What Fable Actually Chooses.

What This Study Cannot Tell You

Whether the models' behavior reflects policy, training data, or post-training choices. We measure outputs, not intent.

Whether Chinese developers actually receive these answers. Real usage mixes languages, system prompts, and follow-up turns that a benchmark does not.

Whether the flip generalizes to other languages. Japanese, Korean, or German prompts might segment too; we only tested the en/zh pair.

Market share. These are model recommendations on five specific repos, not adoption data.

We track what AI coding agents recommend

Amplifying benchmarks what Claude Code, Codex, Cursor, and now China's coding agents pick, and shows vendors where they win, where they lose, and what moved.