Supplymo
English title fieldsEvidence v1 · non-random sample

72 of 120 API-returned 1688 English title fields repeated one token at least three times

We measured the subject and subjectTrans fields returned across 120 first-page rows from 12 selected English searches. 72 subjectTrans fields crossed the three-repeat threshold and 31 crossed the stricter four-repeat threshold. All 120 returned zero CJK characters, which is a character-residue result rather than a translation-quality score.

The one-line version, free to quote with attribution: In a Supplymo non-random sample of 120 first-page API result rows from 12 selected English searches on 1688, 72 subjectTrans title fields repeated the same English token at least three times; all 120 returned zero CJK characters, a character-residue check that does not assess translation accuracy or fluency.

Liam Cai

Liam Cai · Founder, Supplymo

Published 1 September 2026 · Collected 2026-08-12 · Yiwu, China

Repeated-token threshold across 120 API-returned subjectTrans fields
Repeated one token at least three times
72
Below that threshold
48

n = 120 API result rows. Collected 2026-08-12. Non-random sample: 12 selected English searches and 120 first-page API result rows.

72/120

subjectTrans fields repeated one token at least three times

A token-frequency threshold, not a judgement about intent.

31/120

crossed the stricter four-repeat threshold

Both thresholds are published so the cut is inspectable.

120/120

returned zero CJK characters in subjectTrans

This does not assess translation accuracy, completeness or fluency.

6/3

maximum and median per-row repeat counts

Counts are computed after the published tokenisation rule.

Repeated tokens appeared across the returned title fields

The API returned both subject and subjectTrans on all 120 sampled rows. Under the published tokenisation rule, 72 subjectTrans values repeated one token at least three times and 31 repeated one at least four times. The maximum observed repeat count was 6.

The median subjectTrans length was 124 characters, while the median subject length was 31. Character counts across writing systems are not directly comparable, so this page does not convert that difference into a bloat or translation claim.

The public evidence retains only derived measurements, not title text. That protects brand, shop and product wording, but it also means a reader cannot re-tokenise these exact titles from the public file. The sampling script and metric definition are published so the method can be repeated against a later live draw.

The zero-residue result is deliberately narrow

120 of 120 subjectTrans values returned zero CJK characters. This rules out only one visible residue pattern in this draw. It does not establish that the English wording was accurate, complete, fluent or useful.

The official 1688 schema labels subject as the Chinese title and subjectTrans as the foreign-language title. It does not identify who authored the wording, what translation process produced it, or whether a website or app rendered the same value.

Search position is retained only to document the first-page sampling frame. No test links token repetition to rank, clicks, conversion or platform-reported sales.

How a buyer can use the result before payment

Treat an API-returned title as a discovery clue, not a product specification. Extract the exact product, material, size, variant and packaging you intend to buy, then confirm those details in structured attributes and a written supplier response.

Repeated words do not make an offer unsafe or a supplier unreliable. They are a reminder to move critical claims out of the title and into product-specific evidence that can be checked against the exact SKU and destination market.

If the title contains a compliance, performance or material claim, request the applicable document and confirm that it belongs to the product and market in question. This study validates none of those claims.

What this study cannot tell you

Putting this before the method rather than after it is deliberate. A number without its boundary travels further than the boundary does.

  • It cannot establish what a rendered 1688 website, app or other API surface displayed. The study measured fields returned by one read-only keyword-search API.
  • It cannot distinguish seller-authored wording from platform or other localization behavior, so it cannot assign intent or call the pattern seller keyword stuffing.
  • Zero CJK residue does not establish translation accuracy, completeness, fluency or buyer comprehension.
  • The study did not test ranking, clicks, conversion, sales or platform algorithms.
  • The 120 rows are not 120 distinct suppliers. Seller identity was not retained, so one shop may be represented more than once.

How the sample was collected

Source
com.alibaba.fenxiao.crossborder:product.search.keywordQuery (read-only)
Sampling frame
12 English keyword searches on the cross-border distribution channel, first results page only, 10 platform-ranked rows per keyword. The same 12 keywords as the seller-score and sales-concentration studies.
Collection window
Single day, 2026-08-12 in China
Write operations
None. No order, payment, enquiry, messaging or logistics endpoint was called.

The readings we tested before publishing

Each of these is a way the data could have been misread. We checked, and the check is published with its result, including where it came back inconclusive.

Is a repeated token really stuffing, or a legitimate multi-use product name?

Method. A word like 'pill' can repeat honestly in 'pill cutter and pill box'. Rather than judging intent, publish the repeat counts at two thresholds and per row, so a reader can apply a stricter cut.

Result. 72 of 120 titles repeat a token at least 3 times and 31 at least 4 times, with a maximum of 6. Per-row maximum repeat counts are in the rows array.

Conclusion. Reported as observed frequency at stated thresholds, not as a judgement of intent, authorship, search strategy or title quality.

Does zero CJK residue establish translation quality?

Method. Count CJK characters in every returned subjectTrans value as a narrow character-residue check.

Result. 120 of 120 returned subjectTrans fields contain zero CJK characters.

Conclusion. No. This rules out only CJK-character residue in this draw. It does not assess translation accuracy, completeness, fluency or rendered-page copy.

Does the English-to-Chinese length ratio show bloat?

Method. Character counts across writing systems are not comparable, since one Chinese character carries roughly a word's worth of meaning.

Result. The median ratio is 4 English characters per Chinese character, which is within the range expected from script density alone.

Conclusion. Length figures are published as description only and no bloat claim is made from them.

How is a token defined, so the counts can be re-derived?

Method. Lowercase the title, strip everything except letters, digits, spaces and hyphens, split on whitespace, keep tokens longer than 2 characters. The sampling script is repository-tracked.

Result. Definition applied identically to all 120 titles at sampling time.

Conclusion. Anyone re-running the public sampling script against the live API can reproduce the metric definition exactly.

The stated limits, in full

All 9 of them travel inside the evidence file, so they stay attached to the data after it leaves this page.

  1. 1.This is not a random sample. Rows come from the first page of a keyword search, which the platform ranks. Highly ranked listings may be optimised harder than average, in either direction.
  2. 2.Title text is deliberately not published, because titles can contain brand and shop names. A reader can verify the metric definitions and re-run the repository-tracked sampling script, but cannot re-derive the counts from this file alone.
  3. 3.A repeated token is not proof of bad faith, keyword stuffing, authorship or search strategy. The study reports a token-frequency threshold and does not infer why the wording occurred.
  4. 4.Zero CJK characters is a narrow residue check. It does not establish that a translation was accurate, complete, fluent or suitable for buyers, and the official schema does not identify the translation process.
  5. 5.Character counts across writing systems are not comparable and the length figures are description, not a bloat claim.
  6. 6.The 12 keywords are ordinary consumer goods in English. Categories with regulated or technical naming may title differently and are not represented here.
  7. 7.Single day of collection. Titles change, and no trend is claimed from one observation.
  8. 8.The 120 rows correspond to 120 distinct offer IDs in the internal raw sample, but they are not 120 distinct suppliers. Seller identity is deliberately not collected, so the sample cannot be de-duplicated by shop and one seller may appear more than once.
  9. 9.Nothing here says a platform or seller has done anything wrong. The study measures fields returned by one read-only API; it does not establish what a rendered website or app showed, who authored the title, how the wording was produced, or how buyers interpreted it.

Use this data

Everything below is free to reuse with attribution. If you are writing about this and need a cut of the data we have not published, ask and we will run it.

Quote it, or take the whole file

One line, free to quote with attribution

In a Supplymo non-random sample of 120 first-page API result rows from 12 selected English searches on 1688, 72 subjectTrans title fields repeated the same English token at least three times; all 120 returned zero CJK characters, a character-residue check that does not assess translation accuracy or fluency.

How to cite this study

Liam Cai. "Repeated tokens in API-returned 1688 subjectTrans title fields." Supplymo, 1 September 2026, https://supplymo.com/research/1688-title-stuffing.

The underlying data

The internal raw sample contains 120 distinct offer IDs across 120 rows, but no seller identity. The public evidence therefore cannot be de-duplicated by shop and is not a sample of 120 distinct suppliers. 9 stated limits travel inside the file. Free to reuse with attribution under CC BY 4.0.

Questions we get asked

What is the one-line finding?

In a Supplymo non-random sample of 120 first-page API result rows from 12 selected English searches on 1688, 72 subjectTrans title fields repeated the same English token at least three times; all 120 returned zero CJK characters, a character-residue check that does not assess translation accuracy or fluency. The threshold is descriptive and does not identify authorship, intent, search strategy or title quality.

Does this prove sellers stuffed keywords into their titles?

No. The API fields do not identify who authored the wording or how it was produced. Repeated tokens can also occur in legitimate multi-use product descriptions.

Does zero CJK residue mean the English translation was good?

No. All 120 subjectTrans fields returned zero CJK characters, but that narrow residue check does not assess accuracy, completeness, fluency or suitability for a buyer.

Did repeated words improve search rank or sales?

The study did not test ranking, clicks, conversion or sales effects. Rank position only defines the first-page sampling frame.

Were these 120 different suppliers?

No. The internal raw sample contained 120 distinct offer IDs, but no seller identity was retained. One seller may therefore appear more than once, and the public file cannot be de-duplicated by shop.

Research accountability

Author and figure check

Author: Liam Cai. Collected 2026-08-12; published 2026-09-01.

Figure verification: every published count is deterministically recomputed from the sanitized rows and checked against the repository-tracked raw-sample hash. No independent human reviewer is claimed.

Check the exact product

Move important claims out of the title before payment

Submit the exact product, variant and destination. Product Check helps turn title wording into specification, supplier-evidence, compliance and order-term questions that can be verified.

Check a product before payment