Supplymo
Original research · August 2026Evidence v1 · non-random sample

A supplier score of 5.0 can sit on top of a 1.0

We read 120 platform-ranked 1688 listings through the marketplace's own API. 65 showed the maximum headline supplier score, and 38 of them carried a component score at or below 2.5 underneath it.

The buyer sees one number. The platform holds five. This page shows how far apart they can be, and states in full what the sample cannot support.

Liam Cai

Liam Cai · Founder, Supplymo

Published August 4, 2026 · Collected 2026-08-04 · Yiwu, China

What we measured

The headline is almost always high. The components are not.

Every one of the 120 listings returned both a headline score and its five components, so nothing below rests on a partial reading.

Listings measured

120

Twelve consumer categories, one platform, one day.

Showed the maximum 5.0

65 of 120

109 scored 4.5 or above.

High headline, low component

38 of 120

Headline of 4 or more with a component at or below 2.5. 21 of those showed the maximum 5.0.

Median gap

1.9 points

Between the headline score and that listing's own weakest component. The largest gap was 4.9.

Weakest area

23 listings

Scored at or below 2.5 on pre-sale responsiveness, the most of any component.

Not the same scale

2 values

Dispute handling returned only 4 or 5 across all 120 listings, while pre-sale responsiveness returned ten distinct values from 0 to 5.

Which component gives way

Four of the five components move. One of them barely does.

Each row shows how many listings scored at or below 2.5, and the range that component actually occupied across the sample. Counts are shown next to every bar so the pattern does not depend on colour.

Number of listings scoring at or below 2.5 on each component
  1. Pre-sale responsiveness· Getting an answer before you buy23 of 120 · 19.2%range 05, 10 values
  2. Logistics experience· Getting the goods moved13 of 120 · 10.8%range 1.75, 15 values
  3. Listing experience· Finding the listing accurate11 of 120 · 9.2%range 25, 10 values
  4. After-sales experience· Getting help once it has shipped3 of 120 · 2.5%range 1.65, 8 values
  5. Dispute handling· The formal complaints process0 of 120 · 0.0%range 45, 2 values

Dispute handling did not fall below 2.5 on any listing, but it also never returned anything other than a 4 or a 5. Reading that as strong performance would be wrong. It is a component that did not vary in this sample, so it cannot be compared with the four that did. That matters for the headline score, because the headline averages all five.

Observed distribution

How far the headline sits above the weakest component

Each listing is placed by the distance between its own headline score and its own lowest component.

Distribution of the gap between headline score and weakest component score
  1. Headline below the component5 of 120 · 4.2%
  2. Under 1 point7 of 120 · 5.8%
  3. 1 to under 2 points49 of 120 · 40.8%
  4. 2 to under 3 points41 of 120 · 34.2%
  5. 3 points or more18 of 120 · 15.0%

The median gap is 1.9 points and the widest is 4.9. A gap is not a fault. It only means the single number shown is not a summary of the five underneath it. The top band is where the headline sits below its own components: all 5 of those listings returned a headline score of zero, which we read as new or inactive rather than poor. They are left in the median, which would be 2.0 without them.

Method

How the sample was collected

Everything here is reproducible from the evidence file and the sampling script.

1

Twelve keyword searches

Twelve English keyword searches on the cross-border distribution channel, covering everyday consumer categories.

2

First page only, ten rows each

Ten platform-ranked rows were taken from each first results page, giving 120 listings. Nothing was filtered or reordered by us.

3

Detail call per listing

Each listing was then read through the detail endpoint, which is where the five component scores are returned.

4

Read-only throughout

No order, payment, enquiry, messaging or logistics endpoint was called at any point.

Sanitised evidence v1

Check the figures yourself

The evidence file carries every listing row, the full method and all ten stated limits. No supplier identity exists at any stage: the sampling script records the category keyword and the numeric scores, and never writes an offer ID, seller ID, shop name, title or price.

Rows
120
Collected
2026-08-04
Stated limits
10

Boundaries

What this research does not say

Stated plainly, because a number travels further than the caveat attached to it.

  • That these suppliers are bad, or that a low component predicts a bad order.
  • That the platform hides anything. It publishes the components. They are simply not the number read first.
  • That this is what 1688 looks like overall. The sample is ranked, not random.
  • Anything at all about product quality, factory capability or export compliance.
  • That the five components can be compared with each other. One of them did not vary in this sample.

Limits

The ten limits of this sample

The first one is the one that matters most, so it is first.

  1. 1

    This is not a random sample. Rows come from the first page of a keyword search, which the platform ranks. Listings that rank highly are more likely to be promoted, high-volume or well-rated, so the headline scores here are probably higher than the platform-wide distribution.

  2. 2

    The five components are not on the same observed scale. Dispute handling returned only two values across all 120 listings, 4 and 5, while pre-sale responsiveness returned ten distinct values spanning 0 to 5. A component that never falls low in this sample may be a component that cannot fall low, and no conclusion should be drawn about supplier behaviour on dispute handling from these figures.

  3. 3

    Rows were not deduplicated by seller. A supplier with several listings on the first page of a search appears more than once, so 120 listings does not mean 120 distinct suppliers. We cannot say how many distinct suppliers are represented, because no seller identity was retained at any stage.

  4. 4

    One platform, one day, one snapshot. Scores move over time and no trend can be read from a single collection.

  5. 5

    Twelve consumer-goods keywords. Categories outside that range are not represented.

  6. 6

    120 listings is enough to show that a pattern exists and not enough to size it precisely for any single category.

  7. 7

    The platform does not publish how the headline score is computed from the components, so the relationship described here is observed, not derived.

  8. 8

    Five listings returned a headline score of zero. These are counted in the totals but are likely to be new or inactive sellers rather than poor performers.

  9. 9

    Component scores were read from the listing detail response. Where a component was absent it was excluded from that listing's weakest-component calculation rather than treated as zero.

  10. 10

    Nothing here measures product quality, factory capability or export compliance. It measures what a buyer is shown on a marketplace page.

Questions

Questions a reader should ask about this

What is the headline score and what are the components?

A 1688 listing shows one supplier score. The platform also returns five component scores through its API: logistics experience, listing experience, pre-sale responsiveness, after-sales experience and dispute handling. A buyer reads the first number. All five come back in the same response.

What does “38 of 120” actually describe?

In this non-random sample, 38 listings carried a headline score of 4 or above together with at least one component score at or below 2.5. It describes these 120 listings only. It is not an estimate for all suppliers on 1688.

Does a low component score mean the supplier is bad?

No. A component score records past buyer experience on one dimension. It does not predict how a specific order will go, and it says nothing about product quality, factory capability or export compliance.

Dispute handling never scored low. Does that mean suppliers handle disputes well?

No, and it would be a mistake to read it that way. Across all 120 listings dispute handling returned only two values, 4 and 5, while pre-sale responsiveness returned ten distinct values spanning 0 to 5. A component that never falls low may simply be one that cannot fall low in this data. We report the observation and draw no conclusion about supplier behaviour from it.

Is this a random sample of 1688?

No. Rows come from the first page of a keyword search, which the platform ranks. Highly ranked listings are more likely to be promoted, high-volume or well-rated, so the headline scores here are probably higher than the platform-wide distribution.

How should a buyer use this?

Ask for the component scores before treating the headline as a summary of them, and decide which component matters for the order in front of you. A first order with specification questions rests on pre-sale responsiveness. A repeat order rests on logistics.

Before you pay

A score is a starting point, not a check

We run supplier and product checks from Yiwu before money moves. We do not promise a lowest price or a guaranteed customs outcome.

Check a supplier