ShopRank-Bench

A contamination-limited benchmark for judged e-commerce preference

Overview

Where the LookBench retrieval track ranks a gallery against a query image, ShopRank-Bench asks a different question: given a shopping query and two candidate products, which one would a shopper prefer? It is a pairwise preference benchmark for text rerankers.

The pairs come from live shopping traffic and are labelled by a panel of three LLM judges from different model families, each seeing both presentation orders. A judge counts as committing only when both orders agree; one that flips is counted as abstaining. Pairs are kept when every committed judge chose the same product, and tiered by how many families committed — so the tier is a measure of label strength, not of difficulty alone.

Because the queries are private traffic rather than a public catalogue dump, the benchmark is contamination-limited: these query strings were not published before release. Every pair ships in two formats — a structured attribute schema and a natural-language rendering of the same attributes, carrying the same label — so a per-format gap measures sensitivity to serialization rather than a difference in labels.

PropertyValue
Preference pairs10,511
Distinct queries2,991
Gold tier (3/3 judges committed)1,843
Silver tier (2/3)4,445
Bronze tier (1/3)4,223
Formats per pairstructured + natural language
Metricpairwise accuracy (chance = 50%)

Leaderboard

Structured format, all 10,511 pairs. Accuracy in percent; intervals are 95% query-clustered bootstrap CIs, since pairs sharing a query are not independent.

ModelOverall [95% CI]GoldSilverBronze
ZooWork-ShopRanker-8B83.4 [82.5, 84.2]96.586.174.8
ZooWork-ShopRanker-4B81.2 [80.3, 82.0]96.384.371.4
Qwen3-Reranker-8B79.2 [78.2, 80.1]93.282.669.4
Jina-Reranker-m0 (2.4B)79.2 [78.3, 80.1]90.882.970.2
Qwen3-Reranker-4B77.5 [76.5, 78.5]92.781.466.7
ZooWork-ShopRanker-0.6B76.6 [75.7, 77.6]91.279.767.1
Qwen3-Reranker-0.6B73.5 [72.4, 74.5]85.777.663.7
BGE-Reranker-v2-m3 (0.6B)73.3 [72.3, 74.3]84.977.863.5
BGE-Reranker-large (0.6B)71.6 [70.5, 72.7]82.476.062.2
BM2556.2 [54.9, 57.4]69.060.446.2
Reference: zero-shot reasoning LLMs, not rerankers
Qwen3.5-27B92.1 [91.5, 92.7]99.996.084.6
GLM-5.391.9 [91.2, 92.4]99.695.984.3

The reasoning LLMs are listed as a reference point on how much of the benchmark's preference is recoverable at all, not as competing systems: they take seconds per decision against milliseconds for a reranker, and are not deployed this way.

Evaluating a reranker

The dataset repo ships evaluate.py, which runs the LookBench reranking harness on all 10,511 preference pairs in both formats and prints accuracy per tier with a query-clustered 95% CI:

pip install torch transformers peft huggingface_hub
curl -LO https://huggingface.co/datasets/srpone/zoowork-shoprank-bench/resolve/main/evaluate.py
python evaluate.py --model srpone/zoowork-shopranker-4b   # or any Qwen3-Reranker

The script clones LookBench on first run. For a ZooWork-ShopRanker repo it loads the base weights and the LoRA adapter from that repo; add --adapter to evaluate your own adapter on a Qwen3-Reranker base.

A model is correct on a pair when it scores the preferred candidate above the rejected one. Ties are reported rather than silently resolved.