Structured format, all 10,511 pairs. Accuracy in percent; intervals are 95% query-clustered bootstrap CIs, since pairs sharing a query are not independent.
Where the LookBench retrieval track ranks a gallery against a query image, ShopRank-Bench asks a different question: given a shopping query and two candidate products, which one would a shopper prefer? It is a pairwise preference benchmark for text rerankers.
The pairs come from live shopping traffic and are labelled by a panel of three LLM judges from different model families, each seeing both presentation orders. A judge counts as committing only when both orders agree; one that flips is counted as abstaining. Pairs are kept when every committed judge chose the same product, and tiered by how many families committed — so the tier is a measure of label strength, not of difficulty alone.
Because the queries are private traffic rather than a public catalogue dump, the benchmark is contamination-limited: these query strings were not published before release. Every pair ships in two formats — a structured attribute schema and a natural-language rendering of the same attributes, carrying the same label — so a per-format gap measures sensitivity to serialization rather than a difference in labels.
| Property | Value |
|---|---|
| Preference pairs | 10,511 |
| Distinct queries | 2,991 |
| Gold tier (3/3 judges committed) | 1,843 |
| Silver tier (2/3) | 4,445 |
| Bronze tier (1/3) | 4,223 |
| Formats per pair | structured + natural language |
| Metric | pairwise accuracy (chance = 50%) |
Structured format, all 10,511 pairs. Accuracy in percent; intervals are 95% query-clustered bootstrap CIs, since pairs sharing a query are not independent.
| Model | Overall [95% CI] | Gold | Silver | Bronze |
|---|---|---|---|---|
| ZooWork-ShopRanker-8B | 83.4 [82.5, 84.2] | 96.5 | 86.1 | 74.8 |
| ZooWork-ShopRanker-4B | 81.2 [80.3, 82.0] | 96.3 | 84.3 | 71.4 |
| Qwen3-Reranker-8B | 79.2 [78.2, 80.1] | 93.2 | 82.6 | 69.4 |
| Jina-Reranker-m0 (2.4B) | 79.2 [78.3, 80.1] | 90.8 | 82.9 | 70.2 |
| Qwen3-Reranker-4B | 77.5 [76.5, 78.5] | 92.7 | 81.4 | 66.7 |
| ZooWork-ShopRanker-0.6B | 76.6 [75.7, 77.6] | 91.2 | 79.7 | 67.1 |
| Qwen3-Reranker-0.6B | 73.5 [72.4, 74.5] | 85.7 | 77.6 | 63.7 |
| BGE-Reranker-v2-m3 (0.6B) | 73.3 [72.3, 74.3] | 84.9 | 77.8 | 63.5 |
| BGE-Reranker-large (0.6B) | 71.6 [70.5, 72.7] | 82.4 | 76.0 | 62.2 |
| BM25 | 56.2 [54.9, 57.4] | 69.0 | 60.4 | 46.2 |
| Reference: zero-shot reasoning LLMs, not rerankers | ||||
| Qwen3.5-27B | 92.1 [91.5, 92.7] | 99.9 | 96.0 | 84.6 |
| GLM-5.3 | 91.9 [91.2, 92.4] | 99.6 | 95.9 | 84.3 |
The reasoning LLMs are listed as a reference point on how much of the benchmark's preference is recoverable at all, not as competing systems: they take seconds per decision against milliseconds for a reranker, and are not deployed this way.
The dataset repo ships evaluate.py, which runs the
LookBench reranking harness
on all 10,511 preference pairs in both formats and prints accuracy per tier with a
query-clustered 95% CI:
pip install torch transformers peft huggingface_hub
curl -LO https://huggingface.co/datasets/srpone/zoowork-shoprank-bench/resolve/main/evaluate.py
python evaluate.py --model srpone/zoowork-shopranker-4b # or any Qwen3-Reranker
The script clones LookBench on first run. For a ZooWork-ShopRanker repo it loads the
base weights and the LoRA adapter from that repo; add --adapter to evaluate
your own adapter on a Qwen3-Reranker base.
A model is correct on a pair when it scores the preferred candidate above the rejected one. Ties are reported rather than silently resolved.