Gensmo Logo
logo

LookBench

A Live and Holistic Open Benchmark for Fashion Image Retrieval
 

🔔 News

Visitor count Last updated: June 30, 2026

📄 [2026-06-30]: New submission from ZooClaw.ai: ZooClaw-FashionSigLIP2 — a SigLIP2-base model fine-tuned with knowledge distillation for robust fashion retrieval, evaluated on text-image retrieval tasks (ZooClaw-Fashion long/short and H&M). Paper · Model 🎉

🎯 [2026-06-04]: New submission from Kuaishou Team: Tianmu-MERE — a multimodal embedding model for product understanding and retrieval in e-commerce, achieving strong results across all subtasks! Model weights are now open-sourced on HuggingFace. 🏆

🚀 [2026-01-20]: Initial release of LookBench v2601 with 4 diverse subtasks across AI-generated and real-world scenarios, and a new open-source model GR-Lite! 🌟

🔥 [2026-01-20]: Release of GR-Pro (proprietary) and GR-Lite (open-source) models achieving state-of-the-art performance! 🎉

Introduction

LookBench is a live, holistic, and challenging benchmark for fashion image retrieval in real e-commerce settings. Unlike static benchmarks that are vulnerable to data contamination, LookBench features continuously refreshing samples, diverse retrieval intents across multiple difficulty levels, and attribute-supervised evaluation with over 100 visually grounded properties.

The current release (v2601) comprises approximately 2,300 queries, each evaluated against a carefully curated retrieval corpus of about 60,000 images per task:

Dataset Image Source # Retro Items Difficulty # Queries / Corpus
RealStudioFlat Real studio flat-lay product photos Single Easy 1,011 / 62,226
AIGen-Studio AI-generated lifestyle studio images Single Medium 192 / 59,254
RealStreetLook Real street outfit photos Multi Hard 1,000 / 61,553
AIGen-StreetLook AI-generated street outfit compositions Multi Hard 160 / 58,846

Each evaluation set is assessed using Coarse Recall, Fine Recall, and nDCG at @1, @5, @10, and @20.

Leaderboard

Open-Source Proprietary
Released on 2026-01-20

LookBench — Multi-attribute Image Retrieval

Real Studio
Real StreetLook
AI-Gen StreetLook
AI-Gen Studio

Text-Image Retrieval

ZooClaw-Fashion (long)
ZooClaw-Fashion (short)
H&M

Results of different models. The best-performing model in each metric is in-bold, and the second best is underlined.
GR-Pro and GR-Lite are our proprietary and open-source models respectively.
ZooClaw-Fashion and H&M numbers are from Xue & Xu, 2026. ZooClaw-Fashion reports R@1 and R@10 (long & short query); H&M reports R@10 + MRR@10. The ZooClaw-Fashion evaluation dataset will be released shortly at srpone/zooclaw-fashion-eval.

BibTeX

@article{gao2026lookbench,
      title={LookBench: A Live and Holistic Open Benchmark for Fashion Image Retrieval},
      author={Chao Gao and Siqiao Xue and Yimin Peng and Jiwen Fu and Tingyi Gu and Shanshan Li and Fan Zhou},
      year={2026},
      url={https://arxiv.org/abs/2601.14706},
      journal={arXiv preprint arXiv:2601.14706},
}

@article{xue2026zooclaw,
      title={ZooClaw-FashionSigLIP2: Distilled Fine-tuning for Robust Fashion Retrieval},
      author={Siqiao Xue and Chunxue Xu},
      year={2026},
      url={https://arxiv.org/abs/2606.27708},
      journal={arXiv preprint arXiv:2606.27708},
}

Visitor Map

Visitor map