πŸ“Š
Dataset

datacomp200m

by Adams Story adams-story/datacomp200m
Free2AITools Nexus Index
27.0
S: Semantic 50

Query-time baseline · scored live at search

A: Authority 33
P: Popularity 62
R: Recency 0
Q: Quality 50
Tech Context
Vital Performance
Data Integrity 27 FNI Score
- Size
- Rows
- Tokens
Dataset Information Summary
Entity Passport
Registry ID adams-story/datacomp200m
Provider huggingface
πŸ“œ

Cite this dataset

Academic & Research Attribution

BibTeX
@misc{hf_dataset_adams_story_datacomp200m,
  author = {Adams Story},
  title = {datacomp200m Dataset},
  year = {2026},
  howpublished = {\url{https://huggingface.co/datasets/adams-story/datacomp200m}},
  note = {Accessed via Free2AITools.}
}
APA Style
Adams Story. (2026). datacomp200m [Dataset]. Free2AITools. https://huggingface.co/datasets/adams-story/datacomp200m

πŸ”¬Technical Deep Dive

Full Specifications [+]

βš–οΈ Free2AITools Nexus Index V2.0

Semantic (S) 50

Query-time baseline · scored live at search

Authority (A) 33
Popularity (P) 62
Recency (R) 0
Quality (Q) 50

πŸ’¬ Index Insight

FNI V2.0 for datacomp200m: Authority (A:33), Popularity (P:62), Recency (R:0), Quality (Q:50). Semantic (S) is a query-time baseline scored live at search.

Free2AITools Nexus Index

Data Sources / Provenance

Open data Updated: Live data
⬇️
Downloads
248,940

πŸ‘οΈ Data Preview

πŸ“Š

Row-level preview not available for this dataset.

Schema structure is shown in the Field Logic panel when available.

πŸ”— Explore Full Dataset β†—

🧬 Field Logic

🧬

Schema not yet indexed for this dataset.

Dataset Specification

Datacomp200m

This is a smaller version of the datacomp_1b dataset.

Filtering was done by taking all rows that had self similarity (inner product) above 0.32. This resulted in 213009083 (213 million) rows.

The results of the datacomp paper suggest that filtering by CLIP score is better than random sampling.

Included in this repo are search indices created using autofaiss, over the text and image embeddings. There are two ways to access metadata, either in .parquet files in the ./metadata directory, or the ./index/metadata.hdf5 hdf5 file.

I would suggest using embedding-reader to load the text and image embeddings.

πŸ“Š Structured Schema (Zero-Fabrication)

Feature Key Data Type
uid string
url string
text string
original_width int64
original_height int64
clip_b32_similarity_score float32
clip_l14_similarity_score float32
face_bboxes Sequence[Sequence]
sha256 string
__index_level_0__ int64

Estimated Rows: 213,009,083

Social Proof

HuggingFace Hub
248.9KDownloads
πŸ”„ Updated daily

Source summary: Based on Hugging Face metadata. Not a recommendation.

πŸ“Š FNI Methodology πŸ“š Knowledge Baseℹ️ Verify with original source

πŸ›‘οΈ Dataset Transparency Report

Technical metadata sourced from upstream repositories.

Open Metadata

πŸ†” Identity & Source

id
hf-dataset--adams-story--datacomp200m
slug
adams-story--datacomp200m
source
huggingface
author
Adams Story
license
tags
size_categories:100m<n<1b, format:parquet, modality:image, modality:tabular, modality:text, library:datasets, library:dask, library:mlcroissant, library:polars, region:us

βš™οΈ Technical Specs

architecture
null
params billions
0.2
context length
null
pipeline tag

πŸ“Š Engagement & Metrics

downloads
248,940
stars
0
forks
0

Data indexed from public sources. Updated daily.