Exploration

ResearchFeatured

Discovery

DiscoverSourcesQuality

Analysis

Working setReviews
Flow StudioTeamConcept
Settings

Partners

  • AI AlliancePrime
  • BrightQueryBuilds Meridian
  • OpenMinedFunded partner
  • MLCommonsFunded partner
  • Hugging FaceDeployment platform
See the full consortium and what each partner wires

Meridian is the discovery layer for research data, built by BrightQuery within the AI Alliance.

hybrid · semantic + lexical · 1286 datasets ranked · 3.54s

Structuretabular385
Depthcataloged901measured385
Licenseunknown1286
Accessopen1286
Sourcehuggingface1286
clear
21-40 of 1286sortrelevancemeasured firstqualitysize
tabular

bezzam/coraal

0.00

63 rows × 58 cols

58 unknown

Corpus of Regional African American Language (CORAAL) Dataset link: https://oraal.github.io/coraal Hugging Face Hub preparation scripts: https://github.com/ebezzam/prepare_coraal Citation Kendall, Tyler and Charlie Farrington. 2023. The Corpus of Regional African American Language. Version 2023.06. Eugene, OR: The Online Resources for African American Language Project. https://doi.org/10.7264/1ad5-6t35.

open·-·huggingface·0% null·completeSource
tabular

Narsil/image_dummy

0.00

3 rows × 2 cols

2 text

\

open·-·huggingface·0% null·completeSource
tabular

google-research-datasets/mbpp

0.00

500 rows × 6 cols

3 text · 2 unknown · 1 numeric

Dataset Card for Mostly Basic Python Problems (mbpp) Dataset Summary The benchmark consists of around 1,000 crowd-sourced Python programming problems, designed to be solvable by entry level programmers, covering programming fundamentals, standard library functionality, and so on. Each problem consists of a task description, code solution and 3 automated test cases. As described in the paper, a subset of the data has been hand-verified by us. Released here as part of… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/mbpp.

open·-·huggingface·16% null·completeSource
tabular

aps/super_glue

0.00

100,730 rows × 9 cols

9 unknown

Dataset Card for "super_glue" Dataset Summary SuperGLUE (https://super.gluebenchmark.com/) is a new benchmark styled after GLUE with a new set of more difficult language understanding tasks, improved resources, and a new public leaderboard. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure Data Instances axb Size of downloaded dataset files: 0.03 MB Size of… See the full description on the dataset page: https://huggingface.co/datasets/aps/super_glue.

open·-·huggingface·0% null·completeSource
tabular

TIGER-Lab/MMLU-Pro

0.00

12,032 rows × 8 cols

5 text · 2 numeric · 1 unknown

MMLU-Pro Dataset MMLU-Pro dataset is a more robust and challenging massive multi-task understanding dataset tailored to more rigorously benchmark large language models' capabilities. This dataset contains 12K complex questions across various disciplines. |Github | 🏆Leaderboard | 📖Paper | 🚀 What's New [2026.03.11] Added more cutting-edge frontier models to the leaderboard, including the Claude-4.6 series, Seed2.0 series, Qwen3.5 series, and Gemini-3.1-Pro, among… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/MMLU-Pro.

open·-·huggingface·0% null·completeSource
tabular

uoft-cs/cifar10

0.00

50,000 rows × 3 cols

3 unknown

Dataset Card for CIFAR-10 Dataset Summary The CIFAR-10 dataset consists of 60000 32x32 colour images in 10 classes, with 6000 images per class. There are 50000 training images and 10000 test images. The dataset is divided into five training batches and one test batch, each with 10000 images. The test batch contains exactly 1000 randomly-selected images from each class. The training batches contain the remaining images in random order, but some training batches may contain… See the full description on the dataset page: https://huggingface.co/datasets/uoft-cs/cifar10.

open·-·huggingface·33% null·completeSource
tabular

huggingface/badges

0.00

1 rows × 2 cols

2 unknown

Badges A set of badges you can use anywhere. Just update the anchor URL to point to the correct action for your Space. Light or dark background with 4 sizes available: small, medium, large, and extra large. How to use? With markdown, just copy the badge from: https://huggingface.co/datasets/huggingface/badges/blob/main/README.md?code=true With HTML, inspect this page with your web browser and copy the outer html. Available sizes Small Medium Large Extra… See the full description on the dataset page: https://huggingface.co/datasets/huggingface/badges.

open·-·huggingface·50% null·completeSource
tabular

allenai/openbookqa

0.00

4,957 rows × 9 cols

9 unknown

Dataset Card for OpenBookQA Dataset Summary OpenBookQA aims to promote research in advanced question-answering, probing a deeper understanding of both the topic (with salient facts summarized as an open book, also provided with the dataset) and the language it is expressed in. In particular, it contains questions that require multi-step reasoning, use of additional common and commonsense knowledge, and rich text comprehension. OpenBookQA is a new kind of… See the full description on the dataset page: https://huggingface.co/datasets/allenai/openbookqa.

open·-·huggingface·0% null·completeSource
tabular

ylecun/mnist

0.00

60,000 rows × 3 cols

3 unknown

Dataset Card for MNIST Dataset Summary The MNIST dataset consists of 70,000 28x28 black-and-white images of handwritten digits extracted from two NIST databases. There are 60,000 images in the training dataset and 10,000 images in the validation dataset, one class per digit so a total of 10 classes, with 7,000 images (6,000 train images and 1,000 test images) per class. Half of the image were drawn by Census Bureau employees and the other half by high school students… See the full description on the dataset page: https://huggingface.co/datasets/ylecun/mnist.

open·-·huggingface·33% null·completeSource
tabular

pwc-archive/evaluation-tables

0.00

2,254 rows × 9167 cols

9167 unknown

[!CAUTION] This dataset will not be updated. It corresponds to the last available public snapshot of the data, retrieved on July 28th, 2025.

open·-·huggingface·2218% null·completeSource
tabular

rajpurkar/squad

0.00

87,599 rows × 6 cols

6 unknown

Dataset Card for SQuAD Dataset Summary Stanford Question Answering Dataset (SQuAD) is a reading comprehension dataset, consisting of questions posed by crowdworkers on a set of Wikipedia articles, where the answer to every question is a segment of text, or span, from the corresponding reading passage, or the question might be unanswerable. SQuAD 1.1 contains 100,000+ question-answer pairs on 500+ articles. Supported Tasks and Leaderboards Question Answering.… See the full description on the dataset page: https://huggingface.co/datasets/rajpurkar/squad.

open·-·huggingface·0% null·completeSource
tabular

xlangai/DS-1000

0.00

1,000 rows × 9 cols

9 unknown

DS-1000 in simplified format 🔥 Check the leaderboard from Eval-Arena on our project page. See testing code and more information (also the original fill-in-the-middle/Insertion format) in the DS-1000 repo. Reformatting credits: Yuhang Lai, Sida Wang

open·-·huggingface·0% null·completeSource
tabular

EleutherAI/hendrycks_math

0.00

1,295 rows × 4 cols

4 text

Dataset Summary MATH dataset from https://github.com/hendrycks/math Citation Information @article{hendrycksmath2021, title={Measuring Mathematical Problem Solving With the MATH Dataset}, author={Dan Hendrycks and Collin Burns and Saurav Kadavath and Akul Arora and Steven Basart and Eric Tang and Dawn Song and Jacob Steinhardt}, journal={NeurIPS}, year={2021} }

open·-·huggingface·0% null·completeSource
tabular

rtrm/debug

0.00

1 rows × 1 cols

1 text

test3

open·-·huggingface·0% null·completeSource
tabular

HuggingFaceH4/MATH-500

0.00

500 rows × 6 cols

5 text · 1 numeric

Dataset Card for MATH-500 This dataset contains a subset of 500 problems from the MATH benchmark that OpenAI created in their Let's Verify Step by Step paper. See their GitHub repo for the source file: https://github.com/openai/prm800k/tree/main?tab=readme-ov-file#math-splits

open·-·huggingface·0% null·completeSource
tabular

princeton-nlp/SWE-bench_Lite

0.00

300 rows × 12 cols

12 text

Dataset Summary SWE-bench Lite is subset of SWE-bench, a dataset that tests systems’ ability to solve GitHub issues automatically. The dataset collects 300 test Issue-Pull Request pairs from 11 popular Python. Evaluation is performed by unit test verification using post-PR behavior as the reference solution. The dataset was released as part of SWE-bench: Can Language Models Resolve Real-World GitHub Issues? Want to run inference now? This dataset only contains the… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/SWE-bench_Lite.

open·-·huggingface·0% null·completeSource
tabular

VLM2Vec/MSR-VTT

0.00

9,000 rows × 9 cols

4 text · 4 numeric · 1 unknown

Clone from "friedrichor/MSR-VTT". MSRVTT contains 10K video clips and 200K captions. We adopt the standard 1K-A split protocol, which was introduced in JSFusion and has since become the de facto benchmark split in the Text-Video Retrieval field. Train: train_7k: 7,010 videos, 140,200 captions train_9k: 9,000 videos, 180,000 captions Test: test_1k: 1,000 videos, 1,000 captions 🌟 Citation @inproceedings{xu2016msrvtt, title={Msr-vtt: A large video description dataset… See the full description on the dataset page: https://huggingface.co/datasets/VLM2Vec/MSR-VTT.

open·-·huggingface·0% null·completeSource
tabular

fancyzhx/ag_news

0.00

120,000 rows × 2 cols

1 text · 1 numeric

Dataset Card for "ag_news" Dataset Summary AG is a collection of more than 1 million news articles. News articles have been gathered from more than 2000 news sources by ComeToMyHead in more than 1 year of activity. ComeToMyHead is an academic news search engine which has been running since July, 2004. The dataset is provided by the academic comunity for research purposes in data mining (clustering, classification, etc), information retrieval (ranking, search, etc), xml… See the full description on the dataset page: https://huggingface.co/datasets/fancyzhx/ag_news.

open·-·huggingface·0% null·completeSource
tabular

akasheroor/American-Sign-Language-Dataset

0.00

25 rows × 2 cols

2 unknown

American Sign Language (ASL) Dataset Description:This dataset contains 108,618 videos representing 2,208 ASL words, with each word having a minimum of 30 videos. The videos were scraped, collected from multiple sources, and preprocessed to ensure consistency, quality, and usability for machine learning and gesture recognition tasks. Each video is ≤10 MB, optimized for storage and model training.The dataset can be used for ASL gesture recognition, video-based ML tasks, and model… See the full description on the dataset page: https://huggingface.co/datasets/akasheroor/American-Sign-Language-Dataset.

open·-·huggingface·50% null·completeSource
tabular

jamesqijingsong/zidian

0.00

8,614 rows × 3 cols

3 unknown

时间线: 2018年搭建成网站 https://zidian.18dao.net 2024年使用AI技術為《國語字典》生成配圖。 2025年上傳到Hugging Face做成數據集。 数据集中的文件: 目录 "image/" 下的文件数量: 4307,文生圖原始png圖片 目录 "image-zidian/" 下的文件数量: 4307,加字後的jpg圖片 目录 "text-zidian/" 下的文件数量: 4307,圖片解釋文字 目录 "pinyin/" 下的文件数量: 1702,拼音mp3文件

open·-·huggingface·0% null·completeSource
← previouspage 2next →

Select a result to see its full details here: the measured structure, quality, and the loader, without leaving your search.