Exploration

ResearchFeatured

Discovery

DiscoverSourcesQuality

Analysis

Working setReviews
Flow StudioTeamConcept
Settings

Partners

  • AI AlliancePrime
  • BrightQueryBuilds Meridian
  • OpenMinedFunded partner
  • MLCommonsFunded partner
  • Hugging FaceDeployment platform
See the full consortium and what each partner wires

Meridian is the discovery layer for research data, built by BrightQuery within the AI Alliance.

hybrid · semantic + lexical · 901 datasets ranked · 5.13s

Depthcataloged901
Licenseunknown901
Accessopen901
Sourcehuggingface901
clear
1-20 of 901sortrelevancemeasured firstqualitysize
declared

HirakoSan/bio-lens

0.00

🌿 iNaturalist Bronze Dataset (Research-Grade, Deduplicated) Description This dataset contains research-grade observations from iNaturalist, processed through a bronze-layer pipeline that includes: Research-grade only observations (community-verified) One photo per observation (deduplicated by observation_uuid, keeping the first photo) AVIF encoded images stored as binary in Parquet ~5TB total size across over 41,316 shards Note: This dataset started as all… See the full description on the dataset page: https://huggingface.co/datasets/HirakoSan/bio-lens.

open·-·huggingface·completeSource
declared

axentx/surrogate-2-business-pipeline

0.00

axentx Surrogate-2 Business Pipeline Continuous business synthesis output from the axentx burn-loop daemon. Schema Column Type Description id int Sequential record id timestamp str (ISO 8601) UTC generation time category str Business vertical (15 categories, weighted) text str (markdown) Flat human-readable serialization — primary view idea dict Structured BMC: name, value_prop, target_customer, problem, differentiator, revenue_model… See the full description on the dataset page: https://huggingface.co/datasets/axentx/surrogate-2-business-pipeline.

open·-·huggingface·completeSource
declared

KakologArchives/KakologArchives

0.00

ニコニコ実況 過去ログアーカイブ ニコニコ実況 過去ログアーカイブは、ニコニコ実況 のサービス開始から現在までのすべての過去ログコメントを収集したデータセットです。 去る2020年12月、ニコニコ実況は ニコニコ生放送内の一公式チャンネルとしてリニューアル されました。これに伴い、2009年11月から運用されてきた旧システムは提供終了となり(事実上のサービス終了)、torne や BRAVIA などの家電への対応が軒並み終了する中、当時の生の声が詰まった約11年分の過去ログも同時に失われることとなってしまいました。 そこで 5ch の DTV 板の住民が中心となり、旧ニコニコ実況が終了するまでに11年分の全チャンネルの過去ログをアーカイブする計画が立ち上がりました。紆余曲折あり Nekopanda 氏が約11年分のラジオや BS も含めた全チャンネルの過去ログを完璧に取得してくださったおかげで、11年分の過去ログが電子の海に消えていく事態は回避できました。しかし、旧 API が廃止されてしまったため過去ログを API… See the full description on the dataset page: https://huggingface.co/datasets/KakologArchives/KakologArchives.

open·-·huggingface·completeSource
declared

xlangai/ubuntu_osworld_file_cache

0.00

OSWorld File Cache This repository serves as a file cache for the OSWorld project, providing reliable and fast access to evaluation files that were previously hosted on Google Drive. Overview OSWorld is a scalable, real computer environment for multimodal agents, supporting task setup, execution-based evaluation, and interactive learning across various operating systems and applications. This cache repository ensures that all evaluation files are consistently accessible… See the full description on the dataset page: https://huggingface.co/datasets/xlangai/ubuntu_osworld_file_cache.

open·-·huggingface·completeSource
declared

osv5m/osv5m

0.00

OpenStreetView-5M The Many Roads to Global Visual Geolocation 📍🌍 First authors: Guillaume Astruc, Nicolas Dufour, Ioannis SiglidisSecond authors: Constantin Aronssohn, Nacim Bouia, Stephanie Fu, Romain Loiseau, Van Nguyen Nguyen, Charles Raude, Elliot Vincent, Lintao XU, Hongyu ZhouLast author: Loic LandrieuResearch Institute: Imagine, LIGM, Ecole des Ponts, Univ Gustave Eiffel, CNRS, Marne-la-Vallée, France Introduction 🌍 OpenStreetView-5M is the first large-scale… See the full description on the dataset page: https://huggingface.co/datasets/osv5m/osv5m.

open·-·huggingface·completeSource
declared

IPEC-COMMUNITY/bridge_orig_lerobot

0.00

This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "widowx", "total_episodes": 53192, "total_frames": 1893026, "total_tasks": 19974, "total_videos": 212768, "total_chunks": 54, "chunks_size": 1000, "fps": 5, "splits": { "train": "0:53192" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/IPEC-COMMUNITY/bridge_orig_lerobot.

open·-·huggingface·completeSource
declared

hf-doc-build/doc-build-dev

0.00

This is a dataset which contains the docs from all the PRs that are updating one of the docs from https://huggingface.co/docs. It is automatically updated by this github action from the doc-buider repo.

open·-·huggingface·completeSource
declared

Benjy/typed_digital_signatures

0.00

Typed Digital Signatures Dataset This comprehensive dataset contains synthetic digital signatures rendered across 30 different Google Fonts, specifically selected for their handwriting and signature-style characteristics. Each font contributes unique stylistic elements, making this dataset ideal for robust signature analysis and font recognition tasks. Dataset Overview Total Fonts: 30 different Google Fonts Images per Font: 3,000 signatures Total Dataset Size: ~90,000… See the full description on the dataset page: https://huggingface.co/datasets/Benjy/typed_digital_signatures.

open·-·huggingface·completeSource
declared

mlfoundations/dclm-baseline-1.0

0.00

DCLM-baseline DCLM-baseline is a 4T token / 3B document pretraining dataset that achieves strong performance on language model benchmarks. Below are comparisions of model trained on DCLM-baseline with other models in the 7B regime. Model Params Tokens Open dataset? CORE MMLU EXTENDED Open weights, closed datasets Llama2 7B 2T ✗ 49.2 45.8 34.1 DeepSeek 7B 2T ✗ 50.7 48.5 35.3 Mistral-0.3 7B ? ✗ 57.0 62.7 45.1 QWEN-2 7B ? ✗ 57.5 71.9 50.5 Llama3 8B 15T ✗ 57.6… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0.

open·-·huggingface·completeSource
declared

genrobot2025/10Kh-RealOmin-OpenData

0.00

Boasting over 13,000 hours of cumulative data and 5 million+ clips, it ranks as the largest open-source embodied intelligence dataset in the industry. Update Notes:Stage 3 data upload completed. 13,000+ hours of pure dual-hand data with frame-level alignment latency < 1ms Full high-precision trajectory reconstruction, breaking the limit of superficial open source, fully ready-to-use 3,000+ contributors and 10,000+ real household scenarios with exceptional diversity Comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/genrobot2025/10Kh-RealOmin-OpenData.

open·-·huggingface·completeSource
declared

IPEC-COMMUNITY/droid_lerobot

0.00

This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "franka", "total_episodes": 92233, "total_frames": 27044326, "total_tasks": 31308, "total_videos": 276699, "total_chunks": 93, "chunks_size": 1000, "fps": 15, "splits": { "train": "0:92233" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/IPEC-COMMUNITY/droid_lerobot.

open·-·huggingface·completeSource
declared

HuggingFaceFW/finephrase

0.00

Dataset Card for HuggingFaceFW/finephrase Dataset Summary Synthetic data generated by DataTrove: Model: HuggingFaceTB/SmolLM2-1.7B-Instruct (main) Source dataset: HuggingFaceFW/fineweb-edu, config sample-350BT, split train Generation config: temperature=1.0, top_p=1.0, top_k=50, max_tokens=2048, model_max_context=8192 Speculative decoding: {"method":"suffix","num_speculative_tokens":32} System prompt: None Input column: text Prompt families: faq prompt Rewrite the… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finephrase.

open·-·huggingface·completeSource
declared

jat-project/jat-dataset-tokenized

0.00

Dataset Card for "jat-dataset-tokenized" More Information needed

open·-·huggingface·completeSource
declared

just-me7ss/American-Sign-Language-Dataset

0.00

American Sign Language (ASL) Dataset Description:This dataset contains 108,618 videos representing 2,208 ASL words, with each word having a minimum of 30 videos. The videos were scraped, collected from multiple sources, and preprocessed to ensure consistency, quality, and usability for machine learning and gesture recognition tasks. Each video is ≤10 MB, optimized for storage and model training.The dataset can be used for ASL gesture recognition, video-based ML tasks, and model… See the full description on the dataset page: https://huggingface.co/datasets/just-me7ss/American-Sign-Language-Dataset.

open·-·huggingface·completeSource
declared

permutans/arxiv-papers-by-subject

0.00

arXiv Papers by Subject A reorganised version of the nick007x/arxiv-papers dataset, partitioned by subject code, year, and month for efficient selective access. Dataset Description This dataset contains metadata for over 2.5 million arXiv papers, organised into a hierarchical directory structure that allows users to download only the specific subjects and time periods they need, rather than the entire dataset. Motivation The original nick007x/arxiv-papers… See the full description on the dataset page: https://huggingface.co/datasets/permutans/arxiv-papers-by-subject.

open·-·huggingface·completeSource
declared

HuggingFaceFW/fineweb

0.00

🍷 FineWeb 15 trillion tokens of the finest data the 🌐 web has to offer What is it? The 🍷 FineWeb dataset consists of more than 18.5T tokens (originally 15T tokens) of cleaned and deduplicated english web data from CommonCrawl. The data processing pipeline is optimized for LLM performance and ran on the 🏭 datatrove library, our large scale data processing library. 🍷 FineWeb was originally meant to be a fully open replication of 🦅 RefinedWeb, with a… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/fineweb.

open·-·huggingface·completeSource
declared

nvidia/PhysicalAI-Robotics-GR00T-X-Embodiment-Sim

0.00

PhysicalAI-Robotics-GR00T-X-Embodiment-Sim Github Repo: Isaac GR00T N1 We provide a set of datasets used for post-training of GR00T N1. Each dataset is a collection of trajectories from different robot embodiments and tasks. Cross-embodied bimanual manipulation: 9k trajectories Dataset Name #trajectories bimanual_panda_gripper.Threading 1000 bimanual_panda_hand.LiftTray 1000 bimanual_panda_gripper.ThreePieceAssembly 1000… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-GR00T-X-Embodiment-Sim.

open·-·huggingface·completeSource
declared

WINGNUS/ACL-OCL

0.00

Dataset Card for ACL Anthology Corpus This repository provides full-text and metadata to the ACL anthology collection (80k articles/posters as of September 2022) also including .pdf files and grobid extractions of the pdfs. How is this different from what ACL anthology provides and what already exists? We provide pdfs, full-text, references and other details extracted by grobid from the PDFs while ACL Anthology only provides abstracts. There exists a similar corpus… See the full description on the dataset page: https://huggingface.co/datasets/WINGNUS/ACL-OCL.

open·-·huggingface·completeSource
declared

vyokky/GUI-360

0.00

GUI-360°: A Comprehensive Dataset And Benchmark For Computer-Using Agents Paper | Code GUI-360° is a large-scale, comprehensive dataset and benchmark suite designed to advance Computer-Using Agents (CUAs). 🎯 Key Features 🔢 1.2M+ executed action steps across thousands of trajectories 💼 Popular Windows office applications (Word, Excel, PowerPoint) 📸 Full-resolution screenshots with accessibility metadata 🎨 Multi-modal trajectories with reasoning traces ✅ Both… See the full description on the dataset page: https://huggingface.co/datasets/vyokky/GUI-360.

open·-·huggingface·completeSource
declared

HuggingFaceM4/the_cauldron

0.00

Dataset Card for The Cauldron Dataset description The Cauldron is part of the Idefics2 release. It is a massive collection of 50 vision-language datasets (training sets only) that were used for the fine-tuning of the vision-language model Idefics2. Load the dataset To load the dataset, install the library datasets with pip install datasets. Then, from datasets import load_dataset ds = load_dataset("HuggingFaceM4/the_cauldron", "ai2d") to download and load the… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceM4/the_cauldron.

open·-·huggingface·completeSource
page 1next →

Select a result to see its full details here: the measured structure, quality, and the loader, without leaving your search.