Exploration

ResearchFeatured

Discovery

DiscoverSourcesQuality

Analysis

Working setReviews
Flow StudioTeamConcept
Settings

Partners

  • AI AlliancePrime
  • BrightQueryBuilds Meridian
  • OpenMinedFunded partner
  • MLCommonsFunded partner
  • Hugging FaceDeployment platform
See the full consortium and what each partner wires

Meridian is the discovery layer for research data, built by BrightQuery within the AI Alliance.

hybrid · semantic + lexical · 115 datasets ranked · 2.69s

Depthcataloged115
Licenseopen104unknown7share alike3non commercial1
Accessopen115
Formatgzip106zip15csv11parquet8tsv4
Sourcezenodo115
clear
1-20 of 115sortrelevancemeasured firstqualitysize
declared

Genome assembly and annotation of Aegilops speltoides with B chromosomes

0.00

Chen, Jianyong

6 files · 1.5 GB · gzipdeclared

docx3
xlsx3
netcdf2
npy2
pdf2
vcf2
fasta1
fits1
hdf51
rar1
sqlite1
tiff1
torch1

A first sequence of the Ae. speltoides B chromosome has been available since 2020 (RUBAN et al. 2020b), but this assembly represents ~16% of the B only and thus cannot be used to uncover B-encoded genes. Now, high-molecular-weight DNA from +B leaf tissue of the clonally propagated +B plant was used to generate a chromosome-scale genome assembly. A total of 154 Gb of PacBio HiFi reads and 36.4 Gb of Nanopore reads (>25 kb) were generated for primary assembly using hifiasm (Cheng et al. 2026). The resulting 5.47 Gb assembly achieved 93.7% BUSCO completeness (contig N50 = 17.1 Mb). Approximately 104 Gb of Hi-C sequencing data derived from leaf tissue of the same +B plant were employed to scaffold the primary contigs. The assembly yielded eight large scaffolds (398-835 Mb), each displaying a characteristic Rabl configuration. Alignment of these scaffolds to the reference genome of Ae. speltoides accession AEG-9674-1 without B chromosome (Avni et al. 2022) revealed that seven of the eight scaffolds showed strong synteny with the standard A chromosomes 1S-7S. To determine whether the remaining large scaffold corresponded to the B chromosome, we generated ~60 Gb of whole-genome sequencing (WGS) data from 0B AR-derived lateral root tissue of the same plant, as well as approximately 36 Gb of WGS data from +B leaf tissue. Comparative read-mapping analyses showed that the eighth scaffold exhibited normal sequencing coverage in +B leaf-derived data but substantially reduced coverage in 0B AR-derived data. Thus, the 398 Mb scaffold represents the B chromosome, accounting for 69% of its size as estimated by flow cytometry. Additionally, 4.81 Gb of contigs were assigned to the seven pairs of A chromosomes, representing 91% of their estimated size (1C=5.27 Gb). Consequently, we produced a high-quality chromosome-scale assembly of Ae. speltoides carrying B chromosomes. To identify genes associated with the B chromosome elimination process, RNA-seq was performed across developmental stages and different tissues in which B chromosome behavior differs. Using all +B RNA-seq datasets, we annotated the Ae. speltoides genome assembly containing the B chromosome. This annotation identified 59,981 transcripts and 47792 protein-coding genes on the seven A chromosomes and 5,940 transcripts and 4,196 protein-coding genes on the B chromosome.

open·CC-BY-4.0·Zenodo·completeSource
declared

Fine tuning an LLM with a domain a specific data set

0.00

Madhusudan, Gujral

6 files · 29 MB · parquetdeclared

Large language models (LLMs) are trained on massive, publicly available text datasets comprising trillions of tokens, enabling them to excel at general language tasks like next-token prediction. However, LLMs often struggle with domain-specific prompts, exhibiting reduced accuracy or generating inaccurate information (hallucinations). This is because they lack sufficient subject matter expertise. Two primary approaches exist to address this limitation for augmenting LLMs knowledge: Retrieval-Augmented Generation (RAG) and fine-tuning. This presentation focuses on fine-tuning smaller LLMs with domain-specific instruct datasets using the LoRA (Low-Rank Adaptation) technique on Gaudi hardware. We will leverage publicly available LLMs and datasets from the Hugging Face Hub for this demonstration. Though it is possible to fine tune LLMs with plain text data - sourced from documents, articles, and other materials.

open·CC-BY-4.0·Zenodo·completeSource
declared

Model checkpoints for Regional climate risk assessment from climate models using probabilistic machine learning

0.00

Lopez-Gomez, Ignacio

2 files · 45 GB · gzipdeclared

Trained GenFocal model checkpoints. This record contains checkpoints for GenFocal: Checkpoints for the Diffusion-based Super-Resolution model (~8.15 GiB) CHeckpoints for the Flow Matching Debiasing model (~36.45 GiB)

open·CC-BY-4.0·Zenodo·completeSource
declared

Contig-AMR Linkage

0.00

Bakare, Akeem

1 files · 34 KB · gzipdeclared

Contig-ARG-Linkage Reproducible Snakemake pipeline that links antibiotic resistance genes to bacterial hosts in metagenomic data. For each sample, it assembles reads (MEGAHIT), detects ARGs (BLASTn vs MEGARes), classifies contigs (Kraken2), and joins by contig ID. Pinned envs, checksummed databases, container builds, and CI-tested on synthetic data.

open·CC-BY-4.0·Zenodo·completeSource
declared

Calibration-Conditioned FiLM Decoders for Low-Latency Decoding of Quantum Error Correction Evaluated on IBM Repetition-Code Experiments - Datasets

0.00

Stein, Samuel

5 files · 139 MB · gzipdeclared

Raw experimental dataset accompanying the paper "Calibration-Conditioned FiLM Decoders for Low-Latency Decoding of Quantum Error Correction Evaluated on IBM Repetition-Code Experiments." This dataset contains repetition-code experiments executed on three IBM Quantum processors -- ibm_kingston, ibm_pittsburgh, and ibm_fez. It comprises 352 hardware snapshots spanning code distances d = 3, 5, 7, 9, 11, syndrome-round counts r = 1 to 11, and both the X and Z logical bases. Each snapshot runs several repetition-code chains in parallel and carries the device calibration data captured at execution time, totalling several million measurement shots. All data is provided raw, exactly as returned by the hardware -- no machine-learning processing or sparsification is applied -- and is anonymized (no IBM Runtime job identifiers are released). CONTENTS - ibm_kingston.tar.gz, ibm_pittsburgh.tar.gz, ibm_fez.tar.gz per-device archives, each unpacking to <device>/d<D>_r<R>/job_<n>/ - index.csv one row per snapshot: path, backend, d, rounds, basis, logical_states, n_chains, shots - README.md full description of the layout and field definitions Each job_<n>/ directory contains: - info.json experiment parameters (device, d, rounds, basis, states, shots, n_chains) - calibration.json the device calibration snapshot at execution time (T1, T2, gate and readout error rates, coupling map) - circuit_state0.qasm, circuit_state1.qasm transpiled circuits as executed - bitstrings.json raw per-shot measurement records for every parallel chain USAGE Because each snapshot stores its own calibration data, the per-device and per-(d, r, basis) structure used in the paper is recoverable by filtering index.csv. The same calibration is consumed by both the FiLM decoder (via its calibration-graph encoder) and the modified MWPM baseline (via its detector-graph edge weights). Github to be updated for corresponding experiments on approval. Please contact Samuel Stein (samuel.stein@pnnl.gov) if you need any information or source docs before hand.

open·CC-BY-4.0·Zenodo·completeSource
declared

EuroFlood: a queryable cloud-native index for the CEMS-EFAS Satellite-Derived Flood Depth Maps

0.00

Hackl, Jürgen

6 files · 132 MB · parquet, tiffdeclared

EuroFlood is an open, cloud-native index over the JRC/Copernicus CEMS-EFAS Satellite-Derived Flood Depth Maps for Europe (Betterle & Salamon, 2025; CC-BY-4.0) - ~3,280 satellite-derived observed flood-depth maps across Europe, 2015-2024. The bundle is a sparse Cloud-Optimized GeoTIFF encoding, per pixel, the set of flood events that inundated it, plus a combo_id -sorted GeoParquet dictionary and a small events table. Query by region and time via HTTP range reads (GDAL /vsicurl + DuckDB) to retrieve matching events, then fetch only the source depth rasters needed. Built with the open-source EuroFlood Python package ( pip install euroflood ).

open·CC-BY-4.0·Zenodo·completeSource
declared

Structural models and molecular dynamics data for cathepsin B, S, and K complexes with glycosaminoglycan-like ligands relevant to allosteric regulation

0.00

Bojarski, Krzysztof Kamil

1 files · 66 MB · gzipdeclared

This repository contains structural models and molecular dynamics input files for complexes of cathepsins B, S, and K with three glycosaminoglycan-like ligands: naphthalene-1,3,6-trisulfonate (NTS), amide-linked bis(naphthalene disulfonate) (BNS), and suramin. Initial protein and ligand structures used for molecular docking are provided in AutoDock3 format. For each cathepsin-ligand complex, the three most populated docking clusters are included, with three representative binding poses per cluster. The repository further contains system topologies and initial coordinates for molecular dynamics simulations in AMBER format ( .parm7 and .rst7 ), ligand library files required for tleap, molecular dynamics input files, and scripts used for post-processing and analysis of the MD trajectories.

open·CC-BY-4.0·Zenodo·completeSource
declared

Multiple myeloma and therapy reshape the bone marrow niche to durably constrain immune reconstitution and vaccine responsiveness

0.00

Chander, Aishwarya

2 files · 8.7 GB · gzip, zipdeclared

Analysis code and processed data for the manuscript: "Multiple myeloma and therapy reshape the bone marrow niche to durably constrain immune reconstitution and vaccine responsiveness." This repository accompanies a longitudinal multi-omic study of immune dysfunction in multiple myeloma (MM). Matched bone marrow and peripheral blood from MM patients were profiled across diagnosis, induction, autologous stem cell transplant (ASCT), and recovery, together with matched healthy donors and vaccine response subcohorts. The analyses show that the tumor imposes a compartment specific immune program; marrow-restricted metabolic, inflammatory, and cytotoxic-effector changes not mirrored in blood; and, that adaptive immune reconstitution remains impaired up to two years post-ASCT. Half of patients failed to mount IgG responses to a high dose nonadjuvanted influenza vaccine, a defect overcome by the LNP adjuvanted COVID mRNA vaccine. Contents: Single cell RNA-seq (BMMC and PBMC), flow cytometry, Olink proteomics, and MSD cytokine analysis pipelines, plus the notebooks generating all manuscript figures. Processed/derived data needed to reproduce the figures are included; raw sequencing data are deposited in GEO.

open·CC-BY-4.0·Zenodo·completeSource
declared

VerNFR v1.0 for ISoLA 2026

0.00

Amilon, Jesper

1 files · 38 KB · gzipdeclared

Version 1.0 of VerNFR, a Frama-C plugin for verifying non-functional requirements of C code and interface contracts. Release to accompany submission to ISoLA2026 See the Readme for further details and instructions.

open·GPL-2.0-only·Zenodo·completeSource
declared

GROND

0.00

Walsh, Calum · Srinivas, Meghana · Stinear, Timothy · et al.

100 files · 2.7 GB · gzip, tsvdeclared

GROND (Genome-derived Ribosomal OperoN Database) A quality-checked and publicly-available database of 16S-ITS-23S RRNA operon sequences and their constituent 16S and 23S genes. Based on GTDB release R232.

open·CC-BY-4.0·Zenodo·completeSource
declared

HG_JUVENILE - Juvenile herring gulls (Larus argentatus, Laridae) hatched at the southern North Sea coast (Belgium)

0.00

Allaert, Reinoud A. · Stienen, Eric W.M. · Lens, Luc · et al.

6 files · 125 MB · csv, gzipdeclared

HG_JUVENILE - Juvenile herring gulls (Larus argentatus, Laridae) hatched at the southern North Sea coast (Belgium) is a bird tracking dataset published by the Centre for Research on Ecology, Cognition and Behaviour of Birds at Ghent University and the Research Institute for Nature and Forest (INBO) . It contains animal tracking data for the project/study HG_JUVENILE , using trackers developed by Interrex ( http://www.interrex-tracking.com ). The study has been operational since 2022. In total 204 individuals of European herring gull ( Larus argentatus ) have been tagged. 150 individuals were raised from egg by Ghent University researchers at the Wildlife Rescue Center in Ostend, completed several cognitive and behavioural tests when approximately three weeks old, and were released in the IJzermonding, Nieuwpoort (Belgium). 54 additional individuals were tagged and released in the wild to collect baseline data. The main goal of the study is to link cognitive performance in the lab to behaviour in the wild. Data are automatically synced with Movebank and from there periodically archived on Zenodo (see https://github.com/inbo/bird-tracking ). Files Data in this package are exported from Movebank study 2217728245 . Fields in the data follow the Movebank Attribute Dictionary and are described in datapackage.json . Files are structured as a Frictionless Data Package . You can access all data in R via https://zenodo.org/records/21279427/files/datapackage.json using frictionless . datapackage.json : technical description of the data files. HG_JUVENILE-reference-data.csv : reference data about the animals, tags and deployments. HG_JUVENILE-gps-yyyy.csv.gz : GPS data recorded by the tags, grouped by year. Acknowledgements This dataset was collected using infrastructure provided by the ERC and Ghent University.

open·CC0-1.0·Zenodo·completeSource
declared

Constrained Heat Index and Five-Level Heat Classification Framework

0.00

Liu, Ping

1 files · 13 KB · gzipdeclared

This record provides code, documentation, and small example data for applying the constrained heat index and five-level heat classification framework described in Liu (2026), "Assessing and Refining the Heat Index for Subdaily Heat Conditions," Journal of Applied Meteorology and Climatology, https://doi.org/10.1175/JAMC-D-25-0250.1. The package includes scripts to apply the HI ≥ T constraint, classify heat index values into five levels (L1-L5), and run point-based examples based on the 1995 Chicago O'Hare and 2024 Furnace Creek cases. The included sample data are intended for demonstration and testing only. Large ERA5 reanalysis files are not redistributed; users should obtain ERA5 data from the Copernicus Climate Data Store or NCAR Research Data Archive as described in the paper. Version 1.0.0 corresponds to the initial public release prepared for the JAMC article.

open·MIT·Zenodo·completeSource
declared

Reference data bundle for PacificBiosciences/Kinnex-IsoSeq-WDL

0.00

Mokveld, Tom · Bruand, Jocelyne · Rowell, William

1 files · 958 MB · gzipdeclared

Static input files to support demultiplexing, primer/barcode handling, alignment, annotation, and classification for Kinnex Iso-Seq using the GRCh38 reference. https://github.com/PacificBiosciences/Kinnex-IsoSeq-WDL kinnex-isoseq-wdl-resources-v0.1.0 ├── GRCh38 │ ├── annotation │ │ └── gencode.v49.annotation.gtf.gz │ ├── classification │ │ ├── intropolis.v1.hg19_with_liftover_to_hg38.tsv.min_count_10.modified2.sorted.tsv │ │ ├── polyA.list.txt │ │ └── refTSS_v3.3_human_coordinate.hg38.sorted.bed │ ├── human_GRCh38_no_alt_analysis_set.fasta │ └── human_GRCh38_no_alt_analysis_set.fasta.fai ├── GRCh38.ref_map.v0p1p0.template.tsv └── kinnex ├── isoseq_v2_barcoded_primers │ └── IsoSeq_v2_primers_12.fasta ├── kinnex_hifi_barcodes │ └── kinnex_hifi_barcodes.fasta └── kinnex_primers ├── kinnex_12fold_primers.fasta ├── kinnex_16fold_primers.fasta └── kinnex_8fold_primers.fasta

open·CC-BY-4.0·Zenodo·completeSource
declared

Supplementary material for "Susceptibility Haplotypes in Non-Cystic Fibrosis Newborns with Elevated Immunoreactive Trypsinogen"

0.00

Uva, Paolo · Rosamilia, Francesca

1 files · 376 KB · gzipdeclared

Introduction Cystic fibrosis (CF) represents the most prevalent life-threatening autosomal recessive disorder in Europe, primarily affecting the respiratory tract, exocrine pancreatic function, and lipid metabolism. Immunoreactive trypsinogen (IRT) quantification constitutes the first-tier test in neonatal CF screening programs; however, elevated IRT levels may also be detected in infants who do not develop CF, generating false-positive (FP) results. The biological basis underlying increased IRT in these cases remains poorly understood and may involve genetic determinants independent of CFTR. We hypothesized that a subset of infants with false-positive CF NBS may harbor genetic variants associated with pancreatic enzyme regulation or exocrine pancreatic biology, potentially contributing to elevated IRT despite the absence of CF Description This dataset comprises aggregated whole-exome sequencing (WES) data from 212 samples, aligned to the genome assembly GRCh38 (hg38). Variants were identified using GATK HaplotypeCaller (v4.1.9.0), normalized and annotated using Variant Effect Predictor (VEP, v113.0). Genotypes with low Depth of Coverage (DP < 8) or low genotype quality (GQ < 20) were set to missing. Variants were subsequently excluded if they showed a missingness rate (Fmiss) ≥20% in either group or a between-group difference in Fmiss exceeding 5%. Samples with high genotype missingness, excessive identity-by-descent (IBD), or ancestry discordant with the reference cohort based on principal component analysis (PCA) were exluded. Analyses were restricted to genes already known to be implicated in pancreatic diseases: CASR, CCL2, CEL, CELA3B, CFTR, CLDN2, CPA1, CTRB1, CTRB2, CTRC, CTSB, CXCL8, KRT18, KRT8, MORC4, PRSS1, PRSS2, RIPPLY1, SBDS, SPINK1, TRPV6. The annotated VCF file includes the following metrics, calculated separately for samples classified as false positives (FP; n = 94) and true negatives (TN; n = 118): AN_FP: Total number of alleles in called genotypes in FP AN_TN: Total number of alleles in called genotypes in TN AC_FP: Allele count in genotypes in FP AC_TN: Allele count in genotypes in TN AC_Hom_FP: Allele counts in homozygous genotypes in FP AC_Hom_TN: Allele counts in homozygous genotypes in TN AC_Het_FP: Allele counts in heterozygous genotypes in FP AC_Het_TN: Allele counts in heterozygous genotypes in TN AF_FP: Allele frequency in FP AF_TN: Allele frequency in TN

open·CC-BY-4.0·Zenodo·completeSource
declared

Dataset for performance of universal machine learning potentials in global optimization of inorganic crystal structures

0.00

Marcial, Edan · Chaudhary, Laxman · Gorbunova, Olesya · et al.

1 files · 18 MB · gzipdeclared

This dataset accompanies the manuscript "Performance of universal machine learning potentials in global optimization of inorganic crystal structures." It provides structural records for merged candidate pools collected from evolutionary optimization workflows together with energetic and metadata records for perturbation tests, including Zn c/a distortions, MB4 phase-stability tests, and Li-B formation-energy calculations. The main database is provided as data/dat.json.gz . The included json query script can read this compressed file directly, so manual extraction is not required. The script can be used to query systems, methods, available energies, atomic positions, POSCAR-format structures, merged-pool structure sets, and Wyckoff information. The archive also contains checksums, the collection manifest, and a README file with usage notes for inspecting and reproducing the collected records.

open·CC-BY-4.0·Zenodo·completeSource
declared

cellassign: Lightweight marker-based assignment of cell categories in AnnData objects.

0.00

Ascensión, Alex M.

2 files · 23 KB · gzipdeclared

cellassign assigns cell-group or cluster-level labels using user-defined marker gene sets. It is designed for single-cell workflows where cells have already been clustered, and where a simple marker-based annotation layer is useful.

open·MIT·Zenodo·completeSource
declared

Dataset for: "Local Distortions and B-site-Resolved Environments in Ca2Mn1-xTixO4 Solid Solutions"

0.00

Cesário, A. N. · Silva dos Santos, Samuel · Rodrigues, Pedro · et al.

3 files · 7.7 MB · gzipdeclared

A series of Ruddlesden-Popper perovskite solid solutions, Ca$_{2}$Mn$_{1-x}$Ti$_{x}$O$_{4}$, is examined by combining computational calculations with local scale experimental studies conducted at ISOLDE-CERN. Perturbed Angular Correlation (PAC) spectroscopy measurements, X-Ray Diffraction (XRD) data and Density Functional Theory (DFT) reuslts are given. Publication DOI: https://doi.org/10.1103/b18r-2d5k

open·CC-BY-4.0·Zenodo·completeSource
declared

A Multilingual Telegram Corpus Partially Annotated for Malicious-Content Taxonomy

0.00

De Souza, Debora F · Beltrao, Gabriela · Chulvi, Berta · et al.

4 files · 1.2 GB · csv, gzipdeclared

This dataset contains 6,805,925 messages collected from 490 public Telegram channels associated with conspiracy theories, anti-vaccine movements, and far-right/nationalist narratives across multiple languages (including English, Spanish, Turkish, Lithuanian, French and German). It was compiled as part of the MOISES project's research on the identification and analysis of malicious/extremist content on social platforms. A subset of 20,941 messages (466 of 490 channels) was manually annotated by human experts against a hierarchical taxonomy of five independent dimensions: Role , Tactic , Feature , Target and Vulnerability . This human-labeled subset supports the evaluation of automated (LLM-based) content-labeling systems against the taxonomy. Files tg_moises_full_corpus.jsonl.gz - 6,805,925 messages from 490 Telegram channels, one JSON object per line. tg_moises_human_labeled_subset.jsonl.gz - 20,941 messages with a human-assigned taxonomy_human label (466 channels contribute at least one labeled message). channel_index.csv - per-channel message counts and labeling coverage. channel_metadata.json - per-channel aggregate language distribution and topic-model output (from the _META collection), if present. Record fields (*.jsonl.gz) Field Description _id Original message id within its channel (MongoDB _id , not globally unique - combine with channel ) channel Telegram channel name (added at export time; corresponds to the original MongoDB collection) message_clean Cleaned message text message_language Detected language of the message text_cleaned Whether text cleaning was applied language_processed Whether language detection was applied repeated Whether the message was flagged as a repeat/duplicate taxonomy_human Human annotation across 5 dimensions: Role, Tactic, Feature, Target, Vulnerability (empty lists = not annotated / not applicable for that dimension) topic_id Topic-model cluster id, where computed (Field set observed by sampling; not every field is present on every record - this is a heterogeneous scrape.)

open·CC-BY-4.0·Zenodo·completeSource
declared

Genome-wide SNP Genotype Dataset and BLUP Phenotypic Data for 105 Indian Mungbean (Vigna radiata L. Wilczek) Accessions

0.00

Sarma, R N

2 files · 88 MB · csv, vcfdeclared

This dataset contains the genotype and phenotype data generated for a genome-wide association study (GWAS) of agronomic traits in a diverse panel of 105 Indian mungbean ( Vigna radiata L. Wilczek) accessions . The dataset comprises a filtered genome-wide SNP dataset in Variant Call Format (VCF) and the corresponding Best Linear Unbiased Predictor (BLUP) values for the measured agronomic traits.

open·CC-BY-4.0·Zenodo·completeSource
declared

Quantifying the Risk of Software Attacks Using Formal Verification – Artifact

0.00

Lanzinger, Florian · Reiche, Frederik · Dörre, Felix · et al.

1 files · 790 MB · gzipdeclared

This is the artifact for our (yet unpublished) paper "Quantifying the Risk of Sofware Attacks Using Formal Verification." It includes the versions of KeY and Palladio that we used as well as all data for the CoRReCt case study. See the included README for more information on how to run the artifact.

open·CC-BY-4.0·Zenodo·completeSource
page 1next →

Select a result to see its full details here: the measured structure, quality, and the loader, without leaving your search.