Exploration

ResearchFeatured

Discovery

DiscoverSourcesQuality

Analysis

Working setReviews
Flow StudioTeamConcept
Settings

Partners

  • AI AlliancePrime
  • BrightQueryBuilds Meridian
  • OpenMinedFunded partner
  • MLCommonsFunded partner
  • Hugging FaceDeployment platform
See the full consortium and what each partner wires

Meridian is the discovery layer for research data, built by BrightQuery within the AI Alliance.

hybrid · semantic + lexical · 10 datasets ranked · 0.55s

Structuresequence4tabular1
Depthcataloged5measured5
Licenseopen10
Accessopen10
Formatfasta10csv2zip2gff1pdf1
Sourcezenodo5zenodo-bio5
clear
1-10 of 10sortrelevancemeasured firstqualitysize
sequence

Database of virus genomes from ultra-deep sequencing of wastewater (WVDB)

0.00

Kantor, Rose · Shakya, Migun · Ruth, Nelson · et al.

2,095 rows · 907 KB · fasta, tsv

A virus genome database representing 21,015 near-complete virus genomes collected from untargeted ultra-deep RNA/DNA combined sequencing of wastewater. Sequence data was provided by the CASPER consortium and raw data may be found on NCBI SRA under bioprojects PRJNA1247874 and PRJNA1198001. Data underwent read trimming, rRNA and human read removal, de novo assembly, and selection of high-quality viral contigs. Contigs were clustered at 95% identity and 85% query coverage to dereplicate. Chimera-checking required at least two independent assemblies of the same viral genome or presence of the genome in another reference database. Annotation made use of RdRpCATCH, geNomad, checkV, BLASTN against NCBI core-nt, and RNAVirHost. The RdRp fasta files contain representative RdRp sequences identified through homology to major RdRp reference databases and clustered at 90% sequence identity over 75% sequence coverage. Included sequences contain all three conserved RdRp motifs (A, B, and C) arranged in either the canonical ABC configuration or the permuted CAB configuration.

sqlite1
tsv1
xlsx1
open·CC-BY-4.0·zenodo-bio·completeSource
sequence

Panel Information Files for "PvGAP: Development of a Globally Applicable, Highly Multiplexed Microhaplotype Amplicon Panel for Plasmodium vivax"

0.00

Hubbard, Alfred · Solares, Edwin · Hemming-Schroeder, Elizabeth

88 rows · 18 KB · fasta

These are the files needed to run the Broad Institute's malaria amplicon pipeline for the PvGAP Plasmodium vivax panel, described in detail here . They consist of FASTA files containing the forward and reverse primers and another FASTA file containing reference sequences for each target, derived from the PvP01 reference genome.

open·CC-BY-4.0·zenodo-bio·completeSource
tabular

Root anatomical traits modulate the assembly and nitrogen transformation potential of root-associated microbiomes in a temperate steppe

0.00

Yuan, Guangyuan

72 rows × 13 cols · 6.0 KB · csv, fasta

11 numeric · 2 categorical

This dataset supports the findings of the manuscript "Root anatomical traits modulate the assembly and nitrogen transformation potential of root-associated microbiomes in a temperate steppe" (NPH-MS-2026-55667). It contains root traits data, bacterial 16S rRNA gene absolute abundances, functional genes relative abundances, DNA extraction metadata, and phylogenetic marker sequences for 37 plant species from a temperate steppe ecosystem. The dataset includes the following files: 1. root traits.csv - Root traits including average diameter (AD), specific root length (SRL), specific root area (SRA), root tissue density (RTD), root nitrogen content (RNC), root carbon content (RCC), carbon‑nitrogen ratio (RCN), cortex layer number (CLN), cortex thickness (CT), and the ratio of cortex thickness to root diameter (CTRD). The first column lists plant species names. 2. Absolute abundance of 16S rRNA gene.csv - Quantitative PCR (qPCR) derived absolute abundances of bacterial 16S rRNA gene copies (copies/ng DNA) across different root compartments (rhizosphere, rhizoplane, endosphere), host species, root orders, and cotyledon classes (monocot/dicot). 3. DNA extraction sample weight.csv - Fresh weight (grams) of root material used for DNA extraction for each sample, linked by SampleID to the abundance data. 4. DNA extraction concentration.csv - Qubit‑measured DNA concentrations (ng/μL) and the sample volume (μL) used for quality control, together with sample metadata. 5. 37species.fasta - DNA sequences of two chloroplast markers (matK and rbcL) for the 37 plant species included in the study. The sequences are in FASTA format with headers formatted as ">Species". These were used for host phylogeny construction and Pagel's λ analyses. 6. Quantitative PCR results of functional gene.csv - Quantitative PCR (qPCR) derived relative abundances of bacterial 16S rRNA gene and functional genes across different root compartments (rhizosphere, rhizoplane, endosphere), host species, root orders, and cotyledon classes (monocot/dicot). 7. README.md - A detailed description of each file, column headers, abbreviations, units, and any missing value codings (NA). All data are provided to ensure transparency and reproducibility of the analyses. For methodological details, please refer to the Materials and Methods section of the associated publication. These data are under embargo until the associated research article is published. After that date, they will be freely available under a Creative Commons Attribution 4.0 International (CC BY 4.0) license. During the embargo period, the metadata (title, authors, abstract) and the DOI remain publicly visible, but the data files are not accessible. For access requests before the embargo expires, please contact the corresponding author.

open·CC-BY-4.0·zenodo-bio·6% null·completeSource
sequence

Trypanosoma cruzi (Dm28c) genome

0.00

Requena Rolanía, Jose María · Greif, Gonzalo · ROBELLO, CARLOS

1 files · 8.0 MB · fasta

This dataset contains the genome sequence for Trypanosoma cruzi (strain Dm28c). This genome sequence was de novo assembled using PacBio Hi-Fi and Illumina sequencing platforms by Greif et al (2026. PMID: 41501640). The genome was assembled into 32 contigs, which represent complete chromosomes. The provided Fasta file also contains an additional contig corresponding to the maxicircle (mitochondrial genome) sequence. The Fasta files included in this dataset were downloaded from GenBank (assembly GCA_044048535.1; May 22, 2026). Additional information about the Dm28cT2T genome assembly and gene annotations may be accessed through the link: https://cruzi.pasteur.uy/

open·CC-BY-4.0·zenodo-bio·completeSource
sequence

Multiple sequence alignment, phylogenetic tree, and domain-level annotation of Cas7 homologs

0.00

Burman, Nathaniel · Buyukyoruk, Murat · Wiegand, Tanner · et al.

4 files · 8.0 MB · fasta

This folder contains a multiple sequence alignment of Cas7 homologs in .fasta format, the domain-level annotations from PFAM and CasFinder, and an associated phylogenetic tree in .newick format.

open·CC-BY-4.0·zenodo-bio·completeSource
declared

Software and AMR peptide database for 'PEPTiGEN: a tool for mining antimicrobial resistance PEPTides using GENe data of public available repositories'

0.00

Meekes, Lisa · Tabaro, Francesco · Bexkens, Michiel · et al.

41 files · 8.2 GB · csv, fasta, pdfdeclared

This record contains the Python software for PEPTiGEN, a tool for generating tryptic peptides from prokaryotic gene sequences and their variants, and the associated antimicrobial resistance (AMR) peptide database. The database is provided as an SQL file and a CSV file containing all genes and predicted peptides. The README file contains explanation of the PEPTiGEN tool. The SQL database schema files contains both the database schema of the SQL database used in the PEPTiGEN analysis as the database schema of the AMR peptide datbase.

open·CC-BY-4.0·Zenodo·completeSource
declared

Data set for the manuscript ´´Duplication of the genes coding the proteins that regulate RNA polymerase-III activity and differential transcription in tissues of teleost fish´´

0.00

Bir, Joyanta · Cancio, Ibon · Diaz de cerio, Oihane · et al.

3 files · 100 KB · fasta, xlsxdeclared

This data file contains the data associated with the manuscript entitled "Duplication of the Genes Coding the Proteins That Regulate RNA Polymerase III Activity and Differential Transcription in Tissues of Teleost Fish."

open·CC-BY-4.0·Zenodo·completeSource
declared

Evolutionary Dynamics of the Complete Chemosensory Repertoire in Kissing Bugs of the Genus Rhodnius: Divergent Odorant Receptors Contrast with Conserved Gene Families

0.00

Merle, Marie

18 files · 238 MB · fasta, zipdeclared

This repository contains the complete chemosensory protein and nucleotide sequences, along with the results of evolutionary selection tests for 13 species of the genus Rhodnius . 1. Project Description This dataset supports the study of the chemosensory repertoire (ORs, GRs, IRs, OBPs, and CSPs) across 13 Rhodnius genomes. The study highlights the contrast between the conservation of Gustatory (GRs) and Ionotropic (IRs) receptors and the high dynamic evolution of Odorant Receptors (ORs), particularly in species adapted to human habitats. 2. Repository Structure 2.1 Sequence Data (FASTA) The following files contain all identified chemosensory genes in both amino acid ( .faa ) and nucleotide ( .fna ) formats: Rhodnius_OR_proteins.faa / Rhodnius_OR_CDS.fna : Odorant Receptors. Rhodnius_GR_proteins.faa / Rhodnius_GR_CDS.fna : Gustatory Receptors Rhodnius_IR_proteins.faa / Rhodnius_IR_CDS.fna : Ionotropic Receptors. Rhodnius_OBP_proteins.faa / Rhodnius_OBP_CDS.fna : Odorant-Binding Proteins. Rhodnius_CSP_proteins.faa / Rhodnius_CSP_CDS.fna : Chemosensory Proteins. 2.2 Phylogenetic Trees Archives containing the multiple sequence alignments and the resulting phylogenetic trees (Newick/Treefile format): OR_trees.zip : Alignment ( OR.ali.fasta ) and tree file ( OR.ali.treefile ) for Odorant Receptors. GR_trees.zip : Alignment ( GR.ali.fasta ) and tree file ( GR.ali.treefile ) for Gustatory Receptors. IR_trees.zip : Alignment ( IR.ali.fasta ) and tree file ( IR.ali.treefile ) for Ionotropic Receptors. OBP_trees.zip : Alignment ( OBP.ali.fasta ) and tree file ( OBP.ali.treefile ) for Odorant-Binding Proteins. CSP_trees.zip : Alignment ( CSP.ali.fasta ) and tree file ( CSP.ali.treefile ) for Chemosensory Proteins. 2.3 Evolutionary Selection Tests These archives contain the results of selection pressure analyses (e.g., dN/dS ratios, Likelihood Ratio Tests). Each gene family folder is subdivided by orthologous groups (e.g., GR1, GR2). OR_selection.zip GR_selection.zip IR_selection.zip Inside each selection archive, you will find: *_ali.fasta : Codon-based multiple sequence alignment. *_ali.pml : Codon-based multiple sequence alignment in PAML-friendly format. tree : The phylogenetic tree used for the selection model. LRT_BM.xls / BM_LRT.xls : Results for the Branch Model tests (domiciliary species vs. sylvatic species, see the associated paper). LRT_SM.xls / SM_LRT.xls : Results for the Site Model tests. 3. Methods Brief Genomes: Genomic data were sourced from NCBI (see paper for specific assembly accessions) . Annotation : Initial identification was performed using insectOR and Exonerate , followed by manual curation of gene models. Trees : Alignements was performed using MAFFT and ML trees using IQ-TREE . Selection Tests : Positive selection was assessed using PAML (codeml, EasyCodeML) on codon-aligned sequences. 4. Species Included Rhodnius bretesi Rhodnius colombiensis Rhodnius (=Psammolestes) coreodes Rhodnius domesticus Rhodnius milesi Rhodnius montenegrensis Rhodnius nastutus Rhodnius neglectus Rhodnius neivai Rhodnius pallescens Rhodnius pictipes Rhodnius prolixus Rhodnius robustus 5. Usage and Citation If you use these data, please cite the original publication: Merle, M. et al. (2026). Evolutionary Dynamics of the Complete Chemosensory Repertoire in Kissing Bugs of the Genus Rhodnius: Divergent Odorant Receptors Contrast with Conserved Gene Families. ( in prep ) For the specific dataset version, you can also cite this Zenodo DOI: DOI: 10.5281/zenodo.19064793

open·CC-BY-4.0·Zenodo·completeSource
declared

Genome Draft of Cardita leana (Archiheterodonta; Bivalvia)

0.00

Formaggioni, Alessandro

6 files · 3.4 GB · fasta, gff, zipdeclared

Genome assembly of Cardita leana (Bivalvia: Archiheterodonta) and the associated gene models predicted with AUGUSTUS. The 'Tree' folder contains species trees inferred using different tree reconstruction programs. The MCMC_analysis folder contains inputs and ouputs for every MCMCTree analysis

open·CC-BY-4.0·Zenodo·completeSource
declared

Dataset used for validating the BITSER tool

0.00

Costa Fuganti, Lucas · Lopes, Fabricio M. · Nunes da Rocha, Ulisses · et al.

15 files · 212 MB · fastadeclared

Dataset used for validating the BITSER tool, composed of data from the viruses SARS-CoV-2 (previously tested using the KEVOLVE method, {lebatteux2024machine}), DENV (tested by the GRAMEP method {pimenta2025gramep}), and HBV (extracted from HBVdb {Hayer2012}).

open·CC-BY-4.0·Zenodo·completeSource

Select a result to see its full details here: the measured structure, quality, and the loader, without leaving your search.