Exploration

ResearchFeatured

Discovery

DiscoverSourcesQuality

Analysis

Working setReviews
Flow StudioTeamConcept
Settings

Partners

  • AI AlliancePrime
  • BrightQueryBuilds Meridian
  • OpenMinedFunded partner
  • MLCommonsFunded partner
  • Hugging FaceDeployment platform
See the full consortium and what each partner wires

Meridian is the discovery layer for research data, built by BrightQuery within the AI Alliance.

hybrid · semantic + lexical · 5 datasets ranked · 1.42s

Structuresequence5
Depthmeasured5
Licenseopen5
Accessopen5
Formatfasta4tsv1vcf1
Sourcezenodo-bio5
clear
1-5 of 5sortrelevancemeasured firstqualitysize
sequence

Database of virus genomes from ultra-deep sequencing of wastewater (WVDB)

0.00

Kantor, Rose · Shakya, Migun · Ruth, Nelson · et al.

2,095 rows · 907 KB · fasta, tsv

A virus genome database representing 21,015 near-complete virus genomes collected from untargeted ultra-deep RNA/DNA combined sequencing of wastewater. Sequence data was provided by the CASPER consortium and raw data may be found on NCBI SRA under bioprojects PRJNA1247874 and PRJNA1198001. Data underwent read trimming, rRNA and human read removal, de novo assembly, and selection of high-quality viral contigs. Contigs were clustered at 95% identity and 85% query coverage to dereplicate. Chimera-checking required at least two independent assemblies of the same viral genome or presence of the genome in another reference database. Annotation made use of RdRpCATCH, geNomad, checkV, BLASTN against NCBI core-nt, and RNAVirHost. The RdRp fasta files contain representative RdRp sequences identified through homology to major RdRp reference databases and clustered at 90% sequence identity over 75% sequence coverage. Included sequences contain all three conserved RdRp motifs (A, B, and C) arranged in either the canonical ABC configuration or the permuted CAB configuration.

open·CC-BY-4.0·zenodo-bio·completeSource
sequence

Panel Information Files for "PvGAP: Development of a Globally Applicable, Highly Multiplexed Microhaplotype Amplicon Panel for Plasmodium vivax"

0.00

Hubbard, Alfred · Solares, Edwin · Hemming-Schroeder, Elizabeth

88 rows · 18 KB · fasta

These are the files needed to run the Broad Institute's malaria amplicon pipeline for the PvGAP Plasmodium vivax panel, described in detail here . They consist of FASTA files containing the forward and reverse primers and another FASTA file containing reference sequences for each target, derived from the PvP01 reference genome.

open·CC-BY-4.0·zenodo-bio·completeSource
sequence

Trypanosoma cruzi (Dm28c) genome

0.00

Requena Rolanía, Jose María · Greif, Gonzalo · ROBELLO, CARLOS

1 files · 8.0 MB · fasta

This dataset contains the genome sequence for Trypanosoma cruzi (strain Dm28c). This genome sequence was de novo assembled using PacBio Hi-Fi and Illumina sequencing platforms by Greif et al (2026. PMID: 41501640). The genome was assembled into 32 contigs, which represent complete chromosomes. The provided Fasta file also contains an additional contig corresponding to the maxicircle (mitochondrial genome) sequence. The Fasta files included in this dataset were downloaded from GenBank (assembly GCA_044048535.1; May 22, 2026). Additional information about the Dm28cT2T genome assembly and gene annotations may be accessed through the link: https://cruzi.pasteur.uy/

open·CC-BY-4.0·zenodo-bio·completeSource
sequence

Multiple sequence alignment, phylogenetic tree, and domain-level annotation of Cas7 homologs

0.00

Burman, Nathaniel · Buyukyoruk, Murat · Wiegand, Tanner · et al.

4 files · 8.0 MB · fasta

This folder contains a multiple sequence alignment of Cas7 homologs in .fasta format, the domain-level annotations from PFAM and CasFinder, and an associated phylogenetic tree in .newick format.

open·CC-BY-4.0·zenodo-bio·completeSource
sequence

azure fox cleaned vcf file

0.00

Omukuti, Rodney

1 files · 8.0 MB · vcf

open·CC-BY-4.0·zenodo-bio·completeSource
closeopen full
tabular · zenodo

Datasets for "Refining simulated mineral dust composition through modified size distributions: dual validation with mineral-specific and elemental observations"

Gómez Maqueo Anaya, Sofía · Dos Santos Souza, Eduardo José · Fomba, Khanneh Wadinga · Aryasree, Sudharaj · Kandler, Konrad

measured·open·3 files
Measuredstructure observed by touching the bytes
topology
tabular
records
116
sampled
full dataset
null rate
61.0%
duplicate rate
0.0%
size
7.0 KB
profiler
tabular:stdlib/v1

Measured the full file.

Columns (16)
columntypenullsdistribution
MINERALS MEASUREMENTScategorical string80%23 distinct · Jan-February 1984 - 1985 · Adedokum et al. (1989) · 29 July 2002
coordinatescategorical string79%21 distinct · Santa Cruz de Tenerife, Spain · 28°19'N, 16°30'W · Ile-Ife, Nigeria
mineralscategorical string2%16 distinct · quartz · kaolinite · calcite
measured (%)numeric number2%0.16 - 74.78 · med 8
stdDev(%)
numeric number
82%
0 - 15.77 · med 2.05
model (%)numeric number16%0.6 - 80.5 · med 8.785
model PSD mod (%)numeric number0%0.4 - 43.1 · med 12.15
percentual changenumeric number16%-67.17 - 237.5 · med -3.836
size rangecategorical string0%10 distinct · bulk · clay · silt
col9categorical string100%0 distinct
col10categorical string100%0 distinct
col11categorical string100%0 distinct
col12categorical string100%0 distinct
col13categorical string100%0 distinct
col14categorical string100%0 distinct
col15categorical string100%0 distinct
Columns by kind
Columns with missing values

Measured distributions

measured (%)
p10 1 · median 8 · p90 26.43
stdDev(%)
p10 0.464 · median 2.05 · p90 8.7
model (%)
p10 1.07 · median 8.785 · p90 52.52
model PSD mod (%)
p10 1.021 · median 12.15 · p90 32.98
percentual change
p10 -65.61 · median -3.836 · p90 124.7