Exploration

ResearchFeatured

Discovery

DiscoverSourcesQuality

Analysis

Working setReviews
Flow StudioTeamConcept
Settings

Partners

  • AI AlliancePrime
  • BrightQueryBuilds Meridian
  • OpenMinedFunded partner
  • MLCommonsFunded partner
  • Hugging FaceDeployment platform
See the full consortium and what each partner wires

Meridian is the discovery layer for research data, built by BrightQuery within the AI Alliance.

hybrid · semantic + lexical · 8 datasets ranked · 0.97s

Structurecomposite6sequence1tabular1
Depthmeasured8
Licenseopen8
Accessopen8
Formatzip6tsv2csv1docx1fasta1
Sourcezenodo-bio8
clear
1-8 of 8sortrelevancemeasured firstqualitysize
composite

Transcriptomic reads and mapping-derived coverage for the UG5 (DRT11) genomic island of Sinorhizobium meliloti RMO17

0.00

Toro, Nicolas · Molina-Sánchez, Maria Dolores

gzip1

1 files · 1.9 MB · zip

This dataset contains the sequencing reads and mapping-derived coverage files corresponding to the UG5 genomic island of Sinorhizobium meliloti RMO17. Paired-end reads mapped to the UG5 region were extracted and processed using Bowtie2, Samtools, Bedtools and deepTools. The dataset includes raw FASTQ.gz files (R1, R2 and unpaired), the reference sequence of the UG5 genomic island (FASTA), genomic annotations (BED), genome size file, and all mapping-derived products (sorted BAM/BAl, bedGraph, and bigWig files with raw and CPM-normalized coverage). The dataset is organised in a structured directory (raw_reads, reference, mapping_products, metadata) to facilitate reuse and reproducibility. This resource supports the analyses reported in the associated manuscript.

open·CC-BY-4.0·zenodo-bio·completeSource
sequence

Database of virus genomes from ultra-deep sequencing of wastewater (WVDB)

0.00

Kantor, Rose · Shakya, Migun · Ruth, Nelson · et al.

2,095 rows · 907 KB · fasta, tsv

A virus genome database representing 21,015 near-complete virus genomes collected from untargeted ultra-deep RNA/DNA combined sequencing of wastewater. Sequence data was provided by the CASPER consortium and raw data may be found on NCBI SRA under bioprojects PRJNA1247874 and PRJNA1198001. Data underwent read trimming, rRNA and human read removal, de novo assembly, and selection of high-quality viral contigs. Contigs were clustered at 95% identity and 85% query coverage to dereplicate. Chimera-checking required at least two independent assemblies of the same viral genome or presence of the genome in another reference database. Annotation made use of RdRpCATCH, geNomad, checkV, BLASTN against NCBI core-nt, and RNAVirHost. The RdRp fasta files contain representative RdRp sequences identified through homology to major RdRp reference databases and clustered at 90% sequence identity over 75% sequence coverage. Included sequences contain all three conserved RdRp motifs (A, B, and C) arranged in either the canonical ABC configuration or the permuted CAB configuration.

open·CC-BY-4.0·zenodo-bio·completeSource
tabular

Microbiota study IgG4-RD AG Chang

0.00

Budzinski, Lisa · Beenken, Anne Elisabeth · Sempert, Toni · et al.

9 rows × 1 cols · 743 B · csv, docx, zip

1 categorical

We have investigated an IgG4-RD (IgG4-RD) cohort by our multi-parameter microbiota flow cytometry approach to characterise the microbiota on single-cell level for attributes of the disease. The microbiota is isolated from stool samples and stained according to the published protocol for (a) host immunoglobulins IgA1, IgA2, IgM, IgG and (b) agglutinin binding to mannose, galactose or N-Acetyl-glucosamine surface sugar moieties. For all samples we also determined the microbiome composition by 16S rRNA (V3-V4) sequencing on the illumina MiSeq platform. We provide the raw .fcs and FASTQ files of 40 IgG4-RD patients. For comparison we additionally analysed 36 healthy donors. All .fcs files were generated on BD Influx®. The metadata is collected in the provided meta.csv. The staining parameters are summarized in provided panel.csv.

open·CC-BY-4.0·zenodo-bio·0% null·completeSource
composite

VICMpred: SVM-Based Prediction of Functional Proteins of Gram-Negative Bacteria Using Amino Acid Patterns and Composition

0.00

Saha, Sudipto · Raghava, Gajendra

1 files · 186 KB · zip

VICMpred: SVM-Based Prediction of Functional Proteins of Gram-Negative Bacteria Using Amino Acid Patterns and Composition VICMpred is a computational web server developed for predicting the major functional classes of Gram-negative bacterial proteins from amino acid sequences. The tool classifies Gram-negative bacterial proteins into four broad functional categories: virulence factors, information molecules, cellular process proteins, and metabolism-related proteins. VICMpred uses support vector machine-based models trained on amino acid composition, dipeptide composition, and class-specific tetrapeptide patterns. Web Server: https://webs.iiitd.edu.in/raghava/vicmpred/ Citation Saha, S., and Raghava, G. P. S. VICMpred: An SVM-based method for the prediction of functional proteins of Gram-negative bacteria using amino acid patterns and composition. Genomics, Proteomics & Bioinformatics, 4(1), 42-47, 2006. https://doi.org/10.1016/S1672-0229(06)60015-6 About the Research Functional annotation of proteins is one of the major challenges in the post-genomic era. Due to the rapid growth of protein sequence databases, experimental functional characterization of every newly discovered protein is not practical. Traditional methods such as BLAST, FASTA, and PSI-BLAST depend on sequence similarity. However, proteins with similar functions may show poor sequence similarity, making direct function prediction difficult. VICMpred was developed as a direct function prediction method for Gram-negative bacterial proteins. Instead of only predicting subcellular localization, it predicts broad biological functions directly from protein sequence features. Data Compilation: The final dataset contained 670 non-redundant Gram-negative bacterial proteins. These included 255 cellular process proteins, 60 information molecules, 285 metabolism proteins, and 70 virulence factors. Methodology: VICMpred uses support vector machine-based models trained on amino acid composition, dipeptide composition, PSI-BLAST similarity search, class-specific tetrapeptide patterns, and hybrid combinations of these features.

open·MIT·zenodo-bio·completeSource
composite

HQ MAGs (CheckM comp≥90%, contam≤5%) from the MicroToxBol Bolivian human gut microbiome cohort

0.00

Manghi, Paolo

2 files · 8.0 MB · gzip, tsv

Using gene-level and species-level shotgun metagenomics, we provide the first characterization of the rural, Bolivian microbiome; we identified microbial genes which strongly correlate (rho>0.45) with arsenic in urine, and that overall contribute to substantiate that the gut microbiome helps tolerate arsenic via a evict-out-of-house mechanism. Mediation analysis, followed by phylogenetic investigation of metagenomic-assembled genomes, further strengthens this observation. This study elucidates the role of the microbiome in helping to tolerate arsenic-rich environments, and paves the way for probiotic interventions that may mitigate the effects of this toxic metal. This repository contains 2,478 HQ MAGs from the MicroToxBol cohort in fasta format and a descriptive table comprising taxonomic annotation, quality-checks, and coverage estimation.

open·CC-BY-4.0·zenodo-bio·completeSource
composite

Single nuclear RNA sequencing from human endomyocardial biopsy (IVIG / Placebo treated) - raw/feature barcode matrix

0.00

Sikking, Maurits · Peisker, Fabian · Maatz, Henrike · et al.

1 files · 100 MB · zip

Project description: See related publication Code repository of the related publication: https://github.com/fpeisker303/IVIG_snRNA_project/ Methods use to generate the Single nuclear RNA sequencing data Endomyocardial biopsies (EMB) were taken from the right ventricular septum and collected via the internal jugular vein using a transcatheter bioptome (Cordis, Miami, FL., USA) at baseline before the IVIg treatment and at the standardized six-months follow-up timepoint of the original study (i.e., median 6.4 [5.9-7.3] months). EMB were evaluated regarding viral persistent and immunohistology markers of inflammation and fibrosis. Spare cardiac biopsies were stored at -80°C until preparation of snRNA sequencing. The isolation of cardiac nuclei and the 10x library preparation were performed at the Max Delbrück Center for Molecular Medicine following a published protocol (1) with adaptations to low-sized tissue pieces (2). In brief, 1-4-mg-sized flash-frozen cardiac biopsies were placed in a pre-cooled dish and an equally sized droplet of homogenization buffer (250 mM sucrose, 25 mM KCl, 5 mM MgCl 2 , 10 mM Tris-HCl, 1 μM DTT, 1× protease inhibitor, 0.4 U μl -1 RNaseIn, 0.2 U μl -1 SUPERaseIn and 0.1% Triton X-100 in nuclease-free water) was added. Buffer-encapsulated tissue pieces were sliced with a scalpel. The tissue pieces were then transferred to a 7-ml glass Dounce tissue grinder (Merck), and nuclei were isolated and stained with NucBlue Live ReadyProbes Reagent (Thermo Fisher Scientific). Hoechst + single nuclei were sorted via fluorescence-activated cell sorting (FACS) (BD Biosciences, FACSAria Fusion). Purity and integrity of nuclei were confirmed microscopically, and nuclei numbers were counted using a Countess II (Life Technologies) before processing with the Chromium Controller (10x Genomics) per the manufacturer's protocol. Single-nucleus 3' gene expression libraries were created using version 3.1 Chromium Single Cell Reagent Kits (10x Genomics) following the manufacturer's instructions. cDNA library quality control was performed using Bioanalyzer High Sensitivity DNA Analysis (Agilent Technologies) and a KAPA Library Quantification Kit. cDNA libraries were sequenced on an Illumina NovaSeq with a targeted read number of 30,000-50,000 reads per nucleus. Fastq files with sequencing results were processed using cellranger version 6.1.2 with the GRCh38-2020-A reference provided by 10x Genomics. References 1. Nadelmann ER, Gorham JM, Reichart D, Delaughter DM, Wakimoto H, Lindberg EL, et al. Isolation of Nuclei from Mammalian Cells and Tissues for Single-Nucleus Molecular Profiling. Curr Protoc. 2021;1(5):e132. 2. Maatz H, Lindberg EL, Adami E, López-Anguita N, Perdomo-Sabogal A, Cocera Ortega L, et al. The cellular and molecular cardiac tissue responses in human inflammatory cardiomyopathies after SARS-CoV-2 infection and COVID-19 vaccination. Nat Cardiovasc Res. 2025;4(3):330-45.

open·CC-BY-4.0·zenodo-bio·completeSource
composite

Raw molecular data for "Phylogeny and biogeographic history of monkey moths (Lepidoptera: Eupterotidae)"

0.00

Li, Xuankun · Tu, Yuezheng · Plotkin, David · et al.

5 files · 100 MB · zip

Raw molecular data for 90 newly-sequenced specimens of Eupterotidae (and related families of Lepidoptera) used to generate a molecular phylogeny for the study "Phylogeny and biogeographic history of monkey moths (Lepidoptera: Eupterotidae)" (currently under peer review as of May 2026). File names contain the sequence ID and taxonomic information for each specimen. Data files are in fastq format and have been compressed; there are two files per specimen (labelled "R1" and "R2"). Data files have been organized into five .zip files, based on family-group taxonomy, as follows: Eupterotidae: Eupterotinae (25 specimens, 50 files) Eupterotidae: Ganisa Group (15 specimens, 30 files) Eupterotidae: Janinae (27 specimens, 54 files) Eupterotidae: Striphnopteryginae (15 specimens, 30 files) Other Lepidoptera families (outgroup taxa): Anthelidae, Bombycidae, Lasiocampidae, Saturniidae (8 specimens, 16 files)

open·CC-BY-4.0·zenodo-bio·completeSource
composite

Data and Simulation Files for: "High-Throughput Characterization of Transmembrane Helix Partitioning in Membrane Domains"

0.00

Lolicato, Fabio · Javanainen, Matti

1 files · 100 MB · zip

This repository contains the data, simulation files, and scripts used in the study "High-Throughput Characterization of Transmembrane Helix Partitioning in Membrane Domains." The repository includes: Scripts used to extract FASTA sequences from the Orientations of Proteins in Membranes database (OPM). Scripts and input files used to generate peptide systems and prepare the molecular dynamics simulations. Simulation parameter files, topology files, coordinate files required to reproduce the simulations. Final structure files ( md.gro ) for each simulation system. The deposited material is intended to ensure transparency, reproducibility, and reuse of the computational workflow and simulation datasets associated with this work.

open·CC-BY-4.0·zenodo-bio·completeSource

Select a result to see its full details here: the measured structure, quality, and the loader, without leaving your search.