Exploration

ResearchFeatured

Discovery

DiscoverSourcesQuality

Analysis

Working setReviews
Flow StudioTeamConcept
Settings

Partners

  • AI AlliancePrime
  • BrightQueryBuilds Meridian
  • OpenMinedFunded partner
  • MLCommonsFunded partner
  • Hugging FaceDeployment platform
See the full consortium and what each partner wires

Meridian is the discovery layer for research data, built by BrightQuery within the AI Alliance.

hybrid · semantic + lexical · 24 datasets ranked · 0.68s

Structuremodal15composite9
Depthmeasured24
Licenseopen24
Accessopen24
Formatpdf15gzip9csv1gff1tsv1
Sourcezenodo-bio24
clear
1-20 of 24sortrelevancemeasured firstqualitysize
composite

HSeeker Benchmark FASTA Datasets

0.00

HSeeker Contributors

xlsx1

7 files · 7.5 MB · gzip

Synthetic FASTA files used by the HSeeker benchmark suite (benchmarks/benchmark.py). Generated deterministically with fixed NumPy/Python random seeds documented in SEED_MANIFEST.json. Four sequence profiles: uniform (25 % each ACGT), ga_biased (90 % purine), ct_biased (90 % pyrimidine), realistic (~41 % GC). Two size tiers: small (~30 MB) and medium (~300 MB). Download via: python benchmarks/benchmark.py --from-zenodo

open·CC-BY-4.0·zenodo-bio·completeSource
modal

PD5D long read DNA-seq

0.00

Kim, Kwanho · Lin, Zechuan · Simmons, Sean · et al.

1 files · 37 KB · pdf

This dataset contains CCS corrected HiFi long-read DNA sequencing (lrDNAseq) in FASTQ format for 100 PMDBS samples from Parkinson's patients and healthy controls. It's part of the PD5D atlas, where the same subjects were also profiled with other omics assays including genotyping, single-cell ATACseq, and spatial transcriptomics.

open·CC-BY-4.0·zenodo-bio·completeSource
modal

PD5D midbrain single-nucleus RNA-seq hybrid selection

0.00

Kim, Kwanho · Lin, Zechuan · Simmons, Sean · et al.

1 files · 37 KB · pdf

This dataset contains raw FASTQ files from the midbrain single-nucleus RNA sequencing (snRNAseq) dataset with hybrid selection for the matching PMDBS samples from the PD5D chort. The same subjects were also profiled with other omics assays including genomic DNAseq, genotyping, single-cell ATACseq, and spatial transcriptomics.

open·CC-BY-4.0·zenodo-bio·completeSource
modal

Single nuclei RNA sequencing (10x) of postmortem cingulate cortex and midbrain of healthy donors and Parkinson's disease patients – 10x snRNA-seq.

0.00

Pančíková, Alexandra · Theunis, Koen · Hulselmans, Gert · et al.

1 files · 37 KB · pdf

This dataset consists of raw sequencing snRNA-seq data (10x Genomics Chromium Next GEM Single Cell 3ʹ). The data is part of an overall set of samples derived from postmortem midbrain (n=140), cingulate cortex (n=190) and motor cortex (n=4) of healthy donors (n=114), patients with Parkinson's disease (n=75) or patients with other neurological disorder (n=1). The protocol followed to isolate nuclei from postmortem brain samples and to prepare sequencing libraries can be found below. To increase throughput and to decrease batch effects, several donors have been pooled together into a single sequencing library. To computationally demultiplex the nuclei to their corresponding donors, cellsnp-lite (version commit: aad18644adcde853c313362a856a24245c9b91f7) followed by vireo (https://github.com/single-cell-genetics/vireo/pull/108) has been used. The population VCF with the donor genotypes derived from whole genome sequencing data has been used to assign nuclei back to their donors.

open·CC-BY-4.0·zenodo-bio·completeSource
modal

Single nuclei RNA sequencing (ParseBio) of postmortem cingulate cortex and midbrain of healthy donors and Parkinson's disease patients.

0.00

Pančíková, Alexandra · Theunis, Koen · Hulselmans, Gert · et al.

1 files · 36 KB · pdf

This dataset consists of raw sequencing snRNA-seq data using ParseBio Evercode Whole Transcriptome. The data is part of an overall set of samples derived from postmortem midbrain (n=140), cingulate cortex (n=190) and motor cortex (n=4) of healthy donors (n=114), patients with Parkinson's disease (n=75) or patients with other neurological disorder (n=1). The protocol followed to isolate nuclei from postmortem brain samples and to prepare sequencing libraries can be found below. To increase throughput and to decrease batch effects, several donors have been pooled together into a specific ParseBio barcode. To computationally demultiplex the nuclei to their corresponding donors, cellsnp-lite (version commit: aad18644adcde853c313362a856a24245c9b91f7) followed by vireo (https://github.com/single-cell-genetics/vireo/pull/108) has been used. The population VCF with the donor genotypes derived from whole genome sequencing data has been used to assign nuclei back to their donors.

open·CC-BY-4.0·zenodo-bio·completeSource
modal

Single nuclei ATAC sequencing of postmortem cingulate cortex and midbrain of healthy donors and Parkinson's disease patients – 10x snATAC-seq.

0.00

Pančíková, Alexandra · Theunis, Koen · Hulselmans, Gert · et al.

1 files · 36 KB · pdf

This dataset consists of raw sequencing snATAC-seq data (10x Genomics Chromium Next GEM Single Cell ATAC v2). The data is part of an overall set of samples derived from postmortem midbrain (n=140), cingulate cortex (n=190) and motor cortex (n=4) of healthy donors (n=114), patients with Parkinson's disease (n=75) or patients with other neurological disorder (n=1). The protocol followed to isolate nuclei from postmortem brain samples and to prepare sequencing libraries can be found below. To increase throughput and to decrease batch effects, several donors have been pooled together into a single sequencing library. To computationally demultiplex the nuclei to their corresponding donors, cellsnp-lite (version commit: aad18644adcde853c313362a856a24245c9b91f7) followed by vireo (https://github.com/single-cell-genetics/vireo/pull/108) has been used. The population VCF with the donor genotypes derived from whole genome sequencing data has been used to assign nuclei back to their donors. (edited)

open·CC-BY-4.0·zenodo-bio·completeSource
modal

Single nuclei ATAC sequencing of postmortem cingulate cortex of healthy donors and Parkinson's disease patients – HyDrop-ATAC v2

0.00

Pančíková, Alexandra · Theunis, Koen · Hulselmans, Gert · et al.

1 files · 36 KB · pdf

This dataset consists of raw sequencing snATAC-seq data (HyDrop v2). The data is part of an overall set of samples derived from postmortem midbrain (n=140), cingulate cortex (n=190) and motor cortex (n=4) of healthy donors (n=114), patients with Parkinson's disease (n=75) or patients with other neurological disorder (n=1). The protocol followed to isolate nuclei from postmortem brain samples can be found below. To increase throughput and to decrease batch effects, several donors have been pooled together into a single sequencing library. To computationally demultiplex the nuclei to their corresponding donors, cellsnp-lite (version commit: aad18644adcde853c313362a856a24245c9b91f7) followed by vireo (https://github.com/single-cell-genetics/vireo/pull/108) has been used. The population VCF with the donor genotypes derived from whole genome sequencing data has been used to assign nuclei back to their donors.

open·CC-BY-4.0·zenodo-bio·completeSource
modal

Single nuclei ATAC sequencing of postmortem cingulate cortex and midbrain of healthy donors and Parkinson's disease patients – Scale-ATAC + 10x Genomics.

0.00

Pančíková, Alexandra · Theunis, Koen · Hulselmans, Gert · et al.

1 files · 36 KB · pdf

This dataset consists of raw sequencing ATAC-seq data (Scale-ATAC pre-indexing followed by 10x Genomics snATAC v2). The data is part of an overall set of samples derived from postmortem midbrain (n=140), cingulate cortex (n=190) and motor cortex (n=4) of healthy donors (n=114), patients with Parkinson's disease (n=75) or patients with other neurological disorder (n=1). The protocol followed to isolate nuclei from postmortem brain samples and to prepare sequencing libraries can be found below. To increase throughput and to decrease batch effects, several donors have been pooled together into a single sequencing library. To computationally demultiplex the nuclei to their corresponding donors, cellsnp-lite (version commit: aad18644adcde853c313362a856a24245c9b91f7) followed by vireo (https://github.com/single-cell-genetics/vireo/pull/108) has been used. The population VCF with the donor genotypes derived from whole genome sequencing data has been used to assign nuclei back to their donors.

open·CC-BY-4.0·zenodo-bio·completeSource
modal

Single nuclei ATAC sequencing of postmortem cingulate cortex of healthy donors – Scale-ATAC + HyDrop v2.

0.00

Pančíková, Alexandra · Theunis, Koen · Hulselmans, Gert · et al.

1 files · 36 KB · pdf

This dataset consists of raw sequencing ATAC-seq data (Scale-ATAC + HyDrop v2). The data is part of an overall set of samples derived from postmortem midbrain (n=140), cingulate cortex (n=190) and motor cortex (n=4) of healthy donors (n=114), patients with Parkinson's disease (n=75) or patients with other neurological disorder (n=1). The protocols followed to isolate nuclei from postmortem brain samples and to prepare sequencing libraries can be found below. To increase throughput and to decrease batch effects, several donors have been pooled together into a single sequencing library. To computationally demultiplex the nuclei to their corresponding donors, cellsnp-lite (version commit: aad18644adcde853c313362a856a24245c9b91f7) followed by vireo (https://github.com/single-cell-genetics/vireo/pull/108 ) has been used. The population VCF with the donor genotypes derived from whole genome sequencing data has been used to assign nuclei back to their donors.

open·CC-BY-4.0·zenodo-bio·completeSource
modal

Single nuclei RNA sequencing of postmortem cingulate cortex, midbrain and motor cortex of healthy donors and Parkinson's disease patients – 10x multiome (snRNA-seq and snATAC-seq).

0.00

Pančíková, Alexandra · Theunis, Koen · Hulselmans, Gert · et al.

1 files · 37 KB · pdf

This dataset consists of raw sequencing snRNA-seq data and snATAC-seq data (10x Genomics Chromium Next GEM Multiome ATAC/GEX). The data is part of an overall set of samples derived from postmortem midbrain (n=140), cingulate cortex (n=190) and motor cortex (n=4) of healthy donors (n=114), patients with Parkinson's disease (n=75) or patients with other neurological disorder (n=1). The protocol followed to isolate nuclei from postmortem brain samples and to prepare sequencing libraries can be found below. To increase throughput and to decrease batch effects, several donors have been pooled together into a single sequencing library. To computationally demultiplex the nuclei to their corresponding donors, cellsnp-lite (version commit: aad18644adcde853c313362a856a24245c9b91f7) followed by vireo (https://github.com/single-cell-genetics/vireo/pull/108 ) has been used. The population VCF with the donor genotypes derived from whole genome sequencing data has been used to assign nuclei back to their donors. (edited)

open·CC-BY-4.0·zenodo-bio·completeSource
modal

Single-cell RNAseq of human PBMCs from healthy control, RBD, and PD.

0.00

MacDonald, Adam · Stratton, Jo Anne

1 files · 35 KB · pdf

We performed 10X Genomics single-cell RNAsequencing of human prepheral blood mononuclear cells from healthy control, PD and RBD patients. This dataset contains raw FASTQ files. Sequencing was performed using NovaSeq 6000 S4 PE 100bp. Reads were processed using the 10X Genomics Cell Ranger Single Cell 2.0.0 pipeline. FASTQs generated from sequencing output were aligned to the human GRCh38 reference genome using STAR algorithm 2.7.3a.

open·CC-BY-4.0·zenodo-bio·completeSource
modal

Bulk RNAseq of mouse substantia nigra of wild type and PINK1 KO mice after C.rodentium infection

0.00

Stratton, Jo Anne · Mukherjee, Sriparna · Trudeau, Louis-Eric

1 files · 5.0 KB · pdf

We performed bulk RNAsequencing of substantia nigra from wild type and PINK1 KO mice 26-days post C.rodentium infection. This dataset contains raw fastq files from striatal cells, sorted into 4 groups namely wild type and PINK1 KO uninfected and infected mice. Sequencing was performed using NextSeq500. FASTQs generated from sequencing output were aligned to the mm10 reference genome using STAR aligner.

open·CC-BY-4.0·zenodo-bio·completeSource
modal

Bulk RNAseq of mouse striatum of wild type and PINK1 KO mice after C.rodentium infection

0.00

Stratton, Jo Anne · Mukherjee, Sriparna · Trudeau, Louis-Eric

1 files · 6.1 KB · pdf

We performed bulk RNAsequencing of striatum from wild type and PINK1 KO mice 26-days post C.rodentium infection. This dataset contains raw fastq files from striatal cells, sorted into 4 groups namely wild type and PINK1 KO uninfected and infected mice. Sequencing was performed using NextSeq500. FASTQs generated from sequencing output were aligned to the mm10 reference genome using STAR aligner.

open·CC-BY-4.0·zenodo-bio·completeSource
modal

Single-cell RNAseq of human iPSC-derived wild type and PINK1 KO myeloid cells after lipopolysaccharide and interleukin-1 beta challenge

0.00

Recinto, Sherilyn · Stratton, Jo Anne

1 files · 6.1 KB · pdf

We performed 10X Genomics single-cell RNAsequencing of human iSPC-derived monocytes and macrophages in vitro. Cells were treated with 500 ng/mL lipopolysaccharide (LPS) and 50 ng/mL interleukin-1 beta (IL1b) for 24 hours. This dataset contains raw FASTQ files from myeloid cells, sorted into 4 groups namely monocytes (Mono) and Macrophages (Mac) non-stimulated (NS) and LPS+IL1b-stimulated cells. Sequencing was performed using NovaSeq 6000 S4 PE 100bp. Reads were processed using the 10X Genomics Cell Ranger Single Cell 2.0.0 pipeline. FASTQs generated from sequencing output were aligned to the human GRCh38 reference genome using STAR algorithm 2.7.3a.

open·CC-BY-4.0·zenodo-bio·completeSource
modal

Single-cell RNAseq of lamina propria in wild type and LRRK2 G2019S mice after C. rodentium infection

0.00

Recinto, Sherilyn · Stratton, Jo Anne · Pei, Jessica · et al.

2 files · 6.0 KB · pdf

We performed 10X Genomics single-cell RNAsequencing of colonic lamina propria cells from wild type and LRRK2 G2019S mice following 1-week post C. rodentium infection. The cells were pooled from 3 mice per group of both sexes at 8-12 weeks of age. This dataset contains raw FASTQ files from mouse colonic lamina propria, sorted into 4 groups namely wild type and LRRK2 G2019S uninfected and infected mice. Sequencing was performed using NovaSeq 6000 S4 PE 100bp. Reads were processed using the 10X Genomics Cell Ranger Single Cell 2.0.0 pipeline. FASTQs generated from sequencing output were aligned to the mouse GRCm38 reference genome using STAR algorithm 2.7.3a.

open·CC-BY-4.0·zenodo-bio·completeSource
modal

Single-cell RNAseq of colonic lamina propria in wild type and PINK1 KO mice after C. rodentium infection

0.00

Recinto, Sherilyn · Stratton, Jo Anne

1 files · 5.7 KB · pdf

We performed 10X Genomics single-cell RNA sequencing of colonic lamina propria cells from wild type and PINK1 KO mice following either 1-week or 2-weeks post C. rodentium infection. The cells were pooled from 3 mice per group of both sexes at 8-12 weeks of age. This dataset contains raw FASTQ files from mouse colonic lamina propria, sorted into 8 groups namely wild type and PINK1 KO uninfected and infected mice at 1- or 2-weeks post-infection. Sequencing was performed using NovaSeq 6000 S4 PE 100bp. Reads were processed using the 10X Genomics Cell Ranger Single Cell 2.0.0 pipeline. FASTQs generated from sequencing output were aligned to the mouse GRCm38 reference genome using STAR algorithm 2.7.3a.

open·CC-BY-4.0·zenodo-bio·completeSource
composite

HQ MAGs (CheckM comp≥90%, contam≤5%) from the MicroToxBol Bolivian human gut microbiome cohort

0.00

Manghi, Paolo

2 files · 8.0 MB · gzip, tsv

Using gene-level and species-level shotgun metagenomics, we provide the first characterization of the rural, Bolivian microbiome; we identified microbial genes which strongly correlate (rho>0.45) with arsenic in urine, and that overall contribute to substantiate that the gut microbiome helps tolerate arsenic via a evict-out-of-house mechanism. Mediation analysis, followed by phylogenetic investigation of metagenomic-assembled genomes, further strengthens this observation. This study elucidates the role of the microbiome in helping to tolerate arsenic-rich environments, and paves the way for probiotic interventions that may mitigate the effects of this toxic metal. This repository contains 2,478 HQ MAGs from the MicroToxBol cohort in fasta format and a descriptive table comprising taxonomic annotation, quality-checks, and coverage estimation.

open·CC-BY-4.0·zenodo-bio·completeSource
composite

AmarylOmicBase: An integrated transcriptome dataset for comparative analysis of Amaryllidoideae species

0.00

Goncalves dos Santos, Karen Cristine · Desgagné-Penix, Isabel · Merindol, Natacha

16 files · 8.0 MB · gzip

General description Compiled dataset of transcriptome assemblies, transcriptome annotation and expression quantification for 27 Amaryllidoideae species and 4 hybrid cultivars: Amaryllis belladonna , Clivia miniata , Crinum asiaticum , Crinum x powellii , Galanthus elwesii , Galanthus sp., Hippeastrum cv. Blossom Peacock, Hippeastrum cv. Jewel, Hippeastrum cv. Royal Velvet, Hippeastrum striatum , Hippeastrum vittatum , Leucojum aestivum , Lycoris aurea , Lycoris chinensis , Lycoris incarnata , Lycoris longituba , Lycoris radiata , Lycoris sprengeri , Narcissus aff pseudonarcissus , Narcissus papyraceus , Narcissus pseudonarcissus , Narcissus cv. Tête-à-Tête, Narcissus tazetta , Narcissus viridiflorus , Phycella aff cyrtanthoides , Rhodophiala pratensis , Scadoxus multiflorus , Traubia modesta , Zephyranthes candida , Zephyranthes carinata , Zephyranthes treatiae. Data includes transcriptome assemblies, Transdecoder predictions of peptide sequences (along with predicted coding seqeunces, GFF3 and BED files), expression quantification (count and TPM matrices for genes and trinity isoforms obtained with Kallisto and from Salmon), and annotation results (from EggNOG, Pfam, Uniprot Swissprot and Rfam), as well as signal peptide and transmemberane domain predictions (from SignalP and TmHMM). Annotations were compiled into a report for each species using Trinotate. For Lycoris aurea and Narcissus cv. Tête-à-Tête, there are two assemblies and corresponding files: Lycoris_aurea_PB and Lycoris_aurea_TH, and Narcissus_TaT_PB Narcissus_TaT_TH. PB assemblies were constructed solely with long-read sequencing data (PacBio, PB); while TH assemblies were constructed with short reads using Trinity, with long-read assembly being used for scaffolding step of Trinity (Trinity Hybrid, TH). Expression quantification was published on NCBI GEO (accessions GSE329951 , GSE329957 , GSE330014 and GSE331457 ). File descriptions All unitigs and predicted protein sequences are prefixed with an acronym to identify the species: Species Acronym NCBI TSA accession Amaryllis belladonna Ambel deposited on GSE331457 Clivia miniata Clmin DBNKRK000000000 Crinum asiaticum Crasi DBNIJL000000000 Crinum x powellii Crpow DBNKRS000000000 Galanthus elwesii Gaelw DBNKRO000000000 Galanthus sp. Gasp DBNKRR000000000 Hippeastrum cv. Blossom Peacock HispBP DBNMYE000000000 Hippeastrum cv. Jewel HispJW DBNMYC000000000 Hippeastrum cv. Royal Velvet HispRV DBNMYD000000000 Hippeastrum striatum Histr DBNIJG000000000 Hippeastrum vittatum Hivit DBNKRF000000000 Leucojum aestivum Leaes DBNIJJ000000000 Lycoris aurea (PB) Lyaur DBNFTY000000000 Lycoris aurea (TH) Lyaur DBNKRI000000000 Lycoris chinensis Lychi DBNKRE000000000 Lycoris incarnata Lyinc DBNKRD000000000 Lycoris longituba Lylon DBNKRG000000000 Lycoris radiata Lyrad deposited on GSE331457 Lycoris sprengeri Lyspr DBNKRL000000000 Narcissus aff pseudonarcissus Naafps DBNKRP000000000 Narcissus papyraceus Napap DBNKRN000000000 Narcissus pseudonarcissus Napse DBNIJK000000000 Narcissus Tête-à-Tête (PB) NaspPB DBNNFO000000000 Narcissus Tête-à-Tête (TH) NaspTH DBNUFN000000000 Narcissus tazetta Nataz DBNPME000000000 Narcissus viridiflorus Navir deposited on GSE331457 Phycella sp. Phsp deposited on GSE331457 Rhodophiala pratensis Rhpra deposited on GSE331457 Scadoxus multiflorus Scmul DBNIJH000000000 Traubia modesta Trmod deposited on GSE331457 Zephyranthes candida Zecan DBNKRH000000000 Zephyranthes carinata Zecar DBNIJI000000000 Zephyranthes treatiae Zetre deposited on GSE331457 Files Files are organized by type of analysis/data, meaning all expression quantification data obtained with Kallisto are compressed into a single file, all results from BlastP are in the same file, etc. Compressed file name Individual file type Content Blastp_Uniprot.tar.gz Tabular 33 files (1 per assembly) generated with BLASTP against Uniprot SwissProt release 2024_04. Files in blast output format 6 (standard columns). EggNOG_Emapper.tar.gz Tabular 33 tabular files (1 per assembly). Generated with eggnog.emapper (annotations file format described in eggnog.emapper's wiki ) Expression_Kallisto.tar.gz Tabular 128 files (4 per assembly, 2 for long-read assemblies): read counts and TPM values for both trinity "genes" and trinity "isoforms"). Row names are "gene"/ "isoform" IDs, column names are SRA run IDs. Expression_Salmon.tar.gz Tabular 128 files (4 per assembly, 2 for long-read assemblies): read counts and TPM values for both trinity "genes" and trinity "isoforms"). Row names are "gene"/ "isoform" IDs, column names are SRA run IDs. Final_assemblies.tar.gz Fasta 33 files (1 per assembly). Assemblies generated in this study, after contamination and expression filtering. HMMScan_PfamA.tar.gz Domain hits table 33 files (1 per assembly). Generated with hmmscan (option ‑‑domtblout, domain hits table, explained in hmmer's user guide [119] ) against the Pfam-A database. Infernal_Rfam.tar.gz Target hits table format 2 33 files (1 per assembly) generated with Infernal's cmscan (using the Trinotate wrapper) against the Rfam database. Table format 2 described in Infernal's user guide section 6 ). Metadata_studies.tar.gz Tabular (semi-colon separated columns) 31 tabular files (one per species/cultivar) indicating: Bioproject ID, SRA run ID, Sample name, Biosample ID, tissue, genotype (cultivar, when specified), treatment, batch, original publication citation, and DOI of original publication. Signalp6.tar.gz Tabular or GFF3 99 files (3 per assembly: prediction_results.txt, output.gff3 and region_output.gff3) generated with SignalP6. TmHMM2.tar.gz Tabular 33 files (1 per assembly) generated with TmHMM2 (short format, described in the guide tab of https://services.healthtech.dtu.dk/services/TMHMM-2.0/ Transcriptomes_unfiltered.tar.gz Fasta 33 files (1 per assembly). Unfiltered (prior to contamination and expression screening) assemblies generated in this study. Transdecoder_bed.tar.gz BED 33 BED files (1 per assembly) generated with TransDecoder.Predict (using BLASTp against Uniprot and hmmscan against Pfam) Transdecoder_cds.tar.gz Fasta 33 fasta files (1 per assembly) of predicted coding sequences generated with TransDecoder.Predict (using BLASTp against Uniprot and hmmscan against Pfam) Transdecoder_gff3.tar.gz GFF3 33 GFF3 files (1 per assembly) generated with TransDecoder.Predict (using BLASTp against Uniprot and hmmscan against Pfam). Transdecoder_proteomes.tar.gz Fasta 33 predicted proteome files (.pep, 1 per assembly) generated with TransDecoder.Predict (using BLASTp against Uniprot and hmmscan against Pfam).

open·CC-BY-4.0·zenodo-bio·completeSource
composite

Structural modularity of receptor-binding proteins underlies host-range strategy diversification in Klebsiella pneumoniae phages

0.00

R. Panicker, Vyshakh · J. Smug, Bogna · Klein-Sousa, Victor · et al.

1 files · 8.0 MB · gzip

Data deposit for: Panicker VR, Smug BJ, Klein-Sousa V, Enright MC, Taylor NMI, Drulis-Kawa Z, Mostowy RJ. "Structural modularity of receptor-binding proteins underlies host-range strategy diversification in Klebsiella pneumoniae phages." bioRxiv 2026. https://doi.org/10.64898/2026.05.12.724579 This archive contains all large binary data files required to reproduce the analyses and figures in the paper above. The associated code is available at: https://github.com/VyshakhRP/RBP-div-hostrange Contents: - 01_raw/02_assemblies/ - Raw lysate assemblies for unpublished Klebsiella pneumoniae phages - 01_raw/03_host_genomes/ - Klebsiella pneumoniae host genome sequences (.fasta) - 01_raw/04_phage_genomes/ - Phage genome sequences (.fasta, .gb) - 02_intermediate/09_af3_predictions/ - AlphaFold 3 structure predictions for all receptor-binding proteins (.cif) - 02_intermediate/14_rbps-ecods/ecod.develop288.domains.txt - ECOD domain database (v20230309, develop288) - 02_intermediate/14_rbps-ecods/ecod.develop288.fasta.txt - ECOD domain FASTA (v20230309, develop288) - 05_output/01_genomes/ - Processed phage genome files - 05_output/03_proteins/ - Protein FASTAs and receptor-binding protein structures Usage: Download the archive, extract it, and place each subdirectory into the corresponding location in the cloned GitHub repository. Full instructions are provided in the repository README under Data Availability. Note: The ECOD database files can alternatively be downloaded directly from http://prodata.swmed.edu/ecod/distributions/ (v20230309 / develop288).

open·CC-BY-4.0·zenodo-bio·completeSource
composite

Toy dataset of read files and reference genome for PopFun test run

0.00

Bar, Ido

9 files · 8.0 MB · gff, gzip

These files are downsampled WGS fastq files (250k paired-end reads each) of a fungal pathogen ( Ascochyta rabiei ), generated on MGI DNSeq-T7. The files are intended to be used directly as test datasets in PopFun - a Nextflow pipeline for variant calling in fungal genomes.

open·MIT·zenodo-bio·completeSource
page 1next →

Select a result to see its full details here: the measured structure, quality, and the loader, without leaving your search.