Kantor, Rose · Shakya, Migun · Ruth, Nelson · et al.
hybrid · semantic + lexical · 20 datasets ranked · 2.63s
Kantor, Rose · Shakya, Migun · Ruth, Nelson · et al.
A virus genome database representing 21,015 near-complete virus genomes collected from untargeted ultra-deep RNA/DNA combined sequencing of wastewater. Sequence data was provided by the CASPER consortium and raw data may be found on NCBI SRA under bioprojects PRJNA1247874 and PRJNA1198001. Data underwent read trimming, rRNA and human read removal, de novo assembly, and selection of high-quality viral contigs. Contigs were clustered at 95% identity and 85% query coverage to dereplicate. Chimera-checking required at least two independent assemblies of the same viral genome or presence of the genome in another reference database. Annotation made use of RdRpCATCH, geNomad, checkV, BLASTN against NCBI core-nt, and RNAVirHost. The RdRp fasta files contain representative RdRp sequences identified through homology to major RdRp reference databases and clustered at 90% sequence identity over 75% sequence coverage. Included sequences contain all three conserved RdRp motifs (A, B, and C) arranged in either the canonical ABC configuration or the permuted CAB configuration.
Kim, Kwanho · Lin, Zechuan · Simmons, Sean · et al.
1 files · 37 KB · pdf
This dataset contains CCS corrected HiFi long-read DNA sequencing (lrDNAseq) in FASTQ format for 100 PMDBS samples from Parkinson's patients and healthy controls. It's part of the PD5D atlas, where the same subjects were also profiled with other omics assays including genotyping, single-cell ATACseq, and spatial transcriptomics.
Kim, Kwanho · Lin, Zechuan · Simmons, Sean · et al.
1 files · 37 KB · pdf
This dataset contains raw FASTQ files from the midbrain single-nucleus RNA sequencing (snRNAseq) dataset with hybrid selection for the matching PMDBS samples from the PD5D chort. The same subjects were also profiled with other omics assays including genomic DNAseq, genotyping, single-cell ATACseq, and spatial transcriptomics.
Pančíková, Alexandra · Theunis, Koen · Hulselmans, Gert · et al.
1 files · 37 KB · pdf
This dataset consists of raw sequencing snRNA-seq data (10x Genomics Chromium Next GEM Single Cell 3ʹ). The data is part of an overall set of samples derived from postmortem midbrain (n=140), cingulate cortex (n=190) and motor cortex (n=4) of healthy donors (n=114), patients with Parkinson's disease (n=75) or patients with other neurological disorder (n=1). The protocol followed to isolate nuclei from postmortem brain samples and to prepare sequencing libraries can be found below. To increase throughput and to decrease batch effects, several donors have been pooled together into a single sequencing library. To computationally demultiplex the nuclei to their corresponding donors, cellsnp-lite (version commit: aad18644adcde853c313362a856a24245c9b91f7) followed by vireo (https://github.com/single-cell-genetics/vireo/pull/108) has been used. The population VCF with the donor genotypes derived from whole genome sequencing data has been used to assign nuclei back to their donors.
Pančíková, Alexandra · Theunis, Koen · Hulselmans, Gert · et al.
1 files · 36 KB · pdf
This dataset consists of raw sequencing snRNA-seq data using ParseBio Evercode Whole Transcriptome. The data is part of an overall set of samples derived from postmortem midbrain (n=140), cingulate cortex (n=190) and motor cortex (n=4) of healthy donors (n=114), patients with Parkinson's disease (n=75) or patients with other neurological disorder (n=1). The protocol followed to isolate nuclei from postmortem brain samples and to prepare sequencing libraries can be found below. To increase throughput and to decrease batch effects, several donors have been pooled together into a specific ParseBio barcode. To computationally demultiplex the nuclei to their corresponding donors, cellsnp-lite (version commit: aad18644adcde853c313362a856a24245c9b91f7) followed by vireo (https://github.com/single-cell-genetics/vireo/pull/108) has been used. The population VCF with the donor genotypes derived from whole genome sequencing data has been used to assign nuclei back to their donors.
Pančíková, Alexandra · Theunis, Koen · Hulselmans, Gert · et al.
1 files · 36 KB · pdf
This dataset consists of raw sequencing snATAC-seq data (10x Genomics Chromium Next GEM Single Cell ATAC v2). The data is part of an overall set of samples derived from postmortem midbrain (n=140), cingulate cortex (n=190) and motor cortex (n=4) of healthy donors (n=114), patients with Parkinson's disease (n=75) or patients with other neurological disorder (n=1). The protocol followed to isolate nuclei from postmortem brain samples and to prepare sequencing libraries can be found below. To increase throughput and to decrease batch effects, several donors have been pooled together into a single sequencing library. To computationally demultiplex the nuclei to their corresponding donors, cellsnp-lite (version commit: aad18644adcde853c313362a856a24245c9b91f7) followed by vireo (https://github.com/single-cell-genetics/vireo/pull/108) has been used. The population VCF with the donor genotypes derived from whole genome sequencing data has been used to assign nuclei back to their donors. (edited)
Pančíková, Alexandra · Theunis, Koen · Hulselmans, Gert · et al.
1 files · 36 KB · pdf
This dataset consists of raw sequencing snATAC-seq data (HyDrop v2). The data is part of an overall set of samples derived from postmortem midbrain (n=140), cingulate cortex (n=190) and motor cortex (n=4) of healthy donors (n=114), patients with Parkinson's disease (n=75) or patients with other neurological disorder (n=1). The protocol followed to isolate nuclei from postmortem brain samples can be found below. To increase throughput and to decrease batch effects, several donors have been pooled together into a single sequencing library. To computationally demultiplex the nuclei to their corresponding donors, cellsnp-lite (version commit: aad18644adcde853c313362a856a24245c9b91f7) followed by vireo (https://github.com/single-cell-genetics/vireo/pull/108) has been used. The population VCF with the donor genotypes derived from whole genome sequencing data has been used to assign nuclei back to their donors.
Pančíková, Alexandra · Theunis, Koen · Hulselmans, Gert · et al.
1 files · 36 KB · pdf
This dataset consists of raw sequencing ATAC-seq data (Scale-ATAC pre-indexing followed by 10x Genomics snATAC v2). The data is part of an overall set of samples derived from postmortem midbrain (n=140), cingulate cortex (n=190) and motor cortex (n=4) of healthy donors (n=114), patients with Parkinson's disease (n=75) or patients with other neurological disorder (n=1). The protocol followed to isolate nuclei from postmortem brain samples and to prepare sequencing libraries can be found below. To increase throughput and to decrease batch effects, several donors have been pooled together into a single sequencing library. To computationally demultiplex the nuclei to their corresponding donors, cellsnp-lite (version commit: aad18644adcde853c313362a856a24245c9b91f7) followed by vireo (https://github.com/single-cell-genetics/vireo/pull/108) has been used. The population VCF with the donor genotypes derived from whole genome sequencing data has been used to assign nuclei back to their donors.
Pančíková, Alexandra · Theunis, Koen · Hulselmans, Gert · et al.
1 files · 36 KB · pdf
This dataset consists of raw sequencing ATAC-seq data (Scale-ATAC + HyDrop v2). The data is part of an overall set of samples derived from postmortem midbrain (n=140), cingulate cortex (n=190) and motor cortex (n=4) of healthy donors (n=114), patients with Parkinson's disease (n=75) or patients with other neurological disorder (n=1). The protocols followed to isolate nuclei from postmortem brain samples and to prepare sequencing libraries can be found below. To increase throughput and to decrease batch effects, several donors have been pooled together into a single sequencing library. To computationally demultiplex the nuclei to their corresponding donors, cellsnp-lite (version commit: aad18644adcde853c313362a856a24245c9b91f7) followed by vireo (https://github.com/single-cell-genetics/vireo/pull/108 ) has been used. The population VCF with the donor genotypes derived from whole genome sequencing data has been used to assign nuclei back to their donors.
Pančíková, Alexandra · Theunis, Koen · Hulselmans, Gert · et al.
1 files · 37 KB · pdf
This dataset consists of raw sequencing snRNA-seq data and snATAC-seq data (10x Genomics Chromium Next GEM Multiome ATAC/GEX). The data is part of an overall set of samples derived from postmortem midbrain (n=140), cingulate cortex (n=190) and motor cortex (n=4) of healthy donors (n=114), patients with Parkinson's disease (n=75) or patients with other neurological disorder (n=1). The protocol followed to isolate nuclei from postmortem brain samples and to prepare sequencing libraries can be found below. To increase throughput and to decrease batch effects, several donors have been pooled together into a single sequencing library. To computationally demultiplex the nuclei to their corresponding donors, cellsnp-lite (version commit: aad18644adcde853c313362a856a24245c9b91f7) followed by vireo (https://github.com/single-cell-genetics/vireo/pull/108 ) has been used. The population VCF with the donor genotypes derived from whole genome sequencing data has been used to assign nuclei back to their donors. (edited)
Yuan, Guangyuan
72 rows × 13 cols · 6.0 KB · csv, fasta
11 numeric · 2 categorical
This dataset supports the findings of the manuscript "Root anatomical traits modulate the assembly and nitrogen transformation potential of root-associated microbiomes in a temperate steppe" (NPH-MS-2026-55667). It contains root traits data, bacterial 16S rRNA gene absolute abundances, functional genes relative abundances, DNA extraction metadata, and phylogenetic marker sequences for 37 plant species from a temperate steppe ecosystem. The dataset includes the following files: 1. root traits.csv - Root traits including average diameter (AD), specific root length (SRL), specific root area (SRA), root tissue density (RTD), root nitrogen content (RNC), root carbon content (RCC), carbon‑nitrogen ratio (RCN), cortex layer number (CLN), cortex thickness (CT), and the ratio of cortex thickness to root diameter (CTRD). The first column lists plant species names. 2. Absolute abundance of 16S rRNA gene.csv - Quantitative PCR (qPCR) derived absolute abundances of bacterial 16S rRNA gene copies (copies/ng DNA) across different root compartments (rhizosphere, rhizoplane, endosphere), host species, root orders, and cotyledon classes (monocot/dicot). 3. DNA extraction sample weight.csv - Fresh weight (grams) of root material used for DNA extraction for each sample, linked by SampleID to the abundance data. 4. DNA extraction concentration.csv - Qubit‑measured DNA concentrations (ng/μL) and the sample volume (μL) used for quality control, together with sample metadata. 5. 37species.fasta - DNA sequences of two chloroplast markers (matK and rbcL) for the 37 plant species included in the study. The sequences are in FASTA format with headers formatted as ">Species". These were used for host phylogeny construction and Pagel's λ analyses. 6. Quantitative PCR results of functional gene.csv - Quantitative PCR (qPCR) derived relative abundances of bacterial 16S rRNA gene and functional genes across different root compartments (rhizosphere, rhizoplane, endosphere), host species, root orders, and cotyledon classes (monocot/dicot). 7. README.md - A detailed description of each file, column headers, abbreviations, units, and any missing value codings (NA). All data are provided to ensure transparency and reproducibility of the analyses. For methodological details, please refer to the Materials and Methods section of the associated publication. These data are under embargo until the associated research article is published. After that date, they will be freely available under a Creative Commons Attribution 4.0 International (CC BY 4.0) license. During the embargo period, the metadata (title, authors, abstract) and the DOI remain publicly visible, but the data files are not accessible. For access requests before the embargo expires, please contact the corresponding author.
MacDonald, Adam · Stratton, Jo Anne
1 files · 35 KB · pdf
We performed 10X Genomics single-cell RNAsequencing of human prepheral blood mononuclear cells from healthy control, PD and RBD patients. This dataset contains raw FASTQ files. Sequencing was performed using NovaSeq 6000 S4 PE 100bp. Reads were processed using the 10X Genomics Cell Ranger Single Cell 2.0.0 pipeline. FASTQs generated from sequencing output were aligned to the human GRCh38 reference genome using STAR algorithm 2.7.3a.
Stratton, Jo Anne · Mukherjee, Sriparna · Trudeau, Louis-Eric
1 files · 5.0 KB · pdf
We performed bulk RNAsequencing of substantia nigra from wild type and PINK1 KO mice 26-days post C.rodentium infection. This dataset contains raw fastq files from striatal cells, sorted into 4 groups namely wild type and PINK1 KO uninfected and infected mice. Sequencing was performed using NextSeq500. FASTQs generated from sequencing output were aligned to the mm10 reference genome using STAR aligner.
Stratton, Jo Anne · Mukherjee, Sriparna · Trudeau, Louis-Eric
1 files · 6.1 KB · pdf
We performed bulk RNAsequencing of striatum from wild type and PINK1 KO mice 26-days post C.rodentium infection. This dataset contains raw fastq files from striatal cells, sorted into 4 groups namely wild type and PINK1 KO uninfected and infected mice. Sequencing was performed using NextSeq500. FASTQs generated from sequencing output were aligned to the mm10 reference genome using STAR aligner.
Recinto, Sherilyn · Stratton, Jo Anne
1 files · 6.1 KB · pdf
We performed 10X Genomics single-cell RNAsequencing of human iSPC-derived monocytes and macrophages in vitro. Cells were treated with 500 ng/mL lipopolysaccharide (LPS) and 50 ng/mL interleukin-1 beta (IL1b) for 24 hours. This dataset contains raw FASTQ files from myeloid cells, sorted into 4 groups namely monocytes (Mono) and Macrophages (Mac) non-stimulated (NS) and LPS+IL1b-stimulated cells. Sequencing was performed using NovaSeq 6000 S4 PE 100bp. Reads were processed using the 10X Genomics Cell Ranger Single Cell 2.0.0 pipeline. FASTQs generated from sequencing output were aligned to the human GRCh38 reference genome using STAR algorithm 2.7.3a.
Recinto, Sherilyn · Stratton, Jo Anne · Pei, Jessica · et al.
2 files · 6.0 KB · pdf
We performed 10X Genomics single-cell RNAsequencing of colonic lamina propria cells from wild type and LRRK2 G2019S mice following 1-week post C. rodentium infection. The cells were pooled from 3 mice per group of both sexes at 8-12 weeks of age. This dataset contains raw FASTQ files from mouse colonic lamina propria, sorted into 4 groups namely wild type and LRRK2 G2019S uninfected and infected mice. Sequencing was performed using NovaSeq 6000 S4 PE 100bp. Reads were processed using the 10X Genomics Cell Ranger Single Cell 2.0.0 pipeline. FASTQs generated from sequencing output were aligned to the mouse GRCm38 reference genome using STAR algorithm 2.7.3a.
Recinto, Sherilyn · Stratton, Jo Anne
1 files · 5.7 KB · pdf
We performed 10X Genomics single-cell RNA sequencing of colonic lamina propria cells from wild type and PINK1 KO mice following either 1-week or 2-weeks post C. rodentium infection. The cells were pooled from 3 mice per group of both sexes at 8-12 weeks of age. This dataset contains raw FASTQ files from mouse colonic lamina propria, sorted into 8 groups namely wild type and PINK1 KO uninfected and infected mice at 1- or 2-weeks post-infection. Sequencing was performed using NovaSeq 6000 S4 PE 100bp. Reads were processed using the 10X Genomics Cell Ranger Single Cell 2.0.0 pipeline. FASTQs generated from sequencing output were aligned to the mouse GRCm38 reference genome using STAR algorithm 2.7.3a.
Budzinski, Lisa · Beenken, Anne Elisabeth · Sempert, Toni · et al.
9 rows × 1 cols · 743 B · csv, docx, zip
1 categorical
We have investigated an IgG4-RD (IgG4-RD) cohort by our multi-parameter microbiota flow cytometry approach to characterise the microbiota on single-cell level for attributes of the disease. The microbiota is isolated from stool samples and stained according to the published protocol for (a) host immunoglobulins IgA1, IgA2, IgM, IgG and (b) agglutinin binding to mannose, galactose or N-Acetyl-glucosamine surface sugar moieties. For all samples we also determined the microbiome composition by 16S rRNA (V3-V4) sequencing on the illumina MiSeq platform. We provide the raw .fcs and FASTQ files of 40 IgG4-RD patients. For comparison we additionally analysed 36 healthy donors. All .fcs files were generated on BD Influx®. The metadata is collected in the provided meta.csv. The staining parameters are summarized in provided panel.csv.
Manghi, Paolo
2 files · 8.0 MB · gzip, tsv
Using gene-level and species-level shotgun metagenomics, we provide the first characterization of the rural, Bolivian microbiome; we identified microbial genes which strongly correlate (rho>0.45) with arsenic in urine, and that overall contribute to substantiate that the gut microbiome helps tolerate arsenic via a evict-out-of-house mechanism. Mediation analysis, followed by phylogenetic investigation of metagenomic-assembled genomes, further strengthens this observation. This study elucidates the role of the microbiome in helping to tolerate arsenic-rich environments, and paves the way for probiotic interventions that may mitigate the effects of this toxic metal. This repository contains 2,478 HQ MAGs from the MicroToxBol cohort in fasta format and a descriptive table comprising taxonomic annotation, quality-checks, and coverage estimation.
Aalborg, Trine
7 files · 8.0 MB · csv, gzip
README - Data for "Genomic constraint and hypervariability in tetraploid potatoes" Trine Aalborg, May 2026 The data applied in the study includes phenotypic and genotypic information on the MASPOT panel (768 F1 progeny of an 18-parent diallel cross - property of Danespo A/S). The genotypic data was generated using genotyping-by-sequencing technology as described in the paper. Following genotype calling and filtration (5-60x read depth, < 50 % missing rate, > 1 % MAF), the total SNP set includes 151,164 biallelic SNPs. Coordinates of these SNPs relative to the DMv6.1 potato reference genome are provided. In addition to genotypes across the clones, the estimated GERP score of that SNP from (Wu et al., 2023) is reported. The manuscript analyses only consider markers with reliable GERP scores (MSA alignment depth > 50, and neutral score > 2), which corresponded to 97,815 of the total 151,164 biallelic SNPs. SNPeff annotations of the markers (based on the DMv6.1 reference genome) are also appended. The phenotypes were collected across 1-2 field trials, depending on the traits, and includes a minimum of two replicates per clone from a randomized block design. There are phenotypes for eight traits: dry matter content [%], yield (hkg/ha), senescence [1-9], flesh color [1-9], tubers/plant, tuber length [mm], tuber diameter [mm], and tuber size [mm^3]. Metadata includes phenotyping year and block location of the plot as well as pedigree of the diallel offspring. File descriptions: gt_MASPOT.csv - .csv file of the non-imputed genotypic data of 151,164 SNPs for the 768 F1 clones (those with GERP scores). Columns 1-3 are SNP coordinates and SNP IDs. Column names from column 4 and onwards are clone IDs. gt_MASPOT_imputed.csv - .csv file holding the imputed (random forest, missRanger algorithm) genotypic data of 151,164 SNPs for the 768 F1 clones. gt_MASPOT_recoded.csv - .csv file of the recoded, imputed genotypic data of 151,164 SNPs for the 768 F1 clones. The SNPs are recoded from original MASPOT ref/alt allele (based on AF in the MASPOT panel) to the alternative allele = the derived allele in the 100-Solanaceae panel. pt_MASPOT.csv - .csv file holding the phenotypic data (eight traits) of the 768 F1 clones (Clone_ID). The number of observations varies across traits. Also including metadata: year of phenotyping (Year), block number (Block, Line_in_block), clone parents (Mother, Father, Family). GERP_MASPOT.csv - .csv file holding the GERP scores (including alignment depth, neutral scores, and a marker annotation based on GERP score thresholds (deleterious, neutral, hypervariable, or low quality)), SNPeff annotations, and the derived allele in the 100-Solanaceae panel from (Wu et al., 2023) [MASPOT_Alt_Allele_Is_Sol_Derived_Allele - used for recoding of the genotypes] of the MASPOT SNPs with GERP scores. snps.MASPOT_F1.vcf.gz - zipped .vcf file of the GBS MASPOT genotypic data (both discrete genotype calls and allele frequencies) called to the DMv6.1 potato reference genome. Filtered to read depth 5x, MQ > 30. A total of 160,920 biallelic SNPs. Includes the 768 F1 progeny analyzed. Literature: Wu, Y., Li, D., Hu, Y., Li, H., Ramstein, G. P., Zhou, S., et al. (2023). Phylogenomic discovery of deleterious mutations facilitates hybrid potato breeding. Cell 186, 2313-2328.e15. doi: 10.1016/j.cell.2023.04.008