Kantor, Rose · Shakya, Migun · Ruth, Nelson · et al.
2,095 rows · 907 KB · fasta, tsv
A virus genome database representing 21,015 near-complete virus genomes collected from untargeted ultra-deep RNA/DNA combined sequencing of wastewater. Sequence data was provided by the CASPER consortium and raw data may be found on NCBI SRA under bioprojects PRJNA1247874 and PRJNA1198001. Data underwent read trimming, rRNA and human read removal, de novo assembly, and selection of high-quality viral contigs. Contigs were clustered at 95% identity and 85% query coverage to dereplicate. Chimera-checking required at least two independent assemblies of the same viral genome or presence of the genome in another reference database. Annotation made use of RdRpCATCH, geNomad, checkV, BLASTN against NCBI core-nt, and RNAVirHost. The RdRp fasta files contain representative RdRp sequences identified through homology to major RdRp reference databases and clustered at 90% sequence identity over 75% sequence coverage. Included sequences contain all three conserved RdRp motifs (A, B, and C) arranged in either the canonical ABC configuration or the permuted CAB configuration.
Hubbard, Alfred · Solares, Edwin · Hemming-Schroeder, Elizabeth
88 rows · 18 KB · fasta
These are the files needed to run the Broad Institute's malaria amplicon pipeline for the PvGAP Plasmodium vivax panel, described in detail here . They consist of FASTA files containing the forward and reverse primers and another FASTA file containing reference sequences for each target, derived from the PvP01 reference genome.
This dataset supports the findings of the manuscript "Root anatomical traits modulate the assembly and nitrogen transformation potential of root-associated microbiomes in a temperate steppe" (NPH-MS-2026-55667). It contains root traits data, bacterial 16S rRNA gene absolute abundances, functional genes relative abundances, DNA extraction metadata, and phylogenetic marker sequences for 37 plant species from a temperate steppe ecosystem. The dataset includes the following files: 1. root traits.csv - Root traits including average diameter (AD), specific root length (SRL), specific root area (SRA), root tissue density (RTD), root nitrogen content (RNC), root carbon content (RCC), carbon‑nitrogen ratio (RCN), cortex layer number (CLN), cortex thickness (CT), and the ratio of cortex thickness to root diameter (CTRD). The first column lists plant species names. 2. Absolute abundance of 16S rRNA gene.csv - Quantitative PCR (qPCR) derived absolute abundances of bacterial 16S rRNA gene copies (copies/ng DNA) across different root compartments (rhizosphere, rhizoplane, endosphere), host species, root orders, and cotyledon classes (monocot/dicot). 3. DNA extraction sample weight.csv - Fresh weight (grams) of root material used for DNA extraction for each sample, linked by SampleID to the abundance data. 4. DNA extraction concentration.csv - Qubit‑measured DNA concentrations (ng/μL) and the sample volume (μL) used for quality control, together with sample metadata. 5. 37species.fasta - DNA sequences of two chloroplast markers (matK and rbcL) for the 37 plant species included in the study. The sequences are in FASTA format with headers formatted as ">Species". These were used for host phylogeny construction and Pagel's λ analyses. 6. Quantitative PCR results of functional gene.csv - Quantitative PCR (qPCR) derived relative abundances of bacterial 16S rRNA gene and functional genes across different root compartments (rhizosphere, rhizoplane, endosphere), host species, root orders, and cotyledon classes (monocot/dicot). 7. README.md - A detailed description of each file, column headers, abbreviations, units, and any missing value codings (NA). All data are provided to ensure transparency and reproducibility of the analyses. For methodological details, please refer to the Materials and Methods section of the associated publication. These data are under embargo until the associated research article is published. After that date, they will be freely available under a Creative Commons Attribution 4.0 International (CC BY 4.0) license. During the embargo period, the metadata (title, authors, abstract) and the DOI remain publicly visible, but the data files are not accessible. For access requests before the embargo expires, please contact the corresponding author.
OMAmer - tree-driven and alignment-free protein assignment to subfamilies OMAmer is an alignment-free protein family assignment method designed to avoid overly specific subfamily predictions and to scale efficiently to phylogenomic databases containing thousands of genomes. It relies on an innovative approach that uses evolutionarily informed k-mers for alignment-free mapping to ancestral protein subfamilies. This dataset provides precomputed OMAmer databases derived from the Hierarchical Orthologous Groups in the OMA Browser . We aim to update these databases with every new OMA Browser release. Each OMAmer database is built using the latest version of the OMAmer package available at the time of the corresponding OMA Browser release. The dataset includes databases for different subsets of the species taxonomy. In most cases, we recommend using the LUCA.h5 database, which contains information from all species in the OMA database. The subset-specific databases are mainly useful when disk space is limited. The release May2026 is based on the OMA Browser release May 2026 which comprises 2983 species. We used OMAmer version 2.1.0 to build these databases.
Requena Rolanía, Jose María · Greif, Gonzalo · ROBELLO, CARLOS
1 files · 8.0 MB · fasta
This dataset contains the genome sequence for Trypanosoma cruzi (strain Dm28c). This genome sequence was de novo assembled using PacBio Hi-Fi and Illumina sequencing platforms by Greif et al (2026. PMID: 41501640). The genome was assembled into 32 contigs, which represent complete chromosomes. The provided Fasta file also contains an additional contig corresponding to the maxicircle (mitochondrial genome) sequence. The Fasta files included in this dataset were downloaded from GenBank (assembly GCA_044048535.1; May 22, 2026). Additional information about the Dm28cT2T genome assembly and gene annotations may be accessed through the link: https://cruzi.pasteur.uy/
Burman, Nathaniel · Buyukyoruk, Murat · Wiegand, Tanner · et al.
4 files · 8.0 MB · fasta
This folder contains a multiple sequence alignment of Cas7 homologs in .fasta format, the domain-level annotations from PFAM and CasFinder, and an associated phylogenetic tree in .newick format.
Tiskus, Edvinas · Tiškuvienė, Rūta · Bučas, Martynas · et al.
37 files · 1.6 GB · hdf5, tiff, torchdeclared
This record contains the labeled data, trained models, and analysis code supporting the article "Comparing a Vision Foundation Model (DINOv3) and a Task-Specific U-Net for Mapping Emergent Aquatic Vegetation from Fused UAV Multispectral and LiDAR Data" (Remote Sensing in Ecology and Conservation). Contents: - masks/ : georeferenced ground-truth segmentation masks (five classes: aquatic vegetation, water, sand, other objects, background), aligned to the fused UAV orthomosaics and spanning 13 sites across nine Lithuanian waterbodies surveyed between May and August 2024. - models/ : the two final trained segmentation models, a Keras/HDF5 U-Net and a PyTorch DINOv3 model. - code/ : Python scripts for training, evaluation, the label-efficiency experiment, and full-scene prediction. The fused 9-band orthomosaics (five-band multispectral, RGB, and a LiDAR canopy height model; approximately 62 GB) are archived separately because of their size and are available from the corresponding author on request. The DINOv3 SAT-493M pretrained backbone is distributed by Meta under its own license and is not redistributed here; obtain it from the official DINOv3 release.
Pre-computed, resolution-broadened (R=1500) stellar atmosphere spectra cache for the astroARIADNE SED fitting package. Contains 7 model grids (Phoenix v2, BT-Settl, BT-NextGen, BT-Cond, Castelli & Kurucz 2004, Kurucz 1993, Coelho 2014) resampled to a common logarithmic wavelength grid (0.125-4.629 µm). This cache eliminates the need to download the full ~770 GB model libraries for SED plotting.
This record contains the trained autoencoder, extracted encoder, and processed world/real-river latent-space reference cloud associated with the manuscript *A data-driven approach to discern the curvature spectral complexity of compound meander bends*. The full autoencoder is provided to support reconstruction-based validation and reproducibility of the learned representation. The extracted encoder is provided for inference and future software tools. It maps preprocessed 64 × 64 single-channel curvature-spectrum images to the two-dimensional latent space used to analyse meander shape complexity and skewness. The file `world_latent_cloud.npy` contains the two-dimensional latent coordinates of the world/real-river meander dataset used as the reference background cloud in the manuscript latent-space figures. This file is a processed latent-coordinate dataset only; it does not contain raw satellite imagery, raw centreline geometries, or training images. The release includes model weights, architecture files, model summaries, export metadata, the world/real-river latent cloud, example inference scripts, a validation script, environment files, and a minimal example input. The models should only be applied to curvature-spectrum images generated consistently with the preprocessing workflow described in the associated manuscript. Main files included in this release are: - trained_autoencoder.h5: full trained autoencoder. - encoder_only.h5: extracted encoder in HDF5/Keras format. - encoder_only.keras: extracted encoder in native Keras format. - model_architecture.json: full autoencoder architecture. - encoder_architecture.json: encoder architecture. - model_summary.tx and encoder_summary.txt: layer summaries. - world_latent_cloud.npy: world/real-river reference latent-space cloud. - world_latent_cloud_metadata.json: metadata for the world/real-river latent-space cloud. - model_card.md: intended use, inputs, outputs, limitations, and citation guidance.
Meekes, Lisa · Tabaro, Francesco · Bexkens, Michiel · et al.
41 files · 8.2 GB · csv, fasta, pdfdeclared
This record contains the Python software for PEPTiGEN, a tool for generating tryptic peptides from prokaryotic gene sequences and their variants, and the associated antimicrobial resistance (AMR) peptide database. The database is provided as an SQL file and a CSV file containing all genes and predicted peptides. The README file contains explanation of the PEPTiGEN tool. The SQL database schema files contains both the database schema of the SQL database used in the PEPTiGEN analysis as the database schema of the AMR peptide datbase.
Bir, Joyanta · Cancio, Ibon · Diaz de cerio, Oihane · et al.
3 files · 100 KB · fasta, xlsxdeclared
This data file contains the data associated with the manuscript entitled "Duplication of the Genes Coding the Proteins That Regulate RNA Polymerase III Activity and Differential Transcription in Tissues of Teleost Fish."
1. May24_CIRBE-GOES_flux.h5 contains the necessary flux data to recreate the flux plots. 2. Oct24_CIRBE-GOES_flux.h5 contains the necessary flux data to recreate the flux plots.
This repository contains the complete chemosensory protein and nucleotide sequences, along with the results of evolutionary selection tests for 13 species of the genus Rhodnius . 1. Project Description This dataset supports the study of the chemosensory repertoire (ORs, GRs, IRs, OBPs, and CSPs) across 13 Rhodnius genomes. The study highlights the contrast between the conservation of Gustatory (GRs) and Ionotropic (IRs) receptors and the high dynamic evolution of Odorant Receptors (ORs), particularly in species adapted to human habitats. 2. Repository Structure 2.1 Sequence Data (FASTA) The following files contain all identified chemosensory genes in both amino acid ( .faa ) and nucleotide ( .fna ) formats: Rhodnius_OR_proteins.faa / Rhodnius_OR_CDS.fna : Odorant Receptors. Rhodnius_GR_proteins.faa / Rhodnius_GR_CDS.fna : Gustatory Receptors Rhodnius_IR_proteins.faa / Rhodnius_IR_CDS.fna : Ionotropic Receptors. Rhodnius_OBP_proteins.faa / Rhodnius_OBP_CDS.fna : Odorant-Binding Proteins. Rhodnius_CSP_proteins.faa / Rhodnius_CSP_CDS.fna : Chemosensory Proteins. 2.2 Phylogenetic Trees Archives containing the multiple sequence alignments and the resulting phylogenetic trees (Newick/Treefile format): OR_trees.zip : Alignment ( OR.ali.fasta ) and tree file ( OR.ali.treefile ) for Odorant Receptors. GR_trees.zip : Alignment ( GR.ali.fasta ) and tree file ( GR.ali.treefile ) for Gustatory Receptors. IR_trees.zip : Alignment ( IR.ali.fasta ) and tree file ( IR.ali.treefile ) for Ionotropic Receptors. OBP_trees.zip : Alignment ( OBP.ali.fasta ) and tree file ( OBP.ali.treefile ) for Odorant-Binding Proteins. CSP_trees.zip : Alignment ( CSP.ali.fasta ) and tree file ( CSP.ali.treefile ) for Chemosensory Proteins. 2.3 Evolutionary Selection Tests These archives contain the results of selection pressure analyses (e.g., dN/dS ratios, Likelihood Ratio Tests). Each gene family folder is subdivided by orthologous groups (e.g., GR1, GR2). OR_selection.zip GR_selection.zip IR_selection.zip Inside each selection archive, you will find: *_ali.fasta : Codon-based multiple sequence alignment. *_ali.pml : Codon-based multiple sequence alignment in PAML-friendly format. tree : The phylogenetic tree used for the selection model. LRT_BM.xls / BM_LRT.xls : Results for the Branch Model tests (domiciliary species vs. sylvatic species, see the associated paper). LRT_SM.xls / SM_LRT.xls : Results for the Site Model tests. 3. Methods Brief Genomes: Genomic data were sourced from NCBI (see paper for specific assembly accessions) . Annotation : Initial identification was performed using insectOR and Exonerate , followed by manual curation of gene models. Trees : Alignements was performed using MAFFT and ML trees using IQ-TREE . Selection Tests : Positive selection was assessed using PAML (codeml, EasyCodeML) on codon-aligned sequences. 4. Species Included Rhodnius bretesi Rhodnius colombiensis Rhodnius (=Psammolestes) coreodes Rhodnius domesticus Rhodnius milesi Rhodnius montenegrensis Rhodnius nastutus Rhodnius neglectus Rhodnius neivai Rhodnius pallescens Rhodnius pictipes Rhodnius prolixus Rhodnius robustus 5. Usage and Citation If you use these data, please cite the original publication: Merle, M. et al. (2026). Evolutionary Dynamics of the Complete Chemosensory Repertoire in Kissing Bugs of the Genus Rhodnius: Divergent Odorant Receptors Contrast with Conserved Gene Families. ( in prep ) For the specific dataset version, you can also cite this Zenodo DOI: DOI: 10.5281/zenodo.19064793
This includes the source code, background plasma conditions, and simulation results that produced the simulation data reported in Green et al., "Evaluating the Impact of Multiscale E-Region Turbulence on HF/VHF Scintillation."
Genome assembly of Cardita leana (Bivalvia: Archiheterodonta) and the associated gene models predicted with AUGUSTUS. The 'Tree' folder contains species trees inferred using different tree reconstruction programs. The MCMC_analysis folder contains inputs and ouputs for every MCMCTree analysis
Wedig, Bryce · Daylan, Tansu · Huang, Alan · et al.
2 files · 6.5 GB · hdf5declared
Data Challenge Overview The Roman Space Telescope is expected to observe O(10^5) galaxy-galaxy strong gravitational lenses, providing high angular resolution images of galaxy-galaxy strong gravitational lenses that can be used to probe the nature of dark matter at sub-galactic scales ( Daylan and Birrer 2023 , Wedig et al. 2025 ). The Roman Data Challenge for Dark Matter Substructure with Galaxy-Galaxy Strong Gravitational Lenses provides realistic simulated Roman images of strong lenses with various dark matter substructure populations and challenges the community to test out substructure detection and characterization pipelines. Dataset Description The goal of this rung is to distinguish between mass distributions with Cold Dark Matter subhalos and no subhalos. In this rung, you will train a binary classifier to determine whether subhalos are present. This is the unlabeled dataset. It does not include the boolean substructure flag and a few other related parameters that were included in the labeled dataset. Rung 1 submissions will be scored for this dataset. Changelog v2.0: Fixes a bug where SNRs were calculated from 601 second exposures but images were simulated with exposure time of 610 seconds. The difference in SNR is approximately 1%. New major version because the systems are different from v1.0 v1.0: Initial version
Gupte, Nihar · Miller, M. Coleman · Udall, Rhiannon · et al.
36 files · 36 GB · hdf5, torchdeclared
Data release for the DINGO O4a eccentricity paper. It contains the per-event parameter-estimation products, population selection function, and hierarchical-inference posteriors needed to reproduce every figure, table, and number in the paper, plus the trained DINGO neural networks used for the analyses. Event data : eccentric, quasicircular, and precessing per-event posterior samples (posteriors_eccentric.h5, posteriors_quasicircular.h5, posteriors_precessing.h5); slimmed log-uniform-eccentricity-prior posteriors used as the hierarchical-likelihood input (posteriors_log_uniform_eccentric.h5); per-event posteriors reweighted by the population-informed posterior (posteriors_population_reweighted.h5); per-event summary statistics with pre-computed Bayes factors (summary_statistics.h5); e_gw conversions (egw_conversions.h5); and the eccentricity-mean-anomaly prior hull (e_zeta_prior_hull.h5). Selection function : the injection p_draw dataframe with detection probabilities including the analysis-window factor (injection_p_draw.h5), a fixed-injection eccentricity sweep (fixed_injection_ecc_sweep.h5), and matched-filter survival-function data (survival_function.h5). Hierarchical inference : the selection-corrected velocity-dispersion posterior marginalized over the GWTC-4 mass/spin/redshift hyperposterior (sigma_posterior.h5), the capture-eccentricity lookup table (capture_ecc_table.h5), the external GWTC-4 hyperposterior fit (gwtc4_hyperposterior.h5), and the GC/NSC branching-fraction posterior (branching_fraction_posterior.h5). Glitch analyses : glitch-marginalized posteriors for GW190701, GW231114_043211, and GW231223_032836. Networks : trained DINGO networks (SEOBNRv5EHM, SEOBNRv5HM, SEOBNRv5PHM) with their training settings; see MODEL_MANIFEST.md. Zenodo stores files flat; the companion code maps each file into the foldered layout the notebooks expect. Code to download the data and reproduce all figures: github.com/nihargupte-ph/o4a-eccentricity , archived at doi:10.5281/zenodo.21221948 .