Kantor, Rose · Shakya, Migun · Ruth, Nelson · et al.
2,095 rows · 907 KB · fasta, tsv
hybrid · semantic + lexical · 57 datasets ranked · 2.66s
Kantor, Rose · Shakya, Migun · Ruth, Nelson · et al.
2,095 rows · 907 KB · fasta, tsv
A virus genome database representing 21,015 near-complete virus genomes collected from untargeted ultra-deep RNA/DNA combined sequencing of wastewater. Sequence data was provided by the CASPER consortium and raw data may be found on NCBI SRA under bioprojects PRJNA1247874 and PRJNA1198001. Data underwent read trimming, rRNA and human read removal, de novo assembly, and selection of high-quality viral contigs. Contigs were clustered at 95% identity and 85% query coverage to dereplicate. Chimera-checking required at least two independent assemblies of the same viral genome or presence of the genome in another reference database. Annotation made use of RdRpCATCH, geNomad, checkV, BLASTN against NCBI core-nt, and RNAVirHost. The RdRp fasta files contain representative RdRp sequences identified through homology to major RdRp reference databases and clustered at 90% sequence identity over 75% sequence coverage. Included sequences contain all three conserved RdRp motifs (A, B, and C) arranged in either the canonical ABC configuration or the permuted CAB configuration.
Hubbard, Alfred · Solares, Edwin · Hemming-Schroeder, Elizabeth
88 rows · 18 KB · fasta
These are the files needed to run the Broad Institute's malaria amplicon pipeline for the PvGAP Plasmodium vivax panel, described in detail here . They consist of FASTA files containing the forward and reverse primers and another FASTA file containing reference sequences for each target, derived from the PvP01 reference genome.
Yuan, Guangyuan
72 rows × 13 cols · 6.0 KB · csv, fasta
11 numeric · 2 categorical
This dataset supports the findings of the manuscript "Root anatomical traits modulate the assembly and nitrogen transformation potential of root-associated microbiomes in a temperate steppe" (NPH-MS-2026-55667). It contains root traits data, bacterial 16S rRNA gene absolute abundances, functional genes relative abundances, DNA extraction metadata, and phylogenetic marker sequences for 37 plant species from a temperate steppe ecosystem. The dataset includes the following files: 1. root traits.csv - Root traits including average diameter (AD), specific root length (SRL), specific root area (SRA), root tissue density (RTD), root nitrogen content (RNC), root carbon content (RCC), carbon‑nitrogen ratio (RCN), cortex layer number (CLN), cortex thickness (CT), and the ratio of cortex thickness to root diameter (CTRD). The first column lists plant species names. 2. Absolute abundance of 16S rRNA gene.csv - Quantitative PCR (qPCR) derived absolute abundances of bacterial 16S rRNA gene copies (copies/ng DNA) across different root compartments (rhizosphere, rhizoplane, endosphere), host species, root orders, and cotyledon classes (monocot/dicot). 3. DNA extraction sample weight.csv - Fresh weight (grams) of root material used for DNA extraction for each sample, linked by SampleID to the abundance data. 4. DNA extraction concentration.csv - Qubit‑measured DNA concentrations (ng/μL) and the sample volume (μL) used for quality control, together with sample metadata. 5. 37species.fasta - DNA sequences of two chloroplast markers (matK and rbcL) for the 37 plant species included in the study. The sequences are in FASTA format with headers formatted as ">Species". These were used for host phylogeny construction and Pagel's λ analyses. 6. Quantitative PCR results of functional gene.csv - Quantitative PCR (qPCR) derived relative abundances of bacterial 16S rRNA gene and functional genes across different root compartments (rhizosphere, rhizoplane, endosphere), host species, root orders, and cotyledon classes (monocot/dicot). 7. README.md - A detailed description of each file, column headers, abbreviations, units, and any missing value codings (NA). All data are provided to ensure transparency and reproducibility of the analyses. For methodological details, please refer to the Materials and Methods section of the associated publication. These data are under embargo until the associated research article is published. After that date, they will be freely available under a Creative Commons Attribution 4.0 International (CC BY 4.0) license. During the embargo period, the metadata (title, authors, abstract) and the DOI remain publicly visible, but the data files are not accessible. For access requests before the embargo expires, please contact the corresponding author.
Zahid, R · Lázaro, A · Moreno‐Alcántar, G · et al.
1 files · 82 KB · tar
Synthetic nanozymes have emerged as promising alternatives to natural enzymes for catalytic and therapeutic applications, yet their limited stability, aqueous compatibility, and catalytic scope impede broader utilization. Here, we report a mild, one-step sol-gel synthesis that yields ultrasmall, water-stable octa-amino silsesquioxanes functioning as metal-free nanozymes. These minimalistic nanostructures exhibit aldolase-like organocatalytic activity in water and enable dynamic, stimuli-responsive modulation of catalysis through reversible supramolecular aggregation and disaggregation triggered by specific chemical inputs, thus forming a multifunctional platform for tunable catalysis and biomedical applications. Structural simplicity, stability, and functional versatility together permit tunable, enzyme-like catalysis in water without auxiliary surfactants or phase-transfer additives. Furthermore, the nanozymes display high biocompatibility and efficient cellular internalization, enabling their use in living cells, for instance, as intracellular prodrug activators via retro-aldol activation of a doxorubicin prodrug in human glioblastoma and metastatic melanoma cells, resulting in selective cytotoxicity. This system provides a cost-effective, sustainable, and scalable platform for water-compatible, metal-free organocatalysis that bridges abiotic catalysis and biological function. These findings demonstrate how rationally designed silsesquioxane frameworks can emulate natural enzyme reactivity while integrating adaptive, stimuli-responsive behavior, broadening the applicability of synthetic nanozymes to catalytic and therapeutic contexts.
Requena Rolanía, Jose María · Greif, Gonzalo · ROBELLO, CARLOS
1 files · 8.0 MB · fasta
This dataset contains the genome sequence for Trypanosoma cruzi (strain Dm28c). This genome sequence was de novo assembled using PacBio Hi-Fi and Illumina sequencing platforms by Greif et al (2026. PMID: 41501640). The genome was assembled into 32 contigs, which represent complete chromosomes. The provided Fasta file also contains an additional contig corresponding to the maxicircle (mitochondrial genome) sequence. The Fasta files included in this dataset were downloaded from GenBank (assembly GCA_044048535.1; May 22, 2026). Additional information about the Dm28cT2T genome assembly and gene annotations may be accessed through the link: https://cruzi.pasteur.uy/
Burman, Nathaniel · Buyukyoruk, Murat · Wiegand, Tanner · et al.
4 files · 8.0 MB · fasta
This folder contains a multiple sequence alignment of Cas7 homologs in .fasta format, the domain-level annotations from PFAM and CasFinder, and an associated phylogenetic tree in .newick format.
Yang, Qingliu
6 files · 8.0 MB · bzip2
Dataset Description This dataset contains 3D lightning location results, DALMA and FALMA waveform for two Energetic Compact Stroke (ECS) events. location results are included: HF3D_1732785151.dat - 3D lightning locations for the ECS leader A flash. HF3D_1734785454.dat - 3D lightning locations for the ECS leader B flash. The timestamp 1734785454 and 1732785151 corresponds to the occurrence time of the lightning flash in Japan Standard Time. File format and parameters The first row contains the lightning occurrence time. Column descriptions: Time (ms) - time relative to the lightning source. X, Y, Z (m) - 3D spatial coordinates relative to ground level. The origin (0, 0, 0) corresponds to latitude 36.76°N and longitude 136.76°E. FALMA and DALMA waveform ECSLeaderA_DALMA_waveform.bz2 is DALMA waveform of Leader A. ECSLeaderA_FALMA_waveform.bz2 is FALMA waveform of Leader A. ECSLeaderB_DALMA_waveform.bz2 is DALMA waveform of Leader B. ECSLeaderB_FALMA_waveform.bz2 is FALMA waveform of Leader B. This dataset allows analysis of the spatial and temporal development of these two ECS flashes.
Modiri, Ehsan · Shrestha, Pallav Kumar · Samaniego Eguiguren, Luis Eduardo
100 files · 49 GB · netcdf, tardeclared
Historical Hydrological Simulations over the South African Domain (1990-2024) The mHM's simulations of the Planet4Health project This dataset contains historical hydrological simulations for the South African domain (domain 1020011530) conducted with the Mesoscale Hydrological Model (mHM) at a spatial resolution of 0.015625°. The simulation period spans 1990-2024 and was part of the Planet4Health (P4H) project, utilising the ERA5 meteorological forcing. This archive is prepared for DOI assignment and ensures long-term reproducibility. It includes relevant clipped NetCDF components for soil moisture layers 4, 5 and 6, consistent with the infrastructure provided within the Helmholtz Centre for Environmental Research (UFZ). The simulations were executed using a specific version of the mHM model with the SCC method for gauges, paired with the mRMv1.0 routing configuration. 🛰️ Simulation Details Model: Mesoscale Hydrological Model (mHM) Codebase: scc_for_gauges branch (https://git.ufz.de/shresthp/mhm/-/tree/scc_for_gauges?ref_type=heads) Spatial resolution: 0.015625° Temporal resolution: Daily Simulation period: 1990-2024 Simulation type: Historical simulation Spin-up: 30-year spin-up using 1990-2019 ERA5 climatology Model version: v1.0 Setup Scope: Model run for domain 1020011530, post-processed and clipped. Configuration & Modules The configuration utilises standard structural components with the SCC methodology. Modules included: Snow processes: Degree-day method Soil moisture: Feddes equation for evapotranspiration reduction Infiltration: Multi-layer Brooks-Corey-like approach Direct runoff: Linear reservoir exceedance method Potential evapotranspiration: Hargreaves-Samani method Interflow: Storage reservoir with nonlinear outflow Groundwater: Linear reservoir Routing: Adaptive time-step routing with mRMv1.0 mechanisms 📥 Input Datasets Meteorological Forcing: ERA5 (Hersbach et al., 2020) at a native input meteorological resolution of 0.25°, dynamically downscaled/mapped to model requirements. Processing Infrastructure: Tracked, processed, and validated under the Planet4Health deployment pipeline (https://git.ufz.de/planet4health/mhm_production/-/tree/main/postproc?ref_type=heads). Data Interfaces: Climate Data Interface version 2.2.4 (CDI) | Climate Data Operators version 2.2.2 (CDO) | NetCDF Operators version 5.1.7 (NCO). 📤 Output Variables sm_l04: Volumetric soil moisture layer 4 (300-500 mm depth) [mm mm-1, fraction between 0 and 1] sm_l05: Volumetric soil moisture layer 5 (500-1000 mm depth) [mm mm-1, fraction between 0 and 1] sm_l06: Volumetric soil moisture layer 6 (1000-2000 mm depth) [mm mm-1, fraction between 0 and 1] 📫 Contact Ehsan Modiri - ehsan.modiri@ufz.de Pallav Kumar Shrestha - pallav-kumar.shrestha@ufz.de Institution Helmholtz Centre for Environmental Research - UFZ, Department of Computational Hydrosystems
Geisler, Jan · Rakhimberdiev, Eldar · Boom, Michiel P. · et al.
29 files · 124 MB · csv, shapefiledeclared
1. Many migratory birds now reach their Arctic breeding grounds earlier in order to keep pace with advancing springs and shifting nutrient peaks, either by departing earlier from non-breeding grounds or by travelling faster. For dark-bellied brent geese, there is limited potential to travel faster, as their migration to the Siberian breeding grounds is already among the fastest of Arctic geese and swans. Earlier departure would require reaching departure body mass earlier, either through a faster accumulation of energy stores during spring staging or via adjustments earlier in the annual cycle. 2. We examined long-term shifts in spring staging phenology and changes in winter and spring body mass trajectories of brent geese at the population level, with particular emphasis on the effects of winter temperature on body mass and spring body mass on departure timing. 3. We used more than five decades of body mass measurements from individuals caught in the United Kingdom and France, and in the Dutch Wadden Sea to reconstruct changes in spring and winter mass trajectories, respectively. These data were combined with over five decades of migration counts in the Netherlands and more than two decades of counts in Denmark to quantify changes in spring staging phenology. 4. We found that brent geese have not shifted their spring arrival in the Wadden Sea but have advanced departure timing. Furthermore, brent geese were heavier during and after milder winters, and have changed mass trajectories over recent decades. They no longer lose mass during winter and the second spring staging phase, and fuelling rates in the first spring staging phase have declined. Annual variation in body mass was not related to annual departure timing. 5. These results suggest that milder winters have relaxed energetic constraints and improved body condition in brent geese throughout the non-breeding season. Our findings highlight the importance of considering the full annual cycle when assessing how animals with limited capacity to adjust migration timing or speed respond to global change.
Kavanagh, Jack · Anthony, Patrick
6 files · 11 MB · geojson, geopackage, shapefiledeclared
A historical map showing a segment of the boundaries of New Spain in c. 1800. This map is fully open to fellow researchers and is available in multiple open source formats (SHP, GPKG, GeoJSON).
Steinhoff, Daniel · Monaghan, Andrew
1 files · 3.5 MB · tardeclared
WHATCH'EM (Water Height and Temperature in Container Habitats Energy Model) is a physics-based model that simulates water temperature and water height in containers using an energy balance approach. The model uses meteorological inputs together with container characteristics, shading, rainfall, evaporation, and optional manual water additions to simulate container water dynamics across a range of environmental conditions.
Kavanagh, Jack · Anthony, Patrick
6 files · 9.3 MB · geojson, geopackage, shapefiledeclared
A historical map of the boundaries of Prussia in c. 1795. This map is fully open to fellow researchers and is available in multiple open source formats (SHP, GPKG, GeoJSON).
Francescone, Marco
14 files · 380 KB · shapefile, zipdeclared
This dataset contains original geomorphic mapping of surface fault traces along the Dixie Valley Fault (DVF) range front and piedmont zone, central Nevada, USA. Traces were mapped directly from a 1-m bare-earth lidar digital elevation model (DEM), using hillshade and slope-raster visualizations. This dataset accompanies the manuscript: Francescone, M., et al. (in review), LiDAR-Based Fault-Scarp Analysis and Rupture Hazard Assessment: Earthquake Scenarios of the Dixie Valley Fault System (Nevada, USA). See Section 3.1 ("Fault Trace Mapping") of the manuscript for full methodological details
Meekes, Lisa · Tabaro, Francesco · Bexkens, Michiel · et al.
41 files · 8.2 GB · csv, fasta, pdfdeclared
This record contains the Python software for PEPTiGEN, a tool for generating tryptic peptides from prokaryotic gene sequences and their variants, and the associated antimicrobial resistance (AMR) peptide database. The database is provided as an SQL file and a CSV file containing all genes and predicted peptides. The README file contains explanation of the PEPTiGEN tool. The SQL database schema files contains both the database schema of the SQL database used in the PEPTiGEN analysis as the database schema of the AMR peptide datbase.
Dey, Hemal · Shao, Wanyun
14 files · 5.5 MB · csv, jpeg, shapefiledeclared
Despite the proliferation of social vulnerability assessment methodologies, selecting the most appropriate model remains a critical challenge due to inter-model variability. To explore the inter-model variability, this study systematically investigated inter-algorithmic and inter-classification variability to assess how methodological design influences outcomes.
Bir, Joyanta · Cancio, Ibon · Diaz de cerio, Oihane · et al.
3 files · 100 KB · fasta, xlsxdeclared
This data file contains the data associated with the manuscript entitled "Duplication of the Genes Coding the Proteins That Regulate RNA Polymerase III Activity and Differential Transcription in Tissues of Teleost Fish."
Cheng, Yuhan · Bai, Lubin · Zhang, Xiuyuan · et al.
15 files · 140 MB · shapefile, zipdeclared
Data used in the paper "Annotation-Efficient Building Footprint Updating via Historical Map Reuse and Iterative Instance Segmentation"
Ye, Renhao · Shen, Shiyin
65 files · 46 GB · csv, fits, tardeclared
Part 15 of 20. Synthetic Euclid VIS-band galaxy images generated from DESI Legacy Imaging Surveys r,z-band cutouts using an Image-to-Image Schrödinger Bridge (I2SB) model, over the Euclid Q1 footprint. The full processed dataset covers 2,981,033 FITS cutouts packed into 1,235 tar archives grouped by HEALPix sky pixel (nside=64, NESTED ordering), ~1 GB each. Because Zenodo limits each record to 100 files and 50 GB, the full set is split across 20 records. This record publishes a 60-archive subset (RA 59.1-78.7°, Dec -39.9--24.9°). Integrity and provenance: checksums.tsv and desi2euclid_index.csv record the SHA256 of all 1,235 archives from this processing run, not just the ones published as downloadable files here. These values let this exact version of the data be independently verified in the future -- including before any analysis built on it has been published -- by recomputing an archive's SHA256 and comparing it against the recorded value, to confirm the file has not been modified, corrupted, or substituted since it was originally produced. See README.md for the FITS layout (predicted VIS image, DESI g,r,z input -- of which only r,z were used to generate the prediction -- and prediction-uncertainty map), how to locate the archive for a given sky position, and how to run the integrity check.
Ye, Renhao · Shen, Shiyin
67 files · 46 GB · csv, fits, tardeclared
Part 14 of 20. Synthetic Euclid VIS-band galaxy images generated from DESI Legacy Imaging Surveys r,z-band cutouts using an Image-to-Image Schrödinger Bridge (I2SB) model, over the Euclid Q1 footprint. The full processed dataset covers 2,981,033 FITS cutouts packed into 1,235 tar archives grouped by HEALPix sky pixel (nside=64, NESTED ordering), ~1 GB each. Because Zenodo limits each record to 100 files and 50 GB, the full set is split across 20 records. This record publishes a 62-archive subset (RA 45.0-70.2°, Dec -48.1--30.1°). Integrity and provenance: checksums.tsv and desi2euclid_index.csv record the SHA256 of all 1,235 archives from this processing run, not just the ones published as downloadable files here. These values let this exact version of the data be independently verified in the future -- including before any analysis built on it has been published -- by recomputing an archive's SHA256 and comparing it against the recorded value, to confirm the file has not been modified, corrupted, or substituted since it was originally produced. See README.md for the FITS layout (predicted VIS image, DESI g,r,z input -- of which only r,z were used to generate the prediction -- and prediction-uncertainty map), how to locate the archive for a given sky position, and how to run the integrity check.
Ye, Renhao · Shen, Shiyin
66 files · 46 GB · csv, fits, tardeclared
Part 13 of 20. Synthetic Euclid VIS-band galaxy images generated from DESI Legacy Imaging Surveys r,z-band cutouts using an Image-to-Image Schrödinger Bridge (I2SB) model, over the Euclid Q1 footprint. The full processed dataset covers 2,981,033 FITS cutouts packed into 1,235 tar archives grouped by HEALPix sky pixel (nside=64, NESTED ordering), ~1 GB each. Because Zenodo limits each record to 100 files and 50 GB, the full set is split across 20 records. This record publishes a 61-archive subset (RA 51.5-81.5°, Dec -54.3--30.0°). Integrity and provenance: checksums.tsv and desi2euclid_index.csv record the SHA256 of all 1,235 archives from this processing run, not just the ones published as downloadable files here. These values let this exact version of the data be independently verified in the future -- including before any analysis built on it has been published -- by recomputing an archive's SHA256 and comparing it against the recorded value, to confirm the file has not been modified, corrupted, or substituted since it was originally produced. See README.md for the FITS layout (predicted VIS image, DESI g,r,z input -- of which only r,z were used to generate the prediction -- and prediction-uncertainty map), how to locate the archive for a given sky position, and how to run the integrity check.