OMAmer - tree-driven and alignment-free protein assignment to subfamilies OMAmer is an alignment-free protein family assignment method designed to avoid overly specific subfamily predictions and to scale efficiently to phylogenomic databases containing thousands of genomes. It relies on an innovative approach that uses evolutionarily informed k-mers for alignment-free mapping to ancestral protein subfamilies. This dataset provides precomputed OMAmer databases derived from the Hierarchical Orthologous Groups in the OMA Browser . We aim to update these databases with every new OMA Browser release. Each OMAmer database is built using the latest version of the OMAmer package available at the time of the corresponding OMA Browser release. The dataset includes databases for different subsets of the species taxonomy. In most cases, we recommend using the LUCA.h5 database, which contains information from all species in the OMA database. The subset-specific databases are mainly useful when disk space is limited. The release May2026 is based on the OMA Browser release May 2026 which comprises 2983 species. We used OMAmer version 2.1.0 to build these databases.
Tiskus, Edvinas · Tiškuvienė, Rūta · Bučas, Martynas · et al.
37 files · 1.6 GB · hdf5, tiff, torchdeclared
This record contains the labeled data, trained models, and analysis code supporting the article "Comparing a Vision Foundation Model (DINOv3) and a Task-Specific U-Net for Mapping Emergent Aquatic Vegetation from Fused UAV Multispectral and LiDAR Data" (Remote Sensing in Ecology and Conservation). Contents: - masks/ : georeferenced ground-truth segmentation masks (five classes: aquatic vegetation, water, sand, other objects, background), aligned to the fused UAV orthomosaics and spanning 13 sites across nine Lithuanian waterbodies surveyed between May and August 2024. - models/ : the two final trained segmentation models, a Keras/HDF5 U-Net and a PyTorch DINOv3 model. - code/ : Python scripts for training, evaluation, the label-efficiency experiment, and full-scene prediction. The fused 9-band orthomosaics (five-band multispectral, RGB, and a LiDAR canopy height model; approximately 62 GB) are archived separately because of their size and are available from the corresponding author on request. The DINOv3 SAT-493M pretrained backbone is distributed by Meta under its own license and is not redistributed here; obtain it from the official DINOv3 release.
Pre-computed, resolution-broadened (R=1500) stellar atmosphere spectra cache for the astroARIADNE SED fitting package. Contains 7 model grids (Phoenix v2, BT-Settl, BT-NextGen, BT-Cond, Castelli & Kurucz 2004, Kurucz 1993, Coelho 2014) resampled to a common logarithmic wavelength grid (0.125-4.629 µm). This cache eliminates the need to download the full ~770 GB model libraries for SED plotting.
This record contains the trained autoencoder, extracted encoder, and processed world/real-river latent-space reference cloud associated with the manuscript *A data-driven approach to discern the curvature spectral complexity of compound meander bends*. The full autoencoder is provided to support reconstruction-based validation and reproducibility of the learned representation. The extracted encoder is provided for inference and future software tools. It maps preprocessed 64 × 64 single-channel curvature-spectrum images to the two-dimensional latent space used to analyse meander shape complexity and skewness. The file `world_latent_cloud.npy` contains the two-dimensional latent coordinates of the world/real-river meander dataset used as the reference background cloud in the manuscript latent-space figures. This file is a processed latent-coordinate dataset only; it does not contain raw satellite imagery, raw centreline geometries, or training images. The release includes model weights, architecture files, model summaries, export metadata, the world/real-river latent cloud, example inference scripts, a validation script, environment files, and a minimal example input. The models should only be applied to curvature-spectrum images generated consistently with the preprocessing workflow described in the associated manuscript. Main files included in this release are: - trained_autoencoder.h5: full trained autoencoder. - encoder_only.h5: extracted encoder in HDF5/Keras format. - encoder_only.keras: extracted encoder in native Keras format. - model_architecture.json: full autoencoder architecture. - encoder_architecture.json: encoder architecture. - model_summary.tx and encoder_summary.txt: layer summaries. - world_latent_cloud.npy: world/real-river reference latent-space cloud. - world_latent_cloud_metadata.json: metadata for the world/real-river latent-space cloud. - model_card.md: intended use, inputs, outputs, limitations, and citation guidance.
Cognitive models propose that reading aloud relies on two routes: a lexical route for familiar word recognition, and a sublexical route for grapheme-to-phoneme conversion. Selective disruption of these routes results in surface or phonological dyslexia, respectively. Although the dual route model is well supported behaviorally, its white matter substrates remain incompletely characterized. This study examined the anatomical correlates of acquired lexical and sublexical reading impairments in patients with high-grade glioma in the dominant hemisphere. Thirty-seven patients underwent diffusion tensor imaging (DTI) and a comprehensive reading battery. Reading aloud of regular and irregular words, and pseudowords was used to classify patients as having intact reading, surface dyslexia, or phonological dyslexia, based on normative criteria. Six language-related white matter tracts were reconstructed individually for each patient using subject-specific tractography. Each tract was then systematically segmented along its longitudinal axis, allowing diffusion parameters to be extracted from anatomically defined subsegments. Lesion-tract overlap measures were computed to quantify focal tract involvement. Fourteen patients exhibited impaired sublexical reading, 12 impaired lexical reading, and 11 intact reading. Deficits in the sublexical route were more frequent among patients with arcuate fasciculus lesion overlap, and sublexical error rates correlated with global tract integrity. Deficits in the lexical route were associated with the posterior temporal segment of the inferior longitudinal fasciculus, where surface dyslexia errors correlated with reduced fractional anisotropy. In contrast, uncinate fasciculus involvement was more frequent among patients with intact reading, suggesting a specific association with preserved reading rather than a role in either reading route. These findings provide a fine-grained structure-function mapping of the dual route reading model onto white matter pathways. This deposit contains: 1. The de-identified SPSS dataset ( CORTEX-D-26-00041R1_SPSS_DATA.sav , N=37 post-metastatic exclusion) and accompanying SPSS syntax ( .sps ) underlying the statistical analyses reported in the manuscript 2. A full data codebook describing all variables, coding schemes, and derived measures 3. Per-patient MNI-space tract segmentations ( CORTEX-D-26-00041R1_MNI ) 4. Per-patient, per-tract FA and MD diffusion parameters ( CORTEX-D-26-00041R1_DTI_Reports ), provided as text files reporting percentile-based microstructural values sampled along the longitudinal course of each reconstructed tract Standardized psychometric test materials used in the reading battery are not included in this deposit due to publisher copyright restrictions on redistribution of test items. Raw diffusion-weighted MRI (dMRI) volumes are not included due to re-identifiability risk in this small, clinically-defined tumor cohort. Both might be available upon reasonable request, subject to appropriate licensing and/or institutional ethics approval - see the manuscript's Data Availability Statement for details.
Meekes, Lisa · Tabaro, Francesco · Bexkens, Michiel · et al.
41 files · 8.2 GB · csv, fasta, pdfdeclared
This record contains the Python software for PEPTiGEN, a tool for generating tryptic peptides from prokaryotic gene sequences and their variants, and the associated antimicrobial resistance (AMR) peptide database. The database is provided as an SQL file and a CSV file containing all genes and predicted peptides. The README file contains explanation of the PEPTiGEN tool. The SQL database schema files contains both the database schema of the SQL database used in the PEPTiGEN analysis as the database schema of the AMR peptide datbase.
1. May24_CIRBE-GOES_flux.h5 contains the necessary flux data to recreate the flux plots. 2. Oct24_CIRBE-GOES_flux.h5 contains the necessary flux data to recreate the flux plots.
Code and analysis for the manuscript "Cerebrospinal fluid profile, including α-Synuclein seeding activity, of p.A53T SNCA mutation carriers: Data from the PPMI Study."
This includes the source code, background plasma conditions, and simulation results that produced the simulation data reported in Green et al., "Evaluating the Impact of Multiscale E-Region Turbulence on HF/VHF Scintillation."
Wedig, Bryce · Daylan, Tansu · Huang, Alan · et al.
2 files · 6.5 GB · hdf5declared
Data Challenge Overview The Roman Space Telescope is expected to observe O(10^5) galaxy-galaxy strong gravitational lenses, providing high angular resolution images of galaxy-galaxy strong gravitational lenses that can be used to probe the nature of dark matter at sub-galactic scales ( Daylan and Birrer 2023 , Wedig et al. 2025 ). The Roman Data Challenge for Dark Matter Substructure with Galaxy-Galaxy Strong Gravitational Lenses provides realistic simulated Roman images of strong lenses with various dark matter substructure populations and challenges the community to test out substructure detection and characterization pipelines. Dataset Description The goal of this rung is to distinguish between mass distributions with Cold Dark Matter subhalos and no subhalos. In this rung, you will train a binary classifier to determine whether subhalos are present. This is the unlabeled dataset. It does not include the boolean substructure flag and a few other related parameters that were included in the labeled dataset. Rung 1 submissions will be scored for this dataset. Changelog v2.0: Fixes a bug where SNRs were calculated from 601 second exposures but images were simulated with exposure time of 610 seconds. The difference in SNR is approximately 1%. New major version because the systems are different from v1.0 v1.0: Initial version
Gupte, Nihar · Miller, M. Coleman · Udall, Rhiannon · et al.
36 files · 36 GB · hdf5, torchdeclared
Data release for the DINGO O4a eccentricity paper. It contains the per-event parameter-estimation products, population selection function, and hierarchical-inference posteriors needed to reproduce every figure, table, and number in the paper, plus the trained DINGO neural networks used for the analyses. Event data : eccentric, quasicircular, and precessing per-event posterior samples (posteriors_eccentric.h5, posteriors_quasicircular.h5, posteriors_precessing.h5); slimmed log-uniform-eccentricity-prior posteriors used as the hierarchical-likelihood input (posteriors_log_uniform_eccentric.h5); per-event posteriors reweighted by the population-informed posterior (posteriors_population_reweighted.h5); per-event summary statistics with pre-computed Bayes factors (summary_statistics.h5); e_gw conversions (egw_conversions.h5); and the eccentricity-mean-anomaly prior hull (e_zeta_prior_hull.h5). Selection function : the injection p_draw dataframe with detection probabilities including the analysis-window factor (injection_p_draw.h5), a fixed-injection eccentricity sweep (fixed_injection_ecc_sweep.h5), and matched-filter survival-function data (survival_function.h5). Hierarchical inference : the selection-corrected velocity-dispersion posterior marginalized over the GWTC-4 mass/spin/redshift hyperposterior (sigma_posterior.h5), the capture-eccentricity lookup table (capture_ecc_table.h5), the external GWTC-4 hyperposterior fit (gwtc4_hyperposterior.h5), and the GC/NSC branching-fraction posterior (branching_fraction_posterior.h5). Glitch analyses : glitch-marginalized posteriors for GW190701, GW231114_043211, and GW231223_032836. Networks : trained DINGO networks (SEOBNRv5EHM, SEOBNRv5HM, SEOBNRv5PHM) with their training settings; see MODEL_MANIFEST.md. Zenodo stores files flat; the companion code maps each file into the foldered layout the notebooks expect. Code to download the data and reproduce all figures: github.com/nihargupte-ph/o4a-eccentricity , archived at doi:10.5281/zenodo.21221948 .