Kovaliov, Michael
1 files · 100 MB · parquet
hybrid · semantic + lexical · 23 datasets ranked · 1.05s
Kovaliov, Michael
1 files · 100 MB · parquet
Omukuti, Rodney
1 files · 8.0 MB · vcf
Madhusudan, Gujral
6 files · 29 MB · parquetdeclared
Large language models (LLMs) are trained on massive, publicly available text datasets comprising trillions of tokens, enabling them to excel at general language tasks like next-token prediction. However, LLMs often struggle with domain-specific prompts, exhibiting reduced accuracy or generating inaccurate information (hallucinations). This is because they lack sufficient subject matter expertise. Two primary approaches exist to address this limitation for augmenting LLMs knowledge: Retrieval-Augmented Generation (RAG) and fine-tuning. This presentation focuses on fine-tuning smaller LLMs with domain-specific instruct datasets using the LoRA (Low-Rank Adaptation) technique on Gaudi hardware. We will leverage publicly available LLMs and datasets from the Hugging Face Hub for this demonstration. Though it is possible to fine tune LLMs with plain text data - sourced from documents, articles, and other materials.
Hackl, Jürgen
6 files · 132 MB · parquet, tiffdeclared
EuroFlood is an open, cloud-native index over the JRC/Copernicus CEMS-EFAS Satellite-Derived Flood Depth Maps for Europe (Betterle & Salamon, 2025; CC-BY-4.0) - ~3,280 satellite-derived observed flood-depth maps across Europe, 2015-2024. The bundle is a sparse Cloud-Optimized GeoTIFF encoding, per pixel, the set of flood events that inundated it, plus a combo_id -sorted GeoParquet dictionary and a small events table. Query by region and time via HTTP range reads (GDAL /vsicurl + DuckDB) to retrieve matching events, then fetch only the source depth rasters needed. Built with the open-source EuroFlood Python package ( pip install euroflood ).
Tiskus, Edvinas · Tiškuvienė, Rūta · Bučas, Martynas · et al.
37 files · 1.6 GB · hdf5, tiff, torchdeclared
This record contains the labeled data, trained models, and analysis code supporting the article "Comparing a Vision Foundation Model (DINOv3) and a Task-Specific U-Net for Mapping Emergent Aquatic Vegetation from Fused UAV Multispectral and LiDAR Data" (Remote Sensing in Ecology and Conservation). Contents: - masks/ : georeferenced ground-truth segmentation masks (five classes: aquatic vegetation, water, sand, other objects, background), aligned to the fused UAV orthomosaics and spanning 13 sites across nine Lithuanian waterbodies surveyed between May and August 2024. - models/ : the two final trained segmentation models, a Keras/HDF5 U-Net and a PyTorch DINOv3 model. - code/ : Python scripts for training, evaluation, the label-efficiency experiment, and full-scene prediction. The fused 9-band orthomosaics (five-band multispectral, RGB, and a LiDAR canopy height model; approximately 62 GB) are archived separately because of their size and are available from the corresponding author on request. The DINOv3 SAT-493M pretrained backbone is distributed by Meta under its own license and is not redistributed here; obtain it from the official DINOv3 release.
Sarma, R N
2 files · 88 MB · csv, vcfdeclared
This dataset contains the genotype and phenotype data generated for a genome-wide association study (GWAS) of agronomic traits in a diverse panel of 105 Indian mungbean ( Vigna radiata L. Wilczek) accessions . The dataset comprises a filtered genome-wide SNP dataset in Variant Call Format (VCF) and the corresponding Best Linear Unbiased Predictor (BLUP) values for the measured agronomic traits.
Njie, Adama · Torkayesh, Ali E · Venghaus, Prof. Dr. Sandra
10 files · 1.6 GB · csv, gzip, parquetdeclared
Structured, speaker-attributed corpus of all German Bundestag plenary session transcripts ( Plenarprotokolle ) from the first legislative period to the present (WP01-WP21, September 1949 - April 2026). Every attributed speech is extracted from the official PDFs published by the Deutscher Bundestag under open data policy and linked to the speaker's name, parliamentary role, party affiliation, and gender. Scale: 4,611 sessions · 1,033,723 speeches · 4,205 identified MdBs · 76 years of parliamentary debate Dataset files speeches.parquet - one row per attributed speech: speaker name, role, party, gender, stammdaten_id, full German text (~1 GB) persons.parquet - one row per MdB: cross-session identity linking all name variants via stammdaten_id; canonical name, birth date, career span, total speeches. Use this - not speakers.parquet - for person-level analysis sessions.parquet - one row per plenary session: date, city, Wahlperiode, source PDF hash, extraction engine speakers.parquet - name-string index: one row per unique name string as extracted from the transcripts. Useful for understanding extraction quality; not suitable for person-level aggregation (the same politician often appears under several name variants across sessions) parties.csv - reference table of 31 German parliamentary parties, 1949-present speeches.csv.gz - CSV fallback for Stata and Excel users (same columns as speeches.parquet) datapackage.json - Frictionless Data schema with column descriptions and foreign key constraints Cross-session identity The same politician often appears under different name strings across sessions (e.g. "Schmidt", "Dr. Schmidt", "Frau Dr. Schmidt"). Cross-session person linkage is provided via stammdaten_id , matched against the official Bundestag Stammdaten biographical XML. The persons.parquet table aggregates all name variants for the same MdB into one row with correctly summed speech counts, career span, and birth date. Coverage: ~98.5% of speeches are linked to a stammdaten_id; the remaining ~1.5% are ambiguous surname-only attributions or speakers not in the Stammdaten. Coverage and sources Source PDFs are the official Stenografische Berichte downloaded from the Bundestag open-data portal (bundestag.de). Party-share normalisation in the corpus statistics uses official seat counts per Wahlperiode sourced from the Federal Returning Officer (Bundeswahlleiter, bundeswahlleiter.de). Two PDF generations are covered: scanned and OCR'd documents (WP01-WP09, Bonn era, 1949-1987) and born-digital documents (WP10-WP21, 1987-present). The engine column in sessions.parquet flags whether pdftotext (born-digital) or pdfminer (OCR fallback) was used; this is the primary data-quality indicator for NLP use. Speaker attribution Each speech is attributed using four patterns extracted from the transcript format: presiding officers (Präsident/in, Vizepräsident/in), regular members (name + party), government officials (name + Bundeskanzler/in, Bundesminister/in, etc.), and procedural roles (Berichterstatter/in, etc.). The party field is null for ~60% of speeches - this is expected, as presiding officers and ministers are not identified by party in the transcript. Gender annotation & distribution Gender is derived by matching speaker names against the official Bundestag Stammdaten biographical XML (all MdBs since 1949), with fallbacks for role title, honorific prefix, manually researched overrides, and a gender_guesser first-name heuristic. The gender_source column distinguishes stammdaten (authoritative, 83%), role_title (gendered job title in attribution, 6.4%), title_prefix (Frau/Herr honorific, 0.5%), manual (historically researched, 2.9%), and inferred (name-based heuristic, 4.7%). Gender distribution: Female 26.6% · Male 73.4% · Unknown 0.0%. Data quality All speeches pass automated validation: zero null speaker names, zero sequence gaps, zero CID artefacts, zero party-misclassified-as-Bundesland errors. Eight sessions with conflicting source PDFs were deduplicated (first lexicographic occurrence retained). 252 non-person names incorrectly accepted by the parser (table headers, legislative terms, agenda fragments) are excluded at build time via a curated exclusion list. OCR sessions (WP01-WP09) may contain Unicode replacement characters (U+FFFD); the engine field identifies these sessions. Licence CC BY 4.0. The underlying Plenarprotokolle are official government documents of the Deutscher Bundestag and are in the public domain.
abdulwahab, samaa · aduallah, mahmood z. · Sallomi, Adheed H.
34 files · 2.7 GB · csv, gzip, parquetdeclared
Intrusion-detection research on Internet Protocol version 6 (IPv6) remains bottlenecked by the scarcity of labelled, protocol-aware flow datasets. Existing machine-learning IDS benchmarks are overwhelmingly IPv4-centric, and the few IPv6 corpora that have been released target narrow attack families or rely on small academic testbeds that cannot be re-created by third parties. We present IPv6-CyberBench, a reproducible eight-phase pipeline that constructs a large, protocol-aware translated-flow corpus by harmonising CIC-IDS-2017, CIC-IDS-2018 and CIC-DDoS-2019, applying deterministic IPv4→IPv6 address translation (6to4, NAT64, Teredo, EUI-64), synthesising 27 IPv6-specific flow features grouped in six protocol families, enforcing nineteen RFC-derived constraint categories together with temporal address dynamics, and rebalancing the long-tailed class distribution with a feature-group-conditioned per-class Wasserstein-GAN-GP augmenter and a SMOTE-KDE fallback selected per class by a formal decision rule. We scope the contribution honestly: because the seed corpora are IPv4 captures, the resulting 2,285,774-record benchmark is a translated-flow corpus suitable for training and evaluating flow-level IPv6 IDS classifiers on flooding, brute-force, scan, web-attack and infiltration traffic under IPv6 protocol-header semantics, and for studying IPv6-specific feature engineering and address dynamics in a reproducible setting. It is not a substitute for protocol-native IPv6 attack capture, and we explicitly exclude ICMPv6 Neighbour-Discovery flooding, SEND flooding, NDP exhaustion and extension-header covert-tunnelling from the threat model. The benchmark is evaluated on four axes - fidelity (Kolmogorov-Smirnov, MMD, Fréchet feature distance), utility (stratified 5×5 nested cross-validation over six classifier families including CNN-LSTM and LightGBM), privacy (Shokri-style membership-inference advantage AUC), and external fidelity against a 24 h anonymised CAIDA IPv6 trace (equinix-chicago, US backbone) and a MAWI samplepoint-F trace (WIDE backbone, Tokyo, Japan). The full pipeline, the hyper-parameter manifest, the RFC-constraint manifest, the reproduction scripts, and the 2,285,774-record benchmark are released unconditionally on Zenodo under CC BY 4.0; the dataset and pipeline are openly available at https://doi.org/10.5281/zenodo.19503446 (CC BY 4.0).
Bohl, Michael · Esteban-Medina, Marina · Lenhof, Kerstin · et al.
5 files · 5.0 GB · csv, torch, zipdeclared
This Zenodo record contains all data necessary to reproduce the benchmark results described in the following publication: M. Bohl, M. Esteban-Medina, N. Beerenwinkel, and K. Lenhof, Domain-adaptation deep learning models do not outperform simple baseline models in single-cell anti-cancer drug sensitivity prediction, bioRxiv (2026). Processed bulk and single-cell RNA-Seq datasets with response labels are in processed.zip. scATD model weights are in checkpoint_fold1_epoch_30.pth Full hyperparameter tuning logs/results are in hyperparam_tuning_results.csv A revised version of the source code (without model weights) is in code.zip. If it gets updated in the future, check the latest version at https://github.com/cbg-ethz/SC-Bulk-Domain-Adaptation/
BROCHARD, Pierre
3 files · 141 MB · torchdeclared
YOLOv26 models specialized in text region (TextRegion) and text line (TextLine) segmentation for medieval manuscripts. Sources: The Manicule corpus: https://nakala.fr/collection/10.34847/nkl.e0ef83vx The Alcar-HOME database: https://zenodo.org/record/5600884 The e-NDP corpus: https://zenodo.org/record/7575693 The Himanis project: https://zenodo.org/record/5535306 OCR ground truth for Caroline Miniscule : https://github.com/rescribe/carolineminuscule-groundtruth Ground Truth for ONB-Cod. 3891 : https://zenodo.org/record/7467249 Cremma Medieval : https://zenodo.org/record/7506657 DISTINGUO : https://doi.org/10.34847/NKL.48AD8B8D and synthetic data. HuggingFace Mirror : https://huggingface.co/LaMOP/Yolo-Seg-TextRegion-TextLine-Manuscript
Meekes, Lisa · Tabaro, Francesco · Bexkens, Michiel · et al.
41 files · 8.2 GB · csv, fasta, pdfdeclared
This record contains the Python software for PEPTiGEN, a tool for generating tryptic peptides from prokaryotic gene sequences and their variants, and the associated antimicrobial resistance (AMR) peptide database. The database is provided as an SQL file and a CSV file containing all genes and predicted peptides. The README file contains explanation of the PEPTiGEN tool. The SQL database schema files contains both the database schema of the SQL database used in the PEPTiGEN analysis as the database schema of the AMR peptide datbase.
Barreto, Bruno · Eisencraft, Marcio
3 files · 2.6 GB · parquetdeclared
This dataset provides a fixed benchmark dataset for stellar atmospheric parameter estimation from Sloan Digital Sky Survey Data Release 12 (SDSS DR12) optical stellar spectra. The dataset is organized into three predefined Parquet splits: 30,000 spectra for training, 5,000 spectra for validation, and 15,000 spectra for testing. Each row corresponds to one SDSS stellar spectrum and includes raw spectral arrays, fixed-length processed spectral features, source identifiers, basic metadata, and catalog stellar-parameter labels with their associated uncertainties. The supervised regression targets are the adopted catalog stellar atmospheric parameters: effective temperature (Teff, in K), metallicity ([Fe/H], in dex), and surface gravity (log g, in dex). The dataset also includes relevant observational and catalog information such as SDSS plate, MJD, fiber identifier, sky coordinates, signal-to-noise ratio, adopted radial velocity, raw flux, logarithmic wavelength grid, inverse variance, pixel mask, and processed flux features. This release is intended to support machine-learning research on stellar spectroscopy, including regression models for atmospheric parameter estimation, benchmark comparisons, uncertainty-aware evaluation, and experiments using either processed fixed-length spectra or native observed-frame spectral arrays.
Saldanha, Raphael
8 files · 601 MB · parquet, zipdeclared
Description This deposit contains annual, municipality-level datasets derived from the Brazilian Primary Health Care Information System (SISAB). The files combine two complementary data sources: Public SISAB Saúde report downloads from the Atendimento/Visita production report. CID-10 and CIAP-2 attendance data obtained from SISAB through requests under the Brazilian Access to Information Law (Lei de Acesso à Informação, LAI). The datasets are organized as tidy annual files in CSV (Zipped) and Parquet format. They are intended to support reproducible analysis of primary care production, procedures, evaluated problems/conditions, and CID/CIAP-coded attendances across Brazilian municipalities. The public SISAB report datasets are stratified by competence month, state, municipality, DataSUS age group, SISAB sex category, and the selected report category. For each competence month and report type, the extraction combines 36 stratified SISAB downloads: 18 age groups by 2 sex values. Monthly files are merged into yearly files, completing missing combinations of observed competence, municipality, age group, sex, and category with valor = 0 . The LAI dataset contains yearly CID-10 and CIAP-2 attendance counts by competence month, municipality, code type, and code. When multiple valid LAI files cover the same competence, the processing pipeline selects the file with the largest number of data rows, using file size and request folder order as tie-breakers. Provenance columns identify the selected LAI request and source file. Variables SISAB Saúde Produção Columns: competencia : competence month in YYYYMM format. uf : Brazilian state abbreviation. ibge : municipality IBGE code. municipio : municipality name. faixa_etaria : Age group. sexo : SISAB sex category, Masculino or Feminino . tipo_producao : production type from the SISAB report. valor : count reported by SISAB. SISAB Saúde Procedimento Columns: competencia : competence month in YYYYMM format. uf : Brazilian state abbreviation. ibge : municipality IBGE code. municipio : municipality name. faixa_etaria : age group. sexo : SISAB sex category, Masculino or Feminino . procedimento : procedure from the SISAB report. valor : count reported by SISAB SISAB Saúde Condição Avaliada Columns: competencia : competence month in YYYYMM format. uf : Brazilian state abbreviation. ibge : municipality IBGE code. municipio : municipality name. faixa_etaria : age group. sexo : SISAB sex category, Masculino or Feminino . condicao_avaliada : evaluated problem or condition from the SISAB report. valor : count reported by SISAB. SISAB LAI CID/CIAP Columns: ano_competencia : competence year. competencia : competence month in YYYYMM format. competencia_date : first day of the competence month. co_municipio_ibge : municipality IBGE code. tp_codigo : code type, CID or CIAP . codigo : CID-10 or CIAP-2 code. qt_atendimentos : number of attendances. source_request : selected LAI request folder. source_file : selected source CSV file. Methods The public SISAB report files were generated with the sisab_scrapper processing pipeline. For each month, report type, age group, and sex value, the pipeline downloads the all-Brazil municipality report from SISAB, validates the returned CSV, preserves raw cache files for resumable runs, and writes a sorted monthly tidy dataset. The yearly merge validates required columns, expected age groups, expected sex values, category values, and month gaps unless explicitly allowed. The CID/CIAP files were generated with the sisab_lai processing pipeline. The pipeline imports CSV files received through LAI requests, detects the real CSV header after any SQL*Plus preamble, validates candidate files, resolves overlapping requests by competence, standardizes old and new schemas into one tidy table, and exports annual CSV and Parquet files together with audit reports. Sources - SISAB public reports, Ministry of Health, Brazil: https://sisab.saude.gov.br/ - SISAB LAI files obtained through Brazilian Access to Information Law requests. - Processing code for public SISAB report data: https://github.com/rfsaldanha/sisab_scrapper - Processing code for LAI CID/CIAP data: https://github.com/rfsaldanha/sisab_lai Notes - Counts are aggregated administrative records and should be interpreted in light of SISAB reporting practices, data quality, and changes in municipal reporting coverage. - Municipality boundaries, names, and coding practices may vary over time. - Public SISAB report datasets are completed with zero values only for combinations defined by observed municipalities, observed competencies, all expected age groups, both expected sex values, and observed report categories within the yearly merge. - LAI CID/CIAP data preserves selected source-file provenance through source_request and source_file . - This deposit corresponds to an individual year. Deposits for other years are published separately.
Yang, Jiazhuo · Cai, Hongyan · Xu, Xinliang · et al.
3 files · 408 MB · tiff, torchdeclared
This record contains the 10 m farmland shelterbelt distribution products for Northeast China in 2020 and 2025, together with the pretrained deep learning model weights used for farmland shelterbelt inference. The dataset was generated from spring Sentinel-2 surface reflectance composites using B4, B8, and NDVI features and a ResNet-50/CA deep learning model. The two GeoTIFF files represent binary farmland shelterbelt maps, where 1 indicates farmland shelterbelt and 0 indicates non-shelterbelt. The pretrained model weights are provided as ForestNet50V1.pth to support reproducible inference and local model adaptation. The dataset covers major agricultural regions of Northeast China, including Heilongjiang, Jilin, Liaoning, and eastern Inner Mongolia. It is intended for regional- and landscape-scale analyses of farmland shelterbelt distribution, spatial continuity, fragmentation, stage-based change, and ecological engineering assessment.
Anonymous Authors
15 files · 2.2 GB · parquet, rar, torchdeclared
Economic losses caused by extreme climate are not confined to the locations where events occur, but can propagate across regions through physical and economic linkages. Yet existing climate-impact assessment methods remain poorly suited to tracing how shocks spread across space and reshape the geography of economic loss. Here we develop a mechanistically informed multi-scale spatiotemporal autoregressive graph neural network model to quantify spatially cascading climate impacts. The model couples scale-specific, physically structured spatiotemporal autoregressive processes through an adaptive gating mechanism, allowing heterogeneous cross-scale interactions to be learned from data. Model estimation is achieved through tailored graph convolutional neural networks that are mathematically equivalent to spatiotemporal autoregressive models, enabling scalability while preserving transparent parameter interpretation. Monte Carlo simulation experiments show that the model accurately recovers true parameters and distinguishes between scale-dependent processes. Applying the framework to extreme precipitation, we find that large-scale upwind-to-downwind cascades driven by atmospheric circulations dominate aggregated economic losses. A one-standard-deviation increase in log extreme precipitation is associated with a 0.19 percentage-point decline in economic growth rate at the large scale, with 62.3% of the loss arising from spatial cascades. These findings highlight the need for transboundary risk governance that incorporates spatial cascading into climate-extremes monitoring and early-warning. Description of the uploaded file Monte Carlo simulation code data_generator_factors.py: Data generation script for multi-scale Monte Carlo simulation experiments. sarnn_model.py:Implementation of the proposed MS-STARGNNs model architecture definition. train.py:Training pipeline script for the MS-STARGNNs model. Data and spatial weights matrices for empirical analysis global_panel_1deg_std.parquet:Standardized Large-scale (1°) datase; global_panel_2km_std.parquet:Standardized Small-scale (2 km) dataset. W_global_2km_knn8.pt:Small-scale spatial weights matrix based on 8-nearest neighbors (KNN8). W_Large-scale:A large-scale spatial weights matrix derived from moisture transport pathways (2005-2021) mapping_1deg_to_2km.parquet: Correspondence file mapping large-scale (1°) grid cells to small-scale (2km) grid cells. ipcc_region_mapping_coarse_8regions.parquet: Mapping file linking Large-scale grid cells to the 8 IPCC AR6 reference regions. ipcc_region_mapping_fine_8regions.parquet: Mapping file linking Small-scale grid cells to the 8 IPCC AR6 reference regions. Code model.py: Core architecture definitions for empirical analysis. Provides the base classes and computational layers engineered to handle real-world geospatial complexities. All subsequent training scripts import modules from this file. STARGNNs.py: Implementation of the single-scale baseline. Serves as a reference point for evaluating the efficacy of cross-scale feature fusion. MS-STARGNNs_fixed.py: Configuration script for the MS-STARGNNs model utilizing fixed autoregressive coefficients. MS-STARGNNs.py: Configuration script for the MS-STARGNNs model utilizing annually varying autoregressive coefficients. MS-STARGNNs_8 regions.py: Executes the MS-STARGNNs model with decoupled regional parameters, loading unique autoregressive weights and βvectors for each IPCC region. Implements null value handling for regions lacking observational data (e.g., Antarctica).
zhang, Wenjun
10 files · 459 MB · torchdeclared
This archive contains full training dataset, preprocessing scaler .pkl files and five groups of pre-trained neural network checkpoints (.pth) for reproducing all results in the manuscript. Corresponding source code repository on GitHub: https://github.com/wenjunzhang2020/Prediction-of-combined-effects-of-binary-mixed-systems .
Mutluel, Abdullah Mevlüt
12 files · 449 KB · pdf, torchdeclared
Code and data accompanying the manuscript "Disentangling Electro-osmotic Drag and Back-Diffusion Water Fluxes in Polymer Electrolyte Membranes: A Physics-Informed Neural Network for Net-Flux Inversion" submitted to the Journal of the Electrochemical Society. Contents: - gen_data_lit.py: synthetic benchmark generation. Solves the through-plane membrane water transport boundary-value problem for 24 operating conditions using the Springer drag coefficient and the Nguyen-White diffusion correlation. - pinn_lit.py: physics-informed neural network training. Recovers n_d(lambda) and D_w(lambda) from net-flux data with a hard conservation constraint and a single anchor point. - plots_lit.py: reproduces all manuscript figures (Figs. 1-3). - rev_lib.py: validation experiments, including trend-line baseline comparison, noise robustness, non-monotonic diffusion coefficient recovery, anchor ablation, multi-seed statistics, and held-out condition tests (Fig. 4 and Table 2). - ga_final.py: graphical abstract. - data_lit.json: generated benchmark data (24 operating conditions with internal water-content profiles and net fluxes). - model_lit.pt: trained PyTorch model weights. - manuscript.tex and figure PDFs. Requirements: Python 3 with PyTorch, NumPy, SciPy, and Matplotlib. Run gen_data_lit.py first, then pinn_lit.py, and then plots_lit.py to reproduce the results from scratch, or load model_lit.pt directly with the Model class defined in pinn_lit.py.
David, Cédric
19 files · 370 KB · netcdf, parquetdeclared
RAPID comes along with a set of test files based on a synthetic experiment called the Sandbox. More information is available in SANDBOX.md .
S.S, Ashwin · Minami, Katsuhiko · Nakazato, Kako
11 files · 124 MB · jpeg, torch, zipdeclared
Supplementary code for: Katsuhiko Minami, Kako Nakazato, Sachiko Tamura, S. S. Ashwin, Kazuhiro Maeshima* Machine learning-assisted Repli-Histo labeling reveals distinct transcription-dependent constraints on chromatin motion in living cells (2026).
Gupte, Nihar · Miller, M. Coleman · Udall, Rhiannon · et al.
36 files · 36 GB · hdf5, torchdeclared
Data release for the DINGO O4a eccentricity paper. It contains the per-event parameter-estimation products, population selection function, and hierarchical-inference posteriors needed to reproduce every figure, table, and number in the paper, plus the trained DINGO neural networks used for the analyses. Event data : eccentric, quasicircular, and precessing per-event posterior samples (posteriors_eccentric.h5, posteriors_quasicircular.h5, posteriors_precessing.h5); slimmed log-uniform-eccentricity-prior posteriors used as the hierarchical-likelihood input (posteriors_log_uniform_eccentric.h5); per-event posteriors reweighted by the population-informed posterior (posteriors_population_reweighted.h5); per-event summary statistics with pre-computed Bayes factors (summary_statistics.h5); e_gw conversions (egw_conversions.h5); and the eccentricity-mean-anomaly prior hull (e_zeta_prior_hull.h5). Selection function : the injection p_draw dataframe with detection probabilities including the analysis-window factor (injection_p_draw.h5), a fixed-injection eccentricity sweep (fixed_injection_ecc_sweep.h5), and matched-filter survival-function data (survival_function.h5). Hierarchical inference : the selection-corrected velocity-dispersion posterior marginalized over the GWTC-4 mass/spin/redshift hyperposterior (sigma_posterior.h5), the capture-eccentricity lookup table (capture_ecc_table.h5), the external GWTC-4 hyperposterior fit (gwtc4_hyperposterior.h5), and the GC/NSC branching-fraction posterior (branching_fraction_posterior.h5). Glitch analyses : glitch-marginalized posteriors for GW190701, GW231114_043211, and GW231223_032836. Networks : trained DINGO networks (SEOBNRv5EHM, SEOBNRv5HM, SEOBNRv5PHM) with their training settings; see MODEL_MANIFEST.md. Zenodo stores files flat; the companion code maps each file into the foldered layout the notebooks expect. Code to download the data and reproduce all figures: github.com/nihargupte-ph/o4a-eccentricity , archived at doi:10.5281/zenodo.21221948 .