Kantor, Rose · Shakya, Migun · Ruth, Nelson · et al.
2,095 rows · 907 KB · fasta, tsv
A virus genome database representing 21,015 near-complete virus genomes collected from untargeted ultra-deep RNA/DNA combined sequencing of wastewater. Sequence data was provided by the CASPER consortium and raw data may be found on NCBI SRA under bioprojects PRJNA1247874 and PRJNA1198001. Data underwent read trimming, rRNA and human read removal, de novo assembly, and selection of high-quality viral contigs. Contigs were clustered at 95% identity and 85% query coverage to dereplicate. Chimera-checking required at least two independent assemblies of the same viral genome or presence of the genome in another reference database. Annotation made use of RdRpCATCH, geNomad, checkV, BLASTN against NCBI core-nt, and RNAVirHost. The RdRp fasta files contain representative RdRp sequences identified through homology to major RdRp reference databases and clustered at 90% sequence identity over 75% sequence coverage. Included sequences contain all three conserved RdRp motifs (A, B, and C) arranged in either the canonical ABC configuration or the permuted CAB configuration.
Hubbard, Alfred · Solares, Edwin · Hemming-Schroeder, Elizabeth
88 rows · 18 KB · fasta
These are the files needed to run the Broad Institute's malaria amplicon pipeline for the PvGAP Plasmodium vivax panel, described in detail here . They consist of FASTA files containing the forward and reverse primers and another FASTA file containing reference sequences for each target, derived from the PvP01 reference genome.
This dataset supports the findings of the manuscript "Root anatomical traits modulate the assembly and nitrogen transformation potential of root-associated microbiomes in a temperate steppe" (NPH-MS-2026-55667). It contains root traits data, bacterial 16S rRNA gene absolute abundances, functional genes relative abundances, DNA extraction metadata, and phylogenetic marker sequences for 37 plant species from a temperate steppe ecosystem. The dataset includes the following files: 1. root traits.csv - Root traits including average diameter (AD), specific root length (SRL), specific root area (SRA), root tissue density (RTD), root nitrogen content (RNC), root carbon content (RCC), carbon‑nitrogen ratio (RCN), cortex layer number (CLN), cortex thickness (CT), and the ratio of cortex thickness to root diameter (CTRD). The first column lists plant species names. 2. Absolute abundance of 16S rRNA gene.csv - Quantitative PCR (qPCR) derived absolute abundances of bacterial 16S rRNA gene copies (copies/ng DNA) across different root compartments (rhizosphere, rhizoplane, endosphere), host species, root orders, and cotyledon classes (monocot/dicot). 3. DNA extraction sample weight.csv - Fresh weight (grams) of root material used for DNA extraction for each sample, linked by SampleID to the abundance data. 4. DNA extraction concentration.csv - Qubit‑measured DNA concentrations (ng/μL) and the sample volume (μL) used for quality control, together with sample metadata. 5. 37species.fasta - DNA sequences of two chloroplast markers (matK and rbcL) for the 37 plant species included in the study. The sequences are in FASTA format with headers formatted as ">Species". These were used for host phylogeny construction and Pagel's λ analyses. 6. Quantitative PCR results of functional gene.csv - Quantitative PCR (qPCR) derived relative abundances of bacterial 16S rRNA gene and functional genes across different root compartments (rhizosphere, rhizoplane, endosphere), host species, root orders, and cotyledon classes (monocot/dicot). 7. README.md - A detailed description of each file, column headers, abbreviations, units, and any missing value codings (NA). All data are provided to ensure transparency and reproducibility of the analyses. For methodological details, please refer to the Materials and Methods section of the associated publication. These data are under embargo until the associated research article is published. After that date, they will be freely available under a Creative Commons Attribution 4.0 International (CC BY 4.0) license. During the embargo period, the metadata (title, authors, abstract) and the DOI remain publicly visible, but the data files are not accessible. For access requests before the embargo expires, please contact the corresponding author.
Zahid, R · Lázaro, A · Moreno‐Alcántar, G · et al.
1 files · 82 KB · tar
Synthetic nanozymes have emerged as promising alternatives to natural enzymes for catalytic and therapeutic applications, yet their limited stability, aqueous compatibility, and catalytic scope impede broader utilization. Here, we report a mild, one-step sol-gel synthesis that yields ultrasmall, water-stable octa-amino silsesquioxanes functioning as metal-free nanozymes. These minimalistic nanostructures exhibit aldolase-like organocatalytic activity in water and enable dynamic, stimuli-responsive modulation of catalysis through reversible supramolecular aggregation and disaggregation triggered by specific chemical inputs, thus forming a multifunctional platform for tunable catalysis and biomedical applications. Structural simplicity, stability, and functional versatility together permit tunable, enzyme-like catalysis in water without auxiliary surfactants or phase-transfer additives. Furthermore, the nanozymes display high biocompatibility and efficient cellular internalization, enabling their use in living cells, for instance, as intracellular prodrug activators via retro-aldol activation of a doxorubicin prodrug in human glioblastoma and metastatic melanoma cells, resulting in selective cytotoxicity. This system provides a cost-effective, sustainable, and scalable platform for water-compatible, metal-free organocatalysis that bridges abiotic catalysis and biological function. These findings demonstrate how rationally designed silsesquioxane frameworks can emulate natural enzyme reactivity while integrating adaptive, stimuli-responsive behavior, broadening the applicability of synthetic nanozymes to catalytic and therapeutic contexts.
Requena Rolanía, Jose María · Greif, Gonzalo · ROBELLO, CARLOS
1 files · 8.0 MB · fasta
This dataset contains the genome sequence for Trypanosoma cruzi (strain Dm28c). This genome sequence was de novo assembled using PacBio Hi-Fi and Illumina sequencing platforms by Greif et al (2026. PMID: 41501640). The genome was assembled into 32 contigs, which represent complete chromosomes. The provided Fasta file also contains an additional contig corresponding to the maxicircle (mitochondrial genome) sequence. The Fasta files included in this dataset were downloaded from GenBank (assembly GCA_044048535.1; May 22, 2026). Additional information about the Dm28cT2T genome assembly and gene annotations may be accessed through the link: https://cruzi.pasteur.uy/
Burman, Nathaniel · Buyukyoruk, Murat · Wiegand, Tanner · et al.
4 files · 8.0 MB · fasta
This folder contains a multiple sequence alignment of Cas7 homologs in .fasta format, the domain-level annotations from PFAM and CasFinder, and an associated phylogenetic tree in .newick format.
Large language models (LLMs) are trained on massive, publicly available text datasets comprising trillions of tokens, enabling them to excel at general language tasks like next-token prediction. However, LLMs often struggle with domain-specific prompts, exhibiting reduced accuracy or generating inaccurate information (hallucinations). This is because they lack sufficient subject matter expertise. Two primary approaches exist to address this limitation for augmenting LLMs knowledge: Retrieval-Augmented Generation (RAG) and fine-tuning. This presentation focuses on fine-tuning smaller LLMs with domain-specific instruct datasets using the LoRA (Low-Rank Adaptation) technique on Gaudi hardware. We will leverage publicly available LLMs and datasets from the Hugging Face Hub for this demonstration. Though it is possible to fine tune LLMs with plain text data - sourced from documents, articles, and other materials.
EuroFlood is an open, cloud-native index over the JRC/Copernicus CEMS-EFAS Satellite-Derived Flood Depth Maps for Europe (Betterle & Salamon, 2025; CC-BY-4.0) - ~3,280 satellite-derived observed flood-depth maps across Europe, 2015-2024. The bundle is a sparse Cloud-Optimized GeoTIFF encoding, per pixel, the set of flood events that inundated it, plus a combo_id -sorted GeoParquet dictionary and a small events table. Query by region and time via HTTP range reads (GDAL /vsicurl + DuckDB) to retrieve matching events, then fetch only the source depth rasters needed. Built with the open-source EuroFlood Python package ( pip install euroflood ).
Historical Hydrological Simulations over the South African Domain (1990-2024) The mHM's simulations of the Planet4Health project This dataset contains historical hydrological simulations for the South African domain (domain 1020011530) conducted with the Mesoscale Hydrological Model (mHM) at a spatial resolution of 0.015625°. The simulation period spans 1990-2024 and was part of the Planet4Health (P4H) project, utilising the ERA5 meteorological forcing. This archive is prepared for DOI assignment and ensures long-term reproducibility. It includes relevant clipped NetCDF components for soil moisture layers 4, 5 and 6, consistent with the infrastructure provided within the Helmholtz Centre for Environmental Research (UFZ). The simulations were executed using a specific version of the mHM model with the SCC method for gauges, paired with the mRMv1.0 routing configuration. 🛰️ Simulation Details Model: Mesoscale Hydrological Model (mHM) Codebase: scc_for_gauges branch (https://git.ufz.de/shresthp/mhm/-/tree/scc_for_gauges?ref_type=heads) Spatial resolution: 0.015625° Temporal resolution: Daily Simulation period: 1990-2024 Simulation type: Historical simulation Spin-up: 30-year spin-up using 1990-2019 ERA5 climatology Model version: v1.0 Setup Scope: Model run for domain 1020011530, post-processed and clipped. Configuration & Modules The configuration utilises standard structural components with the SCC methodology. Modules included: Snow processes: Degree-day method Soil moisture: Feddes equation for evapotranspiration reduction Infiltration: Multi-layer Brooks-Corey-like approach Direct runoff: Linear reservoir exceedance method Potential evapotranspiration: Hargreaves-Samani method Interflow: Storage reservoir with nonlinear outflow Groundwater: Linear reservoir Routing: Adaptive time-step routing with mRMv1.0 mechanisms 📥 Input Datasets Meteorological Forcing: ERA5 (Hersbach et al., 2020) at a native input meteorological resolution of 0.25°, dynamically downscaled/mapped to model requirements. Processing Infrastructure: Tracked, processed, and validated under the Planet4Health deployment pipeline (https://git.ufz.de/planet4health/mhm_production/-/tree/main/postproc?ref_type=heads). Data Interfaces: Climate Data Interface version 2.2.4 (CDI) | Climate Data Operators version 2.2.2 (CDO) | NetCDF Operators version 5.1.7 (NCO). 📤 Output Variables sm_l04: Volumetric soil moisture layer 4 (300-500 mm depth) [mm mm-1, fraction between 0 and 1] sm_l05: Volumetric soil moisture layer 5 (500-1000 mm depth) [mm mm-1, fraction between 0 and 1] sm_l06: Volumetric soil moisture layer 6 (1000-2000 mm depth) [mm mm-1, fraction between 0 and 1] 📫 Contact Ehsan Modiri - ehsan.modiri@ufz.de Pallav Kumar Shrestha - pallav-kumar.shrestha@ufz.de Institution Helmholtz Centre for Environmental Research - UFZ, Department of Computational Hydrosystems
Njie, Adama · Torkayesh, Ali E · Venghaus, Prof. Dr. Sandra
10 files · 1.6 GB · csv, gzip, parquetdeclared
Structured, speaker-attributed corpus of all German Bundestag plenary session transcripts ( Plenarprotokolle ) from the first legislative period to the present (WP01-WP21, September 1949 - April 2026). Every attributed speech is extracted from the official PDFs published by the Deutscher Bundestag under open data policy and linked to the speaker's name, parliamentary role, party affiliation, and gender. Scale: 4,611 sessions · 1,033,723 speeches · 4,205 identified MdBs · 76 years of parliamentary debate Dataset files speeches.parquet - one row per attributed speech: speaker name, role, party, gender, stammdaten_id, full German text (~1 GB) persons.parquet - one row per MdB: cross-session identity linking all name variants via stammdaten_id; canonical name, birth date, career span, total speeches. Use this - not speakers.parquet - for person-level analysis sessions.parquet - one row per plenary session: date, city, Wahlperiode, source PDF hash, extraction engine speakers.parquet - name-string index: one row per unique name string as extracted from the transcripts. Useful for understanding extraction quality; not suitable for person-level aggregation (the same politician often appears under several name variants across sessions) parties.csv - reference table of 31 German parliamentary parties, 1949-present speeches.csv.gz - CSV fallback for Stata and Excel users (same columns as speeches.parquet) datapackage.json - Frictionless Data schema with column descriptions and foreign key constraints Cross-session identity The same politician often appears under different name strings across sessions (e.g. "Schmidt", "Dr. Schmidt", "Frau Dr. Schmidt"). Cross-session person linkage is provided via stammdaten_id , matched against the official Bundestag Stammdaten biographical XML. The persons.parquet table aggregates all name variants for the same MdB into one row with correctly summed speech counts, career span, and birth date. Coverage: ~98.5% of speeches are linked to a stammdaten_id; the remaining ~1.5% are ambiguous surname-only attributions or speakers not in the Stammdaten. Coverage and sources Source PDFs are the official Stenografische Berichte downloaded from the Bundestag open-data portal (bundestag.de). Party-share normalisation in the corpus statistics uses official seat counts per Wahlperiode sourced from the Federal Returning Officer (Bundeswahlleiter, bundeswahlleiter.de). Two PDF generations are covered: scanned and OCR'd documents (WP01-WP09, Bonn era, 1949-1987) and born-digital documents (WP10-WP21, 1987-present). The engine column in sessions.parquet flags whether pdftotext (born-digital) or pdfminer (OCR fallback) was used; this is the primary data-quality indicator for NLP use. Speaker attribution Each speech is attributed using four patterns extracted from the transcript format: presiding officers (Präsident/in, Vizepräsident/in), regular members (name + party), government officials (name + Bundeskanzler/in, Bundesminister/in, etc.), and procedural roles (Berichterstatter/in, etc.). The party field is null for ~60% of speeches - this is expected, as presiding officers and ministers are not identified by party in the transcript. Gender annotation & distribution Gender is derived by matching speaker names against the official Bundestag Stammdaten biographical XML (all MdBs since 1949), with fallbacks for role title, honorific prefix, manually researched overrides, and a gender_guesser first-name heuristic. The gender_source column distinguishes stammdaten (authoritative, 83%), role_title (gendered job title in attribution, 6.4%), title_prefix (Frau/Herr honorific, 0.5%), manual (historically researched, 2.9%), and inferred (name-based heuristic, 4.7%). Gender distribution: Female 26.6% · Male 73.4% · Unknown 0.0%. Data quality All speeches pass automated validation: zero null speaker names, zero sequence gaps, zero CID artefacts, zero party-misclassified-as-Bundesland errors. Eight sessions with conflicting source PDFs were deduplicated (first lexicographic occurrence retained). 252 non-person names incorrectly accepted by the parser (table headers, legislative terms, agenda fragments) are excluded at build time via a curated exclusion list. OCR sessions (WP01-WP09) may contain Unicode replacement characters (U+FFFD); the engine field identifies these sessions. Licence CC BY 4.0. The underlying Plenarprotokolle are official government documents of the Deutscher Bundestag and are in the public domain.
abdulwahab, samaa · aduallah, mahmood z. · Sallomi, Adheed H.
34 files · 2.7 GB · csv, gzip, parquetdeclared
Intrusion-detection research on Internet Protocol version 6 (IPv6) remains bottlenecked by the scarcity of labelled, protocol-aware flow datasets. Existing machine-learning IDS benchmarks are overwhelmingly IPv4-centric, and the few IPv6 corpora that have been released target narrow attack families or rely on small academic testbeds that cannot be re-created by third parties. We present IPv6-CyberBench, a reproducible eight-phase pipeline that constructs a large, protocol-aware translated-flow corpus by harmonising CIC-IDS-2017, CIC-IDS-2018 and CIC-DDoS-2019, applying deterministic IPv4→IPv6 address translation (6to4, NAT64, Teredo, EUI-64), synthesising 27 IPv6-specific flow features grouped in six protocol families, enforcing nineteen RFC-derived constraint categories together with temporal address dynamics, and rebalancing the long-tailed class distribution with a feature-group-conditioned per-class Wasserstein-GAN-GP augmenter and a SMOTE-KDE fallback selected per class by a formal decision rule. We scope the contribution honestly: because the seed corpora are IPv4 captures, the resulting 2,285,774-record benchmark is a translated-flow corpus suitable for training and evaluating flow-level IPv6 IDS classifiers on flooding, brute-force, scan, web-attack and infiltration traffic under IPv6 protocol-header semantics, and for studying IPv6-specific feature engineering and address dynamics in a reproducible setting. It is not a substitute for protocol-native IPv6 attack capture, and we explicitly exclude ICMPv6 Neighbour-Discovery flooding, SEND flooding, NDP exhaustion and extension-header covert-tunnelling from the threat model. The benchmark is evaluated on four axes - fidelity (Kolmogorov-Smirnov, MMD, Fréchet feature distance), utility (stratified 5×5 nested cross-validation over six classifier families including CNN-LSTM and LightGBM), privacy (Shokri-style membership-inference advantage AUC), and external fidelity against a 24 h anonymised CAIDA IPv6 trace (equinix-chicago, US backbone) and a MAWI samplepoint-F trace (WIDE backbone, Tokyo, Japan). The full pipeline, the hyper-parameter manifest, the RFC-constraint manifest, the reproduction scripts, and the 2,285,774-record benchmark are released unconditionally on Zenodo under CC BY 4.0; the dataset and pipeline are openly available at https://doi.org/10.5281/zenodo.19503446 (CC BY 4.0).
WHATCH'EM (Water Height and Temperature in Container Habitats Energy Model) is a physics-based model that simulates water temperature and water height in containers using an energy balance approach. The model uses meteorological inputs together with container characteristics, shading, rainfall, evaporation, and optional manual water additions to simulate container water dynamics across a range of environmental conditions.
Meekes, Lisa · Tabaro, Francesco · Bexkens, Michiel · et al.
41 files · 8.2 GB · csv, fasta, pdfdeclared
This record contains the Python software for PEPTiGEN, a tool for generating tryptic peptides from prokaryotic gene sequences and their variants, and the associated antimicrobial resistance (AMR) peptide database. The database is provided as an SQL file and a CSV file containing all genes and predicted peptides. The README file contains explanation of the PEPTiGEN tool. The SQL database schema files contains both the database schema of the SQL database used in the PEPTiGEN analysis as the database schema of the AMR peptide datbase.
This dataset provides a fixed benchmark dataset for stellar atmospheric parameter estimation from Sloan Digital Sky Survey Data Release 12 (SDSS DR12) optical stellar spectra. The dataset is organized into three predefined Parquet splits: 30,000 spectra for training, 5,000 spectra for validation, and 15,000 spectra for testing. Each row corresponds to one SDSS stellar spectrum and includes raw spectral arrays, fixed-length processed spectral features, source identifiers, basic metadata, and catalog stellar-parameter labels with their associated uncertainties. The supervised regression targets are the adopted catalog stellar atmospheric parameters: effective temperature (Teff, in K), metallicity ([Fe/H], in dex), and surface gravity (log g, in dex). The dataset also includes relevant observational and catalog information such as SDSS plate, MJD, fiber identifier, sky coordinates, signal-to-noise ratio, adopted radial velocity, raw flux, logarithmic wavelength grid, inverse variance, pixel mask, and processed flux features. This release is intended to support machine-learning research on stellar spectroscopy, including regression models for atmospheric parameter estimation, benchmark comparisons, uncertainty-aware evaluation, and experiments using either processed fixed-length spectra or native observed-frame spectral arrays.
Description This deposit contains annual, municipality-level datasets derived from the Brazilian Primary Health Care Information System (SISAB). The files combine two complementary data sources: Public SISAB Saúde report downloads from the Atendimento/Visita production report. CID-10 and CIAP-2 attendance data obtained from SISAB through requests under the Brazilian Access to Information Law (Lei de Acesso à Informação, LAI). The datasets are organized as tidy annual files in CSV (Zipped) and Parquet format. They are intended to support reproducible analysis of primary care production, procedures, evaluated problems/conditions, and CID/CIAP-coded attendances across Brazilian municipalities. The public SISAB report datasets are stratified by competence month, state, municipality, DataSUS age group, SISAB sex category, and the selected report category. For each competence month and report type, the extraction combines 36 stratified SISAB downloads: 18 age groups by 2 sex values. Monthly files are merged into yearly files, completing missing combinations of observed competence, municipality, age group, sex, and category with valor = 0 . The LAI dataset contains yearly CID-10 and CIAP-2 attendance counts by competence month, municipality, code type, and code. When multiple valid LAI files cover the same competence, the processing pipeline selects the file with the largest number of data rows, using file size and request folder order as tie-breakers. Provenance columns identify the selected LAI request and source file. Variables SISAB Saúde Produção Columns: competencia : competence month in YYYYMM format. uf : Brazilian state abbreviation. ibge : municipality IBGE code. municipio : municipality name. faixa_etaria : Age group. sexo : SISAB sex category, Masculino or Feminino . tipo_producao : production type from the SISAB report. valor : count reported by SISAB. SISAB Saúde Procedimento Columns: competencia : competence month in YYYYMM format. uf : Brazilian state abbreviation. ibge : municipality IBGE code. municipio : municipality name. faixa_etaria : age group. sexo : SISAB sex category, Masculino or Feminino . procedimento : procedure from the SISAB report. valor : count reported by SISAB SISAB Saúde Condição Avaliada Columns: competencia : competence month in YYYYMM format. uf : Brazilian state abbreviation. ibge : municipality IBGE code. municipio : municipality name. faixa_etaria : age group. sexo : SISAB sex category, Masculino or Feminino . condicao_avaliada : evaluated problem or condition from the SISAB report. valor : count reported by SISAB. SISAB LAI CID/CIAP Columns: ano_competencia : competence year. competencia : competence month in YYYYMM format. competencia_date : first day of the competence month. co_municipio_ibge : municipality IBGE code. tp_codigo : code type, CID or CIAP . codigo : CID-10 or CIAP-2 code. qt_atendimentos : number of attendances. source_request : selected LAI request folder. source_file : selected source CSV file. Methods The public SISAB report files were generated with the sisab_scrapper processing pipeline. For each month, report type, age group, and sex value, the pipeline downloads the all-Brazil municipality report from SISAB, validates the returned CSV, preserves raw cache files for resumable runs, and writes a sorted monthly tidy dataset. The yearly merge validates required columns, expected age groups, expected sex values, category values, and month gaps unless explicitly allowed. The CID/CIAP files were generated with the sisab_lai processing pipeline. The pipeline imports CSV files received through LAI requests, detects the real CSV header after any SQL*Plus preamble, validates candidate files, resolves overlapping requests by competence, standardizes old and new schemas into one tidy table, and exports annual CSV and Parquet files together with audit reports. Sources - SISAB public reports, Ministry of Health, Brazil: https://sisab.saude.gov.br/ - SISAB LAI files obtained through Brazilian Access to Information Law requests. - Processing code for public SISAB report data: https://github.com/rfsaldanha/sisab_scrapper - Processing code for LAI CID/CIAP data: https://github.com/rfsaldanha/sisab_lai Notes - Counts are aggregated administrative records and should be interpreted in light of SISAB reporting practices, data quality, and changes in municipal reporting coverage. - Municipality boundaries, names, and coding practices may vary over time. - Public SISAB report datasets are completed with zero values only for combinations defined by observed municipalities, observed competencies, all expected age groups, both expected sex values, and observed report categories within the yearly merge. - LAI CID/CIAP data preserves selected source-file provenance through source_request and source_file . - This deposit corresponds to an individual year. Deposits for other years are published separately.
Bir, Joyanta · Cancio, Ibon · Diaz de cerio, Oihane · et al.
3 files · 100 KB · fasta, xlsxdeclared
This data file contains the data associated with the manuscript entitled "Duplication of the Genes Coding the Proteins That Regulate RNA Polymerase III Activity and Differential Transcription in Tissues of Teleost Fish."
Part 15 of 20. Synthetic Euclid VIS-band galaxy images generated from DESI Legacy Imaging Surveys r,z-band cutouts using an Image-to-Image Schrödinger Bridge (I2SB) model, over the Euclid Q1 footprint. The full processed dataset covers 2,981,033 FITS cutouts packed into 1,235 tar archives grouped by HEALPix sky pixel (nside=64, NESTED ordering), ~1 GB each. Because Zenodo limits each record to 100 files and 50 GB, the full set is split across 20 records. This record publishes a 60-archive subset (RA 59.1-78.7°, Dec -39.9--24.9°). Integrity and provenance: checksums.tsv and desi2euclid_index.csv record the SHA256 of all 1,235 archives from this processing run, not just the ones published as downloadable files here. These values let this exact version of the data be independently verified in the future -- including before any analysis built on it has been published -- by recomputing an archive's SHA256 and comparing it against the recorded value, to confirm the file has not been modified, corrupted, or substituted since it was originally produced. See README.md for the FITS layout (predicted VIS image, DESI g,r,z input -- of which only r,z were used to generate the prediction -- and prediction-uncertainty map), how to locate the archive for a given sky position, and how to run the integrity check.
Part 14 of 20. Synthetic Euclid VIS-band galaxy images generated from DESI Legacy Imaging Surveys r,z-band cutouts using an Image-to-Image Schrödinger Bridge (I2SB) model, over the Euclid Q1 footprint. The full processed dataset covers 2,981,033 FITS cutouts packed into 1,235 tar archives grouped by HEALPix sky pixel (nside=64, NESTED ordering), ~1 GB each. Because Zenodo limits each record to 100 files and 50 GB, the full set is split across 20 records. This record publishes a 62-archive subset (RA 45.0-70.2°, Dec -48.1--30.1°). Integrity and provenance: checksums.tsv and desi2euclid_index.csv record the SHA256 of all 1,235 archives from this processing run, not just the ones published as downloadable files here. These values let this exact version of the data be independently verified in the future -- including before any analysis built on it has been published -- by recomputing an archive's SHA256 and comparing it against the recorded value, to confirm the file has not been modified, corrupted, or substituted since it was originally produced. See README.md for the FITS layout (predicted VIS image, DESI g,r,z input -- of which only r,z were used to generate the prediction -- and prediction-uncertainty map), how to locate the archive for a given sky position, and how to run the integrity check.
Part 13 of 20. Synthetic Euclid VIS-band galaxy images generated from DESI Legacy Imaging Surveys r,z-band cutouts using an Image-to-Image Schrödinger Bridge (I2SB) model, over the Euclid Q1 footprint. The full processed dataset covers 2,981,033 FITS cutouts packed into 1,235 tar archives grouped by HEALPix sky pixel (nside=64, NESTED ordering), ~1 GB each. Because Zenodo limits each record to 100 files and 50 GB, the full set is split across 20 records. This record publishes a 61-archive subset (RA 51.5-81.5°, Dec -54.3--30.0°). Integrity and provenance: checksums.tsv and desi2euclid_index.csv record the SHA256 of all 1,235 archives from this processing run, not just the ones published as downloadable files here. These values let this exact version of the data be independently verified in the future -- including before any analysis built on it has been published -- by recomputing an archive's SHA256 and comparing it against the recorded value, to confirm the file has not been modified, corrupted, or substituted since it was originally produced. See README.md for the FITS layout (predicted VIS image, DESI g,r,z input -- of which only r,z were used to generate the prediction -- and prediction-uncertainty map), how to locate the archive for a given sky position, and how to run the integrity check.