Exploration

ResearchFeatured

Discovery

DiscoverSourcesQuality

Analysis

Working setReviews
Flow StudioTeamConcept
Settings

Partners

  • AI AlliancePrime
  • BrightQueryBuilds Meridian
  • OpenMinedFunded partner
  • MLCommonsFunded partner
  • Hugging FaceDeployment platform
See the full consortium and what each partner wires

Meridian is the discovery layer for research data, built by BrightQuery within the AI Alliance.

hybrid · semantic + lexical · 46 datasets ranked · 0.96s

Structurecomposite3tabular1
Depthcataloged42measured4
Licenseopen46
Accessopen46
Formatrar38parquet9zip7csv4gzip2
Sourcezenodo46
clear
1-20 of 46sortrelevancemeasured firstqualitysize
composite

ADM_LSIR: a physics-inspired laparoscopic aerosol degradation dataset

0.00

guo, na · pan, jiachen · li, tiantian · et al.

14 files · 100 MB · csv, rar, tsv

ADM_LSIR is a physics-inspired laparoscopic aerosol degradation dataset for aerosol-aware surgical image analysis and image restoration. The v1.0.0 release contains: - 21,916 clean clinical laparoscopic frames (clean/) - 9,562 real intraoperative aerosol-degraded frames (degraded/) - 36,052 simulated aerosol masks, including 19,701 smoke-like masks and 16,351 trajectory masks (mask/) - Blender simulation/cache materials (ADM_LSIR_Blender_simulation_files_v1.0.rar) - metadata_quality_report_v1.0.csv - recommended_splits_v1.0.csv - video_mapping_v1.0.csv - parts_manifest.txt - checksums_v1.0.tsv - release_manifest_v1.0.json All released clinical frames are de-identified and stored as lossless PNG files. Filenames use anonymized video identifiers, e.g., C-V##-####.png for clean frames and D-V##-####.png for degraded frames. The recommended split is defined at the source_video_id/public_video_label level to reduce leakage across frames from the same source video. The public video labels in video_mapping_v1.0.csv provide privacy-safe source-video identifiers (video1-video19). The Blender archive documents the smoke and trajectory mask simulation setup and supports reuse, but it is not a guaranteed exact per-mask reproduction package. The released pre-rendered mask library is the primary reusable dataset component. Source code for synthesis and quality screening is available at: https://github.com/SweetDeathh/ADM_LSIR

xlsx2
docx1
jpeg1
netcdf1
tiff1
torch1
tsv1
open·CC-BY-4.0·Zenodo·completeSource
composite

AI2EMD with Hierarchical Active Learning Enables Accurate and Generalizable Liquid Electrolyte Characterization: Neural network potentials training data

0.00

Xu, Tao

3 files · 100 MB · rar, xlsx

The deepmd_data dataset comprises energy data for 375,781 structures and force data for more than 45 million atoms, generated from six hierarchical active learning iterations. The init directory stores the non-periodic molecular configurations established at the initialization stage. The iter directories contain structures labeled by both AIMD and DPMD-FP methods for each iteration. Each folder is named according to the scheme of solvent molecule index followed by lithium salt designation. The finetuned_data folder contains SEI reaction simulation data used for fine-tuning the MLFFs. Comprehensive details regarding the solvent molecules are available in the solvent_with_smiles.xlsx file.

open·CC-BY-4.0·Zenodo·completeSource
composite

A robust method for microscopic 3D shape restoration via shape-from-focus

0.00

Yuezong Wang · Yu Niu · Jiqiang Chen

2 files · 100 MB · rar, zip

This dataset contains raw experimental images, video sequences of real test samples, simulated microscopic image data from the manuscript " A robust method for microscopic 3D shape restoration via shape-from-focus" , as well as focal volume datasets collected before vibration simulation, after vibration simulation, and post anti-vibration processing. All provided data enable the validation of conclusions and reproducibility of experimental results obtained via the computational pipeline proposed in this paper.

open·CC-BY-4.0·Zenodo·completeSource
tabular

Generated ASO features for the OligoAI dataset

0.00

Kovaliov, Michael

1 files · 100 MB · parquet

open·CC-BY-4.0·Zenodo·completeSource
declared

Fine tuning an LLM with a domain a specific data set

0.00

Madhusudan, Gujral

6 files · 29 MB · parquetdeclared

Large language models (LLMs) are trained on massive, publicly available text datasets comprising trillions of tokens, enabling them to excel at general language tasks like next-token prediction. However, LLMs often struggle with domain-specific prompts, exhibiting reduced accuracy or generating inaccurate information (hallucinations). This is because they lack sufficient subject matter expertise. Two primary approaches exist to address this limitation for augmenting LLMs knowledge: Retrieval-Augmented Generation (RAG) and fine-tuning. This presentation focuses on fine-tuning smaller LLMs with domain-specific instruct datasets using the LoRA (Low-Rank Adaptation) technique on Gaudi hardware. We will leverage publicly available LLMs and datasets from the Hugging Face Hub for this demonstration. Though it is possible to fine tune LLMs with plain text data - sourced from documents, articles, and other materials.

open·CC-BY-4.0·Zenodo·completeSource
declared

Skala Bünte, Süki Vagonu Hallē [2026.07.08.]

0.00

Daugavietis, Jānis

6 files · 3.3 GB · jpeg, rardeclared

Skala Bünte, Süki Vagonu Hallē [2026.07.08.] Koncerta beigu telefona foto/ video. SKALA BÜNTE X SÜKI FACE2FACE https://fb.me/e/7dycfN2gI Details Event by John Dow Vagonu Hall Public · Anyone on or off Facebook 8TH OF JULY VAGONU HALLE TWO BANDS TWO BACKLINES MOSHPIT IN THE MIDDLE ONCE IN A LIFETIME FACE2FACE MASSACRE PROVIDED BY SKALA BÜNTE & SÜKI 🔪🔪🔪 DOORS 19:00 7€

open·CC-BY-4.0·Zenodo·completeSource
declared

EuroFlood: a queryable cloud-native index for the CEMS-EFAS Satellite-Derived Flood Depth Maps

0.00

Hackl, Jürgen

6 files · 132 MB · parquet, tiffdeclared

EuroFlood is an open, cloud-native index over the JRC/Copernicus CEMS-EFAS Satellite-Derived Flood Depth Maps for Europe (Betterle & Salamon, 2025; CC-BY-4.0) - ~3,280 satellite-derived observed flood-depth maps across Europe, 2015-2024. The bundle is a sparse Cloud-Optimized GeoTIFF encoding, per pixel, the set of flood events that inundated it, plus a combo_id -sorted GeoParquet dictionary and a small events table. Query by region and time via HTTP range reads (GDAL /vsicurl + DuckDB) to retrieve matching events, then fetch only the source depth rasters needed. Built with the open-source EuroFlood Python package ( pip install euroflood ).

open·CC-BY-4.0·Zenodo·completeSource
declared

Replication Package for the paper: More Productive, but at What Cost? Understanding How GenAI Shapes Developers' Work Across the SPACE Dimensions

0.00

Murilo Coelho · de Sousa Amâncio, Francisco Dione · Paixao, Matheus · et al.

3 files · 742 KB · rardeclared

This repository contains the replication package for the paper "More Productive, but at What Cost? Understanding How GenAI Shapes Developers' Work Across the SPACE Dimensions", accepted at the 40th Brazilian Symposium on Software Engineering (SBES 2026), São Paulo, Brazil. The package includes: (i) the complete survey instrument and sanitized participant responses; (ii) qualitative coding artifacts, including the codebook, category consolidation, and classification analysis; (iii) inter-rater reliability materials (Cohen's Kappa = 0.81); (iv) quantitative datasets and statistical analysis outputs (SPACE composite scores, Cronbach's alpha, Kruskal-Wallis, Mann-Whitney U, and Dunn's post-hoc tests); and (v) supporting literature review material. All participant data were anonymized prior to disclosure. The survey materials are in Portuguese, the language of data collection.

open·CC-BY-4.0·Zenodo·completeSource
declared

GenixRL reclassification scores for ~ 1.05 million missense Variants of Uncertain Significance (VUS) from ClinVar

0.00

Abbas, Syed Hassan

1 files · 86 MB · rardeclared

Abbas H, et al. "GenixRL- " The VUS were extracted from CLinVar database (downloaded [Date: August 2025]) and scored using the GenixRL framework. This dataset provides the foundation for the VUS reclassification analysis presented in the main manuscript. The dataset is provided as a single compressed CSV file: GenixRL_VUS_Scored.csv.gz Key Coulmn Descriptions: [Variant Identifier Columns, e.g., CHROM, POS, REF, ALT]: Standard genomic coordinates for each variant. SYMBOL: The official gene symbol. GenixRL_Score: The continuous pathogenicity score generated by GenixRL, ranging from 0 (most likely benign) to 1 (most likely pathogenic). GenixRL_Classification: The tiered classification based on the manuscript's thresholds: 'Likely Benign': Score < 0.53 'Likely Pathogenic': Score >= 0.53 and < 0.709 'High-Confidence Pathogenic': Score >= 0.709 [Other relevant columns]: The file also includes intermediate scores from other predictors and allele frequencies used for validation.

open·CC-BY-4.0·Zenodo·completeSource
declared

Supporting Data and Evaluation Outputs for JerseyTrack: A Confidence-Guided Sports Multi-Object Tracking Method via Jersey Semantic Fusion

0.00

Cao, Shiyuan · Li, Yaning

8 files · 18 MB · csv, rardeclared

This repository provides the supporting materials for the manuscript "A Confidence-Guided Sports Multi-Object Tracking Method via Jersey Semantic Fusion." The archived materials include evaluation summaries, final tracking outputs, environment records, dataset mapping files, protocol reproduction materials, intermediate summary files, and revision evidence used to support the reported SportsMOT validation results. The reported formal results were obtained from the real detector, real frame-level semantic feature extraction, confidence partitioning, cascaded association, and TrackEval evaluation pipeline. The original public benchmark datasets, including SportsMOT, TeamTrack, SoccerNet Tracking, MOT17, and MOT20, are not redistributed in this repository and should be obtained from their original providers.

open·CC-BY-4.0·Zenodo·completeSource
declared

Supplementary material 2 from: Nie Y, Huang B (2026) Drechslerosporium cornellii gen. et sp. nov. within the Basidiobolaceae exhibiting unique conidial discharge and digitate chlamydospores. MycoKeys 136: 177-191. https://doi.org/10.3897/mycokeys.136.200461

0.00

Nie, Yong · Huang, Bo

1 files · 1.5 KB · rardeclared

BI tree

open·CC0-1.0·Zenodo·completeSource
declared

Supplementary material 1 from: Nie Y, Huang B (2026) Drechslerosporium cornellii gen. et sp. nov. within the Basidiobolaceae exhibiting unique conidial discharge and digitate chlamydospores. MycoKeys 136: 177-191. https://doi.org/10.3897/mycokeys.136.200461

0.00

Nie, Yong · Huang, Bo

1 files · 9.1 KB · rardeclared

Alignments for phylogen

open·CC0-1.0·Zenodo·completeSource
declared

A Parliamentary Discourse Dataset from the German Bundestag

0.00

Njie, Adama · Torkayesh, Ali E · Venghaus, Prof. Dr. Sandra

10 files · 1.6 GB · csv, gzip, parquetdeclared

Structured, speaker-attributed corpus of all German Bundestag plenary session transcripts ( Plenarprotokolle ) from the first legislative period to the present (WP01-WP21, September 1949 - April 2026). Every attributed speech is extracted from the official PDFs published by the Deutscher Bundestag under open data policy and linked to the speaker's name, parliamentary role, party affiliation, and gender. Scale: 4,611 sessions · 1,033,723 speeches · 4,205 identified MdBs · 76 years of parliamentary debate Dataset files speeches.parquet - one row per attributed speech: speaker name, role, party, gender, stammdaten_id, full German text (~1 GB) persons.parquet - one row per MdB: cross-session identity linking all name variants via stammdaten_id; canonical name, birth date, career span, total speeches. Use this - not speakers.parquet - for person-level analysis sessions.parquet - one row per plenary session: date, city, Wahlperiode, source PDF hash, extraction engine speakers.parquet - name-string index: one row per unique name string as extracted from the transcripts. Useful for understanding extraction quality; not suitable for person-level aggregation (the same politician often appears under several name variants across sessions) parties.csv - reference table of 31 German parliamentary parties, 1949-present speeches.csv.gz - CSV fallback for Stata and Excel users (same columns as speeches.parquet) datapackage.json - Frictionless Data schema with column descriptions and foreign key constraints Cross-session identity The same politician often appears under different name strings across sessions (e.g. "Schmidt", "Dr. Schmidt", "Frau Dr. Schmidt"). Cross-session person linkage is provided via stammdaten_id , matched against the official Bundestag Stammdaten biographical XML. The persons.parquet table aggregates all name variants for the same MdB into one row with correctly summed speech counts, career span, and birth date. Coverage: ~98.5% of speeches are linked to a stammdaten_id; the remaining ~1.5% are ambiguous surname-only attributions or speakers not in the Stammdaten. Coverage and sources Source PDFs are the official Stenografische Berichte downloaded from the Bundestag open-data portal (bundestag.de). Party-share normalisation in the corpus statistics uses official seat counts per Wahlperiode sourced from the Federal Returning Officer (Bundeswahlleiter, bundeswahlleiter.de). Two PDF generations are covered: scanned and OCR'd documents (WP01-WP09, Bonn era, 1949-1987) and born-digital documents (WP10-WP21, 1987-present). The engine column in sessions.parquet flags whether pdftotext (born-digital) or pdfminer (OCR fallback) was used; this is the primary data-quality indicator for NLP use. Speaker attribution Each speech is attributed using four patterns extracted from the transcript format: presiding officers (Präsident/in, Vizepräsident/in), regular members (name + party), government officials (name + Bundeskanzler/in, Bundesminister/in, etc.), and procedural roles (Berichterstatter/in, etc.). The party field is null for ~60% of speeches - this is expected, as presiding officers and ministers are not identified by party in the transcript. Gender annotation & distribution Gender is derived by matching speaker names against the official Bundestag Stammdaten biographical XML (all MdBs since 1949), with fallbacks for role title, honorific prefix, manually researched overrides, and a gender_guesser first-name heuristic. The gender_source column distinguishes stammdaten (authoritative, 83%), role_title (gendered job title in attribution, 6.4%), title_prefix (Frau/Herr honorific, 0.5%), manual (historically researched, 2.9%), and inferred (name-based heuristic, 4.7%). Gender distribution: Female 26.6% · Male 73.4% · Unknown 0.0%. Data quality All speeches pass automated validation: zero null speaker names, zero sequence gaps, zero CID artefacts, zero party-misclassified-as-Bundesland errors. Eight sessions with conflicting source PDFs were deduplicated (first lexicographic occurrence retained). 252 non-person names incorrectly accepted by the parser (table headers, legislative terms, agenda fragments) are excluded at build time via a curated exclusion list. OCR sessions (WP01-WP09) may contain Unicode replacement characters (U+FFFD); the engine field identifies these sessions. Licence CC BY 4.0. The underlying Plenarprotokolle are official government documents of the Deutscher Bundestag and are in the public domain.

open·CC-BY-4.0·Zenodo·completeSource
declared

IPv6-CyberBench: A Synthetic Translated-Flow Stress-Test Corpus for Diagnostic IDS Evaluation under IPv4-to-IPv6 Header Substitution

0.00

abdulwahab, samaa · aduallah, mahmood z. · Sallomi, Adheed H.

34 files · 2.7 GB · csv, gzip, parquetdeclared

Intrusion-detection research on Internet Protocol version 6 (IPv6) remains bottlenecked by the scarcity of labelled, protocol-aware flow datasets. Existing machine-learning IDS benchmarks are overwhelmingly IPv4-centric, and the few IPv6 corpora that have been released target narrow attack families or rely on small academic testbeds that cannot be re-created by third parties. We present IPv6-CyberBench, a reproducible eight-phase pipeline that constructs a large, protocol-aware translated-flow corpus by harmonising CIC-IDS-2017, CIC-IDS-2018 and CIC-DDoS-2019, applying deterministic IPv4→IPv6 address translation (6to4, NAT64, Teredo, EUI-64), synthesising 27 IPv6-specific flow features grouped in six protocol families, enforcing nineteen RFC-derived constraint categories together with temporal address dynamics, and rebalancing the long-tailed class distribution with a feature-group-conditioned per-class Wasserstein-GAN-GP augmenter and a SMOTE-KDE fallback selected per class by a formal decision rule. We scope the contribution honestly: because the seed corpora are IPv4 captures, the resulting 2,285,774-record benchmark is a translated-flow corpus suitable for training and evaluating flow-level IPv6 IDS classifiers on flooding, brute-force, scan, web-attack and infiltration traffic under IPv6 protocol-header semantics, and for studying IPv6-specific feature engineering and address dynamics in a reproducible setting. It is not a substitute for protocol-native IPv6 attack capture, and we explicitly exclude ICMPv6 Neighbour-Discovery flooding, SEND flooding, NDP exhaustion and extension-header covert-tunnelling from the threat model. The benchmark is evaluated on four axes - fidelity (Kolmogorov-Smirnov, MMD, Fréchet feature distance), utility (stratified 5×5 nested cross-validation over six classifier families including CNN-LSTM and LightGBM), privacy (Shokri-style membership-inference advantage AUC), and external fidelity against a 24 h anonymised CAIDA IPv6 trace (equinix-chicago, US backbone) and a MAWI samplepoint-F trace (WIDE backbone, Tokyo, Japan). The full pipeline, the hyper-parameter manifest, the RFC-constraint manifest, the reproduction scripts, and the 2,285,774-record benchmark are released unconditionally on Zenodo under CC BY 4.0; the dataset and pipeline are openly available at https://doi.org/10.5281/zenodo.19503446 (CC BY 4.0).

open·MIT·Zenodo·completeSource
declared

Southeast Asia 6-km Grid Perceived Informality Index Dataset

0.00

Rui, Jin · Cai, Chenfan

1 files · 3.7 MB · rardeclared

This dataset provides a paper-aligned 6-km grid-level perceived informality index for Southeast Asia. It is prepared as the data-release package associated with the manuscript "Unveiling the environmental drivers of urban informality in Southeast Asia through citizen science and explainable machine learning", intended for submission to npj Urban Sustainability. The release retains the ten countries used in the manuscript: Philippines, Indonesia, Vietnam, Cambodia, Myanmar, Thailand, Malaysia, Laos, Brunei and Timor-Leste. Singapore and unassigned grid cells from an earlier draft package were excluded to align the deposited data with the manuscript scope. The primary variable is informality_index, a min-max normalized 0-1 perceived informality index. Higher values indicate stronger predicted perceived urban informality. The original large-valued model output is retained as raw_informality_score for provenance, but informality_index should be used as the main manuscript-ready variable. Normalization was calculated over the retained ten-country grid cells. The package is organized into four main folders. The data/ folder contains the authoritative GeoPackage spatial layer, CSV exports, centroid-coordinate tables, administrative boundaries, country-level summary statistics and whole-dataset summary statistics. The docs/ folder contains a data dictionary, a methodology note and a processing log documenting how the release was generated. The figures/ folder contains quick-look visual summaries, including the informality-index histogram, country-level mean chart and spatial preview map. The checksums/ folder contains SHA-256 checksums for verifying file integrity. The GeoPackage file SEA_Perceived_Informality_6km_Grid.gpkg is the authoritative spatial dataset and includes grid geometries, country assignment, raw_informality_score and informality_index. The CSV files are provided for users who prefer tabular workflows or do not use GIS software. The centroid CSV includes approximate longitude and latitude for each grid cell centroid. Summary-statistics files report country-level and overall distributions of the normalized perceived informality index, with additional raw-score summaries retained for provenance. This release is intended to support reproducibility, spatial visualization and further research on perceived urban informality, environmental exposure and urban sustainability in Southeast Asia. Raw satellite imagery, Google Street View imagery and third-party geospatial layers are not redistributed in this package because they are subject to the access terms and licensing conditions of their original providers.

open·CC-BY-4.0·Zenodo·completeSource
declared

Emission-factor differentiation redistributes global soil nitrous oxide sources and mitigation potential

0.00

Haider, Haroon

1 files · 146 MB · rardeclared

This dataset provides global gridded soil nitrous oxide (N₂O) emissions for the period 1961-2019, distributed as a stack of annual NetCDF files ( N2O_soil_emissions_YYYY.nc ), one per year, giving 59 files in total. Each file resolves emissions on a regular 0.5° × 0.5° longitude-latitude mesh (720 longitude columns by 360 latitude rows) covering the full globe. Grid-cell centres are given by the lat (degrees_north) and lon (degrees_east) coordinate variables, with longitude running from -180° to +180°; the environmental-response field was rolled from its native 0-360° convention to -180-180° so that every layer shares this common grid. Ocean and other missing cells are flagged with a fill value of NaN . Within each annual file, emissions are stored as twelve monthly fields ( n2o_january through n2o_december ) together with an annual total ( n2o_annual ). The monthly fields report soil N₂O emission per grid cell in grams of N₂O-N per cell per month, and the annual field, expressed in grams of N₂O-N per cell per year, is the sum of the twelve monthly layers. Because the stored values are absolute per-cell masses rather than per-area fluxes, a regional or global total is obtained by summing cells directly and converting from grams to teragrams (dividing the summed grams by 1 × 10&sup1;²). Grid-cell emissions were computed as the product of a fertiliser-driven emission field, an environmental scaling term, and a mean precipitation modifier. Each file preserves the full provenance of this calculation through its global attributes, which record the dataset title, the reference year, the input and output units, the calculation string, the longitude-fix note, and a UTC creation timestamp, so that every layer can be traced back to its inputs and processing steps.

open·CC-BY-4.0·Zenodo·completeSource
declared

Dataset and analysis code: Predicting the workability of PCE-modified cement-fly-ash pastes with explainable machine learning

0.00

Gider, Veysel · Ekinci, Serdar · Izci, Davut · et al.

1 files · 96 KB · rardeclared

Experimental dataset (616 Marsh-funnel flow-time and 252 mini-slump measurements; 22 in-house PCE formulations × 4 fly-ash replacement levels × 7 dosages, w/cm = 0.35) and the complete deterministic Python pipeline reproducing every numerical result, table and figure of the associated paper. See README.md for the column dictionary and reproduction instructions.

open·CC-BY-4.0·Zenodo·completeSource
declared

SDSS DR12 Stellar Spectra Dataset for Machine Learning Tasks

0.00

Barreto, Bruno · Eisencraft, Marcio

3 files · 2.6 GB · parquetdeclared

This dataset provides a fixed benchmark dataset for stellar atmospheric parameter estimation from Sloan Digital Sky Survey Data Release 12 (SDSS DR12) optical stellar spectra. The dataset is organized into three predefined Parquet splits: 30,000 spectra for training, 5,000 spectra for validation, and 15,000 spectra for testing. Each row corresponds to one SDSS stellar spectrum and includes raw spectral arrays, fixed-length processed spectral features, source identifiers, basic metadata, and catalog stellar-parameter labels with their associated uncertainties. The supervised regression targets are the adopted catalog stellar atmospheric parameters: effective temperature (Teff, in K), metallicity ([Fe/H], in dex), and surface gravity (log g, in dex). The dataset also includes relevant observational and catalog information such as SDSS plate, MJD, fiber identifier, sky coordinates, signal-to-noise ratio, adopted radial velocity, raw flux, logarithmic wavelength grid, inverse variance, pixel mask, and processed flux features. This release is intended to support machine-learning research on stellar spectroscopy, including regression models for atmospheric parameter estimation, benchmark comparisons, uncertainty-aware evaluation, and experiments using either processed fixed-length spectra or native observed-frame spectral arrays.

open·CC-BY-4.0·Zenodo·completeSource
declared

BAJO PESO AL NACER Y CRECIMIENTO FETAL EN JOVELLANOS, MATANZAS (2024–2025): ANÁLISIS EPIDEMIOLÓGICO

0.00

José Antonio, Alonso Viamonte · María Caridad, Rodríguez Pérez

1 files · 26 MB · rardeclared

Esta base de datos contempla los materiales suplementarios del artículo en evaluación en SciELO Preprints titulado: BAJO PESO AL NACER Y CRECIMIENTO FETAL EN JOVELLANOS, MATANZAS (2024-2025): AN&Aacute;LISIS EPIDEMIOL&Oacute;GICO.

open·CC-BY-4.0·Zenodo·completeSource
declared

SISAB municipal primary care production and CID/CIAP attendance data, Brazil, 2026 (incomplete)

0.00

Saldanha, Raphael

8 files · 601 MB · parquet, zipdeclared

Description This deposit contains annual, municipality-level datasets derived from the Brazilian Primary Health Care Information System (SISAB). The files combine two complementary data sources: Public SISAB Saúde report downloads from the Atendimento/Visita production report. CID-10 and CIAP-2 attendance data obtained from SISAB through requests under the Brazilian Access to Information Law (Lei de Acesso à Informaç&atilde;o, LAI). The datasets are organized as tidy annual files in CSV (Zipped) and Parquet format. They are intended to support reproducible analysis of primary care production, procedures, evaluated problems/conditions, and CID/CIAP-coded attendances across Brazilian municipalities. The public SISAB report datasets are stratified by competence month, state, municipality, DataSUS age group, SISAB sex category, and the selected report category. For each competence month and report type, the extraction combines 36 stratified SISAB downloads: 18 age groups by 2 sex values. Monthly files are merged into yearly files, completing missing combinations of observed competence, municipality, age group, sex, and category with valor = 0 . The LAI dataset contains yearly CID-10 and CIAP-2 attendance counts by competence month, municipality, code type, and code. When multiple valid LAI files cover the same competence, the processing pipeline selects the file with the largest number of data rows, using file size and request folder order as tie-breakers. Provenance columns identify the selected LAI request and source file. Variables SISAB Saúde Produç&atilde;o Columns: competencia : competence month in YYYYMM format. uf : Brazilian state abbreviation. ibge : municipality IBGE code. municipio : municipality name. faixa_etaria : Age group. sexo : SISAB sex category, Masculino or Feminino . tipo_producao : production type from the SISAB report. valor : count reported by SISAB. SISAB Saúde Procedimento Columns: competencia : competence month in YYYYMM format. uf : Brazilian state abbreviation. ibge : municipality IBGE code. municipio : municipality name. faixa_etaria : age group. sexo : SISAB sex category, Masculino or Feminino . procedimento : procedure from the SISAB report. valor : count reported by SISAB SISAB Saúde Condiç&atilde;o Avaliada Columns: competencia : competence month in YYYYMM format. uf : Brazilian state abbreviation. ibge : municipality IBGE code. municipio : municipality name. faixa_etaria : age group. sexo : SISAB sex category, Masculino or Feminino . condicao_avaliada : evaluated problem or condition from the SISAB report. valor : count reported by SISAB. SISAB LAI CID/CIAP Columns: ano_competencia : competence year. competencia : competence month in YYYYMM format. competencia_date : first day of the competence month. co_municipio_ibge : municipality IBGE code. tp_codigo : code type, CID or CIAP . codigo : CID-10 or CIAP-2 code. qt_atendimentos : number of attendances. source_request : selected LAI request folder. source_file : selected source CSV file. Methods The public SISAB report files were generated with the sisab_scrapper processing pipeline. For each month, report type, age group, and sex value, the pipeline downloads the all-Brazil municipality report from SISAB, validates the returned CSV, preserves raw cache files for resumable runs, and writes a sorted monthly tidy dataset. The yearly merge validates required columns, expected age groups, expected sex values, category values, and month gaps unless explicitly allowed. The CID/CIAP files were generated with the sisab_lai processing pipeline. The pipeline imports CSV files received through LAI requests, detects the real CSV header after any SQL*Plus preamble, validates candidate files, resolves overlapping requests by competence, standardizes old and new schemas into one tidy table, and exports annual CSV and Parquet files together with audit reports. Sources - SISAB public reports, Ministry of Health, Brazil: https://sisab.saude.gov.br/ - SISAB LAI files obtained through Brazilian Access to Information Law requests. - Processing code for public SISAB report data: https://github.com/rfsaldanha/sisab_scrapper - Processing code for LAI CID/CIAP data: https://github.com/rfsaldanha/sisab_lai Notes - Counts are aggregated administrative records and should be interpreted in light of SISAB reporting practices, data quality, and changes in municipal reporting coverage. - Municipality boundaries, names, and coding practices may vary over time. - Public SISAB report datasets are completed with zero values only for combinations defined by observed municipalities, observed competencies, all expected age groups, both expected sex values, and observed report categories within the yearly merge. - LAI CID/CIAP data preserves selected source-file provenance through source_request and source_file . - This deposit corresponds to an individual year. Deposits for other years are published separately.

open·CC-BY-4.0·Zenodo·completeSource
page 1next →

Select a result to see its full details here: the measured structure, quality, and the loader, without leaving your search.