Exploration

ResearchFeatured

Discovery

DiscoverSourcesQuality

Analysis

Working setReviews
Flow StudioTeamConcept
Settings

Partners

  • AI AlliancePrime
  • BrightQueryBuilds Meridian
  • OpenMinedFunded partner
  • MLCommonsFunded partner
  • Hugging FaceDeployment platform
See the full consortium and what each partner wires

Meridian is the discovery layer for research data, built by BrightQuery within the AI Alliance.

hybrid · semantic + lexical · 78 datasets ranked · 8.77s

Structurecomposite4tabular1tensor1
Depthcataloged72measured6
Licenseopen78
Accessopen78
Formatrar38sevenzip16hdf515zip13parquet9
Sourcezenodo78
clear
1-20 of 78sortrelevancemeasured firstqualitysize
tensor

Precomputed Databases for OMAmer

0.00

Altenhoff, Adrian

6 files · 100 MB · hdf5

OMAmer - tree-driven and alignment-free protein assignment to subfamilies OMAmer is an alignment-free protein family assignment method designed to avoid overly specific subfamily predictions and to scale efficiently to phylogenomic databases containing thousands of genomes. It relies on an innovative approach that uses evolutionarily informed k-mers for alignment-free mapping to ancestral protein subfamilies. This dataset provides precomputed OMAmer databases derived from the Hierarchical Orthologous Groups in the OMA Browser . We aim to update these databases with every new OMA Browser release. Each OMAmer database is built using the latest version of the OMAmer package available at the time of the corresponding OMA Browser release. The dataset includes databases for different subsets of the species taxonomy. In most cases, we recommend using the LUCA.h5 database, which contains information from all species in the OMA database. The subset-specific databases are mainly useful when disk space is limited. The release May2026 is based on the OMA Browser release May 2026 which comprises 2983 species. We used OMAmer version 2.1.0 to build these databases.

csv6
gzip3
torch3
xlsx3
pdf2
tiff2
docx1
fasta1
jpeg1
netcdf1
npy1
sqlite1
tsv1
open·CC-BY-4.0·Zenodo·completeSource
composite

ADM_LSIR: a physics-inspired laparoscopic aerosol degradation dataset

0.00

guo, na · pan, jiachen · li, tiantian · et al.

14 files · 100 MB · csv, rar, tsv

ADM_LSIR is a physics-inspired laparoscopic aerosol degradation dataset for aerosol-aware surgical image analysis and image restoration. The v1.0.0 release contains: - 21,916 clean clinical laparoscopic frames (clean/) - 9,562 real intraoperative aerosol-degraded frames (degraded/) - 36,052 simulated aerosol masks, including 19,701 smoke-like masks and 16,351 trajectory masks (mask/) - Blender simulation/cache materials (ADM_LSIR_Blender_simulation_files_v1.0.rar) - metadata_quality_report_v1.0.csv - recommended_splits_v1.0.csv - video_mapping_v1.0.csv - parts_manifest.txt - checksums_v1.0.tsv - release_manifest_v1.0.json All released clinical frames are de-identified and stored as lossless PNG files. Filenames use anonymized video identifiers, e.g., C-V##-####.png for clean frames and D-V##-####.png for degraded frames. The recommended split is defined at the source_video_id/public_video_label level to reduce leakage across frames from the same source video. The public video labels in video_mapping_v1.0.csv provide privacy-safe source-video identifiers (video1-video19). The Blender archive documents the smoke and trajectory mask simulation setup and supports reuse, but it is not a guaranteed exact per-mask reproduction package. The released pre-rendered mask library is the primary reusable dataset component. Source code for synthesis and quality screening is available at: https://github.com/SweetDeathh/ADM_LSIR

open·CC-BY-4.0·Zenodo·completeSource
composite

AI2EMD with Hierarchical Active Learning Enables Accurate and Generalizable Liquid Electrolyte Characterization: Neural network potentials training data

0.00

Xu, Tao

3 files · 100 MB · rar, xlsx

The deepmd_data dataset comprises energy data for 375,781 structures and force data for more than 45 million atoms, generated from six hierarchical active learning iterations. The init directory stores the non-periodic molecular configurations established at the initialization stage. The iter directories contain structures labeled by both AIMD and DPMD-FP methods for each iteration. Each folder is named according to the scheme of solvent molecule index followed by lithium salt designation. The finetuned_data folder contains SEI reaction simulation data used for fine-tuning the MLFFs. Comprehensive details regarding the solvent molecules are available in the solvent_with_smiles.xlsx file.

open·CC-BY-4.0·Zenodo·completeSource
composite

A robust method for microscopic 3D shape restoration via shape-from-focus

0.00

Yuezong Wang · Yu Niu · Jiqiang Chen

2 files · 100 MB · rar, zip

This dataset contains raw experimental images, video sequences of real test samples, simulated microscopic image data from the manuscript " A robust method for microscopic 3D shape restoration via shape-from-focus" , as well as focal volume datasets collected before vibration simulation, after vibration simulation, and post anti-vibration processing. All provided data enable the validation of conclusions and reproducibility of experimental results obtained via the computational pipeline proposed in this paper.

open·CC-BY-4.0·Zenodo·completeSource
composite

Additional data for "Trajectory Variance: An Unsupervised Measure of Developmental Vocal Plasticity in Birdsong"

0.00

Lee, Kanghwi

2 files · 100 MB · sevenzip

Trained variational autoencoder (VAE) checkpoints and per-vocalization latent representations for three developing zebra finches (183K-274K vocalizations each, 40-101 days post-hatch), accompanying the Interspeech 2026 paper "Trajectory Variance: An Unsupervised Measure of Developmental Vocal Plasticity in Birdsong." With the code at https://github.com/hwiora/trajectory_variance , these files reproduce the paper's evaluation tables. The raw audio recordings are not included; they are available from the authors on request and will be released in a future study. See ZENODO_README.md for details.

open·CC-BY-4.0·Zenodo·completeSource
tabular

Generated ASO features for the OligoAI dataset

0.00

Kovaliov, Michael

1 files · 100 MB · parquet

open·CC-BY-4.0·Zenodo·completeSource
declared

Fine tuning an LLM with a domain a specific data set

0.00

Madhusudan, Gujral

6 files · 29 MB · parquetdeclared

Large language models (LLMs) are trained on massive, publicly available text datasets comprising trillions of tokens, enabling them to excel at general language tasks like next-token prediction. However, LLMs often struggle with domain-specific prompts, exhibiting reduced accuracy or generating inaccurate information (hallucinations). This is because they lack sufficient subject matter expertise. Two primary approaches exist to address this limitation for augmenting LLMs knowledge: Retrieval-Augmented Generation (RAG) and fine-tuning. This presentation focuses on fine-tuning smaller LLMs with domain-specific instruct datasets using the LoRA (Low-Rank Adaptation) technique on Gaudi hardware. We will leverage publicly available LLMs and datasets from the Hugging Face Hub for this demonstration. Though it is possible to fine tune LLMs with plain text data - sourced from documents, articles, and other materials.

open·CC-BY-4.0·Zenodo·completeSource
declared

Skala Bünte, Süki Vagonu Hallē [2026.07.08.]

0.00

Daugavietis, Jānis

6 files · 3.3 GB · jpeg, rardeclared

Skala Bünte, Süki Vagonu Hallē [2026.07.08.] Koncerta beigu telefona foto/ video. SKALA BÜNTE X SÜKI FACE2FACE https://fb.me/e/7dycfN2gI Details Event by John Dow Vagonu Hall Public · Anyone on or off Facebook 8TH OF JULY VAGONU HALLE TWO BANDS TWO BACKLINES MOSHPIT IN THE MIDDLE ONCE IN A LIFETIME FACE2FACE MASSACRE PROVIDED BY SKALA BÜNTE & SÜKI 🔪🔪🔪 DOORS 19:00 7€

open·CC-BY-4.0·Zenodo·completeSource
declared

EuroFlood: a queryable cloud-native index for the CEMS-EFAS Satellite-Derived Flood Depth Maps

0.00

Hackl, Jürgen

6 files · 132 MB · parquet, tiffdeclared

EuroFlood is an open, cloud-native index over the JRC/Copernicus CEMS-EFAS Satellite-Derived Flood Depth Maps for Europe (Betterle & Salamon, 2025; CC-BY-4.0) - ~3,280 satellite-derived observed flood-depth maps across Europe, 2015-2024. The bundle is a sparse Cloud-Optimized GeoTIFF encoding, per pixel, the set of flood events that inundated it, plus a combo_id -sorted GeoParquet dictionary and a small events table. Query by region and time via HTTP range reads (GDAL /vsicurl + DuckDB) to retrieve matching events, then fetch only the source depth rasters needed. Built with the open-source EuroFlood Python package ( pip install euroflood ).

open·CC-BY-4.0·Zenodo·completeSource
declared

Hydroclimate data to simulate the Lake Ontario – St. Lawrence River system in support of an expedited review of Lake Ontario regulation Plan 2014 (2022-2026)

0.00

Fry, Lauren · Seglenieks, Frank · Shrestha, Narayan Kumar · et al.

3 files · 22 MB · sevenzipdeclared

In response to flooding on Lake Ontario and the St. Lawrence River in 2017 and 2019, the U.S. and Canadian governments directed the Great Lakes - St. Lawrence River Adaptive Management (GLAM) Committee to expedite review of Lake Ontario outflow regulation Plan 2014 ahead of its usual 15-year review cycle. The dataset includes (1) a baseline record based on data from 1961-2020, (2) a set of stochastic scenarios that incorporates extremes that may not be in the historical record but are plausible under current hydroclimate conditions, and (3) a set of climate change scenarios to support evaluation of the plan under plausible decadal-scale hydroclimate changes. All data required to simulate Lake Ontario water levels, outflows, and downstream water levels are provided for each scenario. The data have been compressed into the following files: Baseline.7z - Baseline hydroclimate record from 1961-2020 Climate.7z - 8 future climate hydroclimate records Stochastic.7z - 500 stochastic hydroclimate records Within each compressed file are subdirectories that contain the files for each water supply sequence. The following table lists the files that are provided for each of the water supply sequences. For each file there is a description of the variable contained in the file, the location of the data in the file, and the unit of the data in the file if applicable. Filename Variable Location Units dpmi_flw_cms_qm48_na.csv Flow Des Prairies and Mille Iles River cms rich_flw_cms_qm48_na.csv Flow Richelieu River at Rapides Fryer cms stfr_flw_cms_qm48_na.csv Flow St. Francois River at Chut Hemming cms stmc_flw_cms_qm48_na.csv Flow St. Maurice River at La Gabelle cms ont_nbs_cms_qm48_na.csv Net Basin Supply Lake Ontario cms ont_nbs_cms_qm48_na_spinup.csv Net Basin Supply Lake Ontario cms eri_flw_cms_qm48_na.csv Outflow Lake Erie cms eri_flw_cms_qm48_na_spinup.csv Outflow (one year spinup) Lake Erie cms slon_flw_cms_qm48_na.csv SLON flow St. Lawrence River (Lake St. Louis) cms slon_flw_cms_qm48_na_spinup.csv SLON flow (one year spinup) St. Lawrence River (Lake St. Louis) cms stl_icw_cms_qm48_na.csv Ice/weed retardation St. Lawrence River cms stl_tde_m_qm48_na.csv Tidal Signal St. Lawrence River m ont_mlv_m_qm48_na_spinup.csv Water Level (one year spinup) Lake Ontario m stl_ics_xx_qm48_na.csv Ice status indicator St. Lawrence River N/A bati_rgh_xx_qm48_na.csv Ice/weed roughness factor Batiscan N/A card_rgh_xx_qm48_na.csv Ice/weed roughness factor Cardinal N/A corn_rgh_xx_qm48_na.csv Ice/weed roughness factor Cornwall N/A intw_rgh_xx_qm48_na.csv Ice/weed roughness factor International Tail Water N/A irhw_rgh_xx_qm48_na.csv Ice/weed roughness factor Iroquois Head Water N/A irtw_rgh_xx_qm48_na.csv Ice/weed roughness factor Iroquois Tail Water N/A jty1_rgh_xx_qm48_na.csv Ice/weed roughness factor Montreal Jetty No. 1 N/A lspr_rgh_xx_qm48_na.csv Ice/weed roughness factor Lake St. Pierre N/A lstd_rgh_xx_qm48_na.csv Ice/weed roughness factor Long Sault Dam N/A morr_rgh_xx_qm48_na.csv Ice/weed roughness factor Morrisburg N/A ogde_rgh_xx_qm48_na.csv Ice/weed roughness factor Odgensburg N/A ptcl_rgh_xx_qm48_na.csv Ice/weed roughness factor Pointe - Claire N/A sahw_rgh_xx_qm48_na.csv Ice/weed roughness factor Saunders Head Water N/A sorl_rgh_xx_qm48_na.csv Ice/weed roughness factor Sorel N/A summ_rgh_xx_qm48_na.csv Ice/weed roughness factor Summerstown N/A triv_rgh_xx_qm48_na.csv Ice/weed roughness factor Trois Rivières N/A vare_rgh_xx_qm48_na.csv Ice/weed roughness factor Varennes N/A na_fst_xx_qm48_na.csv Forecast indicator N/A N/A

open·CC-BY-4.0·Zenodo·completeSource
declared

invnet

0.00

Zhang, Yunwei

6 files · 40 GB · hdf5, zipdeclared

Files related to INVNET, a deep learning model for surface wave dispersion spectrum inversion in geophysics.

open·MIT·Zenodo·completeSource
declared

Replication Package for the paper: More Productive, but at What Cost? Understanding How GenAI Shapes Developers' Work Across the SPACE Dimensions

0.00

Murilo Coelho · de Sousa Amâncio, Francisco Dione · Paixao, Matheus · et al.

3 files · 742 KB · rardeclared

This repository contains the replication package for the paper "More Productive, but at What Cost? Understanding How GenAI Shapes Developers' Work Across the SPACE Dimensions", accepted at the 40th Brazilian Symposium on Software Engineering (SBES 2026), São Paulo, Brazil. The package includes: (i) the complete survey instrument and sanitized participant responses; (ii) qualitative coding artifacts, including the codebook, category consolidation, and classification analysis; (iii) inter-rater reliability materials (Cohen's Kappa = 0.81); (iv) quantitative datasets and statistical analysis outputs (SPACE composite scores, Cronbach's alpha, Kruskal-Wallis, Mann-Whitney U, and Dunn's post-hoc tests); and (v) supporting literature review material. All participant data were anonymized prior to disclosure. The survey materials are in Portuguese, the language of data collection.

open·CC-BY-4.0·Zenodo·completeSource
declared

Data and code for: Comparing a Vision Foundation Model (DINOv3) and a Task-Specific U-Net for Mapping Emergent Aquatic Vegetation from Fused UAV Multispectral and LiDAR Data

0.00

Tiskus, Edvinas · Tiškuvienė, Rūta · Bučas, Martynas · et al.

37 files · 1.6 GB · hdf5, tiff, torchdeclared

This record contains the labeled data, trained models, and analysis code supporting the article "Comparing a Vision Foundation Model (DINOv3) and a Task-Specific U-Net for Mapping Emergent Aquatic Vegetation from Fused UAV Multispectral and LiDAR Data" (Remote Sensing in Ecology and Conservation). Contents: - masks/ : georeferenced ground-truth segmentation masks (five classes: aquatic vegetation, water, sand, other objects, background), aligned to the fused UAV orthomosaics and spanning 13 sites across nine Lithuanian waterbodies surveyed between May and August 2024. - models/ : the two final trained segmentation models, a Keras/HDF5 U-Net and a PyTorch DINOv3 model. - code/ : Python scripts for training, evaluation, the label-efficiency experiment, and full-scene prediction. The fused 9-band orthomosaics (five-band multispectral, RGB, and a LiDAR canopy height model; approximately 62 GB) are archived separately because of their size and are available from the corresponding author on request. The DINOv3 SAT-493M pretrained backbone is distributed by Meta under its own license and is not redistributed here; obtain it from the official DINOv3 release.

open·CC-BY-4.0·Zenodo·completeSource
declared

GenixRL reclassification scores for ~ 1.05 million missense Variants of Uncertain Significance (VUS) from ClinVar

0.00

Abbas, Syed Hassan

1 files · 86 MB · rardeclared

Abbas H, et al. "GenixRL- " The VUS were extracted from CLinVar database (downloaded [Date: August 2025]) and scored using the GenixRL framework. This dataset provides the foundation for the VUS reclassification analysis presented in the main manuscript. The dataset is provided as a single compressed CSV file: GenixRL_VUS_Scored.csv.gz Key Coulmn Descriptions: [Variant Identifier Columns, e.g., CHROM, POS, REF, ALT]: Standard genomic coordinates for each variant. SYMBOL: The official gene symbol. GenixRL_Score: The continuous pathogenicity score generated by GenixRL, ranging from 0 (most likely benign) to 1 (most likely pathogenic). GenixRL_Classification: The tiered classification based on the manuscript's thresholds: 'Likely Benign': Score < 0.53 'Likely Pathogenic': Score >= 0.53 and < 0.709 'High-Confidence Pathogenic': Score >= 0.709 [Other relevant columns]: The file also includes intermediate scores from other predictors and allele frequencies used for validation.

open·CC-BY-4.0·Zenodo·completeSource
declared

Supporting Data and Evaluation Outputs for JerseyTrack: A Confidence-Guided Sports Multi-Object Tracking Method via Jersey Semantic Fusion

0.00

Cao, Shiyuan · Li, Yaning

8 files · 18 MB · csv, rardeclared

This repository provides the supporting materials for the manuscript "A Confidence-Guided Sports Multi-Object Tracking Method via Jersey Semantic Fusion." The archived materials include evaluation summaries, final tracking outputs, environment records, dataset mapping files, protocol reproduction materials, intermediate summary files, and revision evidence used to support the reported SportsMOT validation results. The reported formal results were obtained from the real detector, real frame-level semantic feature extraction, confidence partitioning, cascaded association, and TrackEval evaluation pipeline. The original public benchmark datasets, including SportsMOT, TeamTrack, SoccerNet Tracking, MOT17, and MOT20, are not redistributed in this repository and should be obtained from their original providers.

open·CC-BY-4.0·Zenodo·completeSource
declared

Supplementary material 2 from: Nie Y, Huang B (2026) Drechslerosporium cornellii gen. et sp. nov. within the Basidiobolaceae exhibiting unique conidial discharge and digitate chlamydospores. MycoKeys 136: 177-191. https://doi.org/10.3897/mycokeys.136.200461

0.00

Nie, Yong · Huang, Bo

1 files · 1.5 KB · rardeclared

BI tree

open·CC0-1.0·Zenodo·completeSource
declared

Supplementary material 1 from: Nie Y, Huang B (2026) Drechslerosporium cornellii gen. et sp. nov. within the Basidiobolaceae exhibiting unique conidial discharge and digitate chlamydospores. MycoKeys 136: 177-191. https://doi.org/10.3897/mycokeys.136.200461

0.00

Nie, Yong · Huang, Bo

1 files · 9.1 KB · rardeclared

Alignments for phylogen

open·CC0-1.0·Zenodo·completeSource
declared

astroARIADNE pre-computed spectra cache

0.00

Vines, Jose I.

1 files · 3.0 GB · hdf5declared

Pre-computed, resolution-broadened (R=1500) stellar atmosphere spectra cache for the astroARIADNE SED fitting package. Contains 7 model grids (Phoenix v2, BT-Settl, BT-NextGen, BT-Cond, Castelli & Kurucz 2004, Kurucz 1993, Coelho 2014) resampled to a common logarithmic wavelength grid (0.125-4.629 µm). This cache eliminates the need to download the full ~770 GB model libraries for SED plotting.

open·MIT·Zenodo·completeSource
declared

Dataset for Evaluating the Effectiveness of AI-Generated Data for Video-Based Action Recognition

0.00

Gomułka, Kamil · Woźniak, Piotr · Krzeszowski, Tomasz

1 files · 6.4 GB · sevenzipdeclared

======================= License ======================= This dataset is made available under the Creative Commons Attribution 4.0 International (CC BY 4.0) license. ======================= Summary ======================= DATASET FOR EVALUATING THE EFFECTIVENESS OF AI-GENERATED DATA FOR VIDEO-BASED ACTION RECOGNITION This repository contains the dataset accompanying the paper "Evaluating the Effectiveness of AI-Generated Data for Video-Based Action Recognition". The dataset was curated to investigate the impact of incorporating AI-generated synthetic video data into the training process of deep learning models (such as I3D and TimeSformer) for human action recognition. It is particularly focused on evaluating cross-domain generalization and addressing the synthetic-to-real distribution shift using domain adaptation techniques. It is obligatory to cite the following paper in every work that uses the dataset: K. Gomulka, P. Wozniak, T. Krzeszowski: Evaluating the Effectiveness of AI-Generated Data for Video-Based Action Recognition, Electronics, MDPI, 2026. ======================= Data description ======================= The dataset focuses on 15 target action categories from the HMDB51 benchmark: Draw Sword, Fall, Flick Flack, Handstand, Jump, Kick, Pick, Punch, Run, Shoot Bow, Shoot Gun, Sit, Throw, Walk, and Wave. The repository consists of three main data subsets: * Grok Synthetic Dataset (Grok): Contains 1,500 synthetic video sequences (100 videos per class) generated using the Grok Imagine 1.0 model. Resolution is 560x560 pixels with a duration of 6 seconds per clip. * Meta Synthetic Dataset (Meta): Contains 1,480 synthetic video sequences (80 in the 'throw' class, 100 videos per other classes) generated using the Meta AI Vibes model. Resolution is 624x624 pixels with a duration of 5 seconds per clip. * HMDB51 Subset (hmdb51_org): A specific subset of the real-world HMDB51 dataset containing 2,315 video sequences distributed across the 15 target action classes, providing a direct real-world baseline. The repository also includes sample .csv files defining the train, validation, and test splits. In the example configuration, 60% of the HMDB51 dataset is allocated to training, 20% to validation, and 20% to testing, with all synthetic Grok videos appended exclusively to the training set. ======================= Dataset structure ======================= * Grok/ - directory containing 1,500 synthetic videos generated by Grok Imagine 1.0. * Meta/ - directory containing 1,480 synthetic videos generated by Meta AI Vibes. * hmdb51_org/ - directory containing 2,315 real-world videos from the HMDB51 dataset. * ghtrain_hval/ - directory containing example .csv files defining the train, validation, and test splits for the baseline HMDB + Grok setup. * train.csv - training set paths and labels. * val.csv - validation set paths and labels. * test.csv - testing set paths and labels. * README.md - detailed documentation including dataset overview, configuration details, and instructions for training the TimeSformer model using the provided data splits. ======================= Generation Methodology ======================= The synthetic videos were generated using a structured, combinatorial prompt engineering framework resulting in 100 unique prompts per class. Each prompt combined predefined descriptions of the action, environmental context (e.g., vintage film style, urban outdoor), and camera framing to ensure high scene diversity. While the generated sequences preserve essential motion semantics, they also may contain inherent generative artifacts.

open·CC-BY-4.0·Zenodo·completeSource
declared

A Parliamentary Discourse Dataset from the German Bundestag

0.00

Njie, Adama · Torkayesh, Ali E · Venghaus, Prof. Dr. Sandra

10 files · 1.6 GB · csv, gzip, parquetdeclared

Structured, speaker-attributed corpus of all German Bundestag plenary session transcripts ( Plenarprotokolle ) from the first legislative period to the present (WP01-WP21, September 1949 - April 2026). Every attributed speech is extracted from the official PDFs published by the Deutscher Bundestag under open data policy and linked to the speaker's name, parliamentary role, party affiliation, and gender. Scale: 4,611 sessions · 1,033,723 speeches · 4,205 identified MdBs · 76 years of parliamentary debate Dataset files speeches.parquet - one row per attributed speech: speaker name, role, party, gender, stammdaten_id, full German text (~1 GB) persons.parquet - one row per MdB: cross-session identity linking all name variants via stammdaten_id; canonical name, birth date, career span, total speeches. Use this - not speakers.parquet - for person-level analysis sessions.parquet - one row per plenary session: date, city, Wahlperiode, source PDF hash, extraction engine speakers.parquet - name-string index: one row per unique name string as extracted from the transcripts. Useful for understanding extraction quality; not suitable for person-level aggregation (the same politician often appears under several name variants across sessions) parties.csv - reference table of 31 German parliamentary parties, 1949-present speeches.csv.gz - CSV fallback for Stata and Excel users (same columns as speeches.parquet) datapackage.json - Frictionless Data schema with column descriptions and foreign key constraints Cross-session identity The same politician often appears under different name strings across sessions (e.g. "Schmidt", "Dr. Schmidt", "Frau Dr. Schmidt"). Cross-session person linkage is provided via stammdaten_id , matched against the official Bundestag Stammdaten biographical XML. The persons.parquet table aggregates all name variants for the same MdB into one row with correctly summed speech counts, career span, and birth date. Coverage: ~98.5% of speeches are linked to a stammdaten_id; the remaining ~1.5% are ambiguous surname-only attributions or speakers not in the Stammdaten. Coverage and sources Source PDFs are the official Stenografische Berichte downloaded from the Bundestag open-data portal (bundestag.de). Party-share normalisation in the corpus statistics uses official seat counts per Wahlperiode sourced from the Federal Returning Officer (Bundeswahlleiter, bundeswahlleiter.de). Two PDF generations are covered: scanned and OCR'd documents (WP01-WP09, Bonn era, 1949-1987) and born-digital documents (WP10-WP21, 1987-present). The engine column in sessions.parquet flags whether pdftotext (born-digital) or pdfminer (OCR fallback) was used; this is the primary data-quality indicator for NLP use. Speaker attribution Each speech is attributed using four patterns extracted from the transcript format: presiding officers (Präsident/in, Vizepräsident/in), regular members (name + party), government officials (name + Bundeskanzler/in, Bundesminister/in, etc.), and procedural roles (Berichterstatter/in, etc.). The party field is null for ~60% of speeches - this is expected, as presiding officers and ministers are not identified by party in the transcript. Gender annotation & distribution Gender is derived by matching speaker names against the official Bundestag Stammdaten biographical XML (all MdBs since 1949), with fallbacks for role title, honorific prefix, manually researched overrides, and a gender_guesser first-name heuristic. The gender_source column distinguishes stammdaten (authoritative, 83%), role_title (gendered job title in attribution, 6.4%), title_prefix (Frau/Herr honorific, 0.5%), manual (historically researched, 2.9%), and inferred (name-based heuristic, 4.7%). Gender distribution: Female 26.6% · Male 73.4% · Unknown 0.0%. Data quality All speeches pass automated validation: zero null speaker names, zero sequence gaps, zero CID artefacts, zero party-misclassified-as-Bundesland errors. Eight sessions with conflicting source PDFs were deduplicated (first lexicographic occurrence retained). 252 non-person names incorrectly accepted by the parser (table headers, legislative terms, agenda fragments) are excluded at build time via a curated exclusion list. OCR sessions (WP01-WP09) may contain Unicode replacement characters (U+FFFD); the engine field identifies these sessions. Licence CC BY 4.0. The underlying Plenarprotokolle are official government documents of the Deutscher Bundestag and are in the public domain.

open·CC-BY-4.0·Zenodo·completeSource
page 1next →

Select a result to see its full details here: the measured structure, quality, and the loader, without leaving your search.