Exploration

ResearchFeatured

Discovery

DiscoverSourcesQuality

Analysis

Working setReviews
Flow StudioTeamConcept
Settings

Partners

  • AI AlliancePrime
  • BrightQueryBuilds Meridian
  • OpenMinedFunded partner
  • MLCommonsFunded partner
  • Hugging FaceDeployment platform
See the full consortium and what each partner wires

Meridian is the discovery layer for research data, built by BrightQuery within the AI Alliance.

hybrid · semantic + lexical · 7 datasets ranked · 2.88s

Depthcataloged7
Licenseopen6share alike1
Accessopen7
Formatgeopackage6geojson3shapefile3csv1fasta1pdf1sqlite1xlsx1zip1
Sourcezenodo7
clear
1-7 of 7sortrelevancemeasured firstqualitysize
declared

A Layer of Late Victorian London: The Building Footprints from the 1:1,056 Ordnance Survey Map (1891-1896)

0.00

Petitpierre, Remi · di Lenardo, Isabella · Hudson, Polly · et al.

4 files · 1.4 GB · geopackage, zipdeclared

The dataset provides a large-scale vector layer of historical building footprints for the Greater London area at the end of the 19th century. It comprises 1,299,029 individual building footprints, extracted automatically from the Ordnance Survey five-feet-to-the-mile (1:1,056) map. The original maps were surveyed between 1891 and 1895 and published between 1893 and 1896, covering approximately 450 km² of urban and suburban London. The dataset was produced using a deep learning-based semantic segmentation pipeline, achieving approximately 97% precision and 95% recall in building detection. Source The source maps were digitised by the National Library of Scotland at 400 ppi and manually georeferenced in collaboration with the David Rumsey Map Collection. The extraction relies on 753 georeferenced map sheets, provided that six sheets are missing from the original archive. Extraction methodology The dataset was generated through an automated workflow: 262 image patches were annotated manually with three classes: (i) regular buildings; (ii) compound buildings; (iii) building boundaries A Mask2Former [1] semantic segmentation model is used. For training and inference, we followed the approach described by [2]. At inference time, boundary predictions are supplemented by the predictions of a specialist model, trained on cadastral plans [3]. Polygon geometries are extracted, based on predicted building contours, and vectorised. Compound buildings components are merged. Content Two versions of the dataset are provided: london_buildings_1891-96_raw.gpkg Raw output from the automated extraction pipeline (after vectorisation and merging) london_buildings_1891-96_corr_v1.gpkg Minimally corrected version including manual adjustments for major structures (e.g. railway stations, monuments, large buildings) Each geometry has a field buil_class , taking values in ['regular','compound'] , relative to the cartographic representation of building classes and the associated processing approach. In corr_v1, manually added geometries take the value NULL . In addition, we relase the labeled images used for training the segmentation model: annotations.zip ZIP folders containing the training and validation data The archive contains subfolders images , containing the tif image samples and labels , corresponding to the semantic class indexes, in png format. labelmap.txt details the class index codes. Data format Format: Geopackage Coordinate reference system: WGS 84 / Pseudo-Mercator (EPSG:3857) Temporal coverage Survey period: 1891-1895 Publication period: 1893-1896 Data creation: 2024-2026 Descriptive statistics Number of building footprints: 1,300,831 (raw), 1,299,040 (corr_v1) Coverage area: 450 km² Detection performance (raw): Precision: 97% Recall: 95% Use and reuse potential This dataset supports research in: urban history historical GIS urban morphology economic and social history It is particularly suited for studying long-term urban change and fine-grained spatial patterns in industrial-era cities. Related publication This dataset is described in a data paper submitted to the Journal of Open Humanities Data Corresponding author Remi Petitpierre Email: remi.petitpierre@epfl.ch Funding The research was supported by the College of Humanities at EPFL and the European Union Horizon Europe Programme (Grant No. 101233051). License This dataset is licensed under the Creative Commons Attribution 4.0 International (CC BY 4.0) license. Citation If you use this dataset, please cite: @misc{london_footprints_petitpierre_2026, author = {Petitpierre, R{\'{e}}mi and di Lenardo, Isabella and Hudson, Polly and Herold Hendrik and McDonough, Katherine and Hecht, Robert and Vaienti, Beatrice and Fleet, Christopher}, title = {{A Layer of Late Victorian London: The Building Footprints from the 1:1,056 Ordnance Survey Map (1891-1896)}}, year = {2026}, publisher = {EPFL}, url = {https://doi.org/10.5281/zenodo.19497434}} Limitations Lower accuracy for very small buildings Occasional segmentation errors in dense or complex areas Minor inconsistencies near map sheet boundaries Liability The authors assume no liability for the use of this dataset. References Cheng, B., Misra, I., Schwing, A. G., Kirillov, A., & Girdhar, R. (2022). Masked-attention Mask Transformer for Universal Image Segmentation . arXiv. https://doi.org/10.48550/arXiv.2112.01527 Petitpierre R. (2026) Generalizable Multiscale Segmentation of Heterogeneous Map Collections . arXiv. https://doi.org/10.48550/arXiv.2603.05037 Petitpierre, R., di Lenardo, I., & Rappo, L. (2024). Revealing the Structure of Land Ownership through the Automatic Vectorisation of Swiss Cadastral Plans . Digital History Switzerland. https://doi.org/10.13140/RG.2.2.26632.33281

open·CC-BY-4.0·Zenodo·completeSource
declared

Dicionário e Resumo do Banco de Dados Preditor – Recorte São Paulo (SP)

0.00

Marques da Silva, Gabriel · de Freitas Junior, Adirson Maciel · Feitosa, Flávia da Fonseca

2 files · 500 MB · geopackage, xlsxdeclared

Dicionário e Resumo do Banco de Dados Preditor – Recorte São Paulo (SP) Este conjunto de dados contém o recorte territorial do banco de dados preditor consolidado para o município de São Paulo (no formato GeoPackage: 04_final_consolidado_31_completo_sp.gpkg ) e seu respectivo dicionário e resumo de variáveis (no formato Excel: Dicionário Modelo SP.xlsx ). O banco de dados é estruturado em uma grade celular regular de alta resolução com células de 50 metros por 50 metros (50m x 50m) . Cada célula de 2.500 m² representa uma unidade territorial urbana que agrega variáveis morfológicas urbanas, informações socioeconômicas e escores resultantes de um modelo de aprendizado de máquina (Machine Learning) preditor de assentamentos informais/precários. 1. Visão Geral do Recorte de São Paulo (SP) Total de feições (células de análise) : 388.536 Dimensão das células : 50m x 50m (2.500 m² por célula) Área urbana total analisada : Aproximadamente 971,34 km² (composta pela união das células da grade dentro do município) Total de variáveis analisadas : 111 variáveis originais no GeoPackage, sendo 109 colunas ativas com dados (não-nulos) para São Paulo. Qualidade do preenchimento : Excelente cobertura. Das 109 variáveis ativas, 108 apresentam 100% de cobertura espacial para o município de São Paulo. A única variável com preenchimento parcial é a ranking_candidato (com 93,42% de cobertura). 2. Estrutura do Dicionário (Dicionário Modelo SP.xlsx) O arquivo de dicionário e resumo é composto por 5 colunas que descrevem as variáveis do GeoPackage: Variable (Nome da Variável): O nome técnico da coluna na tabela de atributos do GeoPackage (ex: ID , prob_fcu , ibge_mediapopc ). Type (Tipo do Dado): Indica a natureza do dado físico (ex: object , float64 , int16 / int32 , geometry ). Non_Null_Count (Registros Válidos): Quantidade de feições em São Paulo que possuem dados válidos nesta coluna, desconsiderando campos nulos, vazios ou informados como ausentes (como "None" ou "null"). Coverage_Percentage (Taxa de Cobertura): O percentual de feições de São Paulo que possuem a variável preenchida. Sample_Value (Valor de Amostra): Um exemplo real do dado contido na coluna retirado da primeira linha válida encontrada, servindo para referência rápida sobre o formato da informação. 3. Categorias de Variáveis no Banco de Dados As variáveis presentes no banco de dados preditor de São Paulo dividem-se nos seguintes blocos temáticos: A. Variáveis de Localização e Resolução Territorial : Identificam as células geograficamente e associam os recortes às divisões oficiais (ex: ID , id_rg2017_cd_mun , id_rg2017_mun_nome , id_100m , id_200m , id_grid_1km ). B. Métricas de Classificação do Modelo Preditor : Valores resultantes da aplicação do modelo de inteligência artificial sobre o território (ex: prob_fcu , rank_class , ranking_total , ranking_candidato ). C. Variáveis de Importância (Feature Importance) : O modelo registra quais fatores foram os mais determinantes para a classificação de cada polígono específico (ex: top1_feat a top16_feat , top1_score a top16_score , top1_val a top16_val ). D. Variáveis Socioeconômicas e de Infraestrutura (Censo IBGE) : Indicadores estatísticos herdados dos setores censitários do IBGE (ex: ibge_mediapopc , ibge_mediadomc , m_ibge_renddppo , p_ibge_esginade , p_ibge_lixoinade , p_ibge_naocalca , p_ibge_naoilupub ). E. Variáveis Morfológicas Urbanas (GBA) : Métricas físicas de ocupação do solo e edificações (ex: gba_num_edif , gba_edif_por_ha , idade_ocupacao_menor ). 4. Agradecimentos e Financiamento Este estudo foi desenvolvido no âmbito do Centro de Estudos da Favela (CEFAVELA) (projeto FAPESP nº 2022/12259-8). A.M.F.J. agradece à Fundação de Amparo à Pesquisa do Estado de São Paulo (FAPESP) pela bolsa de pós-doutorado vinculada a este projeto (processo 2025/03115-0). G.M.S. agradece à Fundação de Amparo à Pesquisa do Estado de São Paulo (FAPESP) pela bolsa de Treinamento Técnico (TT/FAS) vinculada a este projeto (processo 2025/24201-2). 5. Como Citar (How to Cite) ABNT : SILVA, Gabriel Marques da; FREITAS JUNIOR, Adirson Maciel de; FEITOSA, Flávia da Fonseca. Dicionário e Resumo do Banco de Dados Preditor – Recorte São Paulo (SP) . Zenodo, 2026. DOI: 10.5281/zenodo.21268382. APA : Silva, G. M., Freitas Junior, A. M., & Feitosa, F. F. (2026). Dicionário e Resumo do Banco de Dados Preditor – Recorte São Paulo (SP) . Zenodo. https://doi.org/10.5281/zenodo.21268382 6. Links Relacionados (Related Links) Centro de Estudos da Favela (CEFavela): https://cefavela.ufabc.edu.br/ Projeto Revelando Favelas (Arcabouço Metodológico): https://cefavela.ufabc.edu.br/revelando-favelas-arcabouco-metodologico-para-identificacao-e-caracterizacao-de-favelas/

open·CC-BY-4.0·Zenodo·completeSource
declared

Map of New Spain [Segment], c. 1800

0.00

Kavanagh, Jack · Anthony, Patrick

6 files · 11 MB · geojson, geopackage, shapefiledeclared

A historical map showing a segment of the boundaries of New Spain in c. 1800. This map is fully open to fellow researchers and is available in multiple open source formats (SHP, GPKG, GeoJSON).

open·CC-BY-4.0·Zenodo·completeSource
declared

Map of Prussia, c. 1795

0.00

Kavanagh, Jack · Anthony, Patrick

6 files · 9.3 MB · geojson, geopackage, shapefiledeclared

A historical map of the boundaries of Prussia in c. 1795. This map is fully open to fellow researchers and is available in multiple open source formats (SHP, GPKG, GeoJSON).

open·CC-BY-4.0·Zenodo·completeSource
declared

Software and AMR peptide database for 'PEPTiGEN: a tool for mining antimicrobial resistance PEPTides using GENe data of public available repositories'

0.00

Meekes, Lisa · Tabaro, Francesco · Bexkens, Michiel · et al.

41 files · 8.2 GB · csv, fasta, pdfdeclared

This record contains the Python software for PEPTiGEN, a tool for generating tryptic peptides from prokaryotic gene sequences and their variants, and the associated antimicrobial resistance (AMR) peptide database. The database is provided as an SQL file and a CSV file containing all genes and predicted peptides. The README file contains explanation of the PEPTiGEN tool. The SQL database schema files contains both the database schema of the SQL database used in the PEPTiGEN analysis as the database schema of the AMR peptide datbase.

open·CC-BY-4.0·Zenodo·completeSource
declared

Heliopolis Project: Archaeological Features

0.00

Langermann, Florence · Blaschta, Stephanie · Dietze, Klara · et al.

3 files · 731 KB · geopackagedeclared

Beschreibung (Deutsch) Diese Datenpublikation enthält archäologische Geodaten aus den Forschungsprojekten "The Cultic Centre of the Sun-God of Heliopolis (Egypt)" und „Eclipse and Mutation: The end of the sun temple of Heliopolis" zu den ergrabenen und aufgefundenen Befunden im altägyptischen Heiligtum in Heliopolis im heutigen Kairo. Die Befunde wurden innerhalb einer Reihe von Grabungskampagnen zwischen 2012 bis 2024 ergraben und dokumentiert. Diese Datensammlung stellt eine Zusammenstellug dieser Befunde dar und soll als Grundlage für eine Gesamtkartierung dienen. An der Zusammenstellung der Daten waren maßgeblich die drei verantwortlichen Ausgräberi:innen Stephanie Blaschta , Klara Dietze und Florence Langermann beteiligt. Für die Erstellung der Datenstrukturen und Schemata und dem Zusammenführen der Daten war Michael Schleier verantwortlich. Das Projekt wurde von der Deutschen Forschungsgemeinschaft (DFG) sowie weiteren akademischen und privaten Förderinstitutionen unterstützt. Weitere Informationen: Projektwebseite DAI Datendokumentation des i3mainz, Hochschule Mainz Description (English) This data publication contains archaeological geodata from the research projects "The Cultic Center of the Sun-God of Heliopolis (Egypt)" and "Eclipse and Mutation: The end of the sun temple of Heliopolis" on the excavated and discovered features in the ancient Egyptian sanctuary in Heliopolis in present-day Cairo. The features were excavated and documented during a series of excavation campaigns between 2012 and 2024. This data collection represents a compilation of these features and is intended to serve as the basis for an overall mapping. The three excavators responsible, Stephanie Blaschta , Klara Dietze and Florence Langermann , played a key role in compiling the data. Michael Schleier was responsible for creating the data structures and schemas and merging the data. The dataset is published under a CC BY-SA 4.0 license . The project is funded by the German Research Foundation (DFG) and a wide network of institutional and private sponsors. Further information: Project website at DAI Datendokumentation of the i3mainz, Mainz University of Applied Sciences

open·CC-BY-SA-4.0·Zenodo·completeSource
declared

Map of Kazakh Steppe, c. 1848

0.00

Kavanagh, Jack · Anthony, Patrick

6 files · 1.5 MB · geojson, geopackage, shapefiledeclared

A historical map of the boundaries of the Kazakh Steppe in c. 1848. This map is fully open to fellow researchers and is available in multiple open source formats (SHP, GPKG, GeoJSON).

open·CC-BY-4.0·Zenodo·completeSource

Select a result to see its full details here: the measured structure, quality, and the loader, without leaving your search.