Lemor, Antoine · Pillod, Alizée · Taylor, Matthew · et al.
266,271 rows × 10 cols
7 numeric · 3 text
The Canadian Climate Framing (CCF) Database is a comprehensive, machine-learning-annotated corpus of climate-change media coverage in Canada. It comprises 266,271 articles from 20 major Canadian newspapers (1978-2024) processed into 9,198,958 two-sentence analytical units (82.9% English, 17.1% French). Each unit is annotated across 65 hierarchical categories by 128 BERT and CamemBERT classifiers, with a macro F1 of 0.866 on a 1,000-sentence gold standard double-coded by an independent annotator (Gwet's AC1 = 0.894, Krippendorff's α = 0.698, Cohen's κ = 0.596 on the 400 blind sentences). Each category receives an A/B/C reliability tier summarising annotation quality from classifier performance and inter-coder agreement. The deposit ships six relational tables (bibliographic metadata, sentence-level annotations, named-entity rollups, article-level aggregates, per-category reliability tiers, and 9,462,845 BAAI/bge-m3 sentence-and-title embeddings). Raw newspaper text is excluded for copyright reasons; bibliographic coordinates (media, date, title, author, page_number) are sufficient for any researcher with institutional access to Factiva, Eureka.cc or ProQuest Canadian Major Dailies to recover the original sentences. This deposit accompanies a methodology paper currently under revision at Scientific Data (Nature Portfolio). This deposit is the Apache Parquet mirror of the canonical PostgreSQL edition (cross-referenced in Related identifiers ). Each of the six relational tables is provided as a standalone .parquet file with ZSTD compression; the 1024-dimensional BAAI/bge-m3 embedding column is materialised as a list<float> , and JSONB entity arrays are serialised as UTF-8 JSON strings. The schemas are otherwise identical to the PostgreSQL edition. The Parquet bundle is readable natively by pandas , polars , R/ arrow , DuckDB , and Spark without any database backend: import pandas as pd agg = pd.read_parquet('CCF_article_aggregates.parquet') emb = pd.read_parquet('CCF_sentence_embeddings.parquet') The HNSW index that ships with the PostgreSQL edition is not transferable to Parquet; brute-force cosine similarity remains tractable on the embedding column (≈ 9.46 M × 1024 float16). The full annotation pipeline, training data, manual-annotation JSONL, intercoder-reliability benchmark, methodology manuscript (LaTeX sources + PDF), and reproducibility scripts are bundled with this deposit as ccf_code_and_paper.tar.gz . The same materials are also available on the project's OSF companion deposit ( 10.17605/OSF.IO/Q5W47 ) and on the development mirror at GitHub .
Aimed at year-round recording of the chemical aerosol composition in central Antarctica, an unattended operating aerosol sampler was successfully deployed at the EPICA deep drilling site in Dronning Maud Land (Kohnen Station). Analyses of teflon/nylon filter packs consecutively collected over bi-weekly intervals during the February 2003 to December 2005 period allowed to evaluate seasonal concentration variations of methane sulphonate (MS), Cl-, NO3-, non-sea salt (nss-)SO4**2- and Na+, while NH4+ and mineral dust related ion results remained below detection limits. For MS and nss-SO4**2 distinct late summer maxima around 44 and 200 ng/m**3, respectively, were found, while (total) NO3- showed a broad November maximum of about 52 ng m**-3. In contrast, the highest concentrations of Na+ with peak values of up to 160 ng/m**3 were observed during the winter half year. The seasonality of these species broadly coincided with long-term observations at the coastal Neumayer Station, including surprisingly comparable NO3- levels. However, the biogenic sulphur and sea salt concentrations were lower at Kohnen by typically a factor of 2-3 and 10, respectively. The arrival of sea ice derived sea salt particles at Kohnen could not clearly detected, since even during mid-winter the nss-SO4**2- to Na+ ratio was generally too high to unambiguously identify a sulphur depleted sea salt SO4**2- fraction.
Aerosol samples collected over the North Atlantic from ship were analysed for Sodium, Magnesium, Potassium, Calcium and Chloride. A found dependence of sea salt concentrations from wind velocity is compared with earlier results. The mean of the ratio Cl/Na was close to that for sea water; the Mg-, K- and Ca-concentrations in the aerosol, however, were enriched with respect to sea water. It is shown that continental advection influences the measured aerosol components over the North Atlantic.
Lemor, Antoine · Pillod, Alizée · Taylor, Matthew · et al.
2.2 MB
The Canadian Climate Framing (CCF) Database is a comprehensive, machine-learning-annotated corpus of climate-change media coverage in Canada. It comprises 266,271 articles from 20 major Canadian newspapers (1978-2024) processed into 9,198,958 two-sentence analytical units (82.9% English, 17.1% French). Each unit is annotated across 65 hierarchical categories by 128 BERT and CamemBERT classifiers, with a macro F1 of 0.866 on a 1,000-sentence gold standard double-coded by an independent annotator (Gwet's AC1 = 0.894, Krippendorff's α = 0.698, Cohen's κ = 0.596 on the 400 blind sentences). Each category receives an A/B/C reliability tier summarising annotation quality from classifier performance and inter-coder agreement. The deposit ships six relational tables (bibliographic metadata, sentence-level annotations, named-entity rollups, article-level aggregates, per-category reliability tiers, and 9,462,845 BAAI/bge-m3 sentence-and-title embeddings). Raw newspaper text is excluded for copyright reasons; bibliographic coordinates (media, date, title, author, page_number) are sufficient for any researcher with institutional access to Factiva, Eureka.cc or ProQuest Canadian Major Dailies to recover the original sentences. This deposit accompanies a methodology paper currently under revision at Scientific Data (Nature Portfolio). This deposit is the canonical PostgreSQL edition. It contains a pg_dump -Fd directory archive (compressed into a single .tar file) of the six relational tables, including the pgvector extension and HNSW cosine indexes for sub-second semantic-similarity search. Restoration is a one-liner: tar -xf CCF_Database.tar && createdb CCF_Database && psql -d CCF_Database -c 'CREATE EXTENSION IF NOT EXISTS vector;' && pg_restore -d CCF_Database --no-owner --no-privileges -j 8 CCF_Database_dump A column-oriented Apache Parquet mirror of the same six tables is available as the sister deposit on Zenodo (cross-referenced in Related identifiers ). The Parquet mirror is recommended for users without PostgreSQL access (it is directly readable by pandas, polars, R/arrow, DuckDB, and Spark). The full annotation pipeline, training data, manual-annotation JSONL, intercoder-reliability benchmark, methodology manuscript (LaTeX sources + PDF), and reproducibility scripts are bundled with this deposit as ccf_code_and_paper.tar.gz . The same materials are also available on the project's OSF companion deposit ( 10.17605/OSF.IO/Q5W47 ) and on the development mirror at GitHub . Requirements: PostgreSQL 16 or 17 with pgvector ≥ 0.8.2 (for halfvec(1024) storage of the sentence embeddings).
During the 1965 Atlantic Expedition of the "Meteor“ concentrations of various atmospheric trace gases were measured. The following gases were considered: carbon dioxide (CO2), sulfur dioxide (SO2), nitrogene dioxide (NO2), and nitric oxide (NO). The air whereof these components were measured was sucked in from a height of 14 m above the surface of the sea. The results allow conclusions upon the long term global increase of the atmospheric CO2 content, the meridional distribution of the CO2 on the Atlantic Ocean, and the dependance of its concentration upon the time of the day and the thermal structure of the atmosphere. Attempts at determining concentrations of sulfur dioxide and nitric oxide of non-continental origin failed at large. Concentrations of NO2, however, could succesfully be measured.
Piel, Claudia · Weller, Rolf · Huke, Michael · et al.
98 rows × 14 cols · 13 KB
8 text · 5 numeric · 1 datetime
During three summer campaigns in January/February 2000, 2001, and 2002 the ionic composition of the aerosol at the European Project for Ice Coring in Antarctica (EPICA) deep-drilling site at Kohnen Station was measured in daily resolution. In 2000 and 2002 we observed mean (±std) non-sea-salt sulfate (nss-[SO4]2-) concentrations of 353 ± 100 ng/m**3 and 320 ± 250 ng/m**3, as well as methane sulfonate (MS) concentrations of 59 ± 36 ng/m**3 and 74 ± 80 ng/m**3, respectively. For the summer campaign in 2001, significantly lower nss-[SO4]2- and MS levels of 164 ± 150 ng/m**3 and 19 ± 12 ng/m**3, respectively, were typical. The mean MS/nss-[SO4]2- ratio ranged from about 0.1 to 0.2. MS and nss-[SO4]2- concentrations and their variability were roughly comparable to coastal stations at summer. Supported by air mass back trajectory analyses, this finding documented an efficient long-range transport to Kohnen via the free troposphere. MS/nss-[SO4]2- ratios exhibited a strong dependence on the MS concentration with systematically higher ratios at higher MS concentrations, a peculiarity which is also evident in a firn core drilled at this site.
This dataset quantifies the uncertainty in mapping late-successional and old-growth (LSOG) forest across the approximately 4.2 million hectares of Maine's unorganized townships, and tests whether LSOG is rapidly disappearing. Three to four independent, credible mapping methods are compared on a common 100 m grid: (M1) a reproduction of the Hagan et al. (2026) airborne-LiDAR canopy random forest, rebuilt from their public Zenodo deposit; (M2) a logistic model of the FIA field-structure LSOG class on Potapov (GEDI-calibrated) canopy height; (M3) a direct canopy-height threshold; and (M4) the FIA structural class imputed to every pixel via USFS TreeMap (2016, 2020, 2022). Version 1.2.0 additions. This version adds the materials behind the formal Ecosphere Comment on Hagan et al. (2026): (a) a cross-validated accuracy assessment (AUC) of each mapping approach on the original authors' own training plots, showing that high training accuracy does not transfer to agreement among independent maps; (b) an FIA design-based estimate of older forest with sampling-error confidence intervals, the unbiased ground reference the original analysis lacked, putting older forest at about 3.9 percent (3.3 to 4.6) and rising, including on private commercial timberland; (c) a threshold-sensitivity sweep and a 20-seed reproduction ensemble; (d) an ownership-resolved breakdown (private commercial versus public); (e) a hex-scale (8 km) summary of cross-method disagreement; and (f) the Comment manuscript and Supporting Information. Headline findings. Credible methods disagree by roughly 2.8 times on how much LSOG exists and on the location of most LSOG hectares, while agreeing closely on the rare, well-defined old-growth core. Protecting the top 5 to 20 percent of hectares by one map versus another overlaps on only 16 to 30 percent of the ground, so single-map patch-level prioritization for large expenditures is fragile. The design-based FIA estimate and TreeMap imputation both show older forest stable to increasing rather than rapidly declining; the apparent loss reported elsewhere is a gross harvest flux, not a net stock decline. Contents. Derived 100 m GeoTIFFs (reproduced Hagan class, v5.1-GEDI probability, TreeMap class, a per-cell method-consensus layer), summary tables (area by method, pairwise agreement, concordance, prioritization fragility, AUC by approach, design-based older-forest trend with CIs, ownership breakdown, and FIA validation), the analysis R scripts, quick-look figures, the Ecosphere Comment manuscript and Supporting Information, and a full methods-and-findings report (PDF). Privacy. No FIA plot coordinates are included; all products are derived rasters or aggregate summary tables. Caveats: the robust temporal signal is direction rather than precise rate; FIA stand age is modeled, so a structural large-tree domain is reported alongside the age domain; cross-validated intervals are best read as lower bounds because plots are spatially dispersed but not independent. See the README and report for full methods, provenance, and limitations. Version 1.12.0 additions. The cross-map comparison is refined to independent remote-sensing operationalizations only. (a) A three-map remote-sensing ensemble over Maine on a common 100 m grid: reproduced Hagan airborne-LiDAR (any-LSOG 21.9 percent), an FIA-structure class on Potapov GEDI-calibrated spaceborne canopy height (14.0 percent), and the ORNL/Bruening national old-growth stratum (36.1 percent); the three span a 2.6-fold range and agree on only 2.7 percent of flagged hectares, with the ORNL stratum spatially uncorrelated with the structure maps. The USFS TreeMap imputation is reclassified as a second FIA-anchored accounting, reported with the design-based estimate rather than as an independent map. (b) A design-based estimate of LSOG itself: integrated any-LSOG 14.1 percent (12.9 to 15.3) and strict four-axis true LSOG 3.1 percent (2.5 to 3.7) of Maine forestland, the airborne map exceeding even the inclusive ground estimate. (c) A balanced-model LSOG probability surface at 100 m, with a binary class calibrated to the design-based area to bound over-prediction. (d) A multi-objective support vector regression pilot tracing the Pareto front of total versus systematic (attenuation) error. Derived rasters, the three-map agreement layer, the ORNL stratum reprojected to the study grid, tables, R scripts, the updated Comment, and the companion manuscript are included. No FIA plot coordinates are included. Version 1.13.0. Final consolidated release. Adds: a rare-class remedy menu for the reproduced random forest (default vs class weighting vs balanced sub-sampling vs voting-threshold; old-growth detection 0.24 to 0.82, mapped old-growth area 1.0 to 2.3 percent); the definitive five-model LSOG probability map for Maine with across-model uncertainty and reference reserves (MNAP/TNC network, Baxter, Big Reed) over real state and county boundaries; design-based 95 percent confidence intervals for forest-type and ecoregion representation and for disturbance shares; a full robustness/stress-test matrix; and the copy-edited, sole-authored Comment, companion manuscript, and Maine Forest Products Council technical report. Authorship updated to Aaron R. Weiskittel.
Data archive supporting the manuscript " Multilayer canopy model outperforms big-leaf model for evapotranspiration predictions under high water and heat stress conditions " (Authors: Raghav, Liu, Kumar, Bisht) This archive provides, for each of 28 eddy-covariance sites, the hourly time series of modeled and observed evapotranspiration (ET) together with the meteorological drivers used to force the models. ET is reported as latent heat flux (LE), in W m⁻² , the native model and tower-measurement unit (to convert to a water flux, divide by the latent heat of vaporization λ ≈ 2.45 × 10⁶ J kg⁻¹: 1 W m⁻² ≈ 0.00147 mm h⁻¹). The two model configurations are run from the same CLM-ml code and differ only in the number of within-canopy layers; in this archive both configurations use identical below-ground (root) and soil parameters , so that differences reflect the above-ground canopy representation alone. Contents File Description <SITE>_ET_hourly.csv` Per-site hourly time series (28 files; columns below). all_sites_ET_hourly.csv All 28 sites concatenated (same columns, plus `Site`). sites_metadata.csv Site list with latitude, longitude, number of hours, and date range. model_forcing_netcdf/<SITE>_forcing.nc Complete model forcing for each site (all driver variables and the gap-filled/closure-corrected flux products; see below) Sites (28) CA-Cbo, CA-Gro, CA-TP3, CA-TPD, CH-Lae, CZ-Lnz, CZ-RAJ, CZ-Stn, DE-Hai, FR-Bil, FR-Hes, IT-Cp2, IT-SR2, US-Bar, US-Me2, US-Me6, US-NC1, US-NR1, US-Oho, US-UMB, US-UMd, US-xAB, US-xBR, US-xDL, US-xHA, US-xJE, US-xTA, US-xTR. Coordinates and record lengths are in sites_metadata.csv . Columns in <SITE>_ET_hourly.csv Column Definition Units TIMESTAMP Time at the start of the hour, as provided in the model forcing (site local-standard-time convention) YYYY-MM-DD HH:MM:SS Site Site identifier - ET_obs_LE_gapfilled_Wm2 Observed latent heat flux, gap-filled by the marginal-distribution-sampling (MDS) method (FLUXNET/ONEFlux `LE_F_MDS`). W m⁻² ET_obs_LE_corrected_Wm2 Observed latent heat flux after energy-balance-closure correction (Bowen-ratio-preserving). This is the target used to evaluate the models. W m⁻² ET_MLCAN_Wm2 Modeled latent heat flux from the multilayer canopy (MLCAN) configuration. W m⁻² ET_1L_Wm2 Modeled latent heat flux from the single-layer / big-leaf (1L) configuration. W m⁻² H_obs_gapfilled_Wm2 Observed sensible heat flux, MDS gap-filled (`H_F_MDS`). W m⁻² H_obs_corrected_Wm2 Observed sensible heat flux after energy-balance-closure correction. W m⁻² SW_IN_Wm2 Incoming shortwave radiation (model driver, `FSDS`). W m⁻² TA_degC Air temperature (model driver, `TBOT`, converted from K). °C VPD_kPa Vapor pressure deficit, computed from observed relative humidity and air temperature. kPa SWC_m3m3 Volumetric soil water content (same across all soil layers). m³ m⁻³ LAI_m2m2 Effective leaf area index used to drive the models (`ELAI`). m² m⁻² Note: Missing values are written as empty fields. The fraction of finite, energy-balance-corrected observed ET per site is given in `sites_metadata.csv` (`ET_obs_corrected_pct_finite`). Model forcing NetCDF files (`model_forcing_netcdf/`) Each `<SITE>_forcing.nc` contains the complete set of driver variables used to run both configurations and the full observed-flux products, at the same temporal resolution. Key variables (units as stored): - Meteorology: `TBOT` (K), `RH` (%), `WIND` (m s⁻¹), `FSDS` (incoming shortwave, W m⁻²), `FLDS` (incoming longwave, W m⁻²), `PSRF` (surface pressure, Pa), `PRECTmms` (precipitation, mm s⁻¹), `CO2MF` (CO₂ mole fraction), `ZBOT` (reference height, m). - Vegetation / soil: `ELAI`, `ESAI` (effective leaf/stem area index), `SWC` (volumetric soil water), `H2OSOI`, `TSOI` (soil-profile moisture and temperature). - Observed fluxes: `LE_F_MDS`, `H_F_MDS` (MDS gap-filled), `LE_c`, `H_c` (energy-balance corrected), `GPP_DT`, `GPP_NT` (daytime/nighttime partitioned gross primary productivity). - Temporal coverage: each site's record covers 1 July - 31 December of each year (the study's analysis window). Methods (brief) Eddy-covariance processing . Half-hourly fluxes were computed with EddyPro and quality-controlled; latent and sensible heat were gap-filled using the MDS algorithm and then corrected for the surface-energy-balance closure gap (Bowen-ratio-preserving). Models. CLM-ml (Bonan et al., 2021; https://doi.org/10.1016/j.agrformet.2021.108435) was run in two configurations viz. multilayer (MLCAN) and single-layer (1L) that share identical code, leaf-level formulations, and soil/root parameters and differ only in the number of within-canopy layers. Each configuration was independently calibrated to the energy-balance-corrected observed ET. Provenance, license, and citation The **original** half-hourly eddy-covariance observations for each site are distributed by AmeriFlux, the ICOS Drought-2018 and Warm-Winter-2020 collections, and NEON; the per-site dataset DOIs are listed in the Supplementary Information file of the associated manuscript. The non-gap-filled (raw) observed LE can be obtained from those original datasets. The modeled ET, the processed (gap-filled and corrected) observed ET, and the assembled forcing in this archive are the data generated by this study . - License: Creative Commons Attribution 4.0 (CC-BY-4.0).
Atmospheric trace element concentrations were measured from March 1999 through December 2003 at the Air Chemistry Observatory of the German Antarctic station Neumayer by inductively coupled plasma - quadrupol mass spectrometry (ICP-QMS) and ion chromatogra-phy (IC). This continuous five year long record derived from weekly aerosol sampling re-vealed a distinct seasonal summer maximum for elements linked with mineral dust entry (Al, La, Ce, Nd) and a winter maximum for the mostly sea salt derived elements Li, Na, K, Mg, Ca, and Sr. The relative seasonal amplitude was around 1.7 and 1.4 for mineral dust (La) and sea salt aerosol (Na), respectively. On average a significant deviation regarding mean ocean water composition was apparent for Li, Mg, and Sr which could hardly be explained by mir-abilite precipitation on freshly formed sea ice. In addition we observed all over the year a not clarified high variability of element ratios Li/Na, K/Na, Mg/Na, Ca/Na, and Sr/Na. We found an intriguing co-variation of Se concentrations with biogenic sulfur aerosols (methane sul-fonate and non-sea salt sulfate), indicating a dominant marine biogenic source for this element linked with the marine biogenic sulfur source.
Description: This dataset accompanies the study "Locked in heat: present-day thermal constraints on Summer Olympic host cities and the limits of operational adaptation" (Defrance & Gadais). It provides Wet-Bulb Globe Temperature (WBGT) indicators characterising summer heat-stress constraints on Summer Olympic competition, derived from the ERA5-Land hourly reanalysis. Source data. All indicators are computed from the ERA5-Land hourly reanalysis (Muñoz-Sabater et al., 2021), 0.1° × 0.1° (~9 km) spatial resolution, distributed by the Copernicus Climate Change Service (C3S). Source variables: 2 m air temperature ( t2m ), 2 m dewpoint temperature ( d2m ), 10 m zonal and meridional wind components ( u10 , v10 ), and surface solar radiation downwards ( ssrd , de-accumulated to instantaneous hourly mean flux). Temporal coverage: July and August only, 2006-2025 (20 years), the months encompassing the modern Summer Olympic Games. WBGT computation. WBGT is reconstructed using the Stull (2011) psychrometric approximation for the natural wet-bulb temperature and the Hajizadeh et al. (2017) empirical regression for the black-globe temperature. Two formulations are provided: outdoor (sun-exposed), WBGT = 0.7 Tw + 0.2 Tg + 0.1 Ta (Yaglou & Minard 1957 weighting), with full solar load; and indoor/shaded, following the ISO 7243 shade formulation with incoming shortwave radiation set to zero. Contents. Global gridded indicators ( global_days_gt28_2006-2025.nc , NetCDF): mean number of July-August days per season with at least one hour exceeding WBGT thresholds, on the full ERA5-Land land grid. Variables: days_gt28_outdoor , days_gt28_indoor , days_gt32_outdoor , days_gt32_indoor . Thresholds: 28 °C (high heat-stress risk) and 32 °C (extreme risk), per international sport-federation guidelines. City-scale hourly series ( hourly_<City>.parquet , 8 cities): hourly time series in mean local solar time, July-August 2006-2025, with WBGT (outdoor and indoor) and its components ( Ta , RH , Tw , Tg_sun , ssrd ). City-scale derived indicators (CSV): city_summary.csv (per-city constrained-day counts at both thresholds and formulations, peak-hour WBGT decomposition, 2006-2025 linear trends); diurnal_<City>.csv (percentage of hours exceeding each threshold by local hour); yearly_<City>.csv (annual constrained-day counts). Analysis code is archived together with the data and at [GitHub URL]. Study cities (nearest valid ERA5-Land land cell): Doha (QA), Ahmedabad (IN), Tokyo (JP), Los Angeles (US), Paris (FR), Rio de Janeiro (BR), Brisbane (AU), Cape Town (ZA). Scope. Present, observed climate only; no future projections. WBGT is a first-order indicator at the regional near-surface scale and does not resolve intra-urban or venue microclimates.
WetVegDE is tri-annual wetland land cover mapping (30 m) dataset for Germany from 1985 to 2022. The maps contain detailed information of different wetland vegetations. Overview: Web map Map value Class name Color code (R,G,B) 1 Other land covers (220,220,220) 2 Grassland (143,158,0) 3 Forest (0,150,0) 4 Water (0,0,255) 5 Open swamp (111,244,230) 6 Salt marsh (82,195,220) 7 Cut-over raised bog (213,94,0) 8 Natural or restored raised bog (204,121,167) 9 Natural or restored fen (240,228,66) This dataset contains: WetVegDE_{year}.tif : Tri-annual map data (30 m) in GeoTIFF format (projection ETRS89 / EPSG:3035) legend_EN.qml : Map style to be used in QGIS (English) legend_DE.qml : Map style to be used in QGIS (German)
Chartrand, Allison · MacGregor, Joseph · Morlighem, Mathieu · et al.
12 MB
This repository contains the datasets which accompany the code for the manuscript "A vast valley network beneath the Greenland Ice Sheet". This dataset presents subglacial topography mapped beneath the Greenland Ice Sheet using high-resolution observations of the ice sheet surface and Ice Flow Perturbation Analysis (IFPA) code adapted for Greenland. This repository contains: (1) subglacial topography mapped using IFPA (raw IFPA output and blended+radar-corrected topography), error estimates, and rebounded topography as rasters, (2) a "flow-aware" hillshade model of the ice sheet as a raster, (3) all traced lineations from the blended+radar-corrected subglacial topography as shapefile polylines, (4) automated stream network derived from the rebounded topography as shapefile polylines, (5) automated stream network branching angles as shapefile points, (6) all traced edges of disrupted basal ice where the IFPA method output is disfavored in the blended+radar-corrected topography. The code for pre-processing data, IFPA, and data analysis can be found in the accompanying Github repository, available at DOI:
This dataset contains survey data examining elementary school teachers' perceptions of local wisdom-based digital learning content based on different age groups (junior and senior teachers) in Jakarta, Indonesia. The dataset includes responses to two constructs from the Technology Acceptance Model (TAM), namely Perceived Usefulness (PU) and Perceived Ease of Use (PEOU). The PU construct measures teachers' perceptions of the benefits of using local wisdom-based digital content in improving teaching performance, productivity, effectiveness, instructional quality, and efficiency. The PEOU construct measures teachers' perceptions regarding the ease of learning, operating, understanding, adapting, and using local wisdom-based digital content in various instructional contexts. The instrument was adapted from Davis's (1989) original TAM scales for electronic information systems. It consists of 12 items, including six items measuring Perceived Usefulness (PU1-PU6) and six items measuring Perceived Ease of Use (PEOU1-PEOU6). The items were contextualized to assess the use of local wisdom-based digital learning content in elementary school teaching.
This submission includes publicly available data extracted in its original form. Please reference the Related Publication listed here for source and citation information If you have questions about the underlying data stored here, please contact the EPA at https://www.epa.gov/outdoor-air-quality-data/forms/contact-us-about-outdoor-air-quality-data. If you have questions or recommendations related to this metadata entry and extracted data, please contact the CAFE Data Management team at: climatecafe@bu.edu. The AirData Air Quality Monitors app is a mapping application available on the web and on mobile devices that displays monitor locations and monitor-specific information. It also allows the querying and downloading of data daily and annual summary data. Map layers include: Monitors for all criteria pollutants (CO, Pb, NO2, Ozone, PM10, PM2.5, and SO2) PM2.5 Chemical Speciation Network monitors IMPROVE (Interagency Monitoring of PROtected Visual Environments) monitors NATTS (National Air Toxics Trends Stations) NCORE (Multipollutant Monitoring Network) PAMS (Photochemical Assessment Monitoring Stations) Near road monitors Nonattainment areas for all criteria pollutants Tribal areas Federal Class I areas (national parks and wilderness areas)” Data included are point location layers and supporting documentation used in the EPA Interactive Map of Air Quality Monitors available at https://www.epa.gov/outdoor-air-quality-data/interactive-map-air-quality-monitors and https://epa.maps.arcgis.com/apps/webappviewer/index.html?id=5f239fd3e72f424f98ef3d5def547eb5&extent=-146.2334,13.1913,-46.3896,56.5319. Point location layers were downloaded from the ESRI map services website and saved as Geopackages. CSV lists of individual air quality sites and monitors are included as well. The user interface of the interactive map is documented in PDF format. The individual ESRI map services websites for each map layer which provide some metadata information are documented in PDF format. The point location layers all include the fields “Annual Data Download (Links)” and “Daily Data Download (Links)” with html links listed. The data sourced from these links has previously been archived on the Dataverse website. [Quote from: https://www.epa.gov/outdoor-air-quality-data/interactive-map-air-quality-monitors] on 2026-03-27.
Agreement Between Large Language Models and Humans in Research Proposal Review - Data and Code This repository contains the data and code required to reproduce the analyses, statistical results, and figures presented in the associated manuscript. Files are organized by function and described below. All research proposals are anonymized and labeled with non-identifying identifiers (A, B, C, ...). Reviewer identities were never provided to the authors. To prevent inadvertent disclosure, all free-text review content from both human reviewers and large language models (LLMs) has been removed; only the numerical evaluation data required to reproduce the reported analyses are included. Data files Human_raw_scores.csv Individual numerical scores assigned by human reviewers, one row per reviewer × proposal × criterion. Used to compute panel-level summary statistics and the inter-reviewer reliability metrics reported in the manuscript. Human.csv Proposal-level human reference scores (one row per proposal) used as the human benchmark against which LLM scores and rankings are compared. LLM_data_combined_clean_filtered.csv All numerical scores generated by the evaluated LLMs. Preprocessed to remove evaluations in which a model failed to return one or more required numerical scores (see Data provenance and known limitations for counts). review_criteria.txt The evaluation criteria and rating scales. See the note under Data provenance regarding the 2023 vs. 2024 criterion naming. Data dictionary Human_raw_scores.csv Column Description Applicant Anonymized proposal identifier (A, B, C, ...). Cycle Review cycle the proposal belongs to (2023 or 2024). Label Scoring criterion: Intellectual merit , Potential for impact , Collaborative Potential , or Overall ranking . Reviewer_Seq Reviewer index within a proposal (1, 2, 3, ...). Identities are unknown; this index only links a single reviewer's ratings across criteria for the same proposal, in source-file order. It is not consistent across proposals (reviewer 1 for proposal A is not reviewer 1 for proposal B). Rating Numerical score. Criterion ratings use a 1-5 scale; Overall ranking uses a 1-3 scale (3 = fund, 2 = fund with modifications, 1 = do not fund). Human.csv Column Description Name Anonymized proposal identifier (matches Applicant above). Cycle Review cycle (2023 or 2024). IM , Impact , Collab , Overall Panel-mean scores for the four criteria. Score Weighted composite panel score, computed as 0.4·IM + 0.3·Impact + 0.3·Collab , matching the weighting applied to the LLM composite scores. LLM_data_combined_clean_filtered.csv Column Description Name Anonymized proposal identifier (matches Human.csv ). Type Input given to the model: Abstract or Full_Proposal . Model Model identifier. For models with controllable reasoning depth, the tier is appended as a suffix ( _low , _medium , _high ). The portion before the first underscore is the root model. Prompt Prompting strategy ( OneShot or CoT ). Seed Requested random seed. For models that did not support seed specification at the time of execution (the reasoning models listed in the Methods), the API ignored this value and it functions only as a replicate index; output is not reproducible from it for those models. Temp Sampling temperature (0.1, 0.5, 0.9). IM , Impact , Collab Criterion ratings (1-5 scale). Overall Overall recommendation (1-3 scale). See limitation note on out-of-scale values. Score Weighted composite, 0.4·IM + 0.3·Impact + 0.3·Collab . Category Reasoning/architecture category of the model. Year Evaluation wave in which the run was performed (proposals were re-evaluated as new model generations were released); this is not the proposal's submission cycle. Use Cycle in the human files for submission cycle. Data provenance and known limitations We document the following so that users can interpret the data accurately. Two review cycles, combined. The 28 proposals come from two internal seed grant cycles: 15 from 2023 and 13 from 2024 ( Cycle column). For 2024 proposals, proposal-level means in Human.csv are the official institute panel means; for 2023 proposals they are computed from the individual ratings in Human_raw_scores.csv . Reviewers per proposal. Panels ranged from 3 to 6 reviewers. Reviewer identities were never provided; the human inter-reviewer reliability is therefore estimated with a one-way random-effects model (ICC(1,1)), which is the appropriate model when each proposal is rated by a different, unidentified set of reviewers. Six 2024 reviews not present at the individual level. For six 2024 proposals (G, R, T, W, X, Z), one reviewer's scores were submitted without written comments and are not included in Human_raw_scores.csv . For these proposals, Human.csv carries the official institute panel means, so the panel mean in Human.csv and the mean recomputed from Human_raw_scores.csv differ slightly. The reproducibility check in 06_Human_data.ipynb confirms exact agreement for all proposals with complete individual records and reports the expected small differences for these six. Out-of-scale LLM Overall ratings. A small number of responses (≈1.2% of reviews, almost entirely from gpt-3.5-turbo) rated the overall recommendation on a 1-5 scale rather than the requested 1-3 scale, in a format the parser could not distinguish. The analysis code masks values outside the valid range before computing any Overall -based result; the composite Score does not use Overall and is unaffected. Excluded LLM evaluations. Evaluations in which a model failed to return one or more required numerical scores were removed prior to analysis. The file provided here is the post-exclusion (analyzed) dataset. Code notebooks Two groups of notebooks are provided: (i) the LLM evaluation pipeline and (ii) statistical analysis and figure generation. LLM evaluation pipeline (Notebooks 1-4) Documentation of the methodology used to generate the LLM evaluations. These use synthetic examples and contain no confidential data, API credentials, or real proposal text. 01_pipeline_overview.ipynb - architecture, configuration, criteria, output format, evaluation matrix. 02_prompting_strategies.ipynb - one-shot and chain-of-thought prompting; example selection; text vs. vision input. 03_response_parsing.ipynb - regex extraction of ratings and comments; error handling; decimal ratings. 04_example_evaluation.ipynb - end-to-end workflow on synthetic data. Statistical analysis and figures (Notebooks 5-6) 05_Data_Processing.ipynb - ANOVA and effect sizes; empirical absolute score differences; ICC(2,1) and Spearman correlations vs. the human panel; Monte Carlo comparisons; scatter, slope, and bias figures. 06_Human_data.ipynb - ICC(1,1)/ICC(1,k) for human reviewers with bootstrap confidence intervals; Monte Carlo single-reviewer vs. leave-one-out panel rank agreement; consistency check of Human.csv against Human_raw_scores.csv . Reproducibility Running 05_Data_Processing.ipynb and 06_Human_data.ipynb against the included data reproduces the statistical results and figures in the manuscript. Notebooks 1-4 document the evaluation pipeline using synthetic examples. Requirements pandas numpy scipy statsmodels scikit-learn matplotlib The LLM evaluation pipeline notebooks (1-4) additionally use the packages listed in requirements.txt . Citation If you use this data or code, please cite the associated manuscript and this archive: Gorski, C., Leo, N., Gayah V. Agreement Between Large Language Models and Humans in Research Proposal Review - Data and Scripts. 2026. Zenodo. https://doi.org/10.5281/zenodo.18187034 License Data are released under CC BY 4.0; code is released under the MIT License.
Ahmmed, Imtiaz · Masud, Md Hasan · Sarker, Monosij Kanti · et al.
4.6 MB
This repository contains the reproducibility package associated with the study: "Bridging the Gap Between Awareness and Action: An Empirical Study of Technical Debt Management in Agile Software Development." The package includes: • Survey questionnaire • Survey dataset • Supporting documentation The dataset contains responses from software engineering professionals regarding technical debt awareness, causes, impacts, and management practices in agile software development environments. The materials are provided to facilitate verification, replication, and future research based on the study findings.
General Information This repository contains the comprehensive collection of empirical lacustrine time series, sedimentary core records, standardized data matrices, and execution scripts required to fully replicate the figures, network topologies, and statistical null models presented in the associated manuscript. I. Software and Environment Requirements Python (v3.8+): Required libraries include numpy , pandas , matplotlib , scipy , statsmodels , and pymannkendall . R (v4.2+): Required libraries include wsyn , igraph , zoo , parallel , pbapply , ggplot2 , and cowplot . Geospatial Platform: Esri ArcGIS Desktop (v10.8) or ArcGIS Pro (for cartographic rendering and vector layer manipulation). Computation Note: To mitigate boundary artifacts and edge effects inherent in chronological sliding-window operations, the final 12-24 data points of the generated time series are systematically excluded from the final trend evaluation. II. File Inventory and Component Descriptions 1. Empirical and Core Datasets ( .csv ) Global_Lake_Chlorophyll_a_Time_Series.csv Description: Long-term, multi-decadal monthly gridded chlorophyll- a concentration time series across global limnological cohorts, featuring unique lake identification codes as columns and sequential temporal intervals as rows. Global_Lake_Water_Color_Time_Series.csv Description: Normalized global lake water color time series quantified via the Forel-Ule Index (FUI) framework, structured identically to the chlorophyll dataset for multi-proxy comparison. Global_Lake_Sediment_Pigments.csv Description: Stratigraphic sedimentary core pigment records providing long-term retrospective evidence of historical limnological synchronization and baseline shifts. Figures_data.CSV Description: Consolidated and curated data matrices containing the exact values, coordinates, and regional groupings utilized to plot the core text figures. 2. Statistical Analysis and Mathematical Scripts ( .py & .R ) Pairwise correlation -based synchrony.py Description: Computes macro-scale spatial synchrony trends across lacustrine nodes using the standard pairwise correlation matrix stream following frequency-domain decomposition. Loreau φ Metric for lake synchrony.py Description: Execution script utilizing the classic Loreau-de Mazancourt φ metric to calculate multi-lake population-level synchrony across global and latitudinal cohorts. Sliding window sensitivity.py Description: Explores scale dependency and temporal robustness by executing the analytical data stream across varying sliding temporal windows (e.g., 2, 5, 8, 10, 15 time steps). Network and modularity analysis.R Description: Implements the wsyn continuous signed-power soft-thresholding paradigm and leverages igraph to partition similarity networks into topological communities, calculating decadal modularity . Permutation-based significance of synchrony trends.py Description: Performs Mann-Kendall trend tests on sliding-window synchrony series and runs empirical hypothesis testing to extract true directional shifts. Permutation-based significance of synchrony trends-Null_Model_Generator.py Description: Harnesses multi-core parallel processing to shuffle network weights 1,000 times, constructing empirical null distributions to validate the significance of observed network community dissolution. 3. Geospatial Visualization & Documentation ( .rar & .txt ) Figure 1.rar Description: Compressed archive containing all raw geospatial project databases, vector layer shapefiles ( .shp ), metadata tables, and cartographic layout definitions ( .mxd ) used to generate the global geographic distribution map of sample lakes (Figure 1). Compiled within Esri ArcGIS 10.8. Data_Sources_and_References.txt Description: A comprehensive standalone text file documenting the complete bibliographic literature sources, historical baselines, and corresponding DOI attributions compiled within the empirical datasets. III. Execution and Replication Workflow Data Cleaning & Detrending: Feed raw time series through Python/R scripts to execute STL harmonic regression models and filter low-frequency background signals. Synchrony Calculations: Execute the pairwise and Loreau metric scripts to plot continuous synchrony variations over time. Network Configurations: Run the R network script to output the high-resolution PCA community plots and decadal modularity comparisons. Significance Evaluation: Launch the null model generator to confirm that network configuration shifts significantly exceed random stochastic expectations ( P < 0.001 ). Spatial Reconstruction: Extract Figure 1.rar into your local GIS directory to access, modify, or re-export the multi-layered baseline global sampling maps.
Schlosser, Elisabeth · Reijmer, Carleen H · Oerter, Hans · et al.
342 rows × 5 cols · 15 KB
4 numeric · 1 datetime
The relationship between d18O and air temperature at Neumayer station, Ekströmisen, Antarctica, was investigated using fresh-snow samples from the time period 1981-2000. A trajectory model that calculated 5 day-backward trajectories was used to study the influence of different synoptic weather situations and thus of different moisture sources on this correlation. Generally a high correlation between air temperature and d18O was found, but the quality of the d18O-T relationship varied with the different trajectory classes. Additionally, the sea-ice coverage on the travel path of the moist air was considered. The amount of open ocean water underneath the trajectory has a large influence on the d18O-T relationship. For trajectories that lead completely above open water, no significant correlation between d18O and T was found, because mixing with air masses containing additionally evaporated water vapour from the ocean influences the isotope ratio of precipitation. A very high correlation, however, was found for transports over the completely ice-covered Weddell Sea.
Zolduoarrati, Elijah · Licorish, Sherlock · Stanger, Nigel
3.7 MB
Existing quantitative studies examining how diversity affects Stack Overflow contributions have only captured statistical trends, yet fail to explain how their findings resonate to the actual users' experience. Qualitative investigation is essential to validate these findings. Our study synthesises existing research to yield 42 key outcomes, informing the development of 14 open-ended questions for Stack Overflow users related to participation, contribution utility and value, and quality. 209 responses were accrued that reflected how US-based users interact with the platform, where inductive thematic analysis was conducted to identify emergent themes. This replication package is provided for those interested in further examining our research methodology.
Measurements of atmospheric radioactivity attached to aerosols are described. Fallout was collected in a vessel of large area. Emphasis was on separation of "wet" and "dry" samples. For strontium 90 a ratio of "wet" to "dry" fallout of 5:1 has been found independent of latitude. The total fallout was smaller than comparable values from continents because of very small amounts of rainfall in the equatorial zone. In order to achieve consistency in the global balance a better knowledge not only of radioactivity but also of precipitation over the ocean is required. Fallout of Ra-D clearly shows the ITC as a barrier for the latitudinal movement of near sea-surface air masses. The concentration of short-lived emanation daughters shows large variations according to varying geographic conditions. A variation with time could not be explained. The specific activity of long-lived radioactive substances shows the expected effect of the ITC as well as a seasonal diminuation of average concentration, similar to that measured at Heidelberg.
The present invention is an intelligent gateway which can receive multiple sensor data using sub-1G Hz frequency, analyze data, and transmit processed data to a database server. The intelligent gateway can receive data from up to 100 sensors using sub-1G Hz (433, 868 or 915 MHz) wireless frequency. The received data can be analyzed and the gateway can determine when to transmit data, and which packaged data to transmit to the database server. The intelligent gateway can also receive feedback and instructions from the database server. The process data can be transmitted to the database server with different protocols like WIFI, Ethernet and RS485. The intelligent gateway can also include multiple sensors including temperature and humidity sensors, pressure sensors, air speed sensors and a particulate matter sensor for detecting particulates of less than 2.5 micro meters (PM2.5). These sensors are collect additional indoor environmental quality parameters.
The present invention provides a ventilation system for improving air quality of an indoor space. The system includes sensors for measuring PM2.5 particle level P, CO 2 level C, and TVOC level T in the indoor space. A control circuit is configured to receive P, C and T values and generate an output signal Vout according to a specific algorithm, which in turn controls speed-variable EC motors that drive ventilating fans. The invention exhibits numerous technical merits such as lower energy consumption, programmable operation, high efficiency, and lower noise, among others.
declared
Scholarly works are unavailable this run - the index did not respond. The other tabs are unaffected.