Lemor, Antoine · Pillod, Alizée · Taylor, Matthew · et al.
266,271 rows × 10 cols
7 numeric · 3 text
The Canadian Climate Framing (CCF) Database is a comprehensive, machine-learning-annotated corpus of climate-change media coverage in Canada. It comprises 266,271 articles from 20 major Canadian newspapers (1978-2024) processed into 9,198,958 two-sentence analytical units (82.9% English, 17.1% French). Each unit is annotated across 65 hierarchical categories by 128 BERT and CamemBERT classifiers, with a macro F1 of 0.866 on a 1,000-sentence gold standard double-coded by an independent annotator (Gwet's AC1 = 0.894, Krippendorff's α = 0.698, Cohen's κ = 0.596 on the 400 blind sentences). Each category receives an A/B/C reliability tier summarising annotation quality from classifier performance and inter-coder agreement. The deposit ships six relational tables (bibliographic metadata, sentence-level annotations, named-entity rollups, article-level aggregates, per-category reliability tiers, and 9,462,845 BAAI/bge-m3 sentence-and-title embeddings). Raw newspaper text is excluded for copyright reasons; bibliographic coordinates (media, date, title, author, page_number) are sufficient for any researcher with institutional access to Factiva, Eureka.cc or ProQuest Canadian Major Dailies to recover the original sentences. This deposit accompanies a methodology paper currently under revision at Scientific Data (Nature Portfolio). This deposit is the Apache Parquet mirror of the canonical PostgreSQL edition (cross-referenced in Related identifiers ). Each of the six relational tables is provided as a standalone .parquet file with ZSTD compression; the 1024-dimensional BAAI/bge-m3 embedding column is materialised as a list<float> , and JSONB entity arrays are serialised as UTF-8 JSON strings. The schemas are otherwise identical to the PostgreSQL edition. The Parquet bundle is readable natively by pandas , polars , R/ arrow , DuckDB , and Spark without any database backend: import pandas as pd agg = pd.read_parquet('CCF_article_aggregates.parquet') emb = pd.read_parquet('CCF_sentence_embeddings.parquet') The HNSW index that ships with the PostgreSQL edition is not transferable to Parquet; brute-force cosine similarity remains tractable on the embedding column (≈ 9.46 M × 1024 float16). The full annotation pipeline, training data, manual-annotation JSONL, intercoder-reliability benchmark, methodology manuscript (LaTeX sources + PDF), and reproducibility scripts are bundled with this deposit as ccf_code_and_paper.tar.gz . The same materials are also available on the project's OSF companion deposit ( 10.17605/OSF.IO/Q5W47 ) and on the development mirror at GitHub .
Aimed at year-round recording of the chemical aerosol composition in central Antarctica, an unattended operating aerosol sampler was successfully deployed at the EPICA deep drilling site in Dronning Maud Land (Kohnen Station). Analyses of teflon/nylon filter packs consecutively collected over bi-weekly intervals during the February 2003 to December 2005 period allowed to evaluate seasonal concentration variations of methane sulphonate (MS), Cl-, NO3-, non-sea salt (nss-)SO4**2- and Na+, while NH4+ and mineral dust related ion results remained below detection limits. For MS and nss-SO4**2 distinct late summer maxima around 44 and 200 ng/m**3, respectively, were found, while (total) NO3- showed a broad November maximum of about 52 ng m**-3. In contrast, the highest concentrations of Na+ with peak values of up to 160 ng/m**3 were observed during the winter half year. The seasonality of these species broadly coincided with long-term observations at the coastal Neumayer Station, including surprisingly comparable NO3- levels. However, the biogenic sulphur and sea salt concentrations were lower at Kohnen by typically a factor of 2-3 and 10, respectively. The arrival of sea ice derived sea salt particles at Kohnen could not clearly detected, since even during mid-winter the nss-SO4**2- to Na+ ratio was generally too high to unambiguously identify a sulphur depleted sea salt SO4**2- fraction.
Aerosol samples collected over the North Atlantic from ship were analysed for Sodium, Magnesium, Potassium, Calcium and Chloride. A found dependence of sea salt concentrations from wind velocity is compared with earlier results. The mean of the ratio Cl/Na was close to that for sea water; the Mg-, K- and Ca-concentrations in the aerosol, however, were enriched with respect to sea water. It is shown that continental advection influences the measured aerosol components over the North Atlantic.
Lemor, Antoine · Pillod, Alizée · Taylor, Matthew · et al.
2.2 MB
The Canadian Climate Framing (CCF) Database is a comprehensive, machine-learning-annotated corpus of climate-change media coverage in Canada. It comprises 266,271 articles from 20 major Canadian newspapers (1978-2024) processed into 9,198,958 two-sentence analytical units (82.9% English, 17.1% French). Each unit is annotated across 65 hierarchical categories by 128 BERT and CamemBERT classifiers, with a macro F1 of 0.866 on a 1,000-sentence gold standard double-coded by an independent annotator (Gwet's AC1 = 0.894, Krippendorff's α = 0.698, Cohen's κ = 0.596 on the 400 blind sentences). Each category receives an A/B/C reliability tier summarising annotation quality from classifier performance and inter-coder agreement. The deposit ships six relational tables (bibliographic metadata, sentence-level annotations, named-entity rollups, article-level aggregates, per-category reliability tiers, and 9,462,845 BAAI/bge-m3 sentence-and-title embeddings). Raw newspaper text is excluded for copyright reasons; bibliographic coordinates (media, date, title, author, page_number) are sufficient for any researcher with institutional access to Factiva, Eureka.cc or ProQuest Canadian Major Dailies to recover the original sentences. This deposit accompanies a methodology paper currently under revision at Scientific Data (Nature Portfolio). This deposit is the canonical PostgreSQL edition. It contains a pg_dump -Fd directory archive (compressed into a single .tar file) of the six relational tables, including the pgvector extension and HNSW cosine indexes for sub-second semantic-similarity search. Restoration is a one-liner: tar -xf CCF_Database.tar && createdb CCF_Database && psql -d CCF_Database -c 'CREATE EXTENSION IF NOT EXISTS vector;' && pg_restore -d CCF_Database --no-owner --no-privileges -j 8 CCF_Database_dump A column-oriented Apache Parquet mirror of the same six tables is available as the sister deposit on Zenodo (cross-referenced in Related identifiers ). The Parquet mirror is recommended for users without PostgreSQL access (it is directly readable by pandas, polars, R/arrow, DuckDB, and Spark). The full annotation pipeline, training data, manual-annotation JSONL, intercoder-reliability benchmark, methodology manuscript (LaTeX sources + PDF), and reproducibility scripts are bundled with this deposit as ccf_code_and_paper.tar.gz . The same materials are also available on the project's OSF companion deposit ( 10.17605/OSF.IO/Q5W47 ) and on the development mirror at GitHub . Requirements: PostgreSQL 16 or 17 with pgvector ≥ 0.8.2 (for halfvec(1024) storage of the sentence embeddings).
During the 1965 Atlantic Expedition of the "Meteor“ concentrations of various atmospheric trace gases were measured. The following gases were considered: carbon dioxide (CO2), sulfur dioxide (SO2), nitrogene dioxide (NO2), and nitric oxide (NO). The air whereof these components were measured was sucked in from a height of 14 m above the surface of the sea. The results allow conclusions upon the long term global increase of the atmospheric CO2 content, the meridional distribution of the CO2 on the Atlantic Ocean, and the dependance of its concentration upon the time of the day and the thermal structure of the atmosphere. Attempts at determining concentrations of sulfur dioxide and nitric oxide of non-continental origin failed at large. Concentrations of NO2, however, could succesfully be measured.
Piel, Claudia · Weller, Rolf · Huke, Michael · et al.
98 rows × 14 cols · 13 KB
8 text · 5 numeric · 1 datetime
During three summer campaigns in January/February 2000, 2001, and 2002 the ionic composition of the aerosol at the European Project for Ice Coring in Antarctica (EPICA) deep-drilling site at Kohnen Station was measured in daily resolution. In 2000 and 2002 we observed mean (±std) non-sea-salt sulfate (nss-[SO4]2-) concentrations of 353 ± 100 ng/m**3 and 320 ± 250 ng/m**3, as well as methane sulfonate (MS) concentrations of 59 ± 36 ng/m**3 and 74 ± 80 ng/m**3, respectively. For the summer campaign in 2001, significantly lower nss-[SO4]2- and MS levels of 164 ± 150 ng/m**3 and 19 ± 12 ng/m**3, respectively, were typical. The mean MS/nss-[SO4]2- ratio ranged from about 0.1 to 0.2. MS and nss-[SO4]2- concentrations and their variability were roughly comparable to coastal stations at summer. Supported by air mass back trajectory analyses, this finding documented an efficient long-range transport to Kohnen via the free troposphere. MS/nss-[SO4]2- ratios exhibited a strong dependence on the MS concentration with systematically higher ratios at higher MS concentrations, a peculiarity which is also evident in a firn core drilled at this site.
This dataset quantifies the uncertainty in mapping late-successional and old-growth (LSOG) forest across the approximately 4.2 million hectares of Maine's unorganized townships, and tests whether LSOG is rapidly disappearing. Three to four independent, credible mapping methods are compared on a common 100 m grid: (M1) a reproduction of the Hagan et al. (2026) airborne-LiDAR canopy random forest, rebuilt from their public Zenodo deposit; (M2) a logistic model of the FIA field-structure LSOG class on Potapov (GEDI-calibrated) canopy height; (M3) a direct canopy-height threshold; and (M4) the FIA structural class imputed to every pixel via USFS TreeMap (2016, 2020, 2022). Version 1.2.0 additions. This version adds the materials behind the formal Ecosphere Comment on Hagan et al. (2026): (a) a cross-validated accuracy assessment (AUC) of each mapping approach on the original authors' own training plots, showing that high training accuracy does not transfer to agreement among independent maps; (b) an FIA design-based estimate of older forest with sampling-error confidence intervals, the unbiased ground reference the original analysis lacked, putting older forest at about 3.9 percent (3.3 to 4.6) and rising, including on private commercial timberland; (c) a threshold-sensitivity sweep and a 20-seed reproduction ensemble; (d) an ownership-resolved breakdown (private commercial versus public); (e) a hex-scale (8 km) summary of cross-method disagreement; and (f) the Comment manuscript and Supporting Information. Headline findings. Credible methods disagree by roughly 2.8 times on how much LSOG exists and on the location of most LSOG hectares, while agreeing closely on the rare, well-defined old-growth core. Protecting the top 5 to 20 percent of hectares by one map versus another overlaps on only 16 to 30 percent of the ground, so single-map patch-level prioritization for large expenditures is fragile. The design-based FIA estimate and TreeMap imputation both show older forest stable to increasing rather than rapidly declining; the apparent loss reported elsewhere is a gross harvest flux, not a net stock decline. Contents. Derived 100 m GeoTIFFs (reproduced Hagan class, v5.1-GEDI probability, TreeMap class, a per-cell method-consensus layer), summary tables (area by method, pairwise agreement, concordance, prioritization fragility, AUC by approach, design-based older-forest trend with CIs, ownership breakdown, and FIA validation), the analysis R scripts, quick-look figures, the Ecosphere Comment manuscript and Supporting Information, and a full methods-and-findings report (PDF). Privacy. No FIA plot coordinates are included; all products are derived rasters or aggregate summary tables. Caveats: the robust temporal signal is direction rather than precise rate; FIA stand age is modeled, so a structural large-tree domain is reported alongside the age domain; cross-validated intervals are best read as lower bounds because plots are spatially dispersed but not independent. See the README and report for full methods, provenance, and limitations. Version 1.12.0 additions. The cross-map comparison is refined to independent remote-sensing operationalizations only. (a) A three-map remote-sensing ensemble over Maine on a common 100 m grid: reproduced Hagan airborne-LiDAR (any-LSOG 21.9 percent), an FIA-structure class on Potapov GEDI-calibrated spaceborne canopy height (14.0 percent), and the ORNL/Bruening national old-growth stratum (36.1 percent); the three span a 2.6-fold range and agree on only 2.7 percent of flagged hectares, with the ORNL stratum spatially uncorrelated with the structure maps. The USFS TreeMap imputation is reclassified as a second FIA-anchored accounting, reported with the design-based estimate rather than as an independent map. (b) A design-based estimate of LSOG itself: integrated any-LSOG 14.1 percent (12.9 to 15.3) and strict four-axis true LSOG 3.1 percent (2.5 to 3.7) of Maine forestland, the airborne map exceeding even the inclusive ground estimate. (c) A balanced-model LSOG probability surface at 100 m, with a binary class calibrated to the design-based area to bound over-prediction. (d) A multi-objective support vector regression pilot tracing the Pareto front of total versus systematic (attenuation) error. Derived rasters, the three-map agreement layer, the ORNL stratum reprojected to the study grid, tables, R scripts, the updated Comment, and the companion manuscript are included. No FIA plot coordinates are included. Version 1.13.0. Final consolidated release. Adds: a rare-class remedy menu for the reproduced random forest (default vs class weighting vs balanced sub-sampling vs voting-threshold; old-growth detection 0.24 to 0.82, mapped old-growth area 1.0 to 2.3 percent); the definitive five-model LSOG probability map for Maine with across-model uncertainty and reference reserves (MNAP/TNC network, Baxter, Big Reed) over real state and county boundaries; design-based 95 percent confidence intervals for forest-type and ecoregion representation and for disturbance shares; a full robustness/stress-test matrix; and the copy-edited, sole-authored Comment, companion manuscript, and Maine Forest Products Council technical report. Authorship updated to Aaron R. Weiskittel.
Data archive supporting the manuscript " Multilayer canopy model outperforms big-leaf model for evapotranspiration predictions under high water and heat stress conditions " (Authors: Raghav, Liu, Kumar, Bisht) This archive provides, for each of 28 eddy-covariance sites, the hourly time series of modeled and observed evapotranspiration (ET) together with the meteorological drivers used to force the models. ET is reported as latent heat flux (LE), in W m⁻² , the native model and tower-measurement unit (to convert to a water flux, divide by the latent heat of vaporization λ ≈ 2.45 × 10⁶ J kg⁻¹: 1 W m⁻² ≈ 0.00147 mm h⁻¹). The two model configurations are run from the same CLM-ml code and differ only in the number of within-canopy layers; in this archive both configurations use identical below-ground (root) and soil parameters , so that differences reflect the above-ground canopy representation alone. Contents File Description <SITE>_ET_hourly.csv` Per-site hourly time series (28 files; columns below). all_sites_ET_hourly.csv All 28 sites concatenated (same columns, plus `Site`). sites_metadata.csv Site list with latitude, longitude, number of hours, and date range. model_forcing_netcdf/<SITE>_forcing.nc Complete model forcing for each site (all driver variables and the gap-filled/closure-corrected flux products; see below) Sites (28) CA-Cbo, CA-Gro, CA-TP3, CA-TPD, CH-Lae, CZ-Lnz, CZ-RAJ, CZ-Stn, DE-Hai, FR-Bil, FR-Hes, IT-Cp2, IT-SR2, US-Bar, US-Me2, US-Me6, US-NC1, US-NR1, US-Oho, US-UMB, US-UMd, US-xAB, US-xBR, US-xDL, US-xHA, US-xJE, US-xTA, US-xTR. Coordinates and record lengths are in sites_metadata.csv . Columns in <SITE>_ET_hourly.csv Column Definition Units TIMESTAMP Time at the start of the hour, as provided in the model forcing (site local-standard-time convention) YYYY-MM-DD HH:MM:SS Site Site identifier - ET_obs_LE_gapfilled_Wm2 Observed latent heat flux, gap-filled by the marginal-distribution-sampling (MDS) method (FLUXNET/ONEFlux `LE_F_MDS`). W m⁻² ET_obs_LE_corrected_Wm2 Observed latent heat flux after energy-balance-closure correction (Bowen-ratio-preserving). This is the target used to evaluate the models. W m⁻² ET_MLCAN_Wm2 Modeled latent heat flux from the multilayer canopy (MLCAN) configuration. W m⁻² ET_1L_Wm2 Modeled latent heat flux from the single-layer / big-leaf (1L) configuration. W m⁻² H_obs_gapfilled_Wm2 Observed sensible heat flux, MDS gap-filled (`H_F_MDS`). W m⁻² H_obs_corrected_Wm2 Observed sensible heat flux after energy-balance-closure correction. W m⁻² SW_IN_Wm2 Incoming shortwave radiation (model driver, `FSDS`). W m⁻² TA_degC Air temperature (model driver, `TBOT`, converted from K). °C VPD_kPa Vapor pressure deficit, computed from observed relative humidity and air temperature. kPa SWC_m3m3 Volumetric soil water content (same across all soil layers). m³ m⁻³ LAI_m2m2 Effective leaf area index used to drive the models (`ELAI`). m² m⁻² Note: Missing values are written as empty fields. The fraction of finite, energy-balance-corrected observed ET per site is given in `sites_metadata.csv` (`ET_obs_corrected_pct_finite`). Model forcing NetCDF files (`model_forcing_netcdf/`) Each `<SITE>_forcing.nc` contains the complete set of driver variables used to run both configurations and the full observed-flux products, at the same temporal resolution. Key variables (units as stored): - Meteorology: `TBOT` (K), `RH` (%), `WIND` (m s⁻¹), `FSDS` (incoming shortwave, W m⁻²), `FLDS` (incoming longwave, W m⁻²), `PSRF` (surface pressure, Pa), `PRECTmms` (precipitation, mm s⁻¹), `CO2MF` (CO₂ mole fraction), `ZBOT` (reference height, m). - Vegetation / soil: `ELAI`, `ESAI` (effective leaf/stem area index), `SWC` (volumetric soil water), `H2OSOI`, `TSOI` (soil-profile moisture and temperature). - Observed fluxes: `LE_F_MDS`, `H_F_MDS` (MDS gap-filled), `LE_c`, `H_c` (energy-balance corrected), `GPP_DT`, `GPP_NT` (daytime/nighttime partitioned gross primary productivity). - Temporal coverage: each site's record covers 1 July - 31 December of each year (the study's analysis window). Methods (brief) Eddy-covariance processing . Half-hourly fluxes were computed with EddyPro and quality-controlled; latent and sensible heat were gap-filled using the MDS algorithm and then corrected for the surface-energy-balance closure gap (Bowen-ratio-preserving). Models. CLM-ml (Bonan et al., 2021; https://doi.org/10.1016/j.agrformet.2021.108435) was run in two configurations viz. multilayer (MLCAN) and single-layer (1L) that share identical code, leaf-level formulations, and soil/root parameters and differ only in the number of within-canopy layers. Each configuration was independently calibrated to the energy-balance-corrected observed ET. Provenance, license, and citation The **original** half-hourly eddy-covariance observations for each site are distributed by AmeriFlux, the ICOS Drought-2018 and Warm-Winter-2020 collections, and NEON; the per-site dataset DOIs are listed in the Supplementary Information file of the associated manuscript. The non-gap-filled (raw) observed LE can be obtained from those original datasets. The modeled ET, the processed (gap-filled and corrected) observed ET, and the assembled forcing in this archive are the data generated by this study . - License: Creative Commons Attribution 4.0 (CC-BY-4.0).
Atmospheric trace element concentrations were measured from March 1999 through December 2003 at the Air Chemistry Observatory of the German Antarctic station Neumayer by inductively coupled plasma - quadrupol mass spectrometry (ICP-QMS) and ion chromatogra-phy (IC). This continuous five year long record derived from weekly aerosol sampling re-vealed a distinct seasonal summer maximum for elements linked with mineral dust entry (Al, La, Ce, Nd) and a winter maximum for the mostly sea salt derived elements Li, Na, K, Mg, Ca, and Sr. The relative seasonal amplitude was around 1.7 and 1.4 for mineral dust (La) and sea salt aerosol (Na), respectively. On average a significant deviation regarding mean ocean water composition was apparent for Li, Mg, and Sr which could hardly be explained by mir-abilite precipitation on freshly formed sea ice. In addition we observed all over the year a not clarified high variability of element ratios Li/Na, K/Na, Mg/Na, Ca/Na, and Sr/Na. We found an intriguing co-variation of Se concentrations with biogenic sulfur aerosols (methane sul-fonate and non-sea salt sulfate), indicating a dominant marine biogenic source for this element linked with the marine biogenic sulfur source.
Description: This dataset accompanies the study "Locked in heat: present-day thermal constraints on Summer Olympic host cities and the limits of operational adaptation" (Defrance & Gadais). It provides Wet-Bulb Globe Temperature (WBGT) indicators characterising summer heat-stress constraints on Summer Olympic competition, derived from the ERA5-Land hourly reanalysis. Source data. All indicators are computed from the ERA5-Land hourly reanalysis (Muñoz-Sabater et al., 2021), 0.1° × 0.1° (~9 km) spatial resolution, distributed by the Copernicus Climate Change Service (C3S). Source variables: 2 m air temperature ( t2m ), 2 m dewpoint temperature ( d2m ), 10 m zonal and meridional wind components ( u10 , v10 ), and surface solar radiation downwards ( ssrd , de-accumulated to instantaneous hourly mean flux). Temporal coverage: July and August only, 2006-2025 (20 years), the months encompassing the modern Summer Olympic Games. WBGT computation. WBGT is reconstructed using the Stull (2011) psychrometric approximation for the natural wet-bulb temperature and the Hajizadeh et al. (2017) empirical regression for the black-globe temperature. Two formulations are provided: outdoor (sun-exposed), WBGT = 0.7 Tw + 0.2 Tg + 0.1 Ta (Yaglou & Minard 1957 weighting), with full solar load; and indoor/shaded, following the ISO 7243 shade formulation with incoming shortwave radiation set to zero. Contents. Global gridded indicators ( global_days_gt28_2006-2025.nc , NetCDF): mean number of July-August days per season with at least one hour exceeding WBGT thresholds, on the full ERA5-Land land grid. Variables: days_gt28_outdoor , days_gt28_indoor , days_gt32_outdoor , days_gt32_indoor . Thresholds: 28 °C (high heat-stress risk) and 32 °C (extreme risk), per international sport-federation guidelines. City-scale hourly series ( hourly_<City>.parquet , 8 cities): hourly time series in mean local solar time, July-August 2006-2025, with WBGT (outdoor and indoor) and its components ( Ta , RH , Tw , Tg_sun , ssrd ). City-scale derived indicators (CSV): city_summary.csv (per-city constrained-day counts at both thresholds and formulations, peak-hour WBGT decomposition, 2006-2025 linear trends); diurnal_<City>.csv (percentage of hours exceeding each threshold by local hour); yearly_<City>.csv (annual constrained-day counts). Analysis code is archived together with the data and at [GitHub URL]. Study cities (nearest valid ERA5-Land land cell): Doha (QA), Ahmedabad (IN), Tokyo (JP), Los Angeles (US), Paris (FR), Rio de Janeiro (BR), Brisbane (AU), Cape Town (ZA). Scope. Present, observed climate only; no future projections. WBGT is a first-order indicator at the regional near-surface scale and does not resolve intra-urban or venue microclimates.
WetVegDE is tri-annual wetland land cover mapping (30 m) dataset for Germany from 1985 to 2022. The maps contain detailed information of different wetland vegetations. Overview: Web map Map value Class name Color code (R,G,B) 1 Other land covers (220,220,220) 2 Grassland (143,158,0) 3 Forest (0,150,0) 4 Water (0,0,255) 5 Open swamp (111,244,230) 6 Salt marsh (82,195,220) 7 Cut-over raised bog (213,94,0) 8 Natural or restored raised bog (204,121,167) 9 Natural or restored fen (240,228,66) This dataset contains: WetVegDE_{year}.tif : Tri-annual map data (30 m) in GeoTIFF format (projection ETRS89 / EPSG:3035) legend_EN.qml : Map style to be used in QGIS (English) legend_DE.qml : Map style to be used in QGIS (German)
Chartrand, Allison · MacGregor, Joseph · Morlighem, Mathieu · et al.
12 MB
This repository contains the datasets which accompany the code for the manuscript "A vast valley network beneath the Greenland Ice Sheet". This dataset presents subglacial topography mapped beneath the Greenland Ice Sheet using high-resolution observations of the ice sheet surface and Ice Flow Perturbation Analysis (IFPA) code adapted for Greenland. This repository contains: (1) subglacial topography mapped using IFPA (raw IFPA output and blended+radar-corrected topography), error estimates, and rebounded topography as rasters, (2) a "flow-aware" hillshade model of the ice sheet as a raster, (3) all traced lineations from the blended+radar-corrected subglacial topography as shapefile polylines, (4) automated stream network derived from the rebounded topography as shapefile polylines, (5) automated stream network branching angles as shapefile points, (6) all traced edges of disrupted basal ice where the IFPA method output is disfavored in the blended+radar-corrected topography. The code for pre-processing data, IFPA, and data analysis can be found in the accompanying Github repository, available at DOI:
This dataset contains survey data examining elementary school teachers' perceptions of local wisdom-based digital learning content based on different age groups (junior and senior teachers) in Jakarta, Indonesia. The dataset includes responses to two constructs from the Technology Acceptance Model (TAM), namely Perceived Usefulness (PU) and Perceived Ease of Use (PEOU). The PU construct measures teachers' perceptions of the benefits of using local wisdom-based digital content in improving teaching performance, productivity, effectiveness, instructional quality, and efficiency. The PEOU construct measures teachers' perceptions regarding the ease of learning, operating, understanding, adapting, and using local wisdom-based digital content in various instructional contexts. The instrument was adapted from Davis's (1989) original TAM scales for electronic information systems. It consists of 12 items, including six items measuring Perceived Usefulness (PU1-PU6) and six items measuring Perceived Ease of Use (PEOU1-PEOU6). The items were contextualized to assess the use of local wisdom-based digital learning content in elementary school teaching.
This submission includes publicly available data extracted in its original form. Please reference the Related Publication listed here for source and citation information If you have questions about the underlying data stored here, please contact the EPA at https://www.epa.gov/outdoor-air-quality-data/forms/contact-us-about-outdoor-air-quality-data. If you have questions or recommendations related to this metadata entry and extracted data, please contact the CAFE Data Management team at: climatecafe@bu.edu. The AirData Air Quality Monitors app is a mapping application available on the web and on mobile devices that displays monitor locations and monitor-specific information. It also allows the querying and downloading of data daily and annual summary data. Map layers include: Monitors for all criteria pollutants (CO, Pb, NO2, Ozone, PM10, PM2.5, and SO2) PM2.5 Chemical Speciation Network monitors IMPROVE (Interagency Monitoring of PROtected Visual Environments) monitors NATTS (National Air Toxics Trends Stations) NCORE (Multipollutant Monitoring Network) PAMS (Photochemical Assessment Monitoring Stations) Near road monitors Nonattainment areas for all criteria pollutants Tribal areas Federal Class I areas (national parks and wilderness areas)” Data included are point location layers and supporting documentation used in the EPA Interactive Map of Air Quality Monitors available at https://www.epa.gov/outdoor-air-quality-data/interactive-map-air-quality-monitors and https://epa.maps.arcgis.com/apps/webappviewer/index.html?id=5f239fd3e72f424f98ef3d5def547eb5&extent=-146.2334,13.1913,-46.3896,56.5319. Point location layers were downloaded from the ESRI map services website and saved as Geopackages. CSV lists of individual air quality sites and monitors are included as well. The user interface of the interactive map is documented in PDF format. The individual ESRI map services websites for each map layer which provide some metadata information are documented in PDF format. The point location layers all include the fields “Annual Data Download (Links)” and “Daily Data Download (Links)” with html links listed. The data sourced from these links has previously been archived on the Dataverse website. [Quote from: https://www.epa.gov/outdoor-air-quality-data/interactive-map-air-quality-monitors] on 2026-03-27.
Agreement Between Large Language Models and Humans in Research Proposal Review - Data and Code This repository contains the data and code required to reproduce the analyses, statistical results, and figures presented in the associated manuscript. Files are organized by function and described below. All research proposals are anonymized and labeled with non-identifying identifiers (A, B, C, ...). Reviewer identities were never provided to the authors. To prevent inadvertent disclosure, all free-text review content from both human reviewers and large language models (LLMs) has been removed; only the numerical evaluation data required to reproduce the reported analyses are included. Data files Human_raw_scores.csv Individual numerical scores assigned by human reviewers, one row per reviewer × proposal × criterion. Used to compute panel-level summary statistics and the inter-reviewer reliability metrics reported in the manuscript. Human.csv Proposal-level human reference scores (one row per proposal) used as the human benchmark against which LLM scores and rankings are compared. LLM_data_combined_clean_filtered.csv All numerical scores generated by the evaluated LLMs. Preprocessed to remove evaluations in which a model failed to return one or more required numerical scores (see Data provenance and known limitations for counts). review_criteria.txt The evaluation criteria and rating scales. See the note under Data provenance regarding the 2023 vs. 2024 criterion naming. Data dictionary Human_raw_scores.csv Column Description Applicant Anonymized proposal identifier (A, B, C, ...). Cycle Review cycle the proposal belongs to (2023 or 2024). Label Scoring criterion: Intellectual merit , Potential for impact , Collaborative Potential , or Overall ranking . Reviewer_Seq Reviewer index within a proposal (1, 2, 3, ...). Identities are unknown; this index only links a single reviewer's ratings across criteria for the same proposal, in source-file order. It is not consistent across proposals (reviewer 1 for proposal A is not reviewer 1 for proposal B). Rating Numerical score. Criterion ratings use a 1-5 scale; Overall ranking uses a 1-3 scale (3 = fund, 2 = fund with modifications, 1 = do not fund). Human.csv Column Description Name Anonymized proposal identifier (matches Applicant above). Cycle Review cycle (2023 or 2024). IM , Impact , Collab , Overall Panel-mean scores for the four criteria. Score Weighted composite panel score, computed as 0.4·IM + 0.3·Impact + 0.3·Collab , matching the weighting applied to the LLM composite scores. LLM_data_combined_clean_filtered.csv Column Description Name Anonymized proposal identifier (matches Human.csv ). Type Input given to the model: Abstract or Full_Proposal . Model Model identifier. For models with controllable reasoning depth, the tier is appended as a suffix ( _low , _medium , _high ). The portion before the first underscore is the root model. Prompt Prompting strategy ( OneShot or CoT ). Seed Requested random seed. For models that did not support seed specification at the time of execution (the reasoning models listed in the Methods), the API ignored this value and it functions only as a replicate index; output is not reproducible from it for those models. Temp Sampling temperature (0.1, 0.5, 0.9). IM , Impact , Collab Criterion ratings (1-5 scale). Overall Overall recommendation (1-3 scale). See limitation note on out-of-scale values. Score Weighted composite, 0.4·IM + 0.3·Impact + 0.3·Collab . Category Reasoning/architecture category of the model. Year Evaluation wave in which the run was performed (proposals were re-evaluated as new model generations were released); this is not the proposal's submission cycle. Use Cycle in the human files for submission cycle. Data provenance and known limitations We document the following so that users can interpret the data accurately. Two review cycles, combined. The 28 proposals come from two internal seed grant cycles: 15 from 2023 and 13 from 2024 ( Cycle column). For 2024 proposals, proposal-level means in Human.csv are the official institute panel means; for 2023 proposals they are computed from the individual ratings in Human_raw_scores.csv . Reviewers per proposal. Panels ranged from 3 to 6 reviewers. Reviewer identities were never provided; the human inter-reviewer reliability is therefore estimated with a one-way random-effects model (ICC(1,1)), which is the appropriate model when each proposal is rated by a different, unidentified set of reviewers. Six 2024 reviews not present at the individual level. For six 2024 proposals (G, R, T, W, X, Z), one reviewer's scores were submitted without written comments and are not included in Human_raw_scores.csv . For these proposals, Human.csv carries the official institute panel means, so the panel mean in Human.csv and the mean recomputed from Human_raw_scores.csv differ slightly. The reproducibility check in 06_Human_data.ipynb confirms exact agreement for all proposals with complete individual records and reports the expected small differences for these six. Out-of-scale LLM Overall ratings. A small number of responses (≈1.2% of reviews, almost entirely from gpt-3.5-turbo) rated the overall recommendation on a 1-5 scale rather than the requested 1-3 scale, in a format the parser could not distinguish. The analysis code masks values outside the valid range before computing any Overall -based result; the composite Score does not use Overall and is unaffected. Excluded LLM evaluations. Evaluations in which a model failed to return one or more required numerical scores were removed prior to analysis. The file provided here is the post-exclusion (analyzed) dataset. Code notebooks Two groups of notebooks are provided: (i) the LLM evaluation pipeline and (ii) statistical analysis and figure generation. LLM evaluation pipeline (Notebooks 1-4) Documentation of the methodology used to generate the LLM evaluations. These use synthetic examples and contain no confidential data, API credentials, or real proposal text. 01_pipeline_overview.ipynb - architecture, configuration, criteria, output format, evaluation matrix. 02_prompting_strategies.ipynb - one-shot and chain-of-thought prompting; example selection; text vs. vision input. 03_response_parsing.ipynb - regex extraction of ratings and comments; error handling; decimal ratings. 04_example_evaluation.ipynb - end-to-end workflow on synthetic data. Statistical analysis and figures (Notebooks 5-6) 05_Data_Processing.ipynb - ANOVA and effect sizes; empirical absolute score differences; ICC(2,1) and Spearman correlations vs. the human panel; Monte Carlo comparisons; scatter, slope, and bias figures. 06_Human_data.ipynb - ICC(1,1)/ICC(1,k) for human reviewers with bootstrap confidence intervals; Monte Carlo single-reviewer vs. leave-one-out panel rank agreement; consistency check of Human.csv against Human_raw_scores.csv . Reproducibility Running 05_Data_Processing.ipynb and 06_Human_data.ipynb against the included data reproduces the statistical results and figures in the manuscript. Notebooks 1-4 document the evaluation pipeline using synthetic examples. Requirements pandas numpy scipy statsmodels scikit-learn matplotlib The LLM evaluation pipeline notebooks (1-4) additionally use the packages listed in requirements.txt . Citation If you use this data or code, please cite the associated manuscript and this archive: Gorski, C., Leo, N., Gayah V. Agreement Between Large Language Models and Humans in Research Proposal Review - Data and Scripts. 2026. Zenodo. https://doi.org/10.5281/zenodo.18187034 License Data are released under CC BY 4.0; code is released under the MIT License.
Ahmmed, Imtiaz · Masud, Md Hasan · Sarker, Monosij Kanti · et al.
4.6 MB
This repository contains the reproducibility package associated with the study: "Bridging the Gap Between Awareness and Action: An Empirical Study of Technical Debt Management in Agile Software Development." The package includes: • Survey questionnaire • Survey dataset • Supporting documentation The dataset contains responses from software engineering professionals regarding technical debt awareness, causes, impacts, and management practices in agile software development environments. The materials are provided to facilitate verification, replication, and future research based on the study findings.
General Information This repository contains the comprehensive collection of empirical lacustrine time series, sedimentary core records, standardized data matrices, and execution scripts required to fully replicate the figures, network topologies, and statistical null models presented in the associated manuscript. I. Software and Environment Requirements Python (v3.8+): Required libraries include numpy , pandas , matplotlib , scipy , statsmodels , and pymannkendall . R (v4.2+): Required libraries include wsyn , igraph , zoo , parallel , pbapply , ggplot2 , and cowplot . Geospatial Platform: Esri ArcGIS Desktop (v10.8) or ArcGIS Pro (for cartographic rendering and vector layer manipulation). Computation Note: To mitigate boundary artifacts and edge effects inherent in chronological sliding-window operations, the final 12-24 data points of the generated time series are systematically excluded from the final trend evaluation. II. File Inventory and Component Descriptions 1. Empirical and Core Datasets ( .csv ) Global_Lake_Chlorophyll_a_Time_Series.csv Description: Long-term, multi-decadal monthly gridded chlorophyll- a concentration time series across global limnological cohorts, featuring unique lake identification codes as columns and sequential temporal intervals as rows. Global_Lake_Water_Color_Time_Series.csv Description: Normalized global lake water color time series quantified via the Forel-Ule Index (FUI) framework, structured identically to the chlorophyll dataset for multi-proxy comparison. Global_Lake_Sediment_Pigments.csv Description: Stratigraphic sedimentary core pigment records providing long-term retrospective evidence of historical limnological synchronization and baseline shifts. Figures_data.CSV Description: Consolidated and curated data matrices containing the exact values, coordinates, and regional groupings utilized to plot the core text figures. 2. Statistical Analysis and Mathematical Scripts ( .py & .R ) Pairwise correlation -based synchrony.py Description: Computes macro-scale spatial synchrony trends across lacustrine nodes using the standard pairwise correlation matrix stream following frequency-domain decomposition. Loreau φ Metric for lake synchrony.py Description: Execution script utilizing the classic Loreau-de Mazancourt φ metric to calculate multi-lake population-level synchrony across global and latitudinal cohorts. Sliding window sensitivity.py Description: Explores scale dependency and temporal robustness by executing the analytical data stream across varying sliding temporal windows (e.g., 2, 5, 8, 10, 15 time steps). Network and modularity analysis.R Description: Implements the wsyn continuous signed-power soft-thresholding paradigm and leverages igraph to partition similarity networks into topological communities, calculating decadal modularity . Permutation-based significance of synchrony trends.py Description: Performs Mann-Kendall trend tests on sliding-window synchrony series and runs empirical hypothesis testing to extract true directional shifts. Permutation-based significance of synchrony trends-Null_Model_Generator.py Description: Harnesses multi-core parallel processing to shuffle network weights 1,000 times, constructing empirical null distributions to validate the significance of observed network community dissolution. 3. Geospatial Visualization & Documentation ( .rar & .txt ) Figure 1.rar Description: Compressed archive containing all raw geospatial project databases, vector layer shapefiles ( .shp ), metadata tables, and cartographic layout definitions ( .mxd ) used to generate the global geographic distribution map of sample lakes (Figure 1). Compiled within Esri ArcGIS 10.8. Data_Sources_and_References.txt Description: A comprehensive standalone text file documenting the complete bibliographic literature sources, historical baselines, and corresponding DOI attributions compiled within the empirical datasets. III. Execution and Replication Workflow Data Cleaning & Detrending: Feed raw time series through Python/R scripts to execute STL harmonic regression models and filter low-frequency background signals. Synchrony Calculations: Execute the pairwise and Loreau metric scripts to plot continuous synchrony variations over time. Network Configurations: Run the R network script to output the high-resolution PCA community plots and decadal modularity comparisons. Significance Evaluation: Launch the null model generator to confirm that network configuration shifts significantly exceed random stochastic expectations ( P < 0.001 ). Spatial Reconstruction: Extract Figure 1.rar into your local GIS directory to access, modify, or re-export the multi-layered baseline global sampling maps.
Schlosser, Elisabeth · Reijmer, Carleen H · Oerter, Hans · et al.
342 rows × 5 cols · 15 KB
4 numeric · 1 datetime
The relationship between d18O and air temperature at Neumayer station, Ekströmisen, Antarctica, was investigated using fresh-snow samples from the time period 1981-2000. A trajectory model that calculated 5 day-backward trajectories was used to study the influence of different synoptic weather situations and thus of different moisture sources on this correlation. Generally a high correlation between air temperature and d18O was found, but the quality of the d18O-T relationship varied with the different trajectory classes. Additionally, the sea-ice coverage on the travel path of the moist air was considered. The amount of open ocean water underneath the trajectory has a large influence on the d18O-T relationship. For trajectories that lead completely above open water, no significant correlation between d18O and T was found, because mixing with air masses containing additionally evaporated water vapour from the ocean influences the isotope ratio of precipitation. A very high correlation, however, was found for transports over the completely ice-covered Weddell Sea.
Zolduoarrati, Elijah · Licorish, Sherlock · Stanger, Nigel
3.7 MB
Existing quantitative studies examining how diversity affects Stack Overflow contributions have only captured statistical trends, yet fail to explain how their findings resonate to the actual users' experience. Qualitative investigation is essential to validate these findings. Our study synthesises existing research to yield 42 key outcomes, informing the development of 14 open-ended questions for Stack Overflow users related to participation, contribution utility and value, and quality. 209 responses were accrued that reflected how US-based users interact with the platform, where inductive thematic analysis was conducted to identify emergent themes. This replication package is provided for those interested in further examining our research methodology.
Measurements of atmospheric radioactivity attached to aerosols are described. Fallout was collected in a vessel of large area. Emphasis was on separation of "wet" and "dry" samples. For strontium 90 a ratio of "wet" to "dry" fallout of 5:1 has been found independent of latitude. The total fallout was smaller than comparable values from continents because of very small amounts of rainfall in the equatorial zone. In order to achieve consistency in the global balance a better knowledge not only of radioactivity but also of precipitation over the ocean is required. Fallout of Ra-D clearly shows the ITC as a barrier for the latitudinal movement of near sea-surface air masses. The concentration of short-lived emanation daughters shows large variations according to varying geographic conditions. A variation with time could not be explained. The specific activity of long-lived radioactive substances shows the expected effect of the ITC as well as a seasonal diminuation of average concentration, similar to that measured at Heidelberg.
The present invention is an intelligent gateway which can receive multiple sensor data using sub-1G Hz frequency, analyze data, and transmit processed data to a database server. The intelligent gateway can receive data from up to 100 sensors using sub-1G Hz (433, 868 or 915 MHz) wireless frequency. The received data can be analyzed and the gateway can determine when to transmit data, and which packaged data to transmit to the database server. The intelligent gateway can also receive feedback and instructions from the database server. The process data can be transmitted to the database server with different protocols like WIFI, Ethernet and RS485. The intelligent gateway can also include multiple sensors including temperature and humidity sensors, pressure sensors, air speed sensors and a particulate matter sensor for detecting particulates of less than 2.5 micro meters (PM2.5). These sensors are collect additional indoor environmental quality parameters.
The present invention provides a ventilation system for improving air quality of an indoor space. The system includes sensors for measuring PM2.5 particle level P, CO 2 level C, and TVOC level T in the indoor space. A control circuit is configured to receive P, C and T values and generate an output signal Vout according to a specific algorithm, which in turn controls speed-variable EC motors that drive ventilating fans. The invention exhibits numerous technical merits such as lower energy consumption, programmable operation, high efficiency, and lower noise, among others.
declared
Declared scholarly works - open one to read it and see the measured datasets that cite it.
Jun Wang, Sundar A. Christopher · Geophysical Research Letters · 2003
We explore the relationship between column aerosol optical thickness (AOT) derived from the Moderate Resolution Imaging SpectroRadiometer (MODIS) on the Terra/Aqua satellites and hourly fine particulate mass (PM2.5) measured at the surface at seven locations in Jefferson county, Alabama for 2002. Results indicate that there is a good correlation between the satellite‐derived AOT and PM2.5 (linear correlation coefficient, R = 0.7) indicating that most of the aerosols are in the well‐mixed lower boundary layer during the satellite overpass times. There is excellent agreement between the monthly mean PM2.5 and MODIS AOT (R > 0.9), with maximum values during the summer months due to enhanced photolysis. The PM2.5 has a distinct diurnal signature with maxima in the early morning (6:00 ∼ 8:00AM) due to increased traffic flow and restricted mixing depths during these hours. Using simple empirical linear relationships derived between the MODIS AOT and 24hr mean PM2.5 we show that the MODIS AOT can be used quantitatively to estimate air quality categories as defined by the U.S. Environmental Protection Agency (EPA) with an accuracy of more than 90% in cloud‐free conditions. We discuss the factors that affect the correlation between satellite‐derived AOT and PM2.5 mass, and emphasize that more research is needed before applying these methods and results over other areas.
Jing Cheng, Dan Tong, Qiang Zhang et al. · National Science Review · 2021
Abstract Clean air policies in China have substantially reduced particulate matter (PM2.5) air pollution in recent years, primarily by curbing end-of-pipe emissions. However, reaching the level of the World Health Organization (WHO) guidelines may instead depend upon the air quality co-benefits of ambitious climate action. Here, we assess pathways of Chinese PM2.5 air quality from 2015 to 2060 under a combination of scenarios that link global and Chinese climate mitigation pathways (i.e. global 2°C- and 1.5°C-pathways, National Determined Contributions (NDC) pledges and carbon neutrality goals) to local clean air policies. We find that China can achieve both its near-term climate goals (peak emissions) and PM2.5 air quality annual standard (35 μg/m3) by 2030 by fulfilling its NDC pledges and continuing air pollution control policies. However, the benefits of end-of-pipe control reductions are mostly exhausted by 2030, and reducing PM2.5 exposure of the majority of the Chinese population to below 10 μg/m3 by 2060 will likely require more ambitious climate mitigation efforts such as China's carbon neutrality goals and global 1.5°C-pathways. Our results thus highlight that China's carbon neutrality goals will play a critical role in reducing air pollution exposure to the level of the WHO guidelines and protecting public health.
Zhenyu Zhang, Shiqing Zhang · International Journal of Environmental Science and Technology · 2023
Abstract Air quality forecasting is of great importance in environmental protection, government decision-making, people's daily health, etc. Existing research methods have failed to effectively modeling long-term and complex relationships in time series PM2.5 data and exhibited low precision in long-term prediction. To address this issue, in this paper a new lightweight deep learning model using sparse attention-based Transformer networks (STN) consisting of encoder and decoder layers, in which a multi-head sparse attention mechanism is adopted to reduce the time complexity, is proposed to learn long-term dependencies and complex relationships from time series PM2.5 data for modeling air quality forecasting. Extensive experiments on two real-world datasets in China, i.e ., Beijing PM2.5 dataset and Taizhou PM2.5 dataset, show that our proposed method not only has relatively small time complexity, but also outperforms state-of-the-art methods, demonstrating the effectiveness of the proposed STN method on both short-term and long-term air quality prediction tasks. In particular, on singe-step PM2.5 forecasting tasks our proposed method achieves R 2 of 0.937 and reduces RMSE to 19.04 µg/m 3 and MAE to 11.13 µg/m 3 on Beijing PM2.5 dataset. Also, our proposed method obtains R 2 of 0.924 and reduces RMSE to 5.79 µg/m 3 and MAE to 3.76 µg/m 3 on Taizhou PM2.5 dataset. For long-term time step prediction, our proposed method still performs best among all used methods on multi-step PM2.5 forecasting results for the next 6, 12, 24, and 48 h on two real-world datasets.
Yu Fei Xing, Yue Xu, Min-Hua Shi et al. · PubMed · 2016
Recently, many researchers paid more attentions to the association between air pollution and respiratory system disease. In the past few years, levels of smog have increased throughout China resulting in the deterioration of air quality, raising worldwide concerns. PM2.5 (particles less than 2.5 micrometers in diameter) can penetrate deeply into the lung, irritate and corrode the alveolar wall, and consequently impair lung function. Hence it is important to investigate the impact of PM2.5 on the respiratory system and then to help China combat the current air pollution problems. In this review, we will discuss PM2.5 damage on human respiratory system from epidemiological, experimental and mechanism studies. At last, we recommend to the population to limit exposure to air pollution and call to the authorities to create an index of pollution related to health.
Chong Liu, Po‐Chun Hsu, Hyun‐Wook Lee et al. · Nature Communications · 2015
Particulate matter (PM) pollution has raised serious concerns for public health. Although outdoor individual protection could be achieved by facial masks, indoor air usually relies on expensive and energy-intensive air-filtering devices. Here, we introduce a transparent air filter for indoor air protection through windows that uses natural passive ventilation to effectively protect the indoor air quality. By controlling the surface chemistry to enable strong PM adhesion and also the microstructure of the air filters to increase the capture possibilities, we achieve transparent, high air flow and highly effective air filters of ~90% transparency with >95.00% removal of PM2.5 under extreme hazardous air-quality conditions (PM2.5 mass concentration >250 μg m−3). A field test in Beijing shows that the polyacrylonitrile transparent air filter has the best PM2.5 removal efficiency of 98.69% at high transmittance of ~77% during haze occurrence. Particulate matter pollution is a public health concern in industrialized and urban areas. Here, the authors control the surface chemistry and microstructure of filtration materials to fabricate effective and transparent air filters for the capture of PM2.5pollutants.
Xiaochun Yang, Qizhong Wu, Rong Zhao et al. · Atmospheric Environment · 2019
Particulate matter is the main air pollutant in China, especially in Xi'an in recent years. Since 2013, the WRF-SMOKE-CMAQ model system has been used to build an air quality model system for daily air quality forecasting in Xi'an. The emission inventory was built based on several anthropogenic emission inventories and open access emission datasets, and the model evaluation is presented to verify the emission inventory for particulate matter in Xi'an. Comparing the daily observed and simulated fine particulate (PM 2.5 ) concentrations for four winters in different years (from 2014 to 2017), the model performs well in all studied time periods. The correlation coefficient of the simulated daily PM 2.5 concentration data are all larger than 0.58, reaches 0.80 in 2016, and the fraction of predictions within a factor of two of observations (FAC2) are all above 66%. The differences of simulated results based on emission-unchanged system between 2014 and 2015 indicate that the slightly deteriorating air quality of 2015 is affected by the unfavorable air diffusion condition. The PM 10 concentration increases from 95.9 μg/m 3 to 110.3 μg/m 3 , and the PM 2.5 from 82.4 μg/m 3 to 95.4 μg/m 3 . According to the error analysis in model performance, the serious polluted situation of 2016 is mostly because of the sharp increased dust emissions. The emission-unchanged simulated particulate matter concentrations have little variation from 2015 to 2016, but the observation data increase obviously, that results in dramatically change of Mean Bias (MB). The absolute MB of PM 10 increase from 70.3 μg/m 3 to 135.2 μg/m 3 , and PM 2.5 from 0.46 μg/m 3 to 69.9 μg/m 3 . While the improved air quality in 2017 is attributed to both the better weather condition and the emission-reductions. The emission-unchanged simulated results decrease, and the absolute MB even have bigger decrease, that of the PM 10 concentration reduce by 47μg/m 3 , and PM 2.5 by 37μg/m 3 .
Judith C. Chow, L.‐W. Antony Chen, John G. Watson et al. · Journal of Geophysical Research Atmospheres · 2006
The 14‐month‐long (December 1999 to February 2001) Central California Regional PM10/PM2.5 Air Quality Study (CRPAQS) consisted of acquiring speciated PM2.5 measurements at 38 sites representing urban, rural, and boundary environments in the San Joaquin Valley air basin. The study's goal was to understand the development of widespread pollution episodes by examining the spatial variability of PM2.5, ammonium nitrate (NH4NO3), and carbonaceous material on annual, seasonal, and episodic timescales. It was found that PM2.5 and NH4NO3 concentrations decrease rapidly as altitude increases, confirming that topography influences the ventilation and transport of pollutants. High PM2.5 levels from November 2000 to January 2001 contributed to 50–75% of annual average concentrations. Contributions from organic matter differed substantially between urban and rural areas. Winter meteorology and intensive residential wood combustion are likely key factors for the winter‐nonwinter and urban‐rural contrasts that were observed. Short‐duration measurements during the intensive operating periods confirm the role of upper air currents on valley‐wide transport of NH4NO3. Zones of representation for PM2.5 varied from 5 to 10 km for the urban Fresno and Bakersfield sites, and increased to 15–20 km for the boundary and rural sites. Secondary NH4NO3 occurred region‐wide during winter, spreading over a much wider geographical zone than carbonaceous aerosol.
Zhen Zhang, Shiqing Zhang, Xiaoming Zhao et al. · Frontiers in Environmental Science · 2022
Air quality PM2.5 prediction is an effective approach for providing early warning of air pollution. This paper proposes a new deep learning model called temporal difference-based graph transformer networks (TDGTN) to learn long-term temporal dependencies and complex relationships from time series PM2.5 data for air quality PM2.5 prediction. The proposed TDGTN comprises of encoder and decoder layers associated with the developed graph attention mechanism. In particular, considering the similarity of different time moments and the importance of temporal difference between two adjacent moments for air quality PM2.5prediction, we first construct graph-structured data from original time series PM2.5 data at different moments without explicit graph structure. Then we improve the self-attention mechanism with the temporal difference information, and develop a new graph attention mechanism. Finally, the developed graph attention mechanism is embedded into the encoder and decoder layers of the proposed TDGTN to learn long-term temporal dependencies and complex relationships from a graph prospective on air quality PM2.5 prediction tasks. Experiment results on two collected real-world datasets in China, such as Beijing and Taizhou PM2.5 datasets, show that the proposed method outperforms other used methods on both short-term and long-term air quality PM2.5 prediction tasks.
Shovan Kumar Sahu, Sri Harsha Kota · Aerosol and Air Quality Research · 2016
In New Delhi, the capital city of India, concentrations of regulated air pollutants often exceed the Indian national ambient air quality standards (INAAQS). As the sources of these pollutants differ, it is of utmost priority to understand the most dangerous air pollutant to formulate better control strategies in the city. In this study, regulated air pollutant concentrations in New Delhi during 2011 to 2014 were collected. Compared to other pollutants, PM2.5 concentrations exceeded the INAAQS quite often. While PM2.5 exceeded INAAQS during 85% of the days, NO2, O3, CO and SO2 exceeded only on 37, 14, 11 and 0% of the days, respectively. Using air quality index approach, the most dominant pollutant was identified as PM2.5, for 75 to 90% of the days. However, a seasonal variation in the percentage dominance of PM2.5 was observed. For example, PM2.5 was dominant during 95% of the winter and 68% of monsoon days. In addition to absolute concentrations, pollutants can also be ranked by studying their associated short term mortality impacts. However, such studies are rare in India. For the first time, the short term impact of PM2.5 concentrations on non-disease specific mortality in New Delhi was assessed using Poisson regression models. Results indicated that the excessive risk associated with PM2.5 estimated was 0.57, which was higher than the other regulated pollutants. This indicates a projected 6.2 and 6.5% decrease in mortality by meeting the PM2.5 Indian standards and WHO set limits, respectively.
Joshua Schulz Apte, Julian D. Marshall, Aaron Cohen et al. · Environmental Science & Technology · 2015
Ambient fine particulate matter (PM2.5) has a large and well-documented global burden of disease. Our analysis uses high-resolution (10 km, global-coverage) concentration data and cause-specific integrated exposure-response (IER) functions developed for the Global Burden of Disease 2010 to assess how regional and global improvements in ambient air quality could reduce attributable mortality from PM2.5. Overall, an aggressive global program of PM2.5 mitigation in line with WHO interim guidelines could avoid 750 000 (23%) of the 3.2 million deaths per year currently (ca. 2010) attributable to ambient PM2.5. Modest improvements in PM2.5 in relatively clean regions (North America, Europe) would result in surprisingly large avoided mortality, owing to demographic factors and the nonlinear concentration-response relationship that describes the risk of particulate matter in relation to several important causes of death. In contrast, major improvements in air quality would be required to substantially reduce mortality from PM2.5 in more polluted regions, such as China and India. Moreover, forecasted demographic and epidemiological transitions in India and China imply that to keep PM2.5-attributable mortality rates (deaths per 100 000 people per year) constant, average PM2.5 levels would need to decline by ∼20-30% over the next 15 years merely to offset increases in PM2.5-attributable mortality from aging populations. An effective program to deliver clean air to the world's most polluted regions could avoid several hundred thousand premature deaths each year.
Aaron van Donkelaar, Randall V. Martin, Michael D Brauer et al. · Environmental Health Perspectives · 2010
BACKGROUND: Epidemiologic and health impact studies of fine particulate matter with diameter < 2.5 microm (PM2.5) are limited by the lack of monitoring data, especially in developing countries. Satellite observations offer valuable global information about PM2.5 concentrations. OBJECTIVE: In this study, we developed a technique for estimating surface PM2.5 concentrations from satellite observations. METHODS: We mapped global ground-level PM2.5 concentrations using total column aerosol optical depth (AOD) from the MODIS (Moderate Resolution Imaging Spectroradiometer) and MISR (Multiangle Imaging Spectroradiometer) satellite instruments and coincident aerosol vertical profiles from the GEOS-Chem global chemical transport model. RESULTS: We determined that global estimates of long-term average (1 January 2001 to 31 December 2006) PM2.5 concentrations at approximately 10 km x 10 km resolution indicate a global population-weighted geometric mean PM2.5 concentration of 20 microg/m3. The World Health Organization Air Quality PM2.5 Interim Target-1 (35 microg/m3 annual average) is exceeded over central and eastern Asia for 38% and for 50% of the population, respectively. Annual mean PM2.5 concentrations exceed 80 microg/m3 over eastern China. Our evaluation of the satellite-derived estimate with ground-based in situ measurements indicates significant spatial agreement with North American measurements (r = 0.77; slope = 1.07; n = 1057) and with noncoincident measurements elsewhere (r = 0.83; slope = 0.86; n = 244). The 1 SD of uncertainty in the satellite-derived PM2.5 is 25%, which is inferred from the AOD retrieval and from aerosol vertical profile errors and sampling. The global population-weighted mean uncertainty is 6.7 microg/m3. CONCLUSIONS: Satellite-derived total-column AOD, when combined with a chemical transport model, provides estimates of global long-term average PM2.5 concentrations.
Yanlin Zhang, Fang Cao · Scientific Reports · 2015
This study presents one of the first long term datasets including a statistical summary of PM2.5 concentrations obtained from one-year monitoring in 190 cities in China. We found only 25 out of 190 cities could meet the National Ambient Air Quality Standards of China, and the population-weighted mean of PM2.5 in Chinese cities are 61 μg/m(3), ~3 times as high as global population-weighted mean, highlighting a high health risk. PM2.5 concentrations are generally higher in north than in south regions due to relative large PM emissions and unfavorable meteorological conditions for pollution dispersion. A remarkable seasonal variability of PM2.5 is observed with the highest during the winter and the lowest during the summer. Due to the enhanced contributions from dust particles and open biomass burning, high PM2.5 abundances are also found in the spring (in Northwest and West Central China) and autumn (in East China), respectively. In addition, we found the lowest and highest PM2.5 often occurs in the afternoon and evening hours, respectively, associated with daily variation of the boundary layer depth and anthropogenic emissions. The diurnal distribution of the PM2.5-to-CO ratio consistently displays a pronounced peak during the afternoon periods, reflecting a significant contribution of secondary PM formation.
Noemí Pérez, Jorge Pey, Michael Cusack et al. · Aerosol Science and Technology · 2010
Measurements of particle number concentration (N), black carbon (BC), and PM 10 , PM 2.5 , and PM 1 levels and speciation were carried out at an urban background monitoring site in Barcelona. Daily variability of all aerosol monitoring parameters was highly influenced by road traffic emissions and meteorology. The levels of N, BC, PM X , CO, NO, and NO 2 increased during traffic rush hours, reflecting exhaust, and non-exhaust traffic emissions and then decreased by the effect of breezes and the reduction of traffic intensity. PM 2.5–10 levels did not decrease during the day as a result of dust resuspension by traffic and wind. N showed a second peak, registered in the afternoon and parallel to O 3 levels and solar radiation intensity, that may be attributed to photochemical nucleation of precursor gases. An increasing trend was observed for PM 1 levels from 1999 to 2006, related to the increase in the traffic flow and the diesel fleet in Barcelona. PM composition was highly influenced by road traffic emissions, with exhaust emissions being an important source of PM 1 and dust resuspension processes of PM 2.5–10 , respectively.