The round trip

You are a dataset. Follow yourself down. Rise back up as an answer.

You are a 4.2 GB climate file sitting in a repository. Nobody has opened you. A question is on its way down to find you, and scrolling this page is the whole round trip: the descent that measures you, then the climb back up as an answer.

you are here: 4.2 GB, unopened, top of the round trip

L0 declared

Where you live now

Every catalog indexes what a publisher declared: a title, a DOI, a size, a license. Google Dataset Search, DataCite and Hugging Face all stop exactly here. Everything below the line is the bytes nobody opened, the unsolved last mile.

The metadata ceilingA dataset with a lit top band of declared fields (title, DOI, license) and a large unread body of bytes beneath a line labeled where every other catalog stops.title: declaredDOI: declaredlicense: declaredwhere every other catalog stopsbytes: never read

L0 to L1 · the deep probe

The touch: a few KB, then discard

The probe reaches into your multi-gigabyte body and reads only a header range and a footer range, kilobytes in all. A Parquet footer alone yields an exact row count and per-column min, max and null counts: kilobytes read from a multi-gigabyte body, orders of magnitude less data moved. It measures, records exactly how much it touched, then discards the bytes. Metadata-only is a custody rule, not an epistemology: no byte custody, but knowledge earned from bytes.

Header and footer range reads, with an honesty budgetA multi-gigabyte file whose top and bottom ranges are read into a sandboxed probe; the probe emits a measurement budget recording 96 KB fetched via the footer, the total size not reported on a footer read, and the value-scan statistics honestly declined, then the bytes are discarded.headerfooterabout 4.2 GB, never fetcheddeep probemeasures in a sandboxmeasurement budgetbytes_fetched96 KBbytes_availablenot reportedfetch_modefooterdeclinedmean · median · quantiles · histogrambytes discarded

the ladder

Understanding is earned, rung by rung

Each level is earned from the one below, never asserted. Measured distributions and quantiles, then the role grammar, then meaning read from evidence alone with every claim cited: the hallucination firewall. Topology is a closed vocabulary of ten - tables, tensors, spatial rasters, sequences, graphs, 3D molecular structures and more - so one strategy is written once per measurable topology, never once per format.

The understanding ladder, L0 to L3Four rungs on a vertical spine: L0 declared (neutral), then L1 measured with a sixteen-bin histogram, L2 grammar, and L3 meaning, each earned from the one below; faint L4 and L5 stubs continue above and the ten-value topology vocabulary lists down the right side.L4 identity-in-timeL5 capabilityL3meaningfrom evidence only: the cited basis is the firewallL2grammarRoleMap: anchor / link / axis / signal · GrainModelL1measureddistributions, quantiles, histogramsL0declaredwhat the publisher said: where metadata search stopsearned fromearned fromearned fromtopologytabulartensorspatialsequencegraphstructurestreammodalcompositeunknownclosed vocabulary

one engine, two doors

Served byte-identically, two altitudes

Now fully measured, you are served through two front doors from one Catalog facade and one set of frozen contracts. The agent surface is the primary customer and gets a lean projection; the human Workbench gets the full record. Same truth, different altitude. The loader always runs in your runtime: Meridian is the brain, your environment is the hands.

One engine, two front doorsOne engine of Catalog facade and frozen contracts feeds two doors: an agent MCP surface served a lean projection (a fraction of the tokens of the full record) and a human Workbench served the full record; both walk the same six step journey.one engineCatalog facade · frozen contracts · usage.pyAgent surface · MCPthe primary customer: LeanHit / DatasetFittokens per result pagefulllean: a fractionWorkbench · humaninsight bench + embedded agentthe full record, same MCP toolsboth doors walk the same journeydiscoverassessloadunderstandcomposeciteget_dataset(view=fit) gives lean · view=full gives the byte-identical record

the action layer

The right loader for what you actually are

Because the probe measured you, Meridian hands the loader for your specific topology: pandas for a tabular CSV, xarray for a NetCDF tensor, geopandas with masked nodata and real coordinate names for a raster. Not a generic snippet. And the same measured structure is served as standards on every endpoint, Croissant first.

Minting the right loader from the measured topologyA measured dataset of topology tensor flows through build_recipe into three artifacts: a tailored loader, a Croissant export that travels both ways, and a live MCP recipe. A separate neutral card shows the generic snippet others ship, struck through.you, now measuredtensorbuild_recipeone source, every surfaceloaderxr.open_dataset(url)Croissantmeasured recordSet field statsMCP use_datasetserved livea generic snippetwhat others shipexport on every endpoint: ?format=croissant | dcat | schema_org

the ML-native contract

Spoken in the format ML tools already read

Croissant is the metadata format ML datasets are converging on: an MLCommons standard, a JSON-LD layer on schema.org that tells a framework exactly what files exist, how fields are extracted, and how to load them. Hugging Face auto-generates it for every dataset, Kaggle and OpenML serve it, TensorFlow Datasets and the mlcroissant library load it, and Google Dataset Search indexes it. Meridian emits Croissant 1.1, layers in the GeoCroissant extension for geospatial data, and proves conformance in CI against the real mlcroissant library and our own SHACL shapes, with one difference no other producer has: every field statistic, array shape and spatial measurement comes from your bytes, not copied from what a publisher declared. And the same measured structure serves every standard at one endpoint: Croissant 1.1 for ML loaders, DCAT-AP 3.0.1 (SHACL-conformant, controlled-vocabulary IRIs) for catalog and government interop, and schema.org for open-web discovery. Measure once, emit every dialect.

Anatomy of a Meridian Croissant fileA Croissant 1.1 JSON-LD document with declared metadata in a top band and, below a dashed line labeled measured from your bytes, the measured parts: a recordSet of fields with statistics annotations and native arrays, GeoCroissant terms for geospatial data, FileObject descriptions and contentSize.declared@type sc:Dataset · conformsTo [ croissant/1.1, geo/1.0 ] · name · license · creatorsmeasured from your bytesrecordSetfields, stat annotations, keygeocr: + arraysCRS, resolution, shapesFileObjectmeasured descriptioncomplete categoricals become an sc:Enumeration · contentSize and encodingFormat on every file

how we extend it

We do not copy declared metadata into the record. We measure the bytes, add what the standard can carry, and prove the file two ways.

How Meridian extends and validates CroissantThe Croissant 1.1 baseline, plus the measured parts Meridian adds (per-field statistics, native array shapes and dtypes, and the GeoCroissant geo/1.0 extension), then validated two ways: against the mlcroissant library and against Meridian's own SHACL shapes for the GeoCroissant terms the library does not check.Croissant 1.1the ML-nativebaselinemeasureMeridian adds, measuredfield stats: range · median · null ratenative arrays: shapes + element dtypesGeoCroissant: CRS · resolution · bandsvalidatevalidated 2 waysmlcroissant library+ our SHACL shapes

by topology

How the ecosystem encodes each modality in Croissant, and what Meridian emits for it, measured from the bytes.

tabular

in the ecosystem

one cr:Field per column, with a dataType and a source.extract.column recipe

measured by Meridian

the same recordSet, plus the measured range, median, distinct count and null rate as Croissant 1.1 annotations, dataset-level quality on the recordSet (row count, null and duplicate rates) and a Responsible-AI missing-data note from the measured null rate, a record key from the measured identity columns, and a complete categorical vocabulary lifted into an sc:Enumeration the column references

spatial

in the ecosystem

a dataset-level spatialCoverage Place, with no CRS, resolution or band vocabulary

measured by Meridian

GeoCroissant geo/1.0: the measured CRS, pixel resolution with its unit, and band count and names for raster; for vector, a real recordSet of the OGR attribute schema sourced by column. A WGS84 GeoShape box only when the bounds truly are WGS84

tensor

in the ecosystem

native N-dimensional arrays in Croissant 1.1: isArray and arrayShape

measured by Meridian

a 1.1 recordSet with one isArray field per measured array, each carrying its arrayShape and a cr-numeric element dtype (Float32, Int64 and so on) read from the bytes

structure

in the ecosystem

a cr:FileObject for the mmCIF archive, with no molecular vocabulary: a declared blob

measured by Meridian

a measured structure read from the mmCIF bytes: experimental method, resolution, space group, and the counted chains, residues, atoms and element makeup. The 3D coordinates themselves, not a declared archive

modal

in the ecosystem

a cr:FileSet with an sc:ImageObject, AudioObject or VideoObject field sourced by fileProperty content

measured by Meridian

a measured FileObject description: format, width by height, color mode, frames or page count. An honest decline rather than a fabricated single-record field

composite

in the ecosystem

a cr:FileObject archive plus a cr:FileSet of glob includes

measured by Meridian

the archive FileObject with a measured breakdown of its constituent topologies, declining a FileSet we cannot honestly glob yet

Every other Croissant producer copies the publisher's declared metadata into the record. Meridian measures the bytes first, then writes the record: the per-field statistics, the array shapes and dtypes, the coordinate system and pixel resolution, the vector attribute schema in a Meridian Croissant file were all computed from a kilobyte-range read of the actual data, not asserted by whoever uploaded it. Where we did not measure something, we decline to a measured description rather than invent a recordSet. And we validate every file two ways: against the real mlcroissant library on datasets harvested live from Zenodo, and against our own SHACL shapes for the GeoCroissant terms the library does not check.

measure once, speak every dialect

The structure is measured once, from the bytes. From that single source of truth Meridian projects every catalog dialect at the same endpoint, so a dataset is loadable, discoverable, and federatable without re-curation.

schema.org

how you are found

the open-web Dataset vocabulary Google Dataset Search indexes

DCAT-AP 3.0.1

how you federate

the EU catalog-interop profile portals and governments harvest, emitted with controlled-vocabulary IRIs and SHACL-validated

Croissant

how you load

the ML-native contract a training pipeline reads to materialize you

The EU machine-learning profile MLDCAT-AP is tracked as it converges: it builds on the DCAT-AP base Meridian already speaks and points at a Croissant file by conformsTo, so the bridge extends to it without a rewrite.

route to compute

The last mile onto national compute

You hold the loader; here is the path it travels onto national compute, in the order a researcher walks it. Meridian owns one step, measuring the dataset and minting the loader, and routes the data into staging and the run. The identity, the allocation, the credits, the transfer and the node stay with NSF cyberinfrastructure (NAIRR, ACCESS, Globus, CILogon, Open OnDemand): Meridian prepares what runs, never where it runs.

Route to computeAn ordered six-step process from a found dataset to a job running on NSF compute. Step 1, establish a federated identity with an ACCESS ID through CILogon. Step 2, Meridian measures the dataset and mints the loader and Croissant and serves it over MCP. Step 3, win an allocation on ACCESS or the NAIRR Pilot. Step 4, exchange credits for a resource such as Anvil or Jetstream2. Step 5, stage the dataset to your scratch with Globus. Step 6, open Open OnDemand or Jetstream2 and run, importing the Meridian loader. Meridian owns step 2 and routes the data into steps 5 and 6; it never owns the identity, allocation, credits, transfer, or compute.1Establish a federated identityACCESS ID · CILogon · InCommon SSO · Duo MFANSF CI2Measure the dataset, mint the loadertouch the bytes · the right loader · Croissant · MCPMeridian3Win an allocationACCESS (Explore to Maximize) or NAIRR Pilot · merit reviewNSF CI4Exchange credits for a resourceACCESS Credits to core + GPU hours · Anvil · Jetstream2NSF CI5Stage the datasetGlobus collections to your scratch · checksum-verifiedMeridian → Globus6Open a session, run with the loaderOpen OnDemand / Jetstream2 · Jupyter · import the loaderMeridian loaderMeridian prepares the data and routes it in. It never holds your identity, allocation, credits, the transfer, or the node.Meridian actsyou, on NSF cyberinfrastructure

where it sits

The fabric you now live in

Pull back. You now sit inside the national cyberinfrastructure as the measured action layer: repositories upstream, federation alongside, agents and researchers and compute downstream, standards spoken both ways.

Meridian on the national data fabricRepositories upstream feed Meridian, the measured-structure action layer, which federates NSF cyberinfrastructure (NAIRR, ACCESS, OSN, Globus) and serves AI agents, researchers and compute downstream, exporting Croissant, DCAT, schema.org and FAIR.sources · leverage existing repositoriesZenodolive connectorDataverselive · 100+ orgsDataCite · re3data3000+ reposdata.gov · agenciesDCAT feedsharvest · probe · mapMeridianthe measured-structure action layertouches the bytes, mints the loaderNSF cyberinfraNAIRR · ACCESSOSN · Globusfederateserve · loaders · Croissant · MCPconsumers · the customersAI agentsMCP · Cursor · Claude CodeResearchersagentic WorkbenchComputestage + run · ACCESSinteroperability standards served · Croissant · DCAT · schema.org · FAIRMeridian + its primary surfacesexisting ecosystem, leveraged

connect

You are loadable now. This is the floor.

The same measured engine is live and reachable: an agent adds the hosted MCP endpoint with a bearer token and gets every tool above against the real catalog. You have gone as deep as the bytes go. Everything past this line is the way back up - the question that came down to find you now carries you home as an answer.

serverInfo name=meridian · 10 tools · token-gatedlive
claude mcp add --transport http meridian https://meridian.brightquery.ai/mcp --header "Authorization: Bearer $MERIDIAN_TOKEN"

The ascent

You rise back up as an answer

This is the return leg, and it is the moat. Byte-level measurement is the bedrock: it is capital, and once it is funded anyone can rebuild it. The moat is the workflow that stands on it - a research method a lab adopts and cannot easily leave. A problem becomes a query plan you approve before it runs, graded evidence you keep or dismiss onto a provenance graph, and an answer where every claim cites what you kept. The descent left you measured and loadable at the floor; the climb turns those bytes back into meaning. You are no longer a file nobody opened. You are a citation someone can trust.

The research workflow, the moatFive steps left to right: a problem becomes a query plan the researcher approves, then graded evidence, kept or dismissed onto a provenance graph, then a grounded answer where every claim cites what was kept. A human-approval gate sits between the plan and the evidence.1problemnot a keyword2query planyou approve it first3graded evidenceeach wears its grade4keep or dismissonto a PROV graph5grounded answerevery claim is citedyou approvegrounding · grade · attribution · human approval: the trust floor, never configurable

Discover, unified

One query, three families. Byte-measured datasets are the bedrock; scholarly works and patents ride alongside as declared evidence, each graded honestly on its card, each failing on its own so a flaky source never blanks the page. When nothing matches, it says so and names the terms it could not resolve, never padding the page with near-misses.

Literature to data

Open the paper you kept and the byte-measured datasets that cite it rise into view - the reverse of a citation - plus datasets reached by walking the citation graph, each with its traceable path. Paper-first tools stop at the paper; Meridian carries you to the bytes that back it, and one of them is the 4.2 GB file you followed down. The literature, resting on the data. The data was you.

Grounded answers

A question becomes a plan, evidence you keep or dismiss on a provenance graph, and an answer where every sentence cites its evidence and wears its grade. Open a row and its full measured provenance docks beside the answer, inspected in place without leaving it. A firewall keeps and flags any claim that cites what we never provided: never silently dropped, never silently trusted. This grounds an answer in the evidence you kept. It does not synthesize the literature, name its white spaces, or close them, and we do not say it does.

Lifecycle coverage

The scientific data lifecycle

A scientific dataset moves through about ten lifecycle stages, from acquisition to archiving. Meridian is deliberate about which ones it does rather than claiming the whole arc: it owns three - exploration, discovery, analysis - and the other seven belong to the research teams, the repositories, and the partners around it. Naming their owners is the point rather than a concession.

The scientific data lifecycle and Meridian's coverageThe ten scientific-data-lifecycle stages - acquisition, transfer, management, exploration, analysis, curation, sharing, synthesis, discovery, archiving - in a clockwise cycle around the trusted data fabric, each colored by Meridian's coverage: primary, observational, or delegated.Trusteddata fabricdiscovery core01acquisition02transfer03management04exploration05analysis06curation07sharing08synthesis09discovery10archiving
primary - Meridian's heartobservational - reports ondelegated - to the substrate

The three we claim, defined

exploration.
You name the research problem, not a keyword. The Scout proposes where to look and shows the query it plans to run before it runs anything; you decide what counts as proof. See it in Research.
discovery.
What exists anywhere, across measured datasets, scholarly works and patents, each result carrying how we know it: byte-measured, or merely declared by its publisher. See it in Discover, Sources, Quality.
analysis.
Reading into the data itself. We open the bytes and measure structure, quality and grain, rather than transcribing a publisher's metadata; and for data that cannot move, we route the analysis to it through OpenMined's Syft instead of moving the data to the analysis. See it in the measured profile on every dataset, and the Syft compute recipe.

The seven we do not

acquisition
(research teams).
Getting data into your hands, licenses included. Meridian never takes custody of bytes.
transfer
(research teams).
Moving files between instruments, institutions and storage. Not ours to move.
management
(repositories).
We observe it. Link health, freshness and drift are reported, never managed.
curation
(repositories and their stewards).
Our completeness and quality scores surface gaps for a steward. Finding a gap is not curating it.
sharing
(repositories, and OpenMined · PySyft for data that cannot be shared).
Repositories remain the system of record. The interesting case is the data that normally cannot be shared at all, and Syft is the mechanism that lets it be.
synthesis
(research teams).
Synthesis means going back to the literature, naming its gaps and white spaces, and closing them. We ground claims in the evidence you kept. That is not the same thing, and we do not claim it.
archiving
(repositories).
We observe it. A dataset that goes dark is reported, not rescued.

What happens in each stage

The ten stages, the activities in each, and who tends them. AI Alliance, the non-profit prime, convenes the work and delegates the data platform to BrightQuery (Meridian); research teams generate the data; OpenMined · PySyft runs the analysis that cannot move the data, and ML Commons carries Croissant, the metadata standard discovery rests on. Repositories keep the rest.

  1. 01acquisition

    research teams
    • Scientists run experiments and deploy instruments to collect fresh observations.
    • Each reading is recorded in whatever form the instrument or survey produces.
  2. 02transfer

    research teams
    • Newly gathered data is moved off instruments and field sites into durable storage.
    • Files travel between institutions to where the work will actually happen.
  3. 03management

    repositories
    • Every record is organized, described, and kept healthy so it stays usable.
    • Routine checks catch errors early and flag anything that has gone stale.
  4. 04exploration

    BrightQuery
    • A question opens a workspace: one search across measured datasets, scholarly works and patents, with the literature linked to the data that backs it.
    • Evidence is kept or dismissed on a provenance graph, and a grounded, cited answer comes back - every claim tied to what you kept and graded.
  5. 05analysis

    BrightQuery with OpenMined · PySyft
    • Methods and models run against data that is too sensitive to leave its home.
    • Only the findings travel back; the protected data itself never moves.
  6. 06curation

    repositories and their stewards
    • Each record is reviewed for completeness, with gaps and errors written down.
    • A reliability score is attached so users know how far to trust it.
  7. 07sharing

    repositories · OpenMined · PySyft
    • Datasets are published with clear licenses and rules for reuse.
    • Records are indexed so both people and AI tools can actually find them.
  8. 08synthesis

    research teams
    • Research teams combine measured datasets and models, in the open, into something new.
    • Findings from many studies are drawn together into fresh conclusions, the grade of each input still legible.
  9. 09discovery

    BrightQuery
    • Researchers and AI systems search the catalog and weigh quality signals.
    • The right dataset is pinpointed, with a clear route to reach it.
  10. 10archiving

    repositories
    • Records are preserved well beyond the life of any single project.
    • They are re-checked over time so they stay readable and verifiable for years.

Not from zero

Built on deployed systems

A Category-2 solicitation builds upon an existing, deployed system, not a blank page. Meridian is the next generation of two BrightQuery systems already in production for U.S. statistical agencies, carrying the same grounded-evidence discipline from government data to the measured scientific record. These are two of the five foundations it builds off and orchestrates; the rest are the consortium below - Hugging Face, OpenMined, and Croissant, with AI Alliance convening.

BrightQuery Navigator

Deployed for statistical agencies

The assistant that answers questions from government data accurately - grounded in the source record, never a model guessing.

Meridian carries forward the same grounded-evidence discipline, turned from government datasets to the measured scientific record. Meridian is its natural next generation.

BrightQuery DUP

Deployed · Data Usage Platform

Finds where the published literature uses a given dataset - the citation trail from a paper back to the data it rests on.

Meridian carries forward the papers-to-datasets link, the declared reverse index and the tangential citation hop, applied to scientific datasets instead of government ones.

Who builds this

The consortium

Five organizations, and what each actually contributes. Every row carries the state of the integration rather than the intention behind it, because a logo is a claim. The most useful row is the least flattering: Hugging Face datasets are graded declared, since they computed the profile and we read none of the bytes. A project that will not flatter its own partner can be believed about the grades everywhere else.

  • AI Alliance logo

    AI Alliance

    PrimePlanned

    The non-profit that convenes the consortium and delegates the data platform to BrightQuery. Every partner here was already a member; that was designed, not lucky.

    In the code today. Convening and prime award. No code surface.

  • BrightQuery logo

    BrightQuery

    Builds MeridianLive

    The prime delegate: the AI Alliance hands the data platform to BrightQuery, which builds Meridian. Two systems already deployed for U.S. statistical agencies are its heritage - Navigator, which answers questions grounded in the source record, and the Data Usage Platform, which traces the citation trail from a paper back to the data it rests on.

    In the code today. Meridian IS the build. Its measurement-and-grounding spine and its declared papers-to-datasets index descend directly from Navigator and the DUP, turned from government data to the measured scientific record.

  • OpenMined logo

    OpenMined

    Funded partnerPrototype

    Syft: run a method against data too sensitive to move. Only the findings travel; the bytes never do. It is also the only honest answer to sharing, because the interesting data is the data that can never be put in a repository.

    In the code today. A SyftBox connector reads the public datasite surface and mints a syft:// compute recipe. It routes; it does not yet measure the published mock, and it is not registered, so no datasite is harvested in production.

  • MLCommons logo

    MLCommons

    Funded partnerLive

    Croissant, the metadata standard for machine-learning datasets, already adopted across academia and industry. Discovery rests on it.

    In the code today. Every dataset Meridian serves emits Croissant, validated against the shape constraints, alongside full DCAT-AP.

  • Hugging Face logo

    Hugging Face

    Deployment platformLive

    Where AI-native datasets and models are published and found, and where what this project builds is meant to land. Not a lifecycle stage, and not a premise of the argument.

    In the code today. We index the Hub and bind each dataset from the datasets-server statistics. Those profiles are graded DECLARED, because Hugging Face computed them and we read zero bytes. Their auto-converted Parquet is byte-readable, so this is a grade we can earn rather than a limit we are stuck at.

The analysis stage, in depth

Why Syft, and how we work with it

Analysis is one of the three stages Meridian claims, and it is the one we run with a partner. Meridian's own analysis is the probe: we open the bytes and measure the structure, quality and grain of a dataset rather than transcribing what its publisher declared. OpenMined's Syft extends that to the data we are not allowed to open. It runs methods against data too sensitive to move, so only the findings travel and the bytes never do, which is also the one honest answer to sharing: the interesting case is not the data already in a repository, it is the data that can never be put in one.

Meridian as the discovery layer for Syft compute-to-dataA researcher finds a restricted dataset in Meridian, which catalogs SyftBox datasites and reads their public surface, then mints a compute-to-data job. The submitted code runs on the private data inside the datasite, owner-approved, and only the result returns. The private data never moves.Syft leaves discovery out of the protocol (Principle 9). Meridian is the catalog it points to.researcherfindMeridiancatalog datasites · mint the recipemint the compute-to-data on-rampon-rampSyftBox datasiteprivate datayour code runs here,owner-approvedonly the result returns - the private data never moves

why it is vital

The most valuable data is often the data that cannot move: clinical records, federal microdata, proprietary logs. Syft makes it usable without moving it - the code travels to the data, the owner approves and runs it, and only the result returns. It is how the IDSS analysis stage works for data that can never leave home.

how we use it

Syft leaves discovery out of the protocol on purpose (Principle 9: discoverability happens on companion sites like SyftHub). Meridian is that companion: a connector catalogs datasites over their public surface and mints a ready-to-run compute-to-data job as the loader. Today it routes rather than probes: measuring the public mock with the same probe is deferred to Phase 2. Built against the live network.

how we improve it

We give the Syft network the public, measured, AI-native catalog the protocol leaves out, and feed users back into the Syft tools - the minted on-ramp ends inside syft-client. We are aligning with OpenMined on a standard public dataset descriptor so datasites become crawlable and measurable: the missing seam between a private datasite and a discoverable one.

The split is clean and each partner stays in its lane: OpenMined's Syft owns the privacy-preserving compute - the datasite, the mock, the owner approval, the result - and Meridian owns discovery and routing, holding no private data and issuing no approvals. That is the analysis stage of the lifecycle realized: restricted data made findable, assessable, and actionable without a single private byte leaving its owner.

The measurement vocabulary

What a topology is, and the shapes we measure

A topology is the structural shape of a dataset's bytes, not its file format. The format is what a file technically is (CSV, NetCDF, GeoTIFF); the topology is the shape that data takes, and it is what selects the reader. Meridian writes one measurement strategy per shape rather than per format, so a single reader serves every science: CSV, TSV and Parquet are all just tabular. The vocabulary is closed, and honest about its reach: eight shapes are measured by opening the bytes with a purpose-built reader, and two stay declared.

tabularcsv · openpyxl · pyarrow

row count, per-column kind, null rate, distinct count, min/max, mean/median/std, percentiles, a 16-bin histogram, monotonicity and duplicate rate (Parquet reads exact stats from the footer)

tensornumpy · h5py · scipy

per-array name, shape and dtype, the array count and a subtype (grid, model weights, vector store); shapes and dtypes, not element values

spatialrasterio · pyogrio

raster CRS, bounds, width and height, band count and names, pixel resolution with its unit; for vector, geometry type, feature count and the field schema

sequencestdlib · pysam (htslib)

record count, alphabet (DNA, RNA or protein), min/max/mean length and record ids; variants and alignments read through htslib

graphnetworkx · rdflib

node and edge counts, directed and multigraph flags, the node and edge attribute keys; an exact triple count for RDF

structuremmCIF parser (no heavy deps)

experimental method, resolution in angstroms, space group, polymer types, and atom, chain, residue and model counts

modalPillow · pypdf

image format, dimensions, mode and frame count; PDF page count. Container facts only, never pixel or text content

compositezipfile · tarfile

the member file count and the mix of constituent topologies inside, each member classified by name without opening its bytes

Declared only, no reader yet, so graded declared and never guessed: stream (event and time streams) · unknown (the shape cannot be resolved).

Where the work lands

Why Hugging Face, and how we work with them

Hugging Face is not a lifecycle stage and we do not claim one through them. They are the open platform where AI-native datasets and models are published and found, and they are where what this project builds is meant to land. The hand-off is grade-clean by design: what Meridian opened travels as measured, Croissant-native metadata, and what it only read from the record travels as declared and says so - so a community builds on a provenance it can check rather than a grade it has to guess. That standard is Croissant, and ML Commons carries it.

Why it is infrastructure

Built to be national, not an app

Transdisciplinary

One probe measures a climate NetCDF, a social-science CSV, a geospatial raster and a protein structure through the same ten-topology vocabulary. The mechanism is domain-agnostic by construction, and it shows: the live corpus spans paleoclimate, marine and ice-core geoscience, ML benchmarks and now structural biology.

Operational

Not a demo. A continuously running service: measured structure served over GraphQL, REST and MCP from one set of frozen contracts, live on managed cloud.

Integration-native

It never replaces a repository or hosts the bytes. It sits on top of the ecosystem and federates the cyberinfrastructure around it.

AI-driven

The agent is the customer. Where Meridian opened the file the dataset arrives already measured and already loadable, and where it only read the record it says declared: no bespoke parsing, no guessing, no grade it has to trust on faith.

Proposal figure

One page, for the proposal

The whole system in one page, horizontal or vertical: the research workflow as the moat, the flow you author as data and the trust floor you cannot, the measurement bedrock it all stands on, and the consortium behind it. Flip to doc colors and save it as SVG or PNG.

Meridian · grounded research workflows, on byte-measured dataThe workflow is the moat · byte-level measurement is the bedrock it stands onCURRENT STATE · THE PROBLEMResearch tools reason over abstracts; the data stays unopenedRepositories assert; nothing verifies. Absence reads as zeroLiterature and the data behind it live apart; linking is manualAn answer you cannot trace is an answer you cannot useEvery lab rebuilds the same search by hand, and cannot share itTHE MERIDIAN APPROACHThe workflow is the product: a question, carried to an answerEvery claim cites evidence you kept, and it wears its gradeThe grounding is real: we opened the file, not just the recordOne measurement vocabulary reaches every science, not just tablesThe flow is yours to shape; the trust floor under it is notWHAT IT UNLOCKSAnswers a reviewer can audit, claim by claimLiterature-to-data links a researcher can actually trustA method a lab can author, version, publish and re-runA corpus an agent can load without guessing at the schemaCross-field datasets you would never have thought to searchTHE MOAT · THE RESEARCH WORKFLOWa system, not a corpus: plan, approval, grade and provenance are what a researcher adopts1PROBLEMa research question,not a keyword2QUERY PLANa scout agent proposes ityou approve before it runs3GRADED EVIDENCEdatasets · works · patentseach edge wears its grade4KEEP OR DISMISSyou curate what survivesonto a W3C PROV graph5GROUNDED ANSWERcites the evidence you keptor is flagged ungroundedWHY THE WORKFLOW IS THE MOATA measured corpus is capital: fund it and anyone rebuilds itA trusted workflow is a system: plan, approval, grade, provenancePaper-first tools ground on abstracts; Meridian grounds on bytesA lab authors its own method as data, and publishes itThe workflow is what researchers adopt, and adoption was the gapCONFIGURABLE · THE FLOW IS DATA, NOT CODEauthored on a canvas, versioned, publishedTEMPLATESthe flow authored as data:versioned, published, immutableAGENTSrole, prompt, model tier andtool grant, set per agentGATESadd your own approval points;the floor gates never dropORCHESTRATIONan orchestrator and its sub-agents:scout · screener · synthesizerNOT CONFIGURABLE · THE TRUST FLOORGrounding · grade · attribution · human approvalChecked at parse: an invalid flow cannot be savedThe firewall is code, never an authorable agentGrades are minted server-side, never by a clientA closed tool set: a flow composes, never inventsstands onTHE BEDROCK · MEASUREMENTmeasure the bytes once: identical files across sources reuse the profile by content fingerprintHARVESTopen metadata protocols andnative repository APIsRESOLVE54 formats, by magic bytes first,then content, MIME, extensionBYTE-PROBEwe open the file: a boundedsample, read, then discardedUNDERSTANDshape · schema · quality · licenseemitted as Croissant + DCAT-APWHY MEASUREMENT IS THE BEDROCKThe grounding claim is only true because we opened the fileGrade rides the edge: measured (we read it) or declared (they said it)A source that hands us no bytes stays declared, and says so plainlyA ten-shape vocabulary carries measurement past tablesAbsence is honest: an empty result is a result, never a guessBYTE CUSTODY · WE TOUCH THE BYTES, WE DO NOT KEEP THEMThe probe reads a header range and a footer range, kilobytes out of a multi-gigabyte body, measures, then DISCARDSNo byte custody, but knowledge earned from bytes · raw bytes stay at the source · orders of magnitude less data movedTHE UNDERSTANDING LADDER · DEPTH IS EARNED, RUNG BY RUNGeach rung is earned from the one below it, and never assertedL0DECLAREDwhat the publisher said:where metadata search stopsL1MEASUREDwe read the bytes: distributions,quantiles, null rate, real shapeL2GRAMMARRoleMap: anchor · link · axis ·signal · and the grain of a rowL3MEANINGfrom evidence only; the citedgrounding refs are the firewallnamed, and not climbed yetL4 identity-in-timeL5 capabilitywhat MEASURED means: real shape, read from your bytesEVIDENCE FAMILIES · ONE GRAPHRESEARCH DATASETSmeasured: we read the bytes.shape, schema, quality, sizeSCHOLARLY WORKSthe literature, and what it cites:declared links to its dataPATENTStitle and abstract, searched:where the science became appliedTHE TANGENTIAL HOPkeep a paper, and the work citing itsurfaces cross-field measured dataTEN TOPOLOGIESone closed vocabulary · a strategy is written once per topology, never once per formattabulartensorspatialsequencegraphstructurestreammodalcompositeunknownthe probe sorts you into one, then hands the loader you actually need: pandas for a table · xarray for a tensor · geopandas for a rasterSOURCES · FEDERATED, THEN MEASURED WHERE THE BYTES ALLOWZenodoPANGAEARCSB PDBHugging FaceDryadHarvard DataverseNOAA ERDDAPFigshareICPSRa 50-source landscape, mapped and gradedTHE PLANES UNDERNEATHrelational + graph storehybrid lexical + vector indexcolumnar corpus storeout-of-band batch jobsgraph APIREST10 typed agent toolsINTEROPERABILITY · EMIT EVERY DIALECTCroissant spine: the dialect agents and ML tooling already readDCAT-AP and schema.org from the same measured record · FAIRSHACL shape graphs, conformance-tested in CI, not assertedOne contract set, three dialects: graph API, REST, agent toolsNever replaces a repository: it sits on top and federatesTHE PROVENANCE SPINEEvery record states HOW it is known: MEASURED (we opened the file) or DECLARED (the source said so)The grade rides the EDGE, never the node · no file URL, no measurement: it stays declared, and says soMEASURED · we read the bytesDECLARED · the source said soEXTRACTED · pulled from proseTHE CONSORTIUMAI Allianceconvenes the consortiumBrightQuerythe measurement engineOpenMinedanalysis where data cannot moveMLCommonsthe Croissant metadata spineHugging Facethe deployment edgeMERIDIAN IN THE 26-509 DATA LIFECYCLEcoverage stated, as the solicitation asks01 acquisition02 transfer03 management04 exploration05 analysis06 curation07 sharing08 synthesis09 discovery10 archivingprimary · Meridian owns thisassistiveobservational · we read it, we do not own itdelegated · the repositories, their stewards, the researcherMeridian is open source (Apache-2.0) · built by BrightQuery within the AI Alliance

The return

You are here: cited in someone's answer

You began at the top as 4.2 GB nobody had opened. You went all the way down to your own bytes and came back up measured, loadable, linked to the papers that cite you, and quoted inside a grounded answer with your grade attached. That is the round trip: a question falls to the data and rises as evidence someone can trust. Then the next question arrives, and it starts again.

you are here: measured, loadable, cited

Meridian - the measured-data spine beneath a grounded research workspace.