See how we're accelerating scientific progress by making the world's scientific data accessible for humans and AI.

We rebuilt data storage

problem

Most scientific data sits in opaque binary formats, spread over millions of individual files, with no index and no map. Querying it in the cloud means downloading each one in full.

solution

Icechunk represents that data as infinitely extensible multidimensional arrays — chunked on any axis, compressed, and versioned. Zero-copy ingest turns NetCDF, HDF5, GRIB, and TIFF archives into structured data cubes, and ACID transactions make them safe for operational workloads.

When NASA implemented Icechunk, they saw a 100x speedup. And it's 100% open source (Apache 2.0).

explore icechunk
Icechunk architecture: a repository with snapshot history, branches and tags, over a hierarchy of groups, arrays, and chunks

We bring governance to scientific data

problem

At most organizations, scientific data is essentially invisible, living in random buckets with no discoverability or governance.

solution

Arraylake is the metadata and governance layer of the platform: catalog, search, ownership, permissions, and complete version history over the data you already have.

Bring your own bucket, or use our managed storage. Arraylake works with every public cloud and with S3-compatible on-prem object storage.

explore arraylake
The Arraylake web app: an organization's catalog of Icechunk repositories, storage, and services

We bring your compute to the data

problem

At the petabyte scale, moving data is the slowest and most expensive thing you can do — and every consumer of a dataset expects it in a different shape, over a different protocol.

solution

Compute runs your workloads right next to the data, in the same cloud region, on hardware that scales out to match the size of the problem.

Turn on a service and any repository is queryable over the standards your tools already speak — map tiles, EDR, DAP2, openEO — with only the result crossing the network.

explore compute

We connect the people who create data to the people who use it

problem

Integrating a new dataset means months building a pipeline to copy it into your own systems — and then maintaining that pipeline forever. Now multiply that by every dataset and every consumer.

Meanwhile, data providers build and maintain byzantine data portals and delivery architectures.

solution

What if they could just exchange clean, analysis-ready data in the cloud with the click of a button? That's the Earthmover data marketplace.

explore the marketplace
The Earthmover data marketplace: analysis-ready scientific datasets

Made for humans and agents

Every layer of the platform is reachable the way you already work. A CLI for the terminal and CI, Python and Rust clients for everything else, and data that arrives as plain Xarray, NumPy, or Zarr — so a repository drops into a Jupyter or Marimo notebook, or straight into a PyTorch training loop, without a conversion step.

Agents get first-class access to the same surface. Our hosted MCP server lets an assistant discover repositories, read schemas, and query live data, and agent skills teach it to use the platform correctly instead of guessing at an API.

CLIPythonRustMCPSkills
explore MCP and skills

driven by

Claude
OpenAI
Cursor

ready for

JupyterMarimoXarrayNumPyZarrPyTorch
~/forecasts
$ 

ready to explore earthmover?

Get started now

Want to learn more?

Talk to us