Market data infrastructure for quants

Point at a file.
Get a research‑ready dataset.

Map the columns, choose who reads and writes — MD Store lands, normalizes, documents and secures your market data. No pipelines to build, no infra to babysit.

Early access for research teams. No spam.

mdstore.app / new dataset

Point at your data

s3://desk-raw/binance/trades_2024-03.parquet

✓ Parquet · 38.1M rows · 5 columns detected

ts_mspxqtysymside
170925120012361234.50.012BTCUSDTb
170925120013161234.40.300BTCUSDTs

Map columns to the MD Store schema

  • ts_mstimestamp ms → UTC ns
  • pxprice
  • qtysize
  • syminstrument → BTC-USDT
  • sideside b/s → buy/sell

Dataset

trades.binance.spot

Readers

@quant-research@pm-desk+ add

Writers

@data-eng+ add

Landing…

  • Landed raw → raw/binance/trades
  • Normalized 38.1M rows → clean.trades
  • Deduped 1,204 rows · flagged 3 gaps
  • Data card & lineage generated
  • Granted read ×2 · write ×1

Dataset ready — md.load("trades.binance.spot")

step 1/4

How it works

Three inputs from you. Everything else is automatic.

01

Point at a file

S3, GCS, SFTP, a vendor feed or a local folder. CSV, Parquet, gz — schema is detected for you.

02

Map the columns

Match source columns to one canonical schema. Units, timezones and symbols convert on the fly.

03

Choose roles

Pick readers and writers. Permissions are applied everywhere the data lands.

→ MD Store

Lands it right

Raw kept immutable, clean layer validated, docs and lineage generated, access granted.

Why MD Store

Everything a data team would build.
Already built.

01

Build nothing

No pipelines, no infra, no cleaning scripts.

  • storage & partitioning
  • schemas & versioning
  • schedules & backfills
  • access control

Zero time on infrastructure

Storage, schemas, schedules and permissions are provisioned for you.

03/01/24 09:00 ESTXBTUSD61,234.5
2024-03-01T14:00ZBTC-USD61234.50

Normalization on autopilot

UTC timestamps, one symbol convention, aligned units and sides.

trades.binance.spotcrypto · tick2019 → today@data-eng
trades.cme.esfutures · tick2012 → today@data-eng
trades.okx.perpcrypto · tick2021 → today@quant-research

A data catalog, built in

Every dataset searchable by market, frequency, coverage and owner. Find it before you rebuild it.

02

Know everything

Every dataset explains itself.

trades.binance.spot Tick-level spot trades for microstructure research. @data-eng · 2019 → today · 1.2B rows

Docs that write themselves

A data card per dataset: schema, coverage, owner and the idea behind it.

raw.binance raw.okx bars_1m vol_20d

Lineage for every number

Trace any feature back to its files and transforms. Reproduce any backtest.

03

Research faster

Clean data, one line away.

featuresresearch-ready
cleanvalidated
rawimmutable

A real feature store

Raw data and research features live apart, not in one pile on a shared disk.

import mdstore as md

df = md.load("trades.binance.spot",
             start="2024-03-01")

A Python library for access

Load any dataset into pandas or polars in one line. No paths, no credentials juggling.

trades.binance.spot38.1M rows

timestamppricesizeside
14:00:00.12361234.500.012buy
14:00:00.13161234.400.300sell

A data viewer on the web

Browse schema, sample rows and stats right on the site, before writing any code.

gap

Visualizations out of the box

Coverage, gaps and distributions, ready in the viewer or via md.plot() in Jupyter.

Stop cleaning data. Start researching.

Join the waitlist for early access.