Datasets

Upload historical ticker data and use it in the standard backtesting workflow.

Backtest against a CSV, parquet, or lastra file you upload instead of a managed exchange: create a dataset, PUT the file to a presigned URL, finalize it to trigger ingest, then prepare/execute exactly as normal but with the reserved exchangeId: user.

MethodPathPurpose
POST/datasetsCreate a dataset + first upload session
GET/datasetsList your datasets
GET/datasets/{datasetId}Get one
DELETE/datasets/{datasetId}Delete
POST/datasets/{datasetId}/uploadsOpen a new upload session for an existing dataset
POST/datasets/{datasetId}/uploads/{uploadId}/finalizeTrigger ingest
GET/datasets/{datasetId}/uploads/{uploadId}Poll upload/ingest state
POST/datasets/importsCreate a dataset by fetching history instead of uploading it
GET/datasets/{datasetId}/imports/{importId}Poll fetch/ingest state

v1 is ticker data only — type is always "ticker". instrument must be a plain spot pair (BASE/QUOTE, exactly one /); derivative forms (BTC/USDT:USDT) are rejected.

Creating a dataset

POST /datasets — creates the dataset and its first upload session in one call: a presigned URL your client PUTs the file to directly, no API credentials involved in that PUT.

FieldTypeNotes
namestringrequired, unique among your datasets. 409 if already taken
instrumentstringrequired, plain spot pair
curl -X POST https://api.qtsurfer.net/v1/datasets 
  -H "Authorization: Bearer $TOKEN" 
  -H "Content-Type: application/json" 
  -d '{"name":"My BTC ticks","instrument":"BTC/USDT"}'

DatasetCreated (201) — the metadata available immediately after creation plus the upload session. It is not yet the full Dataset: lifecycle fields such as createdAt, currentVersionId, range, and cadence are obtained from GET /datasets/{datasetId} after the relevant lifecycle stages.

{
  "datasetId": "ds_3f9a1c2e7b0d4a5f", "name": "My BTC ticks",
  "type": "ticker", "instrument": "BTC/USDT",
  "uploadId": "up_1a2b3c4d5e6f7a8b",
  "upload": {
    "url": "https://storage.qtsurfer.com/.../uploads/up_1a2b3c4d5e6f7a8b/raw.csv?X-Amz-...",
    "expiresInMinutes": 15
  }
}

uploadId is what you pass to finalize; upload.url is the presigned target — PUT the raw file there directly, no Authorization header.

Lost this response? Nothing lost — call POST .../uploads on this dataset’s id and you get the very same upload session back, as long as you haven’t finalized it yet.

Errors: 400 invalid request, or instrument isn’t a plain spot pair · 409 dataset name already taken · 429 your tier’s dataset count limit is reached — delete one, or upgrade.

Opening a new upload session

POST /datasets/{datasetId}/uploads — get a fresh upload session for a dataset you already have: a corrected file, or the next chunk of history. Same idempotency contract as POST /datasets’s own upload half: at most one session is open per dataset at a time, so calling this again before finalizing just hands back that same session — safe to retry if a response gets lost. Once a session has been finalized (successfully or not), the next call here opens a genuinely new one.

curl -X POST https://api.qtsurfer.net/v1/datasets/$DATASET_ID/uploads 
  -H "Authorization: Bearer $TOKEN"

201 — the same {uploadId, upload} shape POST /datasets returns, without the dataset metadata around it:

{
  "uploadId": "up_1a2b3c4d5e6f7a8b",
  "upload": {
    "url": "https://storage.qtsurfer.com/.../uploads/up_1a2b3c4d5e6f7a8b/raw.csv?X-Amz-...",
    "expiresInMinutes": 15
  }
}

Errors: 404 no such dataset for this user.

Uploading the file

CSV, parquet, or lastra. A CSV needs a header row; a parquet file carries its columns by name already; a lastra file is our own native columnar format — the same one a dataset’s dataUrl hands you back by default, so downloading a dataset and handing that exact file to another user to upload works with no conversion in between. For CSV/parquet, timestamp (ISO-8601, or numeric epoch seconds / millis / micros — detected from the first row, then enforced for every later row) and close are required; optional: open, high, low, volume, quoteVolume, bid, bidSize, ask, askSize. A lastra upload carries its own fixed column set instead (it’s already a dataUrl download, not a format you construct by hand) and only needs a timestamp series and a close series present. Cadence and timestamp unit are discovered from the data, not declared, for all three.

A CSV upload is converted to our native columnar format (lastra) for storage. A parquet or lastra upload is stored as-is today. Either way, check dataFormat on the ready version for which one you actually get back — don’t assume it from how you uploaded it (a converted CSV and an uploaded lastra file both report dataFormat: "lastra").

The bytes PUT to upload.url may be that file directly, gzipped (.gz), or zipped (.zip, exactly one file inside — a dataset is one file regardless of how it travels). Format is detected from the content itself: there is no filename or Content-Type anywhere in this flow for a client to declare it with, so nothing needs to be sent besides the bytes.

curl -X PUT "$UPLOAD_URL" --data-binary @my-btc-ticks.csv
# or gzip/zip it first -- detected from content, no extra parameter needed
curl -X PUT "$UPLOAD_URL" --data-binary @my-btc-ticks.csv.gz

Finalizing an upload (triggering ingest)

POST /datasets/{datasetId}/uploads/{uploadId}/finalize — call once the PUT above completes. Enqueues ingest and returns immediately; poll GET .../uploads/{uploadId} below. Idempotent while the upload is still open — a repeat finalize before it has produced a version returns the same jobId rather than enqueueing a second ingest. Once it HAS produced a version, uploadId is spent: finalizing it again is a 409, even with different bytes freshly PUT to the same URL — open a new upload session instead of reusing a spent one.

curl -X POST https://api.qtsurfer.net/v1/datasets/$DATASET_ID/uploads/$UPLOAD_ID/finalize 
  -H "Authorization: Bearer $TOKEN"
# → 202 {"jobId": "dataset-upload:.../ds_3f9a1c2e7b0d4a5f:up_1a2b3c4d5e6f7a8b"}

Errors: 404 no such dataset for this user; uploadId wasn’t issued for this dataset (never minted, or minted for a different one); or nothing was PUT to upload.url yet — a finalize with nothing to finalize · 409 uploadId already produced a version (the error message names it) · 413 the uploaded file is many times your tier’s size limit for a dataset (see Size limits: the limit itself applies to the stored size, which is only known after conversion) · 429 your account’s total storage limit (GET /account’s maxTotalStorageBytes) is reached or would be exceeded — delete a dataset to free space, or upgrade.

Size limits

maxDatasetBytes (GET /account) caps a dataset version’s stored size: the bytes its ready version reports, which for a CSV upload is the converted lastra file, not the file you uploaded. Estimate from rows, not from the size of the file. The stored size follows the rows and the columns, and a CSV can come out larger or smaller than the file. Measured on synthetic one-second data: about 92 bytes per row for a timestamp,close CSV, about 55 for timestamp,open,high,low,close,volume. Real data compresses differently, so take these as an order of magnitude and, for a big file, upload a small slice first and scale from its bytes and rows.

The stored size is known only once the file has been converted, so that is where the limit decides: POST .../finalize answers 202, and an upload over the limit ends failed when you poll it, with an error such as Dataset is 196976000 bytes, exceeds the tier's 100000000 byte limit. Poll rather than assume the 202 means it stored. What finalize itself refuses with 413 is a file many times the limit, well beyond anything conversion could bring under it.

Polling ingest

GET /datasets/{datasetId}/uploads/{uploadId} — poll after finalize until status is ready or failed. Also reports uploading (finalize not called yet, but the file was PUT) before you finalize at all. Durably recorded once a version exists, so ready/failed are permanent answers; uploading/ingesting reflect in-flight state that can itself age out (see the 404 case below).

Response — DatasetUploadState

FieldNotes
statusuploading (file PUT, not finalized yet) → ingesting (finalize called, worker parsing/validating) → ready (version carries the result) | failed (e.g. bad CSV contract, mixed timestamp units, a .zip with no file inside or more than one)
jobIdthe ingest job id, while status is ingesting
errorhuman-readable reason, present when status is failed. Durably recorded alongside the failure — stays available however long after the fact you poll
versiona DatasetVersion, present when status is ready or failed

DatasetVersion — one successfully ingested upload

FieldNotes
idthe version id — pass as datasetVersionId on prepare to pin it
bytessize of the stored file (dataUrl) — a converted lastra for a CSV upload (decompressed first, if it arrived as .gz/.zip), or the parquet/lastra file itself, unconverted, for a parquet or lastra upload. Not the size of the bytes originally PUT
rowsnumber of data rows
cadencediscovered from the data’s own timestamps: a fixed grid (1s, 5s, 15s, 1m, 5m, 15m, 30m, 1h, 4h, 1d) when at least half the intervals between rows fall on that step (small clock jitter tolerated), or rt — native data at the rate it was captured, each row at its own timestamp with no fixed step (per-trade on-chain swaps, block-spaced or sub-second ticks, irregular intervals). An rt dataset can be resampled to any fixed cadence at prepare time
timestampUnitiso | s | ms | us — the unit the timestamp column arrived in
gaps, largestGapStepsgap count at the discovered cadence, and the largest one’s size in cadence steps. Always 0 for rt
dataUrlpresigned GET URL to the stored file — see dataFormat. Present once ready
dataFormatlastra (either converted, from a CSV/gzip/zip upload, or unconverted, from a lastra upload — the value alone doesn’t tell you which) | parquet (unconverted, from a parquet upload)
curl https://api.qtsurfer.net/v1/datasets/$DATASET_ID/uploads/$UPLOAD_ID 
  -H "Authorization: Bearer $TOKEN"
{
  "uploadId": "up_1a2b3c4d5e6f7a8b",
  "status": "ready",
  "version": {
    "datasetId": "ds_3f9a1c2e7b0d4a5f", "id": "dsv_8e2b4f19c6a03d7e",
    "bytes": 4831022, "rows": 86400, "cadence": "1s",
    "timestampUnit": "iso", "gaps": 0, "largestGapSteps": 0,
    "dataUrl": "https://storage.qtsurfer.com/.../dsv_8e2b4f19c6a03d7e/ticker_BTC_USDT_....lastra?X-Amz-...",
    "dataFormat": "lastra"
  }
}

Errors: 404 no such dataset for this user, or genuinely nothing known about this uploadId — no version, no in-flight job, nothing was ever PUT to its upload URL.

Importing a dataset instead of uploading one

POST /datasets/imports — a second way to get data into a dataset: instead of PUTting a file yourself, ask the API to go fetch history on your behalf. Creates the dataset and starts the fetch in the same call — there’s no separate upload step, and the result lands as a dataset version indistinguishable from an uploaded one once it’s ready.

type selects the source. dex — history over a pool/pair’s own on-chain market — is the only value today; other source types join this same endpoint later. A dex import has two data shapes, chosen by the top-level cadence:

  • Omitted/blank (default) — on-chain swap history, replayed directly from the pool/pair’s own chain, each swap at its own timestamp (native per-trade cadence).
  • 1s | 1m | 5m — pre-aggregated candles at that width instead of raw trades. The resulting dataset’s type is klines, not ticker.
FieldTypeNotes
namestringrequired, unique among your datasets. 409 if already taken
instrumentstringrequired, plain spot pair — the dataset’s own label, independent of the pool’s on-chain token order
from, tostring (date-time)required, ISO-8601 UTC. from inclusive, to exclusive, from < to. Total span is capped by your tier
cadencestringoptional. Omitted/blank = native per-trade cadence (see below). One of 1s | 1m | 5m asks for pre-aggregated candles at that width on a network that supports it (any other value is 400) — on a network that doesn’t, it’s silently ignored, see below
typestringrequired, "dex" is the only value today
dex.networkstringrequired, one of ethereum | robinhood
dex.idstringrequired unless a candle cadence actually resolves (see below) — a candle cadence ignored on a non-candle network still requires this. "uniswap" is the only value today — which on-chain DEX protocol dex.contract implements
dex.versionstringrequired unless a candle cadence actually resolves (see below). "v2" | "v3"
dex.contractstringrequired, the pool (v3) or pair (v2) contract address
dex.factorystringoptional — omit to auto-discover on-chain from contract; supply only if you already know it or the pool/pair belongs to a non-canonical factory. Either way the pool/pair is validated against whichever factory is used before anything is fetched. Ignored if a candle cadence actually resolves (see below)

On-chain cadence is native, not resampled. A plain dex import (no cadence) keeps the source’s own per-trade event cadence — each swap at the timestamp it happened, so the resulting version’s cadence is rt unless the swaps happen to sit on a fixed grid — rather than bucketing into candles; resample to a coarser cadence afterward as a separate step if you need one from on-chain data. Asking for cadence: "1s"/"1m"/"5m" instead gets you pre-aggregated candles at that width directly — but only on a network that actually has a candle source behind it. On any other network a candle cadence is silently ignored and the import proceeds as if cadence were never sent (native, dex.id/dex.version required) — it does not fail, sync or async. Which networks support candles today isn’t part of this contract and may change; if you need to know before importing, request the native cadence and check the resulting version’s own cadence instead of assuming.

curl -X POST https://api.qtsurfer.net/v1/datasets/imports 
  -H "Authorization: Bearer $TOKEN" 
  -H "Content-Type: application/json" 
  -d '{
    "name": "weth-usdc-week",
    "instrument": "WETH/USDC",
    "from": "2026-08-01T00:00:00Z",
    "to": "2026-08-08T00:00:00Z",
    "type": "dex",
    "dex": {
      "network": "ethereum",
      "id": "uniswap",
      "version": "v3",
      "contract": "0x88e6a0c2ddd26feeb64f039a2c41296fcb3f5640"
    }
  }'
# → 202 {"datasetId":"ds_3f9a1c2e7b0d4a5f","importId":"imp_01j9z...","jobId":"dataset-import:...","status":"fetching"}

Or, for pre-aggregated candles instead of raw on-chain swaps:

curl -X POST https://api.qtsurfer.net/v1/datasets/imports 
  -H "Authorization: Bearer $TOKEN" 
  -H "Content-Type: application/json" 
  -d '{
    "name": "weth-usdc-1s",
    "instrument": "WETH/USDC",
    "from": "2026-08-01T00:00:00Z",
    "to": "2026-08-01T06:00:00Z",
    "cadence": "1s",
    "type": "dex",
    "dex": {
      "network": "ethereum",
      "contract": "0x88e6a0c2ddd26feeb64f039a2c41296fcb3f5640"
    }
  }'

importId is what you poll with, below — there’s no separate “finalize” step the way an upload has.

Errors: 400 invalid request, instrument isn’t a plain spot pair, from >= to, cadence present but not one of its supported values, the range exceeds your tier’s import ceiling, the range’s rough size estimate exceeds your tier’s row limit, dex.network/dex.id not one of their supported values, dex.contract/dex.factory fail basic shape validation, or (when cadence is omitted) dex.id/dex.version missing (whether the pool/pair actually resolves, and for a candle cadence whether that combination is servable on the requested network, is checked later, asynchronously — see failed below) · 409 dataset name already taken · 429 your tier’s dataset count limit is reached, or your account’s total storage limit (GET /account’s maxTotalStorageBytes) is already reached — delete a dataset to free a slot or space, or upgrade.

Polling an import

GET /datasets/{datasetId}/imports/{importId} — poll after POST /datasets/imports until status is ready or failed. An import spends real time fetching from its source before anything is even staged; once fetched, it re-enters the exact same ingest chain an upload uses.

Response — DatasetImportState

FieldNotes
statusfetching (reading from the source, nothing staged yet — the one status only an import ever reports) → ingesting (fetched, staged, worker parsing/validating) → ready (version carries the result) | failed
jobIdthe fetch/ingest job id, while status is fetching or ingesting
errorhuman-readable reason, present when status is failed — an unresolvable pool/pair, no data in the requested range, a range older than the source retains, the fetch exceeding your tier’s time ceiling, or any of the ingest-side reasons DatasetUploadState.error can carry, once fetching hands off to that same chain. Durably recorded, same as on the upload path
versiona DatasetVersion, present when status is ready or failed
curl https://api.qtsurfer.net/v1/datasets/$DATASET_ID/imports/$IMPORT_ID 
  -H "Authorization: Bearer $TOKEN"
{
  "importId": "imp_01j9z1x2y3z4a5b6c7d8e9f0g1",
  "status": "ready",
  "version": {
    "datasetId": "ds_3f9a1c2e7b0d4a5f", "id": "dsv_8e2b4f19c6a03d7e",
    "bytes": 4831022, "rows": 604800, "cadence": "rt",
    "timestampUnit": "us", "gaps": 0, "largestGapSteps": 0,
    "dataUrl": "https://storage.qtsurfer.com/.../dsv_8e2b4f19c6a03d7e/ticker_WETH_USDC_....lastra?X-Amz-...",
    "dataFormat": "lastra"
  }
}

Errors: 404 no such dataset for this user, or genuinely nothing known about this importId.

Dataset shape

Both GET /datasets and GET /datasets/{datasetId} return this — from/to/cadence/timestampUnit mirror the current version’s own discovered range, cadence and timestamp unit, and status, bytes, rows, gaps and largestGapSteps say whether it is usable and how big it is, so you don’t need a second call to see what a dataset covers.

FieldNotes
datasetId, name, type ("ticker" | "klines"), instrument, createdAt, statusalways present. type is "klines" only for a dex import that requested a candle cadence; "ticker" for everything else (uploads, and native-cadence dex imports)
currentVersionIdthe most recently finalized, successfully ingested version. Absent until at least one upload has finished ingesting
statusready — currentVersionId is set and the fields below describe it; failed — the most recent upload or import attempt failed, there is nothing to read yet (see error); pending — nothing has ever been attempted (just created, or an upload was never finalized)
updatedAtwhen currentVersionId last changed; absent until it has a value
from, to, cadence, timestampUnitthe current version’s own range/cadence/timestamp unit (a fixed grid or rt cadence, iso|s|ms|us for timestampUnit, see DatasetVersion), as discovered at ingest. Absent until a version exists
bytes, rows, gaps, largestGapStepsthe current version’s own stored size, row count, gap count and largest gap (in steps of its cadence). Present only when status is ready — see DatasetVersion for what bytes measures
errorwhy the most recent attempt failed. Present only when status is failed

GET /datasets/{datasetId} alone adds dataUrl/dataFormat (same meaning as on DatasetVersion) once the current version is ready, plus _links.self. The bulk listing never mints these — a presigned download URL for every dataset on a screen that renders no chart isn’t worth the exposure.

Listing your datasets

GET /datasets — every dataset you’ve created and not deleted, most recently created first. Never a 404 — an empty array if you have none, same convention as GET /strategies.

curl https://api.qtsurfer.net/v1/datasets -H "Authorization: Bearer $TOKEN"

GET /datasets?includeDeleted=true also lists the datasets you’ve deleted, each with the deletedAt it was deleted at — handy when keeping your own copy of the list in sync, so a deleted dataset shows up as deleted instead of just disappearing.

Getting a dataset

GET /datasets/{datasetId} — a Dataset plus a _links.self.

curl https://api.qtsurfer.net/v1/datasets/$DATASET_ID -H "Authorization: Bearer $TOKEN"

Errors: 404 no such dataset for this user.

Deleting a dataset

DELETE /datasets/{datasetId} — soft delete. Stops appearing in the list/get endpoints and can no longer be prepared from, but its object data is reclaimed later rather than purged inline, so a backtest already running against one of its versions isn’t disrupted.

curl -X DELETE https://api.qtsurfer.net/v1/datasets/$DATASET_ID -H "Authorization: Bearer $TOKEN"
# → {"datasetId": "ds_3f9a1c2e7b0d4a5f", "deleted": true}

Errors: 404 no such dataset for this user, or already deleted.

Backtesting against a dataset

Once a version is ready, prepare and execute exactly as against a managed exchange, but with exchangeId: user and datasetId in place of instrument:

curl -X POST https://api.qtsurfer.net/v1/backtest/user/ticker/prepare 
  -H "Authorization: Bearer $TOKEN" 
  -H "Content-Type: application/json" 
  -d '{"datasetId":"ds_3f9a1c2e7b0d4a5f","from":"2026-03-14","to":"2026-03-15"}'
# → 202 {"jobId":"5ikYAMIO...","datasetId":"ds_3f9a1c2e7b0d4a5f","datasetVersionId":"dsv_8e2b4f19c6a03d7e"}

execute is unchanged — same request body as against a managed exchange, since the instrument and range are recovered from prepareJobId either way. See docs/backtest_execute.md for the full prepare/execute reference, including PrepareRequest’s datasetId/datasetVersionId fields and the dataset-backed coverage shape on PrepareJobState (cadence/gaps/largestGapSteps in place of the hour-walked totalHours/hoursWithData/hoursWithoutData).