Historical coverage is not a single start date. A dataset can cover many seasons while missing the market, bookmaker, observation frequency, or point-in-time fields your analysis requires. Start with the question you want to answer, then check whether the archive can observe it.
Choose the observation you need
| Research question | Data requirement | What a coarse snapshot cannot establish |
|---|---|---|
| How did the prematch price change over a day? | Repeated comparable snapshots with actual observation times | Every intermediate price change |
| What was observable before kickoff? | A snapshot selected using a strict pre-cutoff rule | Availability at your intended decision time if the snapshot is later |
| Which source updated first? | Comparable event-level update observations and trustworthy clocks | Millisecond ordering from five-minute samples |
| How did market availability vary? | Presence, suspension, and missingness records | That a missing row was never offered |
| Did a model generalize? | Point-in-time inputs, outcomes, and versioned evaluation splits | Valid results if later corrections leak into earlier inputs |
These are methodological requirements. They do not imply that any single feed supplies all of them.
Read snapshot semantics literally
An archive describes itself in snapshots: a start date, a sampling interval, and a rule for which snapshot answers a query for a given time. Read each one literally. A query that returns the closest snapshot at or before the requested time behaves differently from one that interpolates, and coverage for a sport, market, or bookmaker usually starts when that item was added to the archive, not at the archive’s first date. Treat stated completeness as a vendor claim until you have audited a sample.
The Pinnacle feed this site covers does not sell an archive. Its REST drops endpoint keeps recent drops for up to about three hours, each with its own timestamp, and polling its markets endpoint with a since cursor records every change from the day you start. If your research can start today, your own collector is an archive whose sampling rule you control.
For any archive, store both the requested timestamp and the returned snapshot timestamp. Repeated requests a minute apart can resolve to the same snapshot; count unique observations rather than treating each request as fresh data.
Check whether an update time belongs to the bookmaker, market, selection, or archive job. A snapshot may contain records with different underlying ages. Preserve the timezone and original precision. RFC 3339 defines a useful timestamp representation, but does not tell you what the feed’s clock measures. RFC 3339.
Write the cutoff rule before sampling
For a study of prices before an event, define the cutoff using the event time known at that point. An event’s final corrected kickoff time may differ from what was scheduled earlier. Keep schedule revisions if the research depends on them.
A defensible extraction rule could be: select the latest eligible observation received no later than a stated cutoff, within a stated maximum age. That is a proposed rule; the age tolerance depends on the question. Report how many events have no eligible observation rather than reaching forward in time to fill them.
Define “closing price” explicitly. Last observed before kickoff, last unsuspended prematch quote, and a feed’s designated closing record may produce different samples. Use one definition consistently and show its limitations.
Audit missingness and revisions
Sample several dates, event types, and market states before buying a large extraction. Count expected events, matched events, usable markets, unknown settlement rules, and missing selections. Keep a reason code where the feed supplies one; do not invent a reason from absence alone.
Ask whether historical corrections overwrite old records, whether original versions are retained, and how canceled or rescheduled events are represented. A corrected archive may help descriptive research while failing to reproduce what a live collector would have seen at the time.
Avoid retaining only completed events with complete prices. That makes the dataset easier to analyze while potentially discarding the exact failures your application must tolerate. Report exclusion counts and whether they differ across feeds or periods.
Estimate extraction and storage
Compute how many unique snapshots the schedule requires, then apply the endpoint’s billing rule. Add requests for pagination or event-specific markets only where required. Reuse an already downloaded permitted snapshot rather than paying to retrieve it repeatedly.
Keep immutable source provenance, extraction time, request scope, mapper version, and a checksum of retained files. Store only what your license permits, and distinguish private analysis from public redistribution or publication of raw rows. API access alone does not answer those permission questions.
Use the cost calculator for volume modeling and the normalization guide for compatible market keys. If you plan to publish findings, prerecord the extraction and exclusion rules using the benchmark methodology.
Sources 2 references
Primary documentation used for this guide. Check dates refer to source review.
- Pinnacle data technical API referenceChecked 26 Sept 2026
- RFC 3339: Date and Time on the InternetChecked 26 Sept 2026