Data dictionary
One page per table: what it holds, every column with its meaning, and which regions carry it, generated from the schema contracts our write path enforces, never hand-edited. data-dictionary.json is the machine-readable twin.
The API and MCP are Pro and Enterprise capabilities; see the API documentation for what the API serves.
public_aggregate 7 tables
Public. These rows are the public layer: aggregate series and reference records such as companies, formations and rig counts. Where a page publishes them you can read them on the website, and a free account can export the aggregate series behind a view (10 a month).
gold_formation_name_mappings- How reported formation names resolve to registry units. The same rock is reported under many names and spellings; each row maps one normalized reported name, within one basin, to its formation_id: the basin scope keeps same-named units in different basins apart. Apply it to your own data to get the same formation resolution we use.
gold_formations- Canonical registry of geological formations: the named rock units wells target and produce from. One row per stratigraphic unit (group, formation, member, or bed, plus broad age buckets where a source reports only a geological age), with its place in the hierarchy and stable links into public stratigraphic references. formation_id columns across our well, completion, and formation-top tables reference this registry, so a name like Wolfcamp means the same unit everywhere.
gold_organization_events- Corporate events that change who owns or operates assets: acquisitions, mergers, name changes, scoped asset transfers, and identity corrections. One row per event with the organizations involved and the date it took effect. This drives ownership-over-time attribution: activity before the effective date stays with the original company, after it with the successor.
gold_organization_name_mappings- How reported operator names resolve to organizations. Regulators publish the same company under many spellings; each row maps one normalized reported name, within one region, to its organization_id. This is the exact lookup our own attribution uses. Apply it to your raw data to get the same resolution.
gold_organization_registrations- The bridge between regulator-issued operator IDs (Texas P-5, New Mexico OGRID, and each region's equivalent) and our canonical company register: normally one row per (region, registration ID), carrying the name as registered, registration status, role flags (operator / gatherer / midstream), and headquarters as filed. organization_id links to gold_organizations once the operator is resolved. Where the company behind an ID was renamed, the ID carries one row per period instead (see valid_from / valid_to) so that filings made before and after the rename are attributed to the right company.
gold_organizations- Master register of the companies behind the wells: operators, service companies, midstream, integrated and downstream groups. One row per organization, deduplicated across regions and name variants; operator id columns across the well, production, and permit tables reference it by organization_id. Where a company was acquired, merged, or renamed, successor_id points to the organization that carries on.
gold_rig_activity- Active drilling-rig counts, a leading indicator of upcoming supply. Parallel series from independent published sources, each at its native breakdown: the Baker Hughes weekly North America count and Mexico's monthly regulator series. Series are never mixed or reconciled, so each row matches the figure its source publishes. Where a source does not break down an axis, that column carries `all`: aggregate only within one source and one period type.
portal 19 tables
Portal. These rows are well-level or producing-unit-level data, or our modelled output at cohort or grid grain. They open with a Portal, Pro or Enterprise plan. The website shows aggregates built from them, never the rows.
gold_directional_surveys- Full wellbore trajectories: one row per survey station (a measured-depth reading along a wellbore leg), with inclination, azimuth, true vertical depth, and position in WGS84. Every filed reading is kept: planned (proposed) surveys as well as as-drilled, and superseded filings alongside the latest, so you can compare plan against actual. For the current as-drilled trajectory, filter is_proposed = false and is_latest_version = true. This is the table for accurate lateral placement, well spacing, and landing-zone work.
gold_organization_event_wells- The specific wells covered by a scoped organization event. When an asset transfer moves only part of a company's assets, this holds the frozen set of wells the deal covered as of its effective date. Most events apply to whole organizations and need no well list, so this table is small and may be empty.
gold_producing_units- The register of producing units coarser than a single well (an FPSO, an oil-sands project, a field, or a whole country), used where a source reports production at that level rather than per well. Each unit carries its grain, jurisdiction, and operator. Volumes live in gold_unit_production; the gold.production view unifies well-grain and unit-grain series into one all-grain read surface.
gold_unit_disposition- Monthly gas and oil disposition for producing units coarser than a single well (an FPSO, a project, a field, or a whole country), one row per unit per month. This is the companion of the well-grain disposition table for sources that report at coarser grain: where the produced volume went, split into sold, flared, vented, reinjected, and fuel. The unit's name, grain, and jurisdiction live on the producing-units register (gold_producing_units); join monthly unit production for the produced denominator. NULL means the source didn't report that split; 0 is a reported zero. Each disposition column carries its own paired _origin column.
gold_unit_production- Monthly production volumes for producing units coarser than a single well (an FPSO, a project, a field, or a whole country), one row per unit per month. This is the companion of the well-grain production table for sources that report at coarser grain; the unit's name, grain, and jurisdiction live on the producing-units register (gold_producing_units). A volume is NULL when that stream isn't reported at this unit's grain (e.g. gas metered nationally while oil is per-facility): NULL means no series at this grain, never zero. The gold.production view unifies this table with well-grain production into one all-grain surface.
gold_well_completions- One row per completion event (the initial completion and each later recompletion) as filed with the regulator. Carries the completion date, the wellbore configuration produced by the event (formation, depths, lateral length, trajectory), and interval details. Plug-and-abandon events are not completions; they live on the well register (plugged_date).
gold_well_disposition- Monthly gas and oil disposition per well: where the produced volume went (sold, flared, vented, reinjected, used as lease or plant fuel), one row per well per month. NULL means the source didn't report that split; 0 is a reported zero. Each disposition column carries its own paired _origin column (reported at well grain vs allocated from lease-level filings). Together with monthly production this gives gross-to-marketed reconciliation and flaring intensity at well grain.
gold_well_forecast_months- Our deployed production forecast: one row per well per future month, oil and gas, starting the month after the well's last reported month and extending up to 120 months (truncated at the economic limit). oil_per_month and gas_per_month are expected values: additive across wells, so portfolio sums are portfolio forecasts. p10/p50/p90 bands appear where an uncertainty model backs them (petroleum convention: P10 optimistic). Which model produced a well's forecast lives in gold_well_forecast_provenance.
gold_well_forecast_provenance- The audit trail for every well's forecast: one row per well per product stream per forecast generation, recording the cohort the well belongs to, the winning forecast method, the training cutoff, and when it was generated. Join from the forecast months when you need to know which model produced a number.
gold_well_formation_surfaces- Gridded formation-top surfaces: one row per (basin, formation, grid point) with the interpolated subsea top depth, its uncertainty, and formation thickness (isopach). Built from our formation-top picks; combine with directional surveys to determine where a wellbore landed.
gold_well_formation_tops- Formation top picks: one row per well per geologic formation penetrated, with the measured top depth (TVD and subsea). Sourced from completion reports, well logs and scout tickets; this is the input data behind our gridded formation surfaces.
gold_well_fracs- One row per hydraulic-fracturing job, reconciled across FracFocus and state filings so one physical treatment appears exactly once. Carries treatment dates, total proppant and water volumes, fluid-system classification, and stage count. Where FracFocus and the state disagree, we keep the better value and say which source supplied it: each key measurement carries its own paired _source column (proppant_lbs_source, total_water_gal_source, frac_stages_source).
gold_well_injection- Monthly volumes injected INTO disposal, enhanced-recovery and storage wells: water (BBL) and gas (MCF), one row per injection well per month, with the injection type. This is fluid going down the hole; produced water lives on the production table. NULL means not reported, never zero.
gold_well_permits- One row per drilling permit, covering every permit on file (pending, approved, rejected, expired, cancelled, and drilled), so the table reads as the forward-looking activity pipeline. Amendments update the permit's row (last_amended_date marks them) rather than creating duplicates. Carries the permitted well's identity and proposed location (including bottom-hole for horizontals), the operator at filing and today, and proposed depth, formation, and purpose. well_id is NULL until a well exists.
gold_well_production- Monthly production volumes for every well, one row per well per month. The series is complete and continuous from first production to the well's reporting frontier: alongside volumes exactly as reported, it includes reliably-inferred rows (interior gap fills, lease-to-well allocations, period-to-month splits), each labeled in record_origin so you can always filter to the raw reported subset. Volumes are monthly totals; per-day rates, cumulatives, and month counts are intentionally not stored. They derive from this table in one SQL expression (see the derived-fields section). Operator columns give both the owner at that month (M&A applied) and the operator exactly as the source reported it.
gold_well_tests- One row per well test - initial potential, deliverability, surveys and retests - with tested oil, gas and water rates, gas-oil ratio as reported, flowing and shut-in pressures, and test duration. Test rates are short-duration measurements; monthly production is the sustained record.
gold_well_transporter- Who moves and buys each well's production: one row per (well, product, role, counterparty, effective interval) linking wells to their gatherer, purchaser, or off-taker, with each counterparty's share of take where reported. Sourced from designation filings (Texas, Colorado, Louisiana) and realized off-take reports (New Mexico).
gold_well_type_curves- Stored type curves: for each cohort (basin-by-formation, plus coarser analog pools), the peak-normalized median decline shape by month-on-production with a p10/p90 donor-dispersion band. Scale a curve by a well's own peak to get that well's expected profile. The basin and region-analog curves are the ones that drive our new-well and total-supply projections; the two sub_region grains (model and descriptive) describe an area's own wells without feeding any projection. curve_grain, n_donors and is_extrapolated tell you the cohort grain, how many wells back the curve, and where the tail is extrapolated rather than donor-backed.
gold_wells- The master well register: one row per well, from the moment a drilling permit is approved through plugging. Carries identity, operators (current owner, as-reported, and the original driller), surface and bottom-hole location, key lifecycle dates (permit, spud, completion, first production, plugged), wellbore geometry (trajectory, depths, lateral length), formation, and status. Many attributes carry a paired _source column telling you where the value came from (reported by the source, measured from a directional survey, or derived). Regional columns add each jurisdiction's own subdivisions (county, district, municipal district, ...).
Looking for how the numbers are built? That is the methodology; region-by-region completeness is on coverage.