Methodology
How wells are classified and production numbers built.
Four geographies, kept honest
Every well carries up to four geographic labels, each with a precise meaning: supply region, basin, sub-basin and play.
The four labels, and the Permian area vs the Permian Basin
- Supply Region (the “supply areas” on this site): the coarse, trader-recognizable producing region (Permian, Appalachia, Eagle Ford, …). Areas are drawn differently: some from a county list, some from a geologic boundary, some from a state or a whole country, some from a country’s onshore or offshore theater. The lens is total: every well belongs to exactly one, with a residual region for wells no named area claims. This is the macro rollup for supply views.
- Basin: the strict geologic province (Permian Basin, Appalachian Basin). Narrower than the market’s loose usage of “basin”, deliberately.
- Sub-basin: the geologic subdivision (Midland, Delaware, …).
- Play: the producing interval × commercial extent (Barnett Shale, Eagle Ford Shale). Assigned only on rock evidence, so the label is honestly sparse. A well whose formation evidence is missing shows “Unassigned: no formation evidence”.
So the Permian supply area and the Permian Basin are two different shapes. The area is the Drilling Productivity Report’s county list plus the wells the basin places outside it, and it is the level supply is traded at. The basin is the geologic province alone.
Play assignment: the well’s own rock evidence, and why location-based inference was rejected
A well gets a play label only when formation evidence on its own record places it in the play’s producing interval, within the play’s extent. Gaps stay empty.
The three evidence kinds, the inference test that failed the bar, and where the formation names come from
That evidence is one of three kinds, each a fact about that wellbore rather than a guess from its neighbors: the formation or pool the operator filed with the official source (most wells); a placement from the well’s own directional survey against its own formation tops, which confirms the filing, corrects it, or fills a missing one; or, failing both, the deepest published stratigraphic pick on the wellbore: a geological-survey pick of the deepest formation penetrated, which is not always the producing interval. The alternative, filling gaps by location-based inference from neighboring wells, was calibrated and measured before being ruled out.
In July 2026 a polygon-scoped inference gate was calibrated against a masked holdout of rock-evidenced wells. The gate combined footprint scope, well profile (horizontal, vintage, fluid), and an estimated formation-depth window.
Best achievable precision was 92.8% on the Barnett, 94.1% on the Haynesville (Texas side only), and 91.7% on the Eagle Ford. That is below the ≥95% publication bar. It reached even that while correctly labeling only 11–22% of the candidate wells, 20–77 per play. Louisiana and Arkansas lack the depth surfaces the test depends on, so the Haynesville test covers the Texas side of the fairway alone.
The failure is structural. After the profile gates, the remaining candidates sit in rock vertically contiguous with the play: Marble Falls directly on the Barnett, Austin Chalk within ~13 ft of the Eagle Ford top, basal Cotton Valley on the Bossier. Depth surfaces carry an honest uncertainty of 174–529 ft, and the contacts they would have to adjudicate sit 0–300 ft apart. The discriminand is smaller than the ruler.
Calibration scope is separate from page coverage. The Haynesville area page covers the full Texas–Louisiana–Arkansas footprint; only this inference test was Texas-scoped.
The recoverable pool is small in any case. In these footprints, 86–99% of formation-null wells are pre-play conventional verticals that fail the profile gates honestly. A bar-clearing gate would add dozens of wells, +0.1–0.3% of the page cohorts. So the area pages publish two numbers side by side: the rock-evidenced well count and the total footprint count. Both are always shown.
Where the formation names come from. Where the lexicon carries the unit, the canonical name we publish for a formation, its rank, its place in the group-formation-member hierarchy, and its position in the stratigraphic column come from the stratigraphic lexicon of Macrostrat (University of Wisconsin–Madison), used under Creative Commons Attribution 4.0 International (CC BY 4.0). We modified it: each unit was matched to the spellings the regulators actually file, units outside our coverage were left out, and units the lexicon does not carry were curated by hand. Macrostrat supplies the vocabulary and its ordering, not the depths: formation tops come from the sources described above.
Full attribution, including the citation form Macrostrat asks for, is on the attribution page.
Footprint definitions
How each supply-area footprint is drawn: the DPR county lists, the Barnett definition, and the Anadarko / Mid-Continent boundary.
The definitions
- DPR county lists. Supply-region footprints for the major plays follow the EIA Drilling Productivity Report RegionCounties list (dpr-data.xlsx, fetched July 12, 2026; 315 counties across 7 regions). The DPR is discontinued, final release May 13, 2024. The EIA still serves the file, and it still parses to exactly that county set, so the map here is reproducible and frozen at that vintage. The current STEO breaks out five regions plus “Rest of Lower 48”, folding Niobrara and Anadarko into the residual. The 7-region set here is a deliberate superset of current EIA practice. We have found no published county crosswalk for the 5-region set.
- Barnett is defined here, not by the DPR. The footprint is 27 counties: the 9 core Barnett Shale counties (Tarrant, Johnson, Wise, Denton, Parker, Hood, Somervell, Ellis, Hill) plus 18 surrounding Fort Worth Basin / Bend Arch producing counties. It deliberately excludes the Central and South Texas counties that over-broad “North Central Texas” province labels sweep in.
- Anadarko vs Mid-Continent. The DPR county list decides. A well in a DPR Anadarko county is Anadarko. A well whose geology says Anadarko Basin but whose county sits outside the DPR list is Mid-Continent. “Mid-Continent” here is the conventional remainder of the Kansas/Oklahoma shelf outside the DPR list, and it excludes Anadarko entirely.
Production numbers: best estimate, revised honestly
Production series are a best estimate: volumes the official source published plus rows derived from them, each tagged with its origin.
EIA reconciliation. Production series are compared against EIA figures region by region, and the difference is tracked over time. The coverage page shows the comparison for every region and stream we benchmark, beside the EIA series month by month. The two numbers are produced differently, so recent months can differ in either direction without either being wrong.
Units are the industry’s own. On every chart, table and sentence on this site the volume prefixes are the oil & gas convention, not SI: M = thousand, MM = million, B = billion, T = trillion. So Mbbl is a thousand barrels, MMbbl a million barrels, MMcf a million cubic feet, Bcf a billion cubic feet and Tcf a trillion cubic feet. Rates: bbl/d = barrels per day; Mbbl/d = thousand barrels per day; MMbbl/d = million barrels per day; Mcf/d = thousand cubic feet per day; MMcf/d = million cubic feet per day; Bcf/d = billion cubic feet per day. Every rate we publish is the volume of a period divided by the calendar days in that period.
How the derived rows are built, compared and revised
- Four derivations produce those rows: a lease or unit total divided among the wells it covers (Texas, Louisiana, Oklahoma, Kansas, Michigan); a multi-month filing spread across its months (New Mexico, Venezuela, Ohio, Trinidad and Tobago, Pennsylvania, Bolivia, Peru, Chile, Cuba, Barbados, Michigan, Guatemala, Suriname); a series adjusted so a well’s own cumulative or a state’s yearly total meets a published figure, the source’s cumulative or the EIA’s annual (Louisiana, Colombia, North Dakota, Mississippi); and an absent month carried at zero, so a series stays continuous through an idle stretch. Two further labels describe the source rather than a derivation: volume a source publishes for a producing facility or unit rather than for each well is labeled a facility total, and figures the official source flags as preliminary keep that flag (California). Where a source’s figures for a stream are unrecoverably corrupt and no reliable estimate exists, we claim no volume for that stream in those months: those zeros mark a filed month whose figures we discarded.
- A redacted month stays absent. Where an official source redacts a well’s figures for a month and publishes nothing in their place, we publish nothing for that well-month: no row, never a zero. The absence is the source’s own statement, and a zero would turn it into a figure nobody filed.
- Figures labeled best estimate include those derived rows; figures labeled reported carry only what the official source published. A figure’s label names the mechanisms that change that figure. The lease split is the one mechanism a wider total can stay quiet about: splitting a lease across its wells leaves any total that already contains the whole lease exactly where it was. A grain that can cut a lease, an area for one, keeps the split named. Every other mechanism is named on every figure it touches.
- Recent periods revise upward as late filings arrive. That is a structural feature of official reporting. Annual figures for the latest year settle over several months.
- Why the two differ. The EIA’s monthly figure is a statistical estimate, published about two months after the production month. It comes from a mandatory survey of the largest operators in the individually surveyed states, scaled to state totals by a model fitted to lagged well-level state data. That model is designed to anticipate the revisions state filings later receive, so the EIA’s estimate stays comparatively stable from release to release. The monthly figures are preliminary all the same, and the EIA revises them on a published schedule. Our series that come from a source’s own records are built the other way, bottom-up, and republished as new records arrive, so our recent months fill in and rise. In settled years the EIA has largely adopted the same state data we ingest, so agreement there is a consistency check. Where a series is the EIA’s own, or is adjusted to the EIA’s annual, agreement is by construction. In recent months, for the series built from a source’s records, the two are genuinely independent estimates of a total neither side has fully observed. Outside the United States the EIA’s international series is one of our own sources, so it is a cross-check rather than an independent one.
- Canadian provenance. Alberta production before 2022 is derived from the AER Well Production Data File, which is available directly from the Alberta Energy Regulator (aer.ca); free AER datasets are reproduced with the AER identified as the source. Alberta volumes from 2022 on and the Saskatchewan volumetric series come from Petrinex (Government of Alberta / Petrinex, Crown copyright acknowledged), the Saskatchewan series being the Saskatchewan Ministry of Energy and Resources’ data. Full source credits: attribution.
- Gas figures are on the produced (wellhead) basis, with one named exception. Produced gas and marketed gas can differ materially where gas is re-injected. Every gas figure on this site (charts, tables and rankings alike) is the produced volume. The exception is Michigan, where the figure is the sold volume: the official source files one gas column, the volume sold, and that is the only gas figure Michigan publishes. The state’s own well rules meter sales gas separately from fuel use, so the sold figure sits a rung below the produced one. Since 2012 the EIA’s own Michigan series carry the same figure for produced and marketed gas, after a few percent of difference in the years before, and since 2012 our figure has stayed within a fraction of a percent of that figure. That comparison reaches the end of 2024, where the EIA’s Michigan series stop; before 2012 our figure and the EIA’s diverge in both directions, from well under to somewhat over. The size of Michigan’s own produced-to-sold wedge stays open, so we publish the figure it files and label the basis on the coverage page. One deliberate consequence of the produced basis elsewhere: Alaska’s produced gas appears in the full area table, on the same stated basis as every other row, and stays out of the named bands on the gas charts, because roughly 89.5% of it is re-injected (2025: 3,562 Bcf produced; after re-injection roughly 373 Bcf remained, an upper bound on what could reach market, before lease fuel, flaring and shrinkage). A headline band on produced gas would present volumes that never reach market as supply. The La Barge area of the Green River basin carries a smaller wedge of the same kind (about 38% of its produced gas stays unsold). A marketed-gas series, built from reported sales volumes, follows.
- Footprint production aggregates cover all wells in the footprint, rock-evidenced play wells and the rest alike: the same two-number honesty as the well counts.
- Producing wells tracked counts a well as producing when it reported oil or gas within the 12 months ending at its region’s last reliable month. Each region’s window is set by its own data, not by a fixed calendar date. A single-region page states its window beside the count (“12 months to Jun 2026”). A page spanning several regions carries “trailing 12 months per region”: 12 trailing months at each region’s last reliable month.
- Type curves. Where a vintage’s type curve has a dashed continuation, the solid part runs through the last month where every one of the vintage’s wells sits inside fully reported data. The dashed curve carries the cohort past that month for up to five years: the average of the full cohort, each well on its own reported figure where it has one, our forecast where it has only that, and zero for the rest. The dashed curve ends at the last month where at least nine in ten of the cohort’s wells are accounted for. A vintage with no dashed continuation is drawn through the months where at least 30% of its wells are old enough for the month to fall inside fully reported data; months below that share would track whichever wells happen to be oldest, rather than the vintage itself, so they are not drawn.
- Colombia’s production series are built from the ANH’s producción fiscalizada open datasets and its well records from the SGC’s POZOS layer, both licensed open data, credited on the attribution page.
Production forecasts
Where a cohort of wells has been scored, the cohort’s forecast method is chosen by backtest against what those wells then produced.
How the forecasts are built
The forward profile is built one well at a time, from the well’s own history, and rebuilt as new volumes arrive.
The method is chosen cohort by cohort. A cohort is the wells of one area, drilled the same way, producing the same stream. A cohort that grows too large splits, as far as it needs to: first by the years the wells first produced, then by a smaller area inside that one, then by the rock they produce from. Where a cohort has been scored, decline-curve and machine-learning methods are fitted on an earlier stretch of its wells’ own history and then measured against the months that followed. That test chooses the method the cohort uses.
Wells long past their early decline are carried forward from their own recent level, at a late-life decline. The same treatment carries marginal wells still producing small volumes, and a well’s second stream. A well too young for a fit of its own is forecast from comparable wells where they exist, and otherwise carried forward from its own recent level. Either way it moves onto its own fit once its history can carry one.
Where our data is held well by well, the totals on this site are built up from those per-well profiles. Where a source keeps restating months it has already published, the totals carry our estimate of that restatement, marked as ours. Where production is published only as a combined total, for a country, an area, a project or a group of leases, that total carries a forward of its own.
How well-level Texas production is built
Texas reports most oil production for a lease rather than for a well. A lease can hold one well or several hundred.
How the division is built, measured and checked
- So around three in five producing Texas wells get all their barrels from our own division of those totals. Texas publishes their volumes only as part of a larger lease total.
- We estimate each well’s share of the lease total, month by month, and the estimate is anchored on measurements of the wells themselves. Texas operators file production tests well by well, and we hold 3,237,611 of them covering 325,638 Texas wells, the earliest from the 1960s and the newest arriving as they are filed. Acquiring those tests, keeping them current and cleaning them is the bulk of the work: a test can be mis-keyed, stale, or simply wrong, and one bad reading would push barrels onto the wrong well for years.
- Where a well has no usable test, its share follows what comparable wells do at the same trajectory, age and stage of decline, and it accounts for whether the well is still producing at all.
- Stacked laterals are handled as a group. Texas lets an operator file one combined test for several bores in a stacked-lateral group, on a designated parent bore, while the other bores file zeros for the same test. Read at face value, that puts the group’s whole rate on the parent bore. We keep a standing register of these groups. Where a lease carries the parent’s combined filing beside a bore’s own zero filing, and that bore has test history of its own, part of the combined value moves onto it; where every bore in the group has that history, the split follows it. Wells in these groups carry 15.40% of the Texas oil produced since 2018.
- The estimate is measured against wells where the true answer is published. A lease holding a single well publishes that well’s own production, and those leases are the yardstick the method is built and checked against.
- The division conserves the reported total exactly. Within every Texas lease and month, the volumes assigned to individual wells sum to the total the operator reported: 19,362,031 lease-months, none deviating by more than a hundredth of a barrel. Conservation fixes the lease total. Which well gets which barrel is the harder question, and it is the one the evidence above is for. When the wells’ individual estimates and the reported lease total differ, the difference falls mostly on the estimates that lean hardest on modeling: a well whose month sits on one of its own measurements moves least, and a well locked to its own reported month stays where it is.
- Essentially all of the production Texas reports reaches a named well (99.74% of liquids and 99.96% of gas, across every month we hold, back to 1993). Of that remainder, 91.63% of the liquids and 87.71% of the gas came from the 1990s and 2000s.
- Every Texas barrel carries its label, reported or allocated.
How we attribute production to a company
Production is attributed to the operator of record for each well in each month. What a company page shows is operated production, gross.
What operated, gross means, and why it is the published basis
When a well changes hands during a year, its volume moves to the new operator from the month of the sale onward, and the months before it stay with the seller. Operated production, gross, is everything the wells and facilities that company operates produced, as the official sources report it. A company’s own financial filings report its net ownership share after royalty, so the two figures answer different questions and will differ, often materially. We publish the operated basis because it is what the sources themselves report, which keeps every figure verifiable well by well.
Primary product labels
Primary product says whether a well is an oil well or a gas well, and the source of that label travels with it.
How labels are read, inferred and gated
Most labels are read from evidence the well already carries. Where a region reports production well by well, a well that has produced is classified from its own first year of production, by its gas-to-oil ratio. Where a well has yet to report production, the official source’s own well-type or permit class supplies the label, and only from a class that has proven precise: a class is used when, checked against production, it is right at least 85% of the time. Where a source states that a well exists to inject, dispose, store, observe, serve or test the rock, that stated purpose stands: the well keeps whatever its own record gives it, and the model leaves it alone. Many sources state no purpose at all. Where we hold a well’s injection record, that record stands in as evidence. An injected volume with no produced hydrocarbons stops the model there. Any “unknown product” share we state is counted over the remaining wells.
Where neither source exists, a per-region model infers the label from what is known about the well: its location and its nearest neighbors’ products, its depth, formation, vintage and trajectory. Those labels are marked inferred wherever they show. A region’s model is used only if it clears an acceptance gate: on a validation set of at least 1,000 non-producing wells whose product is known from the official source’s own classes (with the minority product at least 5% of them), the model must agree with those classes at least 97% of the time after correcting for how the labeled wells differ from the wells being labeled (and do better than simply guessing the commoner product), and a well is labeled only when at least 30 comparable labeled wells exist. That agreement figure states agreement with a source-derived proxy rather than accuracy against ground truth: the proxy classes carry their own measured error, so agreement is an upper bound. Each region’s model also carries a confidence threshold, shown in the table: a label ships only where the model’s confidence for that well reaches it. The gate re-runs with every build; the current table:
| Region | Inferred labels | Confidence threshold | Shift-corrected agreement | Wells labeled |
|---|---|---|---|---|
| Oklahoma | yes | 0.90 | 97.3% | 132,178 |
| Alberta | yes | 0.90 | 98.8% | 123,155 |
| Kansas | yes | 0.90 | 98.5% | 110,441 |
| Texas | yes | 0.90 | 98.0% | 67,795 |
| Ohio | yes | 0.92 | 97.0% | 50,050 |
| Louisiana | yes | 0.95 | 97.1% | 33,830 |
| Saskatchewan | yes | 0.90 | 98.7% | 13,104 |
| Montana | yes | 0.90 | 97.7% | 12,774 |
| Michigan | yes | 0.90 | 98.2% | 273 |
| Argentina | no (minority product under 5% of the validation set; agreement below 97%; does not beat guessing the commoner product) | 0.90 (best attempt) | — | 0 |
| British Columbia | no (minority product under 5% of the validation set) | 0.98 (best attempt) | 98.8% (best attempt) | 0 |
| California | no (minority product under 5% of the validation set) | 0.98 (best attempt) | 99.5% (best attempt) | 0 |
| Colorado | no (minority product under 5% of the validation set) | 0.98 (best attempt) | 99.8% (best attempt) | 0 |
| Gulf of Mexico | no (validation set under 1,000 wells) | 0.98 (best attempt) | 98.9% (best attempt) | 0 |
| Mississippi | no (agreement below 97%) | 0.98 (best attempt) | 95.8% (best attempt) | 0 |
| North Dakota | no (minority product under 5% of the validation set) | 0.98 (best attempt) | 99.5% (best attempt) | 0 |
| Pennsylvania | no (minority product under 5% of the validation set) | 0.98 (best attempt) | 99.7% (best attempt) | 0 |
| West Virginia | no (minority product under 5% of the validation set) | 0.98 (best attempt) | 100.0% (best attempt) | 0 |
9 of 18 regions run so far carry inferred labels (543,600 wells); latest run September 6, 2026. Unit-grain regions, where wells are not produced individually, take the producing unit’s own gas-to-oil ratio instead (source unit production): a derivation, not a model.
Permit counts
A permit counts on the date the official source approved it, in the region that issued it.
What is counted, the classification check, and the windows
- What is counted. Permits the operator withdrew and permits the official source refused are left out. Every type the official source issues counts on an entity’s page: permits to drill a new well, permits to deepen, sidetrack or re-enter existing wells, recompletion permits, permits to plug wells and the rest. Each page states its own mix.
- Permits to drill a new well. Official sources issue different permit types, so the count that compares across supply areas is permits to drill a new well. That is what the ranking on the Americas page counts. Every page states that count with its total, unless the count is zero or the classification check withholds it.
- A check on the classification. Each permit type is read from the official source’s own code list. As a check, a region’s permits to drill a new well over its latest year are compared with the new wells actually appearing there in the same year: wells spudded, or where spud dates are thin, the richest series of wells completed or brought onto production. Where the permits fall below one for every five of those wells, the new-well count is withheld on every page whose latest year that region feeds, and those pages say so. The check judges only where the compared series has at least 100 dated events in the year, so a region with thin activity dates is left unjudged rather than failed.
- Windows. Every window ends at the entity’s own latest approval on record, never at today, so a source that publishes with a lag shows what it holds rather than an empty recent window. The date is stated with every count.
Rig counts
- Source. United States and Canada rig counts are the Baker Hughes North America Rotary Rig Count, published weekly. We serve our own derived series and rollups, never Baker Hughes’s files. Baker Hughes is not affiliated with and does not endorse CommoVision. Mexico’s rig counts come from the country’s official hydrocarbon information system (SIH, monthly) and appear when that series is current.
- Where a count shows. State and province pages count land and inland-water rigs; rigs in US federal offshore waters are counted on the Gulf of Mexico page (Pacific federal waters on the Pacific OCS page), and Alaska keeps its own waters. Supply-area pages roll county-level rig locations up to each county’s dominant supply area in our well data: a county counts toward exactly one area. Weeks with no rigs are real zeros: the source lists only active rigs, so a series that has gone quiet draws at zero rather than disappearing.
How we time an arrival
Each card in the feed shows the elapsed time from the moment new data is first seen to the moment it goes live in the product, where both moments are on record.
The clocks, the steps, and what the badge leaves out
- What the badge measures. The badge covers this end of the chain, on that individual release. The figure is never an average and never a filtered selection. The feed shows every customer-releasable arrival in order, including the slow ones.
- What the badge leaves out. The gap between the source publishing and the change being noticed sits outside the elapsed time. For most sources that gap is unobservable: a publication detected on a poll could have appeared at any point since the previous poll. Only measured latency is published here. The gap is real, and it is not ours to claim away.
- Why the record is kept. These times are measured to be driven down. Step by step, the record shows where time is lost, and shortening those steps is ongoing work.
- Three clocks, always labeled. Most sources are polled, so the clock starts at the poll that saw the change. Some sources are fetched by a query that both finds and downloads new data; there the clock starts when the fetch starts, the earliest moment the change could have been seen. A few releases have neither on record. Those cards say so and show no elapsed time, because borrowing a different clock would make the number mean something else.
- Where it comes from. Every release is collected from the source that published it. That collection is the start of the clock: the elapsed time is measured from it, and the feed’s steps begin after it.
- The steps. Expanding an arrival shows each stage with its own elapsed time. The steps a release passes through depend on the source: a clean data file skips the document parsing, and allocation runs only where a source reports at lease or facility grain:
- Documents parsed. This covers scanned filings too, where the numbers exist only as an image.
- Ingested. If a number is ever questioned, this untouched copy is what we check against.
- Cleaned. Units, well identifiers, dates and operator names are stated on the same terms, source to source.
- Allocated to wells. Where the source reports only a lease or facility total, we estimate each well's share.
- Enriched & verified. We check the numbers before anything publishes, and wells pick up their operator, basin and formation.
- Live. In the product from this moment on. This site follows at its next update.
- Forecasts live. We rebuild forecasts for the wells this release touched. The data is already live while this runs.
- On the website. The pages you are reading rebuilt with this update in them.
- Expected arrivals. Where a source’s own publication rhythm is modeled and the model passes its own backtest, the next arrival gets a named window. The window is deliberately conservative: it is set at the model’s 95th percentile, the point the model expects the real arrival to beat nineteen times in twenty. Only high-confidence sources are named at all.
See it applied
The producing-area pages put this methodology to work: Barnett, Eagle Ford, Haynesville. The coverage matrix lists every published production series with its grain, source, basis and data-through month.
For a question about a specific number, get in touch. Every production figure carries its origin, month by month.