How Is Syndicated Data Collected and Projected?

Updated Aug 20266 min readBy The Sous Team

Syndicated providers collect scan data from a universe of cooperating stores, then project it to represent the full market. The data arrives in provider-defined weeks and periods, and history gets restated when the store universe changes. Those three facts (projection, the calendar, restatements) explain most of the behavior that surprises new readers.

From register to market file

The pipeline starts at the register. Retailers license their scan data to the providers, who pool it across chains, clean it, and standardize it against their product libraries [1]. What lands in your extract is a manufactured dataset: every row has been matched to a product hierarchy, assigned to a market, and slotted into a reporting period.

Not every store participates, and that is the crux for most of the market. The provider observes a universe of cooperating stores and extrapolates to the full market, so many of the numbers you read are estimates built from a large sample, weighted and projected, rather than a census of every register [1].

How scan data becomes a syndicated market file From register to market file Cooperating stores scan and license data Provider cleans, pools, and projects to market Your market file: projected, periodized a sample of stores, weighted and extrapolated, not a census of every register the projection rules live at the provider, not in your file

One exception worth knowing: for some retailers and some providers the read is a census rather than a sample, because the chain supplies data from every store. Where that holds, the numbers are not projections at all.

Projection runs on ACV, not store counts

How do you weight a sample of stores up to a market? By size, and in syndicated data store size means dollars. ACV, all commodity volume, is a store's total sales across every product it carries. Both Circana and Nielsen define it this way and use it as the weighting base for distribution measures [4][5]. A supercenter and a small rural store are not two equal tallies; one may move many times the volume of the other, and the projection math treats it accordingly.

This is why the distribution measures in Chapter 4 are percentages of ACV rather than percentages of stores. The system was built around dollar-weighted store universes from the ground up, and every downstream measure inherits that choice.

One boundary worth knowing: ACV totals conventionally exclude departments that do not scan like packaged goods, such as pharmacy, fuel and lottery [2]. Two stores with equal front-of-store volume can therefore carry different ACV if one runs a large pharmacy business, and a provider's ACV for a retailer will not match that retailer's reported total sales.

A product must scan to count

The subtlest rule in the system: distribution is measured by scans, not by shelf placements. An item counts as in distribution in a store during a period only if it actually scanned there [3]. Sitting in the planogram is not enough.

Play out the consequences. A store that is out of stock on your item all week reads as a store without distribution that week, even though you won the placement. A slow item in a small store might scan zero units some weeks and blink in and out of measured distribution. What looks like a distribution void can be a real void, a chronic stocking problem, or just a slow week. The data reports the scan; it does not tell you which.

Restatements: when history changes under you

The store universe is alive. Chains open and close stores, retailers join and leave the cooperating panel, and providers periodically rebuild their projections to match. When that happens the numbers you already reported can change: a restatement rewrites history so the past is measured with the same universe as the present.

The practical effect: the 52-week trend you pulled in January and the same 52-week trend pulled in July may not match, and neither is wrong. Providers also refresh product hierarchies and category definitions (a reset), which can move items between categories and change your share without a single shopper changing behavior. Teams that do not know this burn days hunting for errors that are actually definitions.

What this means on load day

Everything above lives outside your extract. The projection weights, the universe changes, the hierarchy updates: none of it is visible in the column headers of the file that lands every period. The file just quietly reflects it.

That is the practical argument for validating every new file against the one before it, column by column and measure by measure. It is also exactly what Sous does on load: when a new period's file lands, it checks the structure against the prior period, and anything unexpected (renamed columns, missing measures) stops the load and asks before a single workbook re-runs. How that fits an agent's overall design is covered in how an agent works with syndicated data.

Next: the three companies that build these files, and how SPINS, Circana, and Nielsen actually differ.

Common questions

Is syndicated data accurate? For most of the market it is a projection from a large store sample, refined over decades, and it is the number your retailer and the industry align on. Treat it as reliable for direction and comparison, and remember it is an estimate subject to restatement.

Why did my historical numbers change? Almost certainly a restatement, where the provider rebuilt its store universe or projections, or a reset, where hierarchy and category definitions changed. Both rewrite history by design.

Does shelf presence count as distribution? No. An item must scan in a store during the period to count as distributed there [3]. Out-of-stocks and zero-scan weeks read as missing distribution.

Why don't my distributor shipments match my syndicated sales? Different events. Distributors report cases shipped into warehouses; syndicated POS reports units scanned at registers, projected across a market. Timing, pipeline fill and coverage differences guarantee gaps.