The Four Components: Data, Semantics, Memory, and Rules

Updated Aug 20267 min readBy The Sous Team

A brand context layer has 4 components: data, semantics, memory, and rules. Data establishes what is yours, semantics establish what your words mean, memory retains what has been learned, and rules encode how your business works. The framework was introduced in why an agent needs to know your brand. Knowing which component is missing turns "the AI doesn't get us" into a gap you can name and close.

The four components of a brand context layer Data which items, retailers, markets are yours Semantics what your team's words denote Memory what was asked, corrected, settled Rules calendar, competitive set, legal math An AI that reasons about your brand right items, right scope, right calendar, right math all four are required; each fails in its own recognizable way

Data: what is yours

The data component connects your items, categories, retailers, and markets across sources that never quite share a common ID. A multi-channel CPG brand holds SPINS extracts for the natural channel, Circana or Nielsen files for conventional, retailer portal downloads, and internal spreadsheets, and the same product carries different codes in each. The data component is the mapping that says: these 40 rows are ours, these 12 are the competitor we actually watch, this retailer appears under 3 different market definitions.

How it gets built: mostly at connection time, and then maintained as reality changes. When Sous connects to a brand's extracts, portals, and files, it learns the item universe from the data itself, and the team confirms or corrects the edges (the discontinued flavor, the food-service line that should never count). New items and new retailers join the map as they appear in new periods.

When it is missing, the failure is a blind spot. A total-business read that quietly omits the club channel does not announce itself as incomplete. It just comes back smaller, and nobody in the room can tell.

Semantics: what your words mean

The semantics component resolves your team's language to exact products, markets, and periods. "The core four" is 4 specific UPCs. "The club pack" is the 24-count. "Post-reset" starts on a date. "The natural channel" is a provider-defined universe, and your team may mean SPINS' Natural specifically. The BIRD benchmark's authors called this external knowledge, the unstated facts connecting a question to the database underneath it, and identified it as a core challenge separating real question-to-query work from academic exercises [2].

How it gets built: from use. Nobody sits down and writes a dictionary of their own shorthand, because nobody knows their shorthand is shorthand. The semantics accumulate as the team asks questions in its own words and confirms what came back: the first time someone asks about the core four and adjusts the item list, that phrase is grounded permanently.

When semantics are missing, the failure is a scope error: the right math on the wrong set. The "core four" read that silently includes a discontinued fifth item runs clean, looks right, and answers a question nobody asked. That kind of error survives review, because catching it requires reconstructing the whole item list.

Memory: what has been learned

The memory component retains questions asked, corrections made, and definitions settled, so that nothing needs to be established twice. Anthropic's engineering guidance treats memory as a first-class technique for capable agents: structured note-taking that persists outside the immediate context window lets an agent maintain knowledge across sessions instead of starting from zero [1]. Applied to a brand, that means the correction you made in week 2 ("exclude food service from everything") is still in force in month 9, and the analysis you asked for last quarter is the starting point for this quarter's version.

How it gets built: passively, which is the point. Every session leaves a residue of settled facts. The team was going to ask questions and correct wrong scoping anyway; memory just means the corrections stick.

When memory is missing, the failure is repetition friction. The same clarification, re-supplied every session, until the team stops bothering and the tool goes quiet. Tools rarely get uninstalled for being wrong once. They get abandoned for asking the same question 5 times.

Rules: how the business works

The rules component encodes the standing facts of the operation: the fiscal calendar, the price-pack architecture, which retailer meetings recur, what counts as a win, and, critically for CPG, what math is legal on each provider's measures. Syndicated measures carry aggregation rules (% ACV distribution never averages across markets; weekly velocity never sums into monthly), and those rules differ across SPINS, Circana, and Nielsen. In Sous, provider rulebooks enforce that math on every query, which is covered fully in how an agent works with syndicated data.

How it gets built: two ways. Provider math ships with the system, because it is knowable in advance and non-negotiable. Brand-specific rules arrive like semantics do, through use: the first time the team notes that the fiscal year starts in July, every subsequent period comparison respects it.

When rules are missing, the failure is unquotable numbers: analysis built on the calendar quarter for a team that runs a fiscal one, or a "national average" that a buyer's analyst can dismantle in one question. The work was done, and it cannot be used in the meeting it was built for.

Diagnosing which one is missing

The four failure signatures are distinct enough to work backwards from the symptom.

Symptom you observe Component missing First fix
Whole channels or retailers absent from answers Data Connect the missing source; confirm the item map
Right math on the wrong item set or period Semantics Ground the team's phrases the first time each is used
Same clarification re-asked every session Memory Use a system where corrections persist by default
Numbers your buyer's analyst can take apart Rules Enforce provider rulebooks and the brand's calendar

The table in the seed chapter names these failure modes in summary; the build mechanics above are what closing each gap looks like in practice. Notice that 3 of the 4 components are built by simply using the system and correcting it, which is why the layer gets more complete over time. That dynamic, the compounding, is the subject of the compounding chapter.