What the Lakehouse holds
extracted product data
Structured product records pulled from uploaded documents — names, brands, and
attributes. Enrichment matches these against your catalog products to fill in
missing values.
brand technologies
Named brand technologies and features (materials, cushioning systems, and the
like), with the source document they came from. These enrich the products that
reference them.
the source files
The uploaded files — spec sheets, catalogs, and similar — that products and
technologies are extracted from. Each document keeps a link back to what was
derived from it.
Documents and extraction
You populate the Lakehouse by uploading documents. Extraction then reads each file and proposes the products and technologies it found.1
Upload the file
Upload one or more documents and tag them with the brand and category they
describe. Each upload is tracked as a Job.
2
Spreadsheets pause for column-mapping review first
For a spreadsheet, MerchantOps reads the header row, proposes how each
column maps to a known field, and shows you a preview of the first few rows
— then stops and waits. No rows are extracted yet. You confirm or adjust
the mapping, and extraction only starts once you do. This is a
human-in-the-loop step, distinct from the automatic column mapping used when
uploading products.Every spreadsheet pauses here by default. Ticking Auto-process column
mapping on the upload lets a high-confidence mapping run straight through;
a low-confidence or invalid one still pauses.
3
Extraction runs in the background
With the mapping settled, MerchantOps reads the file and extracts candidate
products and technologies, each with a confidence score, tracked as a Job.
4
A product PDF pauses for extraction review
Before anything from a product PDF is written, MerchantOps applies your
organization’s ingestion settings. Unless
they say otherwise, a document that doesn’t clear the confidence bar stops
and waits for a reviewer to approve it — and the default setting stops a
great many product PDFs here. The document’s Job sits at Awaiting review
until someone resolves it.On the document’s extraction review page you compare each extracted draft
against the page it came from, keep or drop it, correct its values, and then
either Approve & Commit or reject the document. Only what you approve
is written.
5
Extracted data enters the Lakehouse
The approved products and technologies become available to enrichment. A
record arrives flagged Needs Review when its extraction confidence is
low, or when it carries properties MerchantOps couldn’t match to a known
field.
The two pauses are different in kind, and it matters which one you’re looking at.
The column-mapping review (spreadsheets) is a gate in front of extraction:
nothing has been written or even extracted yet, so adjusting a mapping corrects
the instructions rather than undoing a result. The extraction review (product
PDFs) is the opposite — an acceptance step after extraction, where you approve
or reject a result that already exists but has not been committed.
Needs Review: what “Mark reviewed” and “Reject” actually do
A record flagged Needs Review offers two actions, and they are not two flavors of “done”:clears the flag
You checked the record and the data is good. The Needs Review pill goes away
and the record shows as reviewed, with who reviewed it and when.
keeps the flag lit
You checked the record and the data is wrong. Rejecting asks for a reason,
records who rejected it and why — and deliberately leaves the Needs Review
pill lit. A reject means “this record is wrong, someone fix it”, not
“resolved” or “hidden”.
Documents themselves always belong to the organization that uploaded them.
Whether the data extracted from them is shared more broadly is controlled by
your sharing settings, below.
Per-org sharing
The Lakehouse can be shared across organizations, and what you share is configurable. By default, product and technology data is shared, while MAP pricing data is kept private. Changing a sharing setting affects future writes only — data already contributed keeps the sharing decision it was written with. Manage this at Lakehouse sharing settings.How enrichment works
How enrichment consults the Lakehouse before the open web.
MAP policies
The pricing data the Lakehouse can hold, kept private by default.
Ingestion settings
Decide when an extracted PDF pauses for approval before it is written.
Sharing settings
Choose what your organization shares into the Lakehouse.
Uploading products
Adding sellable products to your catalog, as opposed to the Lakehouse.