Skip to main content
The Lakehouse is a shared knowledge base that enrichment reads from when filling in your products. It holds structured product data, brand technologies, and the source documents those were extracted from. Enrichment consults the Lakehouse before scraping the open web, so data your organization has already contributed — or that is shared across organizations — can complete a product without another round trip to a brand site. The Lakehouse is a source, not a destination for your catalog: you don’t add sellable products to it directly. You contribute to it by uploading documents — or by editing the extracted records in place — and it feeds enrichment.

What the Lakehouse holds

extracted product data
Structured product records pulled from uploaded documents — names, brands, and attributes. Enrichment matches these against your catalog products to fill in missing values.
brand technologies
Named brand technologies and features (materials, cushioning systems, and the like), with the source document they came from. These enrich the products that reference them.
the source files
The uploaded files — spec sheets, catalogs, and similar — that products and technologies are extracted from. Each document keeps a link back to what was derived from it.

Documents and extraction

You populate the Lakehouse by uploading documents. Extraction then reads each file and proposes the products and technologies it found.
1

Upload the file

Upload one or more documents and tag them with the brand and category they describe. Each upload is tracked as a Job.
2

Spreadsheets pause for column-mapping review first

For a spreadsheet, MerchantOps reads the header row, proposes how each column maps to a known field, and shows you a preview of the first few rows — then stops and waits. No rows are extracted yet. You confirm or adjust the mapping, and extraction only starts once you do. This is a human-in-the-loop step, distinct from the automatic column mapping used when uploading products.Every spreadsheet pauses here by default. Ticking Auto-process column mapping on the upload lets a high-confidence mapping run straight through; a low-confidence or invalid one still pauses.
3

Extraction runs in the background

With the mapping settled, MerchantOps reads the file and extracts candidate products and technologies, each with a confidence score, tracked as a Job.
4

A product PDF pauses for extraction review

Before anything from a product PDF is written, MerchantOps applies your organization’s ingestion settings. Unless they say otherwise, a document that doesn’t clear the confidence bar stops and waits for a reviewer to approve it — and the default setting stops a great many product PDFs here. The document’s Job sits at Awaiting review until someone resolves it.On the document’s extraction review page you compare each extracted draft against the page it came from, keep or drop it, correct its values, and then either Approve & Commit or reject the document. Only what you approve is written.
5

Extracted data enters the Lakehouse

The approved products and technologies become available to enrichment. A record arrives flagged Needs Review when its extraction confidence is low, or when it carries properties MerchantOps couldn’t match to a known field.
The two pauses are different in kind, and it matters which one you’re looking at. The column-mapping review (spreadsheets) is a gate in front of extraction: nothing has been written or even extracted yet, so adjusting a mapping corrects the instructions rather than undoing a result. The extraction review (product PDFs) is the opposite — an acceptance step after extraction, where you approve or reject a result that already exists but has not been committed.
Uploading documents isn’t the only way to curate the Lakehouse. Lakehouse product records can be edited directly in the product detail view: adjust their properties and technologies, set identifiers such as the vendor ID or catalog key, and attach a catalog link. These actions curate Lakehouse records — the source data enrichment reads from — not the catalog products enrichment fills in.

Needs Review: what “Mark reviewed” and “Reject” actually do

A record flagged Needs Review offers two actions, and they are not two flavors of “done”:
clears the flag
You checked the record and the data is good. The Needs Review pill goes away and the record shows as reviewed, with who reviewed it and when.
keeps the flag lit
You checked the record and the data is wrong. Rejecting asks for a reason, records who rejected it and why — and deliberately leaves the Needs Review pill lit. A reject means “this record is wrong, someone fix it”, not “resolved” or “hidden”.
Rejecting does not remove a record from the Needs Review queue and does not delete it. The record stays in the Lakehouse, still flagged, with your reason attached — visible on the record as “Rejected by … — reason”. Mark reviewed is the only action in the record view that clears the flag, so the usual path after a reject is: someone corrects the record, then marks it reviewed.
Documents themselves always belong to the organization that uploaded them. Whether the data extracted from them is shared more broadly is controlled by your sharing settings, below.

Per-org sharing

The Lakehouse can be shared across organizations, and what you share is configurable. By default, product and technology data is shared, while MAP pricing data is kept private. Changing a sharing setting affects future writes only — data already contributed keeps the sharing decision it was written with. Manage this at Lakehouse sharing settings.

How enrichment works

How enrichment consults the Lakehouse before the open web.

MAP policies

The pricing data the Lakehouse can hold, kept private by default.

Ingestion settings

Decide when an extracted PDF pauses for approval before it is written.

Sharing settings

Choose what your organization shares into the Lakehouse.

Uploading products

Adding sellable products to your catalog, as opposed to the Lakehouse.