The three review modes
every PDF pauses
Every product PDF waits for a reviewer, whatever the extraction looked like.
the platform default
Documents that clear the confidence threshold are ingested automatically;
the rest pause. This is the default, at 80%.
nothing pauses
Extracted data goes straight into the lakehouse. Products still carry a
confidence score and can be reviewed afterwards in the Product Explorer.
What the threshold actually compares
The confidence threshold is a slider from 50% to 100%, and it is used only by threshold mode. It is compared against the document’s confidence, which is the score of its weakest product — not an average.One bad product parks the whole document. A catalog of forty clean products and
one garbled one has a document confidence equal to that one garbled product, so
it pauses.
- it extracted zero products;
- the PDF had blank pages skipped or needed repair before it could be read;
- the document’s brand couldn’t be confirmed; or
- any product is missing its vendor style ID.
Which setting wins for a given upload
A PDF resolves its rule in one order, most specific first:1
The choice made on that upload
The upload dialog offers a per-upload review mode for product PDFs. Leaving
it on Use organization default records no override at all — the dialog
shows you inline what you’re inheriting.
2
This organization's setting
The mode and threshold saved on this page.
3
The platform default
Threshold review at 80%, for an organization that has never saved a setting.
The extraction verification judge
An optional switch runs a verification pass over the extracted products, comparing them against the images of the pages they came from to catch products the extractor invented. It costs one extra vision call per batch of pages and is off by default. Its findings appear as review reasons on the affected products rather than silently changing anything.Reviewing a parked document
A document that parks shows as awaiting review, and its job sits at Awaiting review until someone resolves it — the job cannot finish on its own. Open the document’s extraction review page from the job or from the documents list. You need lakehouse read permission to open it and lakehouse write permission to act on it. The review page puts the extracted drafts beside the PDF itself, so you can check a claim against the page it came from. Opening a product jumps the viewer to the page it was extracted from. From there you can:per draft
Decide product by product what gets committed. Dropped drafts are simply not
written.
an overlay, not a rewrite
Your corrections layer over what the extractor produced — the original
extraction is never altered, and the edits are applied at commit.
re-keys the commit
If the wrong brand was detected, set the right one. This changes the keys the
document writes, so the create-or-merge forecast on screen updates with it.
writes to the lakehouse
Commits the kept products and technologies. The button states exactly how many
of each it will write before you press it.
writes nothing, and is final
Nothing from the extraction is written, the document is marked rejected, and
its job settles. A reason is required and is stored on the document.
Reprocess
The document detail page offers Reprocess on a product PDF that isn’t currently being worked on — it’s the recovery path for a document that failed to extract, extracted badly, or was rejected. You need upload permission to use it. Reprocessing re-extracts the document from scratch and replaces any staged draft, so a review in progress is discarded. It re-uses the same document rather than creating a new one, and the document keeps pointing at the run that produced its current state — so you follow it from the document, not by hunting for a second upload.Related
The lakehouse
What the lakehouse holds and how documents get into it.
Jobs
Why a job sits at Awaiting review, and how to clear it.
Lakehouse sharing
Once data is committed, what your organization shares from it.
How enrichment works
What enrichment does with the lakehouse data you approve.