Documents / Document Intelligence Live

Document Comparison

Compare two versions of a document side by side and list every substantive change, separating what alters meaning from what only alters wording.

About the Agent

Challenges Document Comparison addresses

Done by hand, document intelligence means gathering original version, revised version and comparison settings, working through 4 separate passes over the same material, then producing what changed, substantive changes and worth a second look. None of it is difficult and all of it is exacting, which is the combination people are worst at holding. The errors that matter are the ones a tired reader does not notice, and they surface later — in a reconciliation, or in somebody’s reply. It waits until someone remembers it, which is usually the point at which it has become urgent. As volume grows the work does not get harder, only longer, and the first thing to go is the checking.

Document Comparison runs that same sequence end to end and returns the result as structured artefacts. What it cannot settle it hands over rather than guesses at, and your correction is kept: it asks “Did this catch the changes that matter?” after every run, and those answers become the set it is measured against. Nothing that moves money, alters a contract or reaches a customer executes without human approval, and every action is written to an audit log. The gain is in the volume that no longer has to be read, not in removing the judgement.

How it works

Step 1: Reading the original

First of 4. It works from original version, revised version and comparison settings and feeds the step after it.

Key Tasks:

  • Locating the material: It works from original version, revised version and comparison settings, so nothing has to be forwarded, re-keyed or renamed first.
  • Handling the format it arrives in: Scanned pages, native documents, spreadsheets and message bodies are all read the same way, including layouts where the relevant figure sits inside a table rather than a labelled field.
  • Pulling the fields that matter: Only the fields the rest of the run needs are extracted. What cannot be read confidently is recorded as unread rather than filled in with a best guess.

Outcome:

  • Fields extracted: The fields are available to the steps that follow, with anything unreadable listed rather than silently defaulted — which is what stops a bad extraction becoming a confident wrong answer three steps later.

Step 2: Reading the revision

Step 2 of 4. It takes what step 1 produced and hands its result to step 3.

Key Tasks:

  • Reading the revision specifically: This pass is scoped to the revision rather than to the document as a whole, so a field that appears in more than one place is taken from the one that governs.
  • Keeping the original alongside: Each extracted value stays linked to where it was found, so a figure that looks wrong can be checked against the source rather than re-entered.
  • Locating the material: It works from what step 1 produced, so nothing has to be forwarded, re-keyed or renamed first.

Outcome:

  • The revision captured: The fields are available to the steps that follow, with anything unreadable listed rather than silently defaulted — which is what stops a bad extraction becoming a confident wrong answer three steps later.

Step 3: Comparing the versions

Step 3 of 4. It takes what step 2 produced and hands its result to step 4.

Key Tasks:

  • Finding the counterpart: It searches the connected system for the record this one should correspond to, using the identifiers taken from what step 2 produced.
  • Comparing field by field: Each field is checked against its counterpart rather than the documents being compared as wholes, so a single line that disagrees is reported as that line rather than as a failed match.
  • Applying your tolerances: The variance you accept is configuration. A difference inside it clears; a difference outside it is held, and the amount is stated rather than described as a discrepancy.

Outcome:

  • Everything agrees: The record clears and moves on without anyone reading it.
  • Something does not agree: Each disagreeing field is reported with both values and the size of the gap, so the review starts from the discrepancy rather than from the whole document.
  • No counterpart exists: The record is held and flagged as unmatched rather than passed through as clean, which is the failure mode that costs the most to find later.

Step 4: Writing the review note

Last of 4. It takes what step 3 produced and produces what changed and substantive changes.

Key Tasks:

  • Writing from the run, not from a template: The text is built from what this run actually found, so two document intelligence outputs differ where the underlying records differ.
  • Leading with what needs a decision: The exceptions come first and the routine detail follows, because the reader is deciding rather than reading.
  • Staying inside the evidence: Nothing appears in the text that is not supported by a record the run examined. Gaps are stated as gaps.

Outcome:

  • Artefact ready: A finished artefact, traceable line by line to the records behind it, ready for a person to accept or correct.

Step 5: Your review, and what it changes

The run ends with a person, not with a result being filed.

Key Tasks:

  • Asking a specific question: It asks “Did this catch the changes that matter?” rather than for a rating. A question about this run is answerable; a score out of five is not.
  • Keeping the correction: What you change is recorded against the case that produced it, so the disagreement is retrievable rather than absorbed.
  • Building the evaluation set: Those cases become what the agent is measured on. It is scored against your judgement rather than against a general benchmark.

Outcome:

  • A measured agent, not an assumed one: The cases Document Comparison handles well and the cases it does not are both visible, and the second list is the one that decides what changes. Nothing is retrained silently on the back of a single correction.

Why use Document Comparison?

  • Field mapping you can audit: Every incoming field is shown with what it was mapped to and how confidently. Low-confidence mappings are held rather than applied, which is where silent data corruption otherwise starts.
  • Takes documents as they arrive: Scanned pages, native files and awkward layouts are read as they are. Nothing has to be renamed, re-keyed or converted into a template before a run.
  • Corrected by the people using it: After each run it asks “Did this catch the changes that matter?”. Those answers become the evaluation set, which means it is measured against your judgement rather than ours.
  • Reads and reports, does not act: It returns a result for review rather than writing changes back on its own. Anything that moves money, alters a contract or reaches a customer needs human approval first.
  • Structured results, not prose: All 4 artefacts are structured — what changed, substantive changes and worth a second look — so a result can be scanned, sorted and acted on instead of read end to end.

Oversight

Runs under scoped, least-privilege credentials with every action written to an audit log. Anything that moves money, alters a contract or reaches a customer requires human approval before it executes.

Document Intelligence

Other agents in document intelligence

Extraction and classification across real-world file formats

  • Document Intelligence Live

    Document Summarization

    Upload a document and get a structured summary with key metadata and topics. Long documents are summarized section by section with rolling context.

    View agent Book a call
  • Document Intelligence Live

    Document Translation

    Translate a document section by section, carrying terminology forward so the whole reads as one piece rather than a set of fragments.

    View agent Book a call
  • Data Extraction Live

    OCR Data Extractor

    Read a scanned invoice or receipt and pull out the header fields and line items as structured data.

    View agent Book a call

Next Step

Deploy Document Comparison, or adapt it

It runs as-is. Most deployments diverge — a different source system, a different tolerance, a different approval path. A 30-minute technical call establishes which.

Book a Technical Call
  • No sales script
  • NDA on request
  • Scoping notes sent within 48 hours
Call us Book a call