Documents / Document Intelligence Live

Document Translation

Translate a document section by section, carrying terminology forward so the whole reads as one piece rather than a set of fragments.

About the Agent

Challenges Document Translation addresses

Done by hand, document intelligence means gathering document to translate and translation settings, working through 6 separate passes over the same material, then producing document details, source and translation and translation. None of it is difficult and all of it is exacting, which is the combination people are worst at holding. The errors that matter are the ones a tired reader does not notice, and they surface later — in a reconciliation, or in somebody’s reply. It waits until someone remembers it, which is usually the point at which it has become urgent. As volume grows the work does not get harder, only longer, and the first thing to go is the checking.

Document Translation runs that same sequence end to end and returns the result as structured artefacts. What it cannot settle it hands over rather than guesses at, and your correction is kept: it asks “Does this translation read well?” after every run, and those answers become the set it is measured against. Nothing that moves money, alters a contract or reaches a customer executes without human approval, and every action is written to an audit log. The gain is in the volume that no longer has to be read, not in removing the judgement.

How it works

Step 1: Extracting text

First of 6. It works from document to translate and translation settings and feeds the step after it.

Key Tasks:

  • Locating the material: It works from document to translate and translation settings, so nothing has to be forwarded, re-keyed or renamed first.
  • Handling the format it arrives in: Scanned pages, native documents, spreadsheets and message bodies are all read the same way, including layouts where the relevant figure sits inside a table rather than a labelled field.
  • Pulling the fields that matter: Only the fields the rest of the run needs are extracted. What cannot be read confidently is recorded as unread rather than filled in with a best guess.

Outcome:

  • Fields extracted: The fields are available to the steps that follow, with anything unreadable listed rather than silently defaulted — which is what stops a bad extraction becoming a confident wrong answer three steps later.

Step 2: Building a terminology glossary

Step 2 of 6. It takes what step 1 produced and hands its result to step 3.

Key Tasks:

  • Writing from the run, not from a template: The text is built from what this run actually found, so two document intelligence outputs differ where the underlying records differ.
  • Leading with what needs a decision: The exceptions come first and the routine detail follows, because the reader is deciding rather than reading.
  • Staying inside the evidence: Nothing appears in the text that is not supported by a record the run examined. Gaps are stated as gaps.

Outcome:

  • Artefact ready: A finished artefact, traceable line by line to the records behind it, ready for a person to accept or correct.

Step 3: Splitting into sections

Step 3 of 6. It takes what step 2 produced and hands its result to step 4.

Key Tasks:

  • Working from the run so far: This step takes what step 2 produced and carries it toward step 4.
  • Following the same rules each time: The behaviour is configuration rather than judgement made fresh per run, so document intelligence is handled the same way every time.
  • Surfacing what it cannot settle: Anything ambiguous is passed on as ambiguous rather than resolved silently.

Outcome:

  • Passed on: The result passes to the next step, with anything unresolved carried forward as an open item rather than dropped.

Step 4: Translating each section

Step 4 of 6. It takes what step 3 produced and hands its result to step 5.

Key Tasks:

  • Working from the extracted values: Figures come from what step 3 produced rather than from a re-keyed copy, which removes the transcription step where arithmetic errors usually originate.
  • Applying your rules: Bands, rates and rounding are configuration. The same inputs produce the same figures on every run.
  • Keeping the components: Each total is returned with the parts that produced it, so a figure that looks wrong can be traced rather than recomputed.

Outcome:

  • Figures that reconcile: Figures that reconcile to their own components — the totals shown and the lines above them agree, which is the property that makes a number safe to quote onward.

Step 5: Assembling the translation

Step 5 of 6. It takes what step 4 produced and hands its result to step 6.

Key Tasks:

  • Handling the translation: The work at this step is the translation, scoped to that and not extended to anything the run has already settled.
  • Leaving the record behind it: What this step did and what it decided are written down, so the result can be traced without re-running the step.
  • Working from the run so far: This step takes what step 4 produced and carries it toward step 6.

Outcome:

  • The translation handled: The result passes to the next step, with anything unresolved carried forward as an open item rather than dropped.

Step 6: Preparing the side-by-side view

Last of 6. It takes what step 5 produced and produces document details and source and translation.

Key Tasks:

  • Producing the side-by-side view: What this step assembles is the side-by-side view, in the form the reader actually uses it in rather than as a general summary of the run.
  • Keeping it traceable: Each claim stays linked to the record behind it, so a reviewer can check a line instead of accepting the whole.
  • Writing from the run, not from a template: The text is built from what this run actually found, so two document intelligence outputs differ where the underlying records differ.

Outcome:

  • The side-by-side view assembled: A finished artefact, traceable line by line to the records behind it, ready for a person to accept or correct.

Step 7: Your review, and what it changes

The run ends with a person, not with a result being filed.

Key Tasks:

  • Asking a specific question: It asks “Does this translation read well?” rather than for a rating. A question about this run is answerable; a score out of five is not.
  • Keeping the correction: What you change is recorded against the case that produced it, so the disagreement is retrievable rather than absorbed.
  • Building the evaluation set: Those cases become what the agent is measured on. It is scored against your judgement rather than against a general benchmark.

Outcome:

  • A measured agent, not an assumed one: The cases Document Translation handles well and the cases it does not are both visible, and the second list is the one that decides what changes. Nothing is retrained silently on the back of a single correction.

Why use Document Translation?

  • Takes documents as they arrive: Scanned pages, native files and awkward layouts are read as they are. Nothing has to be renamed, re-keyed or converted into a template before a run.
  • Corrected by the people using it: After each run it asks “Does this translation read well?”. Those answers become the evaluation set, which means it is measured against your judgement rather than ours.
  • Reads and reports, does not act: It returns a result for review rather than writing changes back on its own. Anything that moves money, alters a contract or reaches a customer needs human approval first.
  • Structured results, not prose: All 4 artefacts are structured — document details, source and translation and translation — so a result can be scanned, sorted and acted on instead of read end to end.
  • The same sequence every run: 6 steps in a fixed order, on run one and on run four hundred. The variation that creeps into manual work — a check skipped under time pressure, a threshold applied from memory — has nowhere to enter.

Oversight

Runs under scoped, least-privilege credentials with every action written to an audit log. Anything that moves money, alters a contract or reaches a customer requires human approval before it executes.

Document Intelligence

Other agents in document intelligence

Extraction and classification across real-world file formats

  • Document Intelligence Live

    Document Comparison

    Compare two versions of a document side by side and list every substantive change, separating what alters meaning from what only alters wording.

    View agent Book a call
  • Document Intelligence Live

    Document Summarization

    Upload a document and get a structured summary with key metadata and topics. Long documents are summarized section by section with rolling context.

    View agent Book a call
  • Data Extraction Live

    OCR Data Extractor

    Read a scanned invoice or receipt and pull out the header fields and line items as structured data.

    View agent Book a call

Next Step

Deploy Document Translation, or adapt it

It runs as-is. Most deployments diverge — a different source system, a different tolerance, a different approval path. A 30-minute technical call establishes which.

Book a Technical Call
  • No sales script
  • NDA on request
  • Scoping notes sent within 48 hours
Call us Book a call