Skip to main content
In this tutorial, you build a document processing pipeline that verifies customer onboarding documents against application data. The pipeline receives uploaded files (ID card, proof of address, salary slip), classifies each document, extracts structured data, compares it to what the applicant declared, and routes discrepancies to a human reviewer. What you will build:
  • A file upload UI that accepts multiple documents
  • A document classification workflow that identifies each document type using AI
  • A fan-out extraction pipeline that routes each type to a specialized extractor
  • An AI reconciliation step that compares extracted data against application data
  • Business rules that flag mismatches (name, address, income)
  • A human review task for documents with discrepancies
  • A summary generation step that produces a verification report
AI node types used: Document Understanding, Document Extraction, Text Understanding, Text Generation Patterns demonstrated: Fan-out extraction, AI comparison and reconciliation

Architecture overview

The pipeline processes documents in four phases: upload, classify and extract, reconcile, and review. Workflow breakdown:

Prerequisites

Before starting, make sure you have:
  • Access to a FlowX Designer workspace with AI Platform enabled
  • Familiarity with creating processes, workflows, and UI flows in FlowX
  • A project with the Documents Plugin configured (for file uploads)

Data model

Define the following data model keys in your process. These keys hold the application data submitted by the customer and the results produced by the AI pipeline.
Define these keys under your process data model before building the workflows. The AI nodes and business rules reference these paths at runtime.

Step 1: Build the classification and extraction workflow

Create a workflow named classifyAndExtract. This workflow receives a single document file path, classifies the document type, and then routes to the appropriate extraction branch. This implements the fan-out extraction pattern.

1.1 Add the classification node

Add a Document Understanding node as the first node after the Start node. This node accepts the document file directly and classifies it.
Use Document Understanding here, not Text Understanding — the text nodes accept text input only, while the document nodes work on the uploaded file itself.
Instructions:
Response schema:
Response Key: classificationResult

1.2 Add the Condition node

Add a Condition node after the Document Understanding node. Configure branches based on the document_type value:
The Else branch handles UNKNOWN documents. Use a Script node to return a structured error so the parent process can flag the document for manual classification.

1.3 Configure type-specific extraction nodes

Each branch contains a Document Extraction node with a prompt and schema tailored to that document type.
Use Document Extraction for this step, not Extract Data from File. The two nodes are not interchangeable:
  • Document Extraction takes Instructions and a Response schema, and returns structured fields. It has no extraction-method setting.
  • Extract Data from File takes an Extraction Method (Automatic, LLM Model, OCR Engine, Text Parsing) plus image and signature options, and returns text. It accepts no instructions or response schema.
A single node cannot combine a prompt, a response schema, and an extraction method.
Document Extraction node configuration with Instructions, Document Source, Response Schema, and Response Key
Instructions:
Response schema:
Response Key: extractedData
If you also need the raw text of a document, its embedded images, or signature detection, add a separate Extract Data from File node. Set its Extraction Method to Automatic for a mixed document set like this one, where the format varies per upload and you do not want to pick a strategy per file. See Extract Data from File for a comparison of the methods.
Extract Data from File node with Extraction Method set to Automatic and Response Key documentText
Both Document Extraction and Extract Data from File support the Personal Information Guard, which detects and replaces personal data before it reaches the model. Consider turning it on for this pipeline: ID documents, addresses, and payslips all carry personal data. See Personal Information Guard.

1.4 Add the End Flow node

Add an End Flow node where all branches converge. Set the body to pass results back to the parent process:

Step 2: Build the reconciliation workflow

Create a workflow named reconcileData. This workflow compares the extracted document data against the applicant’s declared data. This implements the AI comparison and reconciliation pattern.

2.1 Add the comparison node

Add a Text Understanding node that receives both the extracted data and the application data.
The Instructions field is static — the node rejects ${...} references inside it. Pass dynamic values through the Context section instead; the node receives them alongside the instructions.
Instructions:
Context: reference the following process data so the node receives it at runtime:
  • ${extraction.classifiedDocs} — the extracted document data
  • ${applicant} — the applicant’s declared data
Response schema:
Response Key: reconciliationResult

2.2 Add the End Flow node


Step 3: Build the summary generation workflow

Create a workflow named generateSummary with a single Text Generation node. Instructions:
Context: reference the following process data:
  • ${applicant} — applicant data
  • ${extraction.classifiedDocs} — extraction results
  • ${reconciliation} — reconciliation results
  • ${review.reviewerNotes} — reviewer notes (if any)
Response Key: summaryReport End Flow body:

Step 4: Build the BPMN process

Create a process named documentVerify that orchestrates the full pipeline using the workflows you built.
1

Add a User Task for file upload

Add a User Task node after the Start Event. This task presents the file upload UI to the user.Configure the task with:
  • Task name: Upload documents
  • Assignment: Assigned to the initiating user
The upload UI is designed directly on this node - you build it in Step 5. A User Task cannot attach a standalone UI Flow; its UI lives on the node.
2

Loop through uploaded documents

For each uploaded document, trigger the classifyAndExtract workflow. Add a Send Message Task node with a Start Integration Workflow action.Input mapping:
Replace [index] with a concrete extraction step for each file: array-indexed ${...} expressions do not resolve in data mappings. Extract the current file into a named object with a business rule or Script node (for example output.currentFile = input.documents.uploadedFiles[0];), then map ${currentFile.filePath} and ${currentFile.fileId}. See Referencing workflow data in node configurations.
Add a Receive Message Task node to capture the extraction output. In the node’s Data Stream, set the Key Name to extraction.classifiedDocs[index].
For multiple documents, repeat the Send/Receive pattern for each file, or use a loop structure with an exclusive gateway that iterates until all files are processed.
3

Trigger the reconciliation workflow

Add another Send Message Task with a Start Integration Workflow action pointing to the reconcileData workflow.Input mapping:
Add a Receive Message Task node. In the node’s Data Stream, set the Key Name to reconciliation.
4

Add a business rule for validation

Add a Service Task with a Business Rule action (JavaScript) to perform deterministic validation checks that supplement the AI reconciliation. Business rules read process data through input. and write results through output..
Business rules provide deterministic, auditable checks. Use them alongside AI reconciliation to catch issues the LLM might miss, such as expired documents or missing required document types.
5

Add the routing gateway

Add an Exclusive Gateway after the business rule. Configure two branches:
A gateway needs its evaluation rule defined in addition to the outgoing branches - without one, the process fails at runtime with “No rules found for gateway node”. Gateway conditions read process data with the input. prefix.
6

Add the human review task

Add a User Task node for manual review. The reviewer sees:
  • Uploaded documents (viewable in a File Preview component)
  • Extracted data side-by-side with declared data
  • The exception report from reconciliation
  • Validation flags from the business rule
The reviewer submits a decision:
  • Approve — continue to summary
  • Reject — end process with rejection status
  • Request re-upload — loop back to the upload step
Store the decision in review.reviewerDecision and any notes in review.reviewerNotes.
7

Route on the reviewer's decision

Add a second Exclusive Gateway after the review task - the three reviewer outcomes each need a branch:
8

Trigger the summary generation workflow

After both the auto-approve and human-review-approve paths converge, add a Send Message Task to trigger the generateSummary workflow.Input mapping:
Add a Receive Message Task. In the node’s Data Stream, set the Key Name to summary.
9

Add the End Event

Add an End Event after the summary is received. The process instance now contains the full verification report at summary.report.

Step 5: Build the upload UI

Design the upload page directly on the Upload documents User Task: select the node and open its UI Designer. A User Task’s UI lives on the node - a standalone UI Flow cannot be attached to a BPMN task.
1

Add an upload component

Add a File Upload component to the node’s UI and restrict the accepted file types to PDF, JPG, and PNG.
2

Add the Upload file action

On a User Task, the File Upload component does not come with an action - create an Upload file action on the node and link the component to it. The action posts the file to the Documents Plugin over Kafka; configure the Address (the document-persist topic) and the Document Type. For the full parameter list and the multi-file behavior, see the Upload file action guide.The AI document nodes read uploaded files with Document Source set to Document Plugin, which resolves paths produced by this upload.
3

Add applicant data display

Add form fields (read-only) that display the applicant’s declared data from applicant. This gives context to the person uploading documents.
4

Add a submit button

Add a Button component labeled Submit documents. Configure it to save the data and advance the User Task.
In a standalone UI Flow (for example, a chat-driven app that triggers workflows directly), the File Upload component behaves differently: it arrives with an Upload action already attached, and the upload result lands under that action’s Response Key - not under the component’s data key. With the default key of response, a successful upload produces:
Pass response.filePath to the workflow in that setup.
File Upload component in a UI Flow with its Upload action open, showing the Response Key field
For the human review step, design a second page the same way - on the Human review User Task node - displaying the extracted data, reconciliation results, and exception report alongside the original documents. Use a side-by-side layout so the reviewer can compare easily.

Step 6: Build the review UI

Design the review page on the Human review User Task node, the same way you built the upload page in Step 5. The review page should include:
Use conditional visibility to highlight rows with MISMATCH or CRITICAL status in the reconciliation table. This draws the reviewer’s attention to the issues that need their judgment.

Testing

1

Test classification in isolation

Open the classifyAndExtract workflow and use Run Workflow with a test file. Upload sample documents one at a time and verify the classification output.
2

Test extraction accuracy

For each document type, compare the extracted fields against the actual document content. Check that:
  • Names are captured correctly (including accented characters)
  • Dates are in the expected YYYY-MM-DD format
  • Numeric values (salary, postal code) are accurate
  • Null is returned for missing fields (not hallucinated values)
3

Test reconciliation with known mismatches

Prepare test data with deliberate mismatches:
Upload an ID card with the name “Jonathan Smith” and a salary slip showing a net salary of 4200. Verify the reconciliation output flags:
  • Name variation as WARNING
  • Income discrepancy (16%) as WARNING
4

Test the business rule

Verify the JavaScript business rule catches:
  • Expired ID documents
  • Missing required document types
  • Income discrepancy above 10%
  • CRITICAL exceptions from reconciliation
Test edge cases: all documents valid (auto-approve path), one missing document (review path), expired ID (review path).
5

Test the full end-to-end flow

Run the complete documentVerify process:
  1. Upload three documents (ID, proof of address, salary slip)
  2. Verify classification and extraction complete
  3. Check the reconciliation report
  4. If routed to review, complete the reviewer task
  5. Verify the summary report is generated
Test both the auto-approve path (all documents match) and the human review path (with discrepancies).

What you learned

In this tutorial, you built a document processing pipeline that demonstrates several key patterns:
  • Fan-out extraction — classifying documents by type and routing each to a specialized extraction node with tailored prompts and schemas
  • AI reconciliation — comparing AI-extracted data against application data with structured exception reports
  • Hybrid AI + business rules — combining AI-driven comparison with deterministic validation (expired documents, missing types, income thresholds)
  • Human-in-the-loop — routing edge cases to a reviewer while auto-approving clean results
  • Workflow composition — building modular workflows for classification, reconciliation, and summary generation, then orchestrating them from a BPMN process

Next steps

Fan-out extraction pattern

Scale the classification and extraction pattern to dozens of document types

AI comparison and reconciliation

Deep-dive into the reconciliation pattern with threshold tuning

Extract Data from File

Configure extraction strategies, image extraction, and signature detection

AI node types

Reference for all AI node types available in Agent Builder
Last modified on August 19, 2026