'From Ingestion to Agents: How AI Teams Build on Document Intelligence - Adit Abraham, Reducto'

Building agents that work in the real world requires solving the unglamorous problem of unstructured data processing; combining traditional CV, VLMs, and age...

By Sean Weldon

From Ingestion to Agents: How AI Teams Build on Document Intelligence

Abstract

This synthesis examines the architectural requirements for building production-grade agentic AI systems, arguing that the primary bottleneck is not model reasoning capability but the quality of unstructured document processing feeding these models. Drawing on operational experience with enterprise-scale document pipelines, the analysis traces a layered architecture combining sub-million-parameter computer vision (CV) models, vision-language models (VLMs), verification-based "agentic OCR," consumer-specific formatting, orchestration, and agent harnesses for tasks like chart extraction and structured extraction. Key findings include frontier models scoring approximately 30% on document-reasoning benchmarks, and structured document representations improving downstream accuracy while reducing token usage and latency. The analysis concludes that granular, stage-by-stage evaluation - rather than reliance on end-to-end benchmarks - is the decisive factor separating deployable systems from demonstrations.

1. Introduction

The first wave of commercially significant large language model (LLM) applications centered on information synthesis: retrieval-augmented generation (RAG), enterprise search, and question-answering chatbots. These systems shared a forgiving operational property - each interaction was largely single-step, with a human reviewing the output before consequential action was taken. The current wave of applications is qualitatively different. Agents increasingly perform end-to-end work: generating and modifying documents, executing multi-step workflows, and making autonomous decisions without per-step human adjudication.

This shift fundamentally changes where risk accumulates in a system. As multi-step agent pipelines compound decisions, errors introduced early - by a mis-parsed table or a misread handwritten annotation - propagate and amplify rather than being caught by a human reviewer at a single checkpoint. The central thesis explored here is that the value of LLM intelligence is bounded by the quality of the data and context to which it is applied. Consequently, the decisive engineering problem for real-world agents is not model selection but unstructured data processing: parsing, formatting, routing, and validating documents, underpinned by rigorous evaluation at every pipeline stage.

Section 2 establishes why document understanding remains an unsolved problem. Section 3 analyzes the core architectural components addressing it. Section 4 distills actionable technical findings, Section 5 discusses broader implications, and Section 6 concludes.

2. Background and Related Work

The PDF (Portable Document Format) is the canonical carrier of enterprise knowledge and simultaneously its principal obstacle. PDFs were specified decades ago to preserve print fidelity, not to support machine reasoning; the format encodes glyph positions rather than semantic structure. Compounding this, humans encode meaning visually - merged table cells express hierarchy, line charts express trends without enumerating underlying values, handwriting annotates printed forms, and creative slide layouts imply nonlinear reading order. Template-based extraction methods perform adequately on narrow, known document formats but fail systematically against the long tail of messy real-world documents.

Empirical support for this difficulty comes from the GDP PDF benchmark (attributed to Serge), on which frontier models score approximately 30% on document-based reasoning tasks. This is notable because these same models often approach or exceed human performance on textual reasoning benchmarks, indicating the bottleneck lies in visual-structural comprehension rather than raw reasoning capacity. A second reference point, a benchmark released by "micro one," characterizes a systematic precision-recall asymmetry between frontier reasoning models and dedicated document-processing services, discussed further in Section 3.

3. Core Analysis

3.1 Traditional Computer Vision and Vision-Language Models: Complementary Tools

Sub-million-parameter CV models remain highly effective for layout detection tasks and are deployable on CPU at scale, offering a low-cost, high-throughput mechanism for identifying document structure before any generative model is invoked. VLMs, by contrast, excel at semantic understanding - handling handwriting and unconventional layouts that defeat traditional OCR. VLMs are described as "fundamentally horizontal in nature," enabling a document to be read "the way that a human would have."

This capability introduces a specific risk: VLMs may "correct" content beyond strict fidelity to the source. For example, models will sometimes recalculate and insert totals into a table that were never present in the original document. To mitigate this, agentic OCR applies token-level edits to a base parse rather than performing full generative rewriting - a correction mechanism analogous to speculative decoding in coding tools such as Cursor. The resulting pipeline architecture is sequential: CV and VLM components produce an initial parse, which passes through a verification/correction layer before yielding a high-confidence output.

3.2 Formatting for Different Consumers

Parsing accuracy alone is insufficient if the resulting representation is poorly suited to downstream consumption. A dynamic table representation strategy renders simple tables in markdown while encoding complex or merged-cell tables in HTML, since markdown cannot express irregular cell structures. However, HTML table representations are token-expensive and perform poorly for embedding-based retrieval. The solution employed is a dual representation: natural language summaries of tables improve embedding-based retrieval accuracy, while the full HTML representation is retained separately for LLM reasoning once a document has been retrieved.

More broadly, providing structured PDF representations rather than raw PDFs improves accuracy and reduces reasoning tokens and latency across Gemini, Anthropic, and OpenAI models - in some cases outperforming the Fable benchmark out-of-box. This indicates that formatting decisions are not cosmetic but materially affect both cost and correctness.

3.3 Orchestration via Classification and Splitting

Dumping full document context into a model causes quality erosion, not merely increased token cost. Classification and splitting serve as an orchestration layer, routing document segments to appropriate downstream pipelines and reducing distraction from irrelevant interleaved content. A representative example involves paper mail packets spanning hundreds of pages with unrelated content interleaved throughout - without splitting, models must contend with distractor content that measurably degrades extraction quality on the target segment.

3.4 Agent Harnesses for Unsolved Extraction Problems

Agent harnesses extend the pipeline to problems that static parsing cannot solve. Extracting precise numerical data from line charts, for instance, requires an agent equipped with a code interpreter and a visualization feedback loop, iteratively reconstructing a data table and rendering it back into chart form until it matches the source image.

For high-cardinality structured extraction (illustrated by CBP forms containing tens of thousands of fields), a parent-agent/sub-agent validation architecture proves effective. This addresses a documented tradeoff: per a benchmark released by "micro one," frontier reasoning models operating at maximum reasoning depth exhibit high precision but poor recall, silently dropping rows, while dedicated document-processing services exhibit the inverse tradeoff. The agent harness approach - decomposing extraction into a parent agent that delegates to validating sub-agents - identifies a local maximum that balances precision and recall better than either approach alone.

4. Technical Insights

Several implementation-level findings recur across the pipeline stages described above:

A key limitation acknowledged in this framework is that static benchmark datasets, including GDP PDF, diverge from production data distributions; consequently, evaluation against live production data is treated as more diagnostic than performance on fixed benchmarks.

5. Discussion

These findings suggest a broader industry pattern: as agentic systems move from demonstration to deployment, the locus of engineering effort shifts away from prompt and model selection toward the unglamorous infrastructure of data preparation. The compounding-risk property of multi-step agents means that even marginal improvements in upstream parsing fidelity can have outsized effects on end-to-end reliability, a dynamic not visible in single-turn RAG evaluations.

An open question raised by this analysis concerns the eventual role of deterministic pipelines altogether. The described trajectory - moving from fixed, deterministic processing pipelines toward agent-navigated file systems, where agents select their own tools and decide which document types to read - suggests that orchestration itself may become an agentic decision rather than an engineered rule. This is operationalized through a content field versus metadata field distinction (e.g., bounding boxes retained as metadata for citation purposes) that allows an agent to navigate documents flexibly rather than following a rigid extraction schema.

This trajectory also implies an expansion beyond extraction toward document generation, indicating that the underlying infrastructure for reading documents accurately is a prerequisite for reliably producing them.

6. Conclusion

This synthesis demonstrates that production-grade agentic systems are constrained less by frontier model reasoning capability than by the fidelity and structure of the data pipelines feeding them. The layered architecture - combining CV models, VLMs, verification-based correction, consumer-aware formatting, orchestration, and agent harnesses - addresses distinct failure modes at each stage, and granular evaluation at every stage is what distinguishes deployable systems from demonstrations. Practically, teams building document-dependent agents should prioritize stage-wise evaluation on production data, adopt dual representations tuned separately for retrieval and reasoning, and consider agent-harness architectures where static extraction pipelines exhibit precision-recall imbalances.


Sources


About the Author

Sean Weldon is an AI engineer and systems architect specializing in autonomous systems, agentic workflows, and applied machine learning. He builds production AI systems that automate complex business operations.

LinkedIn | Website | GitHub