Georgia Tech and J.P. Morgan Develop AI Framework to Fix Errors in Visual Document Parsing
The SlideAgent architecture mimics human reading habits to improve multi-page document accuracy by up to 10 percent.

Researchers from Georgia Tech and J.P. Morgan AI Research have introduced SlideAgent, a new artificial intelligence framework engineered to improve how large language models parse visual, multi-page business documents, as first reported by TechXplore. The system addresses a persistent reliability gap in enterprise automation, where standard visual AI models frequently overlook embedded details, misread complex charts, or misinterpret footnotes across slide decks and corporate reports.
While multimodal models like OpenAI’s GPT, Google’s Gemini, and Anthropic’s Claude have become common workplace tools for document summarization, their architectural limitations pose operational risks. In sectors such as financial services, small errors in data extraction can distort strategic planning and risk evaluation. Yiqiao (Ahren) Jin, a Ph.D. candidate at Georgia Tech’s School of Computational Science and Engineering and the lead author of the study, noted that while these models save time, their imperfections carry significant costs in high-stakes environments where misreading a single metric or footnote can alter executive decisions.
SlideAgent tackles these shortcomings by restructuring how AI processes dense visual data. Standard multimodal software typically evaluates an entire document page as a single image, which often leads to inaccurate visual counting or missed relationships between graphics and accompanying text. SlideAgent instead mimics human cognitive patterns by examining material across three distinct structural tiers: the narrative of the complete document, the context of individual pages, and specific isolated elements such as tables, graphs, and text blocks.
To execute this multi-tier review, SlideAgent deploys a network of specialized AI agents. Each agent operates at a designated structural level before aggregating its findings into a unified, structured analysis. This multi-page contextual reasoning allows the software to pull precise evidence from targeted elements while retaining awareness of the broader document narrative.
During benchmark evaluations across complex datasets—including technical slide decks and corporate financial presentations—SlideAgent demonstrated consistent performance advantages over existing tools. The framework achieved a 7.9% accuracy improvement over its baseline proprietary model and a 9.8% increase over evaluated open-source baseline models, yielding performance gains of up to 10% on complex visual reasoning tasks.
The underlying research grew out of Jin’s internship at J.P. Morgan AI Research, where he collaborated with co-authors Rachneet Kaur, Zhen Zeng, and Sumitra Ganesh. Jin’s academic work at Georgia Tech is supervised by associate professor Srijan Kumar. The team’s findings were detailed in a paper titled "SlideAgent: Hierarchical Agentic Framework for Multi-Page Visual Document Understanding," published on the arXiv preprint server.
The research was accepted for presentation at the 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026), held July 2–7 in San Diego. Moving forward, the research team aims to refine the framework’s operational efficiency, with Jin highlighting the goal of reducing computational demands while preserving the system's accuracy and analytical transparency.
Sources
Written by
The Company Wire
Inside the companies building what’s next. Reporting on startups, technology, funding and the people shaping them.



