Vespper Launches Fine-Tuned Model and MCP Server for Word Document Editing
The Y Combinator-backed startup uses a dedicated reconciler model to translate HTML edits back into Office Open XML, cutting token overhead for autonomous agents.

Vespper, an artificial intelligence startup in Y Combinator’s Fall 2024 cohort, has launched Vespper DOCX MCP, a dedicated tool and fine-tuned model designed to let autonomous agents edit Microsoft Word documents. Announced on Hacker News, the release addresses developer challenges in legal, healthcare, and financial software, where formatted .docx files remain standard working deliverables.
While standard foundation models can generate and revise prose, modifying .docx files programmatically is notoriously difficult. A standard Word document is a compressed ZIP archive containing verbose Office Open XML (OOXML) structures. In these files, a four-sentence paragraph can expand into thousands of XML tokens once styles, split runs, numbering references, and metadata are included. Forcing a primary reasoning agent to parse and write raw OOXML consumes substantial context window capacity on file formatting rather than core domain analysis.
To solve this, Vespper adopted a round-tripping architecture inspired by Infrastructure-as-Code systems such as Terraform. The company converts Word documents into a clean, custom HTML representation that models can edit intuitively. Rather than maintaining a handwritten rule engine to reconcile edited HTML back into valid OOXML, Vespper trained a dedicated language model to handle the translation step.
The reconciler architecture uses a fine-tuned model in the 3 billion to 8 billion parameter range. A deterministic locator maps the agent's edited HTML block back to its original OOXML counterpart. The reconciler model then receives a triplet containing the original HTML, the revised HTML, and the original XML, predicting the updated XML block in a single inference pass without tool calls or reasoning loops. The resulting XML is deterministically diffed against the original block to patch the file and compute tracked changes.
Vespper trained the model using Low-Rank Adaptation (LoRA) via the Unsloth framework on Modal GPUs. The training set comprised roughly 16,000 self-supervised tasks generated from an underlying corpus of 2,046 enterprise documents sourced across government, healthcare, finance, and legal topics.
In internal vendor benchmarks evaluating 279 editing tasks run against OpenAI's GPT 5.6 Sol and GPT 5.6 Terra models, Vespper reported that its MCP approach ran 2.7 to 3.5 times faster and 2.7 to 2.9 times cheaper than existing document-editing skills while maintaining higher accuracy. The startup noted that performance gains were most pronounced on longer files and complex structural elements, including tables, headings, and tracked changes.
Sources
Written by
The Company Wire
Inside the companies building what’s next. Reporting on startups, technology, funding and the people shaping them.



