AppZen introduces ZenLM Plus, finance-specialized language models that outperform frontier models on finance T&E tasks

AppZen introduces ZenLM Plus, finance-specialized language models that outperform frontier models on finance T&E tasks
Bacakan Artikel

EmitenTrust.com New family of finance-specialized large language models led five of six audit control comparisons and reached 97.4 F1 across targeted policy-category cases, 8.7 points above the strongest frontier model tested.

SAN JOSE, Calif., Oct. 1, 2026 /PRNewswire/ -- AppZen, the leader in autonomous finance operations, today introduced ZenLM Plus, a new family of finance-specialized large language models built for the decisions at the center of finance operations.

Finance transformation is entering a new phase, moving beyond AI that assists finance teams to AI agents that can perform work autonomously. Achieving this shift requires specialized intelligence, trusted data, executable workflows, and enterprise governance operating as one integrated system.

AppZen has been building toward this vision through the evolution of its ZenLM models and the Mastermind Platform. Mastermind provides the harness, orchestration, workflows, and controls needed to apply that intelligence across finance operations. It intelligently routes each finance task to the model or method best suited to perform it, whether that is a specialized AppZen ZenLM Plus model, deterministic logic, or a frontier model. This enables finance teams to automate more work at a lower cost, while preserving the accuracy, governance, auditability, and evidence required for every decision.

"Autonomous finance requires more than a powerful model," said Anant Kale, CEO and co-founder of AppZen. "It requires the complete stack, with finance-specialized models, proprietary data, intelligent workflows, and enterprise governance working together to make decisions and take action with accuracy, transparency, and control. ZenLM Plus now outperforms the frontier models we tested on expense audit tasks while operating at a lower cost. Combined with AppZen's agent-first platform, it delivers on the promise of finance transformation, with our AI Agents safely performing more work, autonomously, at enterprise scale."

Benchmark results, tested against seven frontier models

AppZen compared ZenLM Plus with seven frontier models: Opus 5, Sonnet 5, Gemini 3.1 Pro, Gemini 3.5 Flash, GPT-5.6 Sol, GPT-5.6 Terra, and GPT-5.6 Luna. For each audit control, every system received the same relevant expense data, supporting documents, and customer configuration. AppZen scored each system on two measures: precision, the share of flagged expenses that were true violations, and recall, the share of true violations the system caught. It then combined the two into an F1 score of 0 to 100, where 100 indicates perfect accuracy.

ZenLM Plus led five of the six individual audit-control comparisons. It achieved an F1 of 97.4 across targeted policy-category cases, compared with 88.3 for GPT-5.6 Sol, the strongest frontier result. ZenLM Plus led all four consolidated policy groups, including advantages of 10.8 points in premium travel and upgrades, 4.2 points in electronic devices and gifts, and 2.5 points in policy and documentation exceptions.

Among the individual controls, the largest advantage appeared in non-conforming receipt detection. ZenLM Plus achieved an F1 of 92.4, 7.2 points ahead of Opus 5. This task requires more than recognizing a document. A payment slip may prove a card was charged without qualifying as a merchant receipt, while an emailed order confirmation may be valid even if it does not resemble a traditional receipt.

ZenLM Plus also led in catching duplicates across reports, receipt itemization verification, and merchant category matching. Receipt verification was the one control where a frontier model performed slightly better: ZenLM Plus achieved 93.3 F1, compared with 94.1 for Gemini 3.1 Pro and Sonnet 5.

"With ZenLM Plus, we are building the intelligence layer finance teams actually need," said Kunal Verma, CTO and co-founder of AppZen. "Our models are trained on real financial documents, policy decisions, and audit outcomes, then evaluated task by task before they reach production. That approach delivered stronger results on most controls we tested at a fraction of the modeled inference cost. And when a frontier model is better for a specific task, AppZen can use it. Our goal is the best outcome for the customer."

ZenLM Plus also delivered the lowest modeled inference cost of any system evaluated. Even GPT-5.6 Luna, the least expensive frontier model tested, was approximately twice as costly per 1,000 audited expense lines, while Opus 5 was approximately 50 times as costly. This advantage makes it practical to apply specialized AI across high-volume audit workflows while maintaining strong performance.

Purpose-built for the decisions finance teams manage every day 

ZenLM Plus supports a broad range of finance controls used by Travel & Expense (T&E) teams across AppZen Expense Audit. In this evaluation, ZenLM Plus covered:

  • Policy and category decisions: Applies customer-defined rules to premium travel and upgrades, electronic devices and gifts, personal expenses, merchant categories, fuel-service restrictions, and missing-receipt affidavits.
  • Receipt validity and transaction verification: Determines whether supporting evidence qualifies as a valid merchant receipt and whether the merchant, date, amount, expense type, and currency are consistent with the submitted expense.
  • Documentation and itemization requirements: Verifies that hotel and meal receipts contain the required line-item or daily detail, and that expenses are submitted within the organization's configured time limits.
  • Cross-report duplicate detection: Searches across the submitting employee's and other employees' reports to identify the same receipt or underlying purchase submitted more than once.

Depending on the control, ZenLM Plus draws on employee-entered fields, receipt extractions, payment type, a company's configured categories and thresholds, merchant information, card data, and related expense reports. Its judgment reflects how each customer defines risk rather than a general assumption of what a reasonable company would do.

Availability

 ZenLM Plus is currently available for select customers within AppZen's Expense Audit. It will be generally available in December 2026. For more information, please visit https://www.appzen.com/mastermind-ai-automation-platform.

About AppZen 

AppZen is the leader in autonomous finance operations, helping enterprises scale accounts payable, expense auditing, and compliance without adding headcount. Powered by agentic AI, its digital coworkers execute end-to-end finance workflows with built-in governance, full auditability, and enterprise-grade security. Trusted by the Fortune 500, including Amazon, Boeing, Salesforce, Novartis, JPMorgan Chase, VMware, and ServiceNow, the company serves global enterprises across a wide range of industries, including life sciences, manufacturing, financial services, and higher education. With AppZen, CFOs reduce operating costs by up to 50%, achieve automation rates of 80% or more, and realize measurable ROI within weeks. Learn more at www.appzen.com.

Media Contact

Megan Botta

Pitch Public Relations

megan@pitchpublicrelations.com

View original content to download multimedia:https://www.prnewswire.com/news-releases/appzen-introduces-zenlm-plus-finance-specialized-language-models-that-outperform-frontier-models-on-finance-te-tasks-302889866.html

SOURCE AppZen

Artikel ini dipublikasikan ulang secara utuh melalui program kemitraan media dengan prnewswire.com. Seluruh isi materi, data, dan sudut pandang editorial sepenuhnya merupakan tanggung jawab redaksi penerbit asli.