← All research

findings

What Mercor's APEX-Accounting benchmark measures, and what it misses

A single audit case corpus: 12,840 subledger rows, $8.305M

A balance sheet grid with one cell traced back through a chain of source documents; untraced cells fade to outline

Mercor and Ramp released APEX-Accounting on July 31: a benchmark for reconciliations, variance analysis, and closing the books. It’s a serious benchmark. Real workflows, expert-authored pass/fail criteria, open release. It’s also a benchmark that cannot carry evidence integrity, which is what audit work is actually graded on. A reference audit case we examined makes the distinction concrete: a 12,840-row inventory subledger totaling $8.305 million, seeded with planted discrepancies whose combined effect exceeds the engagement’s materiality threshold. Accounting tasks were among the 96.5% without definitive benchmark coverage in June. That gap is now closing. The question is which domains remain open.

Key Takeaways

  • APEX-Accounting brings the APEX playbook to accounting: real workflows, expert-authored pass/fail criteria, open release. It covers the mechanical core of the job.
  • What no reconciliation benchmark can carry is evidence integrity: whether every number traces to a source document and unsupported entries get escalated. That’s what audit work is graded on.
  • Payroll, medical billing, claims adjusting, freight dispatch, and bookkeeping on a real ledger still have no agent environment under any phrasing.

What does APEX-Accounting get right?

APEX-Accounting applies the playbook that made APEX-Agents credible: real workflows rather than isolated calculations, expert-authored pass/fail criteria, and open release. The APEX family is documented research, not just marketing: the original index built expert-written tasks across banking, consulting, law, and medicine (Vidgen et al., 2025), and the agentic follow-on put long-horizon, cross-application tasks on arXiv (Vidgen et al., 2026). APEX-Agents launched with 480 tasks across 33 simulated worlds covering investment banking, law, consulting, medicine, software engineering, and accounting. APEX-Accounting narrows the focus to one domain and goes deeper into reconciliations, variance analysis, and closing the books, the mechanical core of accounting work.

The open release matters. A benchmark that labs can run themselves is a benchmark that labs will cite in system cards. Mercor’s APEX family is becoming part of the frontier measurement layer, and APEX-Accounting extends that layer into accounting.

It’s also a real step past the last generation of finance evals, which were mostly question answering: FinQA tested numerical reasoning over report snippets, and FinanceBench showed retrieval-equipped GPT-4 still hallucinating on open-book questions about real filings. Workflow execution with expert pass/fail criteria is a different, harder category.

What can a benchmark not carry?

Evidence integrity. Real audit work is graded on whether every number traces to a source document: a subledger entry, a journal voucher, a shipping receipt, a receiving log. A reconciliation task that asks “does the GL match the subledger?” tests arithmetic. An audit task that asks “does every GL entry trace to supported source evidence, and are unsupported entries identified and escalated?” tests the actual competency.

The research on LLM auditing finds exactly this split: models can spot statement errors at reasonable rates but fail at explaining them, citing the governing accounting standards, and completing a full audit workflow (Wang et al., 2025). Finding a wrong number is the easy half. Carrying the evidence chain is the job.

Our reference audit case makes it concrete. The case is built on a 12,840-row inventory subledger totaling $8.305 million, seeded with planted discrepancies (unsupported journal entries, cutoff errors, a missed write-down) whose combined overstatement exceeds the engagement’s materiality threshold before any sample projection. A reconciliation benchmark would test whether the agent matches the GL to the subledger. An evidence-integrity environment tests whether the agent identifies every unsupported entry, traces each number to its source, and escalates the control failures. The pass rates on the first tell you nothing about the second.

Which domains are still open?

Accounting tasks were among the 96.5% of tasks without definitive benchmark coverage in the June coverage study. APEX-Accounting closes part of that gap: the reconciliation and closing-workflow part. The evidence-integrity part remains open, and it’s the harder part to build.

The domains that remain genuinely unoccupied:

  • payroll and HRIS
  • medical billing on real software
  • insurance claims adjusting
  • freight dispatch
  • hotel property management
  • warehouse management
  • support ticketing on real software
  • bookkeeping on a real ledger

None has an agent environment under any phrasing. These are the domains where the 96.5% gap is not closing at all.

What this means

APEX-Accounting is a step forward for accounting benchmarks. It is not a substitute for evidence-integrity environments, because a benchmark cannot carry the source-document chain that audit work is graded on. The domains that remain open are the ones where no benchmark exists at all.

FAQ

What is evidence integrity in audit work?

Evidence integrity means every number in a workpaper traces to a source document: a subledger entry, a journal voucher, a shipping receipt. An agent that produces a numerically correct total without source-document support has not completed the audit. The evidence chain is the competency, not the arithmetic.

What is the 96.5% coverage gap?

The June coverage study found that only 3.5% of 202 O*NET occupational tasks have definitive benchmark coverage. The remaining 96.5% are partial-only or uncovered. Accounting tasks were in that gap. APEX-Accounting closes part of it; the evidence-integrity part remains open.

What is APEX-Accounting?

APEX-Accounting is a benchmark released by Mercor and Ramp on July 31, 2026, covering reconciliations, variance analysis, and closing the books. It uses expert-authored pass/fail criteria and is openly released. It is part of Mercor’s APEX benchmark family.