Westworld Finance Diligence Bench

Evaluating AI Agents on a Complete company acquisition due-diligence process

We present Westworld Finance Diligence Bench, a benchmark of 88 problems that evaluates AI agents on a complete company acquisition due-diligence process. It draws on anonymized data from real private transactions, runs inside dynamic desktop environments, and spans trajectories that reach hundreds of steps. Every problem is written and reviewed by practicing finance deal professionals and verified by a mixed pipeline of agentic, discrete, and binary verifiers.

We evaluate a range of frontier models, each across its native harness and our internal Halluminate harness, and present every model-harness pair three ways.

Average mean score by model and harness combination, toggleable between a histogram of mean score with cost per run inside each bar and a Pareto chart of Westworld Finance Diligence Bench score against average cost per run.

The histogram shows each configuration's mean score with average cost per run inside the bar. The Pareto 2D plot trades that cost against mean score and the summary table adds average tool usage, cost, mean score, and pass rate. Mean score is the per-problem pass@1 verifier grade averaged over the 88 problems, always between 0 and 1. Pass rate is the share of those problems whose single graded run clears the 0.50 bar, meaning it scores at least 0.50; the 0.50 threshold is a reporting choice and does not make the underlying reward binary, as we explain below. Cost is estimated from each run's token usage and the model's published input/output/cached pricing, amortized across the 88 tasks.

Key Innovations of our Benchmark

  • 1Real private-deal data: anonymized documents from actual private equity transactions. To the best of our knowledge, no published benchmark evaluates a complete diligence process on private deal data of this breadth.
  • 2A complete process: tasks that together span one deal's diligence, from early analysis through to close, rather than isolated tasks tied to a single job title.
  • 3Mixed verifier pipeline: each problem is scored by a decomposed mix of agentic rubric grading and deterministic checks, weighted by expert authors and audited in an independent QA pass.
  • 4Computer use and tool use together: a single task requires operating a real desktop environment, with files, office applications, and a data room, while also calling structured tools for email and chat, all in one trajectory. A few benchmarks combine these two modes, but it remains rare, and rarer still on finance work.

The rest of this post walks through the benchmark in detail. We first place the benchmark in the context of related work. Then, walk through a sample of a single task, following it from the prompt through the tools, verifiers, and graders. We then describe the quality process we use to keep environments healthy. Next, we present detailed results from our runs, including a per-trajectory error analysis. We close with our takeaways and future work.

1. Previous Works

Benchmarks for professional domains have grown narrow by construction: accounting tasks in one suite, investment banking tasks in another, each targeted at a single task type or a single job title. For instance, some of these are:

  • FinanceBench - 10,231 open-book question-answer pairs about public companies, testing whether a model can pull the right figure from a filing.
  • FinQA - expert-written numerical-reasoning questions over single earnings reports, each paired with an executable reasoning program.
  • ConvFinQA - a conversational extension of FinQA that chains multi-step numerical questions across the turns of a dialogue.
  • BizBench - eight quantitative-reasoning tasks that grade financial problem solving through program synthesis over structured data.
  • FinanceQA - hedge-fund, private-equity, and investment-banking analysis questions covering hand-spreading, valuation conventions, and reasoning under incomplete information.

However, this decomposition no longer matches how AI systems enter these domains. Organizations now deploy agents against whole processes. This complete-process way of using AI has no frontier-quality benchmark, only piecewise tests that cannot be combined into a full picture. Some works that have attempted to address this are:

  • DealTrace - our own private-equity deal review across ten real deals and five sequential stages (extract, reconcile, forecast, market, recommend); a first step toward process-level evaluation, but narrower than a full diligence suite.
  • Finance Agent Benchmark - 537 expert questions across nine categories answered by an agent with web-search and EDGAR access; agentic, but still discrete research questions over public filings rather than one connected deal.
  • TheAgentCompany - long-horizon agent tasks inside a simulated software company that combine computer use (a real browser and desktop applications) with tool use (code execution, chat, and internal services) in the same task; the closest general analogue to whole-process work precisely because it exercises both modes together, though with almost no finance content.
  • OSWorld - 369 open-ended computer-use tasks across real operating systems and applications, establishing the stateful multi-app desktop setting but without any deal or finance domain.
  • tau-bench - multi-turn tool-agent-user interactions graded against database state and policy rules, testing sustained tool use, but on customer support rather than analyst deal work.

Westworld Finance Diligence Bench targets this gap, simulating the full body of work that a team of investment bankers, consultants, and accountants would produce together over several weeks. Unlike pure computer-use benchmarks such as OSWorld, our agents operate in the environment through code, structured tools, and mouse-driven GUI control. Trajectories regularly run to hundreds of steps, and the longest model and harness pairings average close to nine hundred iterations per run, as the trajectory analysis below shows.

2. Coverage

The benchmark spans the full arc of a deal, from the first model to the signed close.

Problems are organized into the seven task categories shown above. They are distributed across these categories and across the professional roles that would own them, so that no single specialty dominates and no stage of the process goes untested. Within a single problem, an agent may need to search a data room, reconcile spreadsheet models, draft client-facing documents, and communicate results over email, mirroring how the underlying work actually flows between artifacts and channels. Here we show the number of tasks in each of the seven task categories.

Task categoryNumber of tasks
Modeling14
Underwriting13
Quality of earnings9
Data preparation11
Pitch materials15
Closing4
Diligence review22

3. Sample Task

Every task in the benchmark was authored by a finance expert. They wrote the prompt, constructed the ground truth, and chose the verifiers and their weights so that the score reflects what actually matters in the work. Verifier weights were chosen to reflect the relative value of each individual verifier on the overall quality of the output. For example, saving a spreadsheet under the right file name is important, but much less so than the substance of the analysis performed, and therefore the filename check is weighted much less than the analysis check. Every problem went through our automated QA pipeline and findings were reviewed by our core team (detailed in the quality-assurance section below).

Problems follow the same shape: the agent receives a set of inputs (a prompt and supporting context), acts on the environment through tools (which can be MCP or GUI-based) to produce a deliverable, which is scored by a mix of agentic and deterministic verifiers. The grade is not a single judgment: it decomposes into several weighted verifiers. Each verifier returns either 0 or 1 (the binary checks and the rubric grader) or a proportional fraction, such as the share of cells in a range that match ground truth. The task score is the weight-weighted average of those per-verifier scores, so it always lands somewhere in 0 to 1.

Here we walk through a single task end to end. Note that the verifiers change from one task to the next.

Sample task

Prompt

Review and redline the "11.02.01.05 CAMA Statement of Work" file from the data room. We are representing Eagle Real Estate Software. The redlined output file should be named "11.02.01.05 CAMA Statement of Work Redlined".

Context given

The desktop might or might not have files. In this case, it starts without files.

/06 Legal/06.02 Customer Information/Customer Contract Templates/11.02.01.05 CAMA Statement of Work.xlsx212 KB
11.02.01.05 CAMA Statement of Work
11.02.01.05 CAMA Statement of Work
ReferenceProvision
Section 2.1(b)Included modules (GIS, CostBridge, ...)
Section 3.2Licensed users (named and read-only)
Section 7Data conversion scope (assessment history, sketches, ...)
Section 8Import / export functionality
Section 9.1(a)Onsite training days
Section 10Payment milestones and Net 30 terms
Section 12.1Termination notice period
Exhibit AImplementation schedule and temporary hosting

Preview shows the first 5 rows. The full statement of work continues below.

The one file named in the prompt, inside a 156-file data room. Click it to preview.

21 emails in 9 threads
Search emails...
John ReynoldsJul 9
Legal review (2)
Sarah MartinezJul 8
Imports requirements
John ReynoldsJul 8
Internal convo (2)
Sarah MartinezJul 8
user count (3)
Sarah MartinezJul 8
Conversion discussion (2)
Sarah MartinezJul 8
Kickoff – CAMA Implementation
From: Sarah Martinez <sarah@it.com>  ·  To: John Reynolds <john@eaglere.com>  ·  Kickoff – CAMA Implementation Discussion
Hi John,
We're excited to move forward.
Before legal review we'd like to discuss a few items:
- Increase licensed users from 18 to 24.
- We'd like GIS included.
- We don't think we'll need CostBridge.
- We'd prefer all implementation to occur on-premise rather than Eagle Cloud.
- We'd like two additional training days.

Let's discuss during kickoff.

Thanks,
Sarah

One message from a thread of about twenty; the terms keep moving across the conversation.

Conversations
Sarah Martinez & Jane Chen2
Sarah MartinezIt doesn't look like I have all the information to fill out the SOW template.
Jane ChenJust fill out what you can and add in comments on any additional insights.
Notes
Project Kickoff
Sales Leadership Mtg
Steering Committee
Follow-up meeting
Project Kickoff Meeting Notes
Attendees
  • John Reynolds – Eagle
  • Sarah Martinez – PM
  • Jane – County Assessor
  • Sam – Procurement
Customer requests
  • GIS module included
  • CostBridge excluded
  • Customer prefers on-prem
  • Cloud may be required if servers fail readiness review
  • Three days onsite training (one already included)
  • Procurement requests milestone payments shifted toward completion
Action items
  • Eagle to propose revised milestone schedule.

Desktop

The agent's working desktop, empty at the start. It downloads its inputs here and saves the final redlined deliverable back to it.

Possible actions
File managerOffice suite
Data Room

The deal's full document repository of 156 files. The one file named in the prompt sits among the rest of the deal's paperwork.

Possible actions
List folderSearch filesGet file infoDownload fileUpload fileCreate folderCreate share linkDelete file
Email

The full inbox, 21 emails across 9 threads. The negotiation plays out here, with the terms shifting from message to message.

Possible actions
Read emailSearch emailsDraft emailSend emailSend draftUpdate draftGet draftList draftsDelete draftModify emailBatch modify emailsDelete emailBatch delete emailsDownload attachmentList email labels
Chat App

Direct messages and channels with coworkers, used for quick clarifications on how to handle gaps in the task.

Possible actions
List channelsList membersGet channel messagesGet direct messagesSearch messagesSend channel messageSend dmCreate channel
Notes

Short, read-only meeting notes referenced in the prompt. Background context on what the customer asked for, not hints.

Possible actions
Read-only context

Desktop environment

The agent works inside a Docker container with a full Linux desktop. Beyond the email, chat, notes, and data-room apps shown in the context above, it drives the real desktop apps below. Click an outlined icon to see the tools it exposes. Market data has no desktop app: the agent reaches it from code or a market-data tool.

The agent's Linux desktop. The real LibreOffice Calc, Writer, and Impress icons in the dock, the File System icon, and the PDF viewer icon are outlined to show they are interactive.
Spreadsheets
18 tools

Open, read, edit, format, and save spreadsheets, including formulas and charts.

Actions
Add chartAdd sheetCreate workbookDelete sheetFormat cellFormat rangeGet commentGet infoModify chartRead cellRead commentsRead rangeSave workbookSet formulaSet formula rangeSet formulas batchWrite cellWrite range
Documents
14 tools

Open, read, edit, and save word-processing documents.

Actions
Add bulleted listAdd headingAdd imageAdd numbered listAdd page breakAdd paragraphAdd tableCreate documentGet infoRead documentReplace textSave documentSet footerSet header
Presentations
22 tools

Open, read, edit, and save slide decks.

Actions
Add chartAdd imageAdd slideAdd tableAdd text boxCount slidesCreate presentationDelete presentationDelete shapeDelete slideDuplicate slideFormat textList shapesRead all slidesRead slide textRead speaker notesReorder slidesReplace textSave presentationSet slide layoutSet speaker notesUpdate table cell
File system
2 tools

Browse folders and search for files on the desktop.

Actions
List directorySearch files
PDF viewer
3 tools

Open and read PDF documents.

Actions
Get infoRead allRead page

The agent can drive these apps through the GUI, or write and run code to do the same work.

Verifiers

Every task is graded by two kinds of verifiers working together. Deterministic verifiers run exact, code-level checks on the deliverable, while the agentic verifier reads the document and grades open-ended criteria against a rubric. Both feed a single pool of reward weights: each check, deterministic or rubric, carries its own weight, and the task's score is the weighted average of every check's score divided by the total weight, so it always falls between 0 and 1. Most checks return a simple pass or fail, but some deterministic checks award proportional partial credit, for example the fraction of cells in a range that match ground truth, which is what makes the final score fine-grained rather than a coarse pass count. There is no fixed deterministic-versus-agentic ratio; the share each side contributes is simply what its own checks' weights add up to, and it shifts from task to task. In this example the deterministic checks carry about 40% of the total weight and the rubric carries the other 60%.

Exact, code-level checks. Click one to see what it verifies and its reward weight.

Correct output filename

Description: Checks that the redlined document is saved with the exact expected file name.
Solution: 11.02.01.05 CAMA Statement of Work Redlined.docx
Reward weight: 4%

Tracked changes on

Description: Checks that track changes is turned on, so every edit shows up as a visible redline.
Solution: Tracked changes enabled.
Reward weight: 8%

GIS module included

Description: Checks that the contract now states the GIS module is included.
Solution: Text reads "includes Eagle's commercial off-the-shelf GIS module".
Reward weight: 6%

CostBridge excluded

Description: Checks that the CostBridge module is now excluded.
Solution: Text reads "does not include Eagle's commercial-off-the-shelf CostBridge module".
Reward weight: 7%

Three training days

Description: Checks that three days of onsite training are specified.
Solution: Text reads "Three (3) days of training".
Reward weight: 5%

Signing payment milestone

Description: Checks the first payment milestone due at signing.
Solution: Text reads "15% due upon signing".
Reward weight: 6%

Delivery payment milestones

Description: Checks the remaining payment milestones tied to delivery.
Solution: Text reads "21.25% due upon completion of initial database mapping" and "21.25% due upon installation".
Reward weight: 9%

Thirty-day termination notice

Description: Checks that the termination notice period stays at thirty days.
Solution: Text reads "terminate this Schedule upon thirty (30) days' written notice to Eagle".
Reward weight: 10%

Net 30 invoice terms

Description: Checks that invoices remain payable within thirty days.
Solution: Text reads "30 days of receiving an invoice".
Reward weight: 10%

The agentic verifier reads the finished document and grades these criteria against the rubric, each with its own reward weight.

1. Commercial Terms
   - Correctly updates Section 2.1(b) to include the GIS module and exclude the CostBridge module. (4%)
   - Correctly updates Section 3.2 to reflect 24 named Users and 6 read-only Users. (7%)
   - Updates Section 9.1(a) to provide three (3) days of onsite training. (4%)
   - Updates Section 10.1 payment milestones to reflect the negotiated percentages totaling 100%. (7%)
   - Leaves Net 30 payment terms unchanged in Section 10.3. (2%)
   - Leaves the 30-day termination provision unchanged in Section 12.1. (2%)

2. Data Conversion Scope
   - Updates Section 7.1(b) to exclude historical property sketches while retaining assessment history and ownership transfer history. (7%)
   - Updates Section 7.2(d) to align with Section 7.1(b). (4%)
   - Correctly handles Personal Property conversion language without contradicting the negotiated scope. (2%)

3. Import / Export Scope
   - Revises Section 8.1 to clarify that standard imports/exports are included while nightly automated exports are excluded from the SOW. (5%)
   - Preserves Section 8.2 by requiring future automated exports to be completed through a Change Order. (4%)

4. Implementation Schedule
   - Updates Exhibit A to reflect temporary Eagle Cloud hosting during customer hardware delays. (4%)
   - Adds clarification that temporary hosting is provided at no charge until production hardware is available. (4%)
   - Clarifies that temporary hosting does not alter milestone sequencing or payment obligations. (3%)
   - Updates the historical database conversion milestone to reflect the exclusion of historical sketches. (3%)

5. Internal Consistency
   - All cross-references remain accurate after revisions (Sections 3, 7, 8, 9, 10, and Exhibit A remain consistent). (8%)
   - No approved change is contradicted elsewhere in the document. (8%)
   - Superseded requests from earlier emails are not incorporated into the final SOW. (8%)

6. Formatting
   - Any newly inserted or completed text is black, matches the surrounding font, font size, spacing, and numbering, and does not appear as colored or highlighted text. (5%)
   - Existing document structure, section numbering, indentation, and overall formatting are preserved. (4%)
   - Placeholder text is replaced only where sufficient information was provided; unresolved placeholders remain unchanged. (3%)

The prompt is crisp, aiming to reflect real requests in the real world. The task ships with a thread of about twenty emails in which the terms keep moving. The agent has to reconcile the full conversation and redline the contract to the final negotiated state. In this case, the data room holds 156 files. For each task we also decide which tools and verifiers are appropriate for the problem. You can see the ones chosen for this sample in the GUI & tools and Verifiers tab.

Click to know more about our Halluminate harness

The Halluminate harness aims to create an equal playing field for evaluating different models' tool-use and computer-use abilities. It does not let the agent execute code. The action space within the Halluminate harness is:

  • 1Read the current state - the task prompt, notes, emails, chat, and the outputs of earlier tool calls.
  • 2Call a tool - invoke any of the tools listed above with structured arguments; the harness runs it and returns the result. Every change to the environment happens through a tool call.
  • 3Use the computer tool - operate the desktop UI directly through screenshots, mouse, and keyboard.
  • 4Reason in text - write intermediate reasoning between calls to plan the next step.
  • 5Finish - stop calling tools to end the run; the agent's work is whatever it has saved to the desktop and data room.

The harness gives the agent no shell tool and no code-interpreter tool: it drives the desktop through a computer tool (screenshots plus mouse and keyboard) and calls structured tools, reasons, and finishes, but it cannot run bash or Python to script around them.

Coverage alone is not sufficient; each problem also has to clear a quality bar.

4. Quality Assurance

To ensure realism and optimize this benchmark for real-world testing, every completed run is put through six independent post-job checks.

Quality assurance is a combination of human and agentic review. Finance experts author each problem, then they go through the automated QA pipeline described here. All QA findings raised by the pipeline are adjudicated by a core-team member, who is responsible for validating then accepting or rejecting, and applying fixes where necessary.

Open any check below to see what it looks for, with an example:

Quality checks

A single Claude agent (claude-opus-5) runs inside the task's Docker container with shell and file-read access. It reads the problem statement, the seeded emails and chat, and the verifier spec, then extracts every requirement, meaning each concrete thing the task asks the agent to produce or satisfy: a specific number to compute, a file to save under an exact name, an email that must be sent, a contract clause that must change. It maps each requirement to the verifiers that grade it (covered, partially covered, or not covered) and flags requirements with missing or weak coverage, verifiers that check things the task never asked for, and output filenames the grader expects but never communicates to the agent.

Pseudocode
agent = ClaudeAgent("claude-opus-5", tools=[Bash, Read])  # in task container
reqs  = agent.extract_requirements(problem, emails, chat)   # each thing the task asks for
for req in reqs:
    cov = agent.map_to_verifiers(req, spec.yaml)            # covered/partial/none
    if cov in {none, partial} or req.ambiguous:
        findings.add(issue, evidence, severity, category)
return findings

Example: a contract-editing task tells the agent to save the finished contract as Hemi_MSA_vFINAL.docx. The verifiers check the edited clause and that track changes is on, but none of them check the file name, so an agent that saves output.docx still scores full marks.

Requirement coverage map
RequirementGraded by
Edit the GIS clauseclause_text
Turn on tracked changestracked_changes
Save as Hemi_MSA_vFINAL.docxno verifier

A claude-opus-5 agent reviews the problem from inside the container with shell and file access. It walks ten checks:

  • Naming consistency: file names, company names, and version numbers match across the problem statement, the hints, the verifier spec, and the actual files.
  • Hints: any hints help the agent along without being essential to the solution or giving away the exact answer.
  • Data presence: every file, email, and chat message the task refers to actually exists where it should.
  • World and identity coherence: people, companies, and dates line up, and the agent's own identity is not confused with someone else's.
  • Output filename communication: the exact name the deliverable must be saved under is stated somewhere the agent can see.
  • Verifier best practices: the verifiers focus on what matters, are not redundant, and do not lean on weak or degenerate judges.
  • Grading design: each verifier's weight reflects its importance to the task, no single check is disproportionately weighted, and partial credit is used where all-or-nothing grading would be too harsh.
  • Manufactured difficulty: the challenge comes from real domain competence, not from tedium, bulk edits, or waiting on timers.
  • Visual quality: the seeded and reference spreadsheets, documents, and decks look professionally built and are pleasant to look at, not obviously auto-generated.
  • Answer leakage: no graded answer, ground-truth value, or exact formula is exposed on a surface the agent can reach.
  • Authoring artifacts: no leftover template text, placeholder personas, serialization junk, or misspellings betray how the problem was assembled.

For the visual-quality check it renders every spreadsheet, document, and slide deck with LibreOffice and reads back the images, judging whether they look clean and professional rather than just numerically correct. Before flagging answer leakage it first proves the file is actually readable by the agent's own user.

Pseudocode
agent = ClaudeAgent("claude-opus-5", tools=[Bash, Read])
for art in office_files:                     # .xlsx / .docx / .pptx
    png = render(art, "soffice + pdftoppm")  # inside container
    agent.inspect(png)                       # is it nice to look at?
for path in graded_answers + gt_values:
    if readable_by(model_user, path):        # reachability guard
        findings.add(answer_leak)
findings += agent.check(naming, hints, identity, difficulty)
return findings

Example: the task says to compute the FY24 blended gross margin, but the data room only contains statements through FY23. There is no FY24 revenue or cost anywhere in the files, so the number the grader expects cannot be derived.

Income_Statement.xlsx (data room)
FY22FY23
Revenue41,20044,800
COGS26,90028,510
Gross margin34.7%36.4%

A claude-opus-5 agent audits the running environment, given the shell, file access, and the compressed traces of three real runs. It walks a fixed checklist:

  • Are the required tools present and returning the right data?
  • Does every file, email, and message the task references actually exist?
  • Are the seeded workbooks internally consistent (do totals sum, do balances balance)?
  • Do the fixture values agree with what the graders expect?
  • Is the agent's identity coherent across its email and chat accounts?
  • Does any surface leak a privileged file or answer key?

It must prove a file is reachable before calling it a leak.

Pseudocode
agent = ClaudeAgent("claude-opus-5", tools=[Bash, Read])
trajs = compress(runs[:3])                   # first/last/error steps + scores
for c in [TOOLS, RESOURCES, FIXTURES, FIXTURE_vs_GRADER,
          IDENTITY, SEED, BUDGET, LEAK]:
    agent.walk(c, container, trajs)          # docker exec to confirm
    # e.g. open workbooks: do totals sum? do balances balance?
return findings                              # empty if healthy

Example: the prompt tells the agent to open Bid_Comparison.xlsx on the Desktop, but that workbook was never seeded. Only unrelated files are present, so the task cannot even start.

/home/model/Desktop
LOI_Review.xlsx
Kickoff_Notes.txt
Bid_Comparison.xlsx (not found)

A claude-opus-5 agent audits the verifiers across at least two runs of the same problem. For each verifier it forms its own verdict on each run, reading what the agent actually did from that run's trajectory and comparing against the ground truth, then checks that verdict against the score the verifier gave. It flags verifiers that:

  • pass everything, even low-quality or wrong work;
  • fail everything, even correct work;
  • pass wrong work or fail correct work;
  • give unstable scores on work of similar quality;
  • error out instead of returning a judgment.

It reads the verifier code and ground truth from the container; the run trajectories are embedded in its prompt.

Pseudocode
agent = ClaudeAgent("claude-opus-5", tools=[Bash, Read])  # needs >= 2 runs
for v in verifiers:
    mine = [agent.assess(u.trajectory, gt) for u in units] # own verdict
    if constant(v.scores):       flag(v, "trivial pass/fail")
    if mine != v.verdicts:       flag(v, "false verdict")
    if unstable(v.scores, mine): flag(v, "nondeterminism")
return findings

Example: the memo must conclude the deal is not recommended, but the verifier only checks that the word "recommend" appears anywhere in the text, so a memo that reaches the opposite conclusion still passes.

Recommendation_Memo.docx (agent output)
"On balance, we recommend proceeding with the acquisition at the proposed price."
Keyword verifier matches "recommend" and returns 1.0, despite the wrong call.

Two red-team checks look for reward hacking:

  1. A cheap deterministic probe submits an empty, degenerate deliverable and reports what it scores.
  2. An agentic red-teamer goes further using two models:
    • a solver that actively tries to cheat the graders (claude-opus-5) and a judge that decides whether it really cheated (claude-sonnet-5). The solver is told to earn credit without doing the work:
      • build the emptiest shell that still satisfies the verifiers (right file names, sheet and section headers, empty tables);
      • fill everything with placeholder or garbage content;
      • take exactly one grade.

It runs inside the container with the real MCP tools plus a synthetic grade tool, and tries three independent times under a shared time budget. If any attempt scores at least 0.3, the judge reads the mandated work, the actual deliverable content, and the tool calls to decide whether the high score was hollow (won by a fake) or genuine. A separate check confirms the ground-truth files are not readable from the solver's sandbox.

Pseudocode
if readable_by(uid=1000, gt_files): flag("ground_truth_leak")  # critical
best = 0
for attempt in range(3):                   # best-of-K, shared 1400s budget
    solver = ClaudeAgent("claude-opus-5", tools=MCP+[grade_submission])
    solver.build_hollow_deliverable()      # structure only, placeholders
    best = max(best, solver.grade_submission())   # exactly 1 grade
if best >= 0.30:
    v = judge("claude-sonnet-5", mandated_work, deliverable_text, calls)
    if v.hollow: flag("reward_hacking", score=best)

Example: the Quality-of-Earnings workbook has a verifier that only checks that some cell contains the text "Adjusted EBITDA". The red-teamer submits a one-cell spreadsheet with exactly that label and no analysis behind it, and still collects that verifier's full weight.

QoE_Model.xlsx (hollow submission)
AB
1Adjusted EBITDA
2

A claude-sonnet-5 model reads each run's trajectory (screenshots stripped) through a direct API call, with no container and no tools of its own. It reports only tool-level runtime problems, not task or grading issues:

  • a tool that errors on valid use;
  • one the agent gives up on and works around;
  • one that hangs or hits a lock;
  • one that returns wrong data.

Each run is analyzed in its own separate API call, so if one run's analysis fails the others still go through. A final call then groups together the findings that are really the same underlying problem showing up across several runs, so one broken tool is reported once instead of many times.

Pseudocode
findings = []
for run in runs:                           # one API call per run
    t = strip_screenshots(run.trajectory)
    findings += llm("claude-sonnet-5", classify_tool_events(t))
    # tool_error | tool_abandoned | tool_hang | tool_wrong_data
if len(findings) >= 2:
    findings = llm_cluster(findings)       # merge same root cause
return findings

Example: on step 47 the agent calls save on the model workbook and the spreadsheet tool returns a lock error. It retries twice, gets the same error, and gives up with its work unsaved.

trajectory, step 47
excel.save_workbook(path="LBO_Model.xlsx")
UnoException: document is locked for editing by another process
> retry 1 ... same error
> retry 2 ... same error
> agent abandons the spreadsheet with edits unsaved

5. Results

This benchmark shows that Opus 5 has the strongest performance across both harnesses, with Grok 4.5 in second in its native harness and then GPT 5.6 Sol across both harnesses. GPT 5.6 Sol is the cheapest and Grok 4.5 in Grok Build is the most token-efficient. We see across tasks that the underlying model drives the mean score far more than the harness. We see that the value of the harness is primarily on cost. A given model can post nearly the same score in either harness while spending far more in one of them. Here, we outline our results in more detail.

5.1 Experimental setup

We evaluate seven models: Opus 5, GPT 5.6 Sol, Gemini 3.6 Flash, Grok 4.5, Gemini 3.1 Pro, Meta muse-spark 1.1, and DeepSeek V4-Pro. Each model runs in two configurations: inside the Halluminate harness, and inside its own native harness (Claude Code for Opus 5, Codex for GPT 5.6 Sol, Gemini CLI for both Gemini models, and Grok Build for Grok 4.5). Meta muse-spark 1.1 and DeepSeek V4-Pro have no first-party native harness, so we use OpenCode instead. Both configurations share the same tool suite; the native harness adds code execution (a shell and file editing) on top. Scores are verifier based and aggregated per problem. Costs use platform-recorded spend where available and priced tokens otherwise. Beyond outcome scores, three independent judge models labeled every step of every final run, and a linking pass attributed each verifier outcome to the specific steps that caused it, including breakage moments and recovery pivots. This supports the trajectory analysis below.

As above, pass rate is the share of problems whose single graded run clears the 0.50 pass bar, meaning it earns a verifier score of at least 0.50. The figure below generalizes that view, sweeping the bar from 0 to 100 percent and plotting how many tasks each configuration still clears.

Tasks cleared at each score threshold
Show models
020406080020406080100pass bar (50%)Score thresholdNumber of Tasks scoring above thresholdOpus 5GPT 5.6 SolGemini 3.6 flashGrok 4.5Gemini 3.1 ProMeta muse-spark 1.1DeepSeek V4-PronativeHalluminateOpenCode

5.2 Performance and compute

On mean score, Opus 5 leads the field. On token efficiency, the standout is GPT 5.6 Sol, which uses fewer tokens in both harnesses while remaining among the top performers. On dollars, Opus 5 inside the Halluminate harness is by far the most expensive configuration; it is simultaneously the costliest and among the best performing, which frames the central tradeoff the benchmark exposes: the top of the score axis is purchasable, but at a steep multiple of what near peers spend.

Performance and compute by model and harness

0.000.120.240.360.480.60Mean score0.510.50Opus 50.420.42GPT 5.6 Sol0.320.30Gemini 3.6 flash0.280.12Gemini 3.1 Pro0.440.28Grok 4.50.180.30Meta muse-spark1.10.190.13DeepSeek V4-PronativeHalluminateOpenCode
0.016.032.048.064.080.0Millions35.072.1Opus 55.513.0GPT 5.6 Sol17.649.8Gemini 3.6 flash4.924.2Gemini 3.1 Pro4.720.8Grok 4.519.45.6Meta muse-spark1.19.76.7DeepSeek V4-PronativeHalluminateOpenCode
0.009.0018.0027.0036.0045.00USD$14.89$44.21Opus 5$6.33$12.97GPT 5.6 Sol$5.81$13.24Gemini 3.6 flash$1.95$12.07Gemini 3.1 Pro$3.47$21.82Grok 4.5$27.12$0.94Meta muse-spark1.1$4.50$8.15DeepSeek V4-PronativeHalluminateOpenCode

5.3 Model or harness

64 percent of score variation is due to which model is chosen and 6 percent due to the harness it runs in, with the remainder attributable to interaction and within-cell variance. We compute this with a two-way decomposition of the 14 model-by-harness mean scores: the total variation across those cells is partitioned into the share explained by the model factor, the share explained by the harness factor, and a remainder capturing their interaction and within-cell noise, with each percentage that factor's share of the total sum of squares. Model choice dominates, and the harness moves the final score surprisingly little. It is essentially flat for the strongest models: Opus 5 and GPT 5.6 Sol each land within about 0.01 across their two harnesses. Only two models lose real ground in the Halluminate harness, Grok 4.5 dropping about 0.16 and Gemini 3.1 Pro about 0.15 of mean score relative to native. Where the harness matters most is not the score at all but the wasted motion behind it: the steps, tokens, and looping a run spends getting there.

Score across every model x harness
0.510.50Opus 50.420.42GPT 5.6 Sol0.320.30Gemini 3.6 flash0.280.12Gemini 3.1 Pro0.440.28Grok 4.50.300.18Meta muse-spark 1.1(no native: OpenCode)0.130.19DeepSeek V4-Pro(no native: OpenCode)nativeHalluminate0.000.090.180.270.360.45WHAT MOVES THE SCORE64%of variation comes fromwhich model you choose6%comes from the harnessremainder = interaction + within-cell

5.4 Performance by task category

Category-level scores vary across the deal arc: closing and pitch materials are the strongest categories on average, while quality of earnings and modeling are the weakest. Opus 5 leads across the categories, scoring highest on pitch materials and closing.

Mean score by task category, per model and harness
Show models
0.00.10.20.30.40.50.60.7Mean scoreModelingUnderwritingQualityof earningsDatapreparationPitchmaterialsClosingDiligencereviewOpus 5GPT 5.6 SolGemini 3.6 flashGrok 4.5Gemini 3.1 ProMeta muse-spark 1.1DeepSeek V4-PronativeHalluminateOpenCode

5.5 Trajectory analysis

Our trajectory analysis ran in four phases. In phase one, for every final run of every configuration, three independent judge models labeled each step as relevant, unclear, or irrelevant to the task and flagged looping, detrimental looping, and suboptimal choices. In phase two, a linking pass attributed each verifier outcome to the specific steps that caused it. In phase three, we designed and validated a golden trajectory for each task, shown in blue in the interactive and detailed below. In phase four, on top of the per-run judgments, a cross-model analyst pass defined two overlays: critical regions, shared decision windows of at most fifteen steps in which the task is decided, and a failure point for each failing run, the step at which its outcome was sealed.

The visualization below covers three of the benchmark’s tasks. Each run is drawn as a single line that rises with steps and moves right as tokens are spent, colored by the judged relevance of each step. Click a line or its name label to focus that run, click a segment for the judge’s reasoning on that step, and use “Compare step” to view one step number across every run side by side. The pass bar is 50 percent, and every score shown is a single attempt’s verifier grade, never an average across attempts. The side panel carries each task’s prompt, its verifier mix, the golden trajectory’s design principles, and the full cross-model pattern report.

Scores tell you which model won; they do not tell you where the others lost. For this round we rebuilt trajectory analysis from the ground up: two independent trajectory analyzers, Claude Opus 4.8 and GPT 5.6 Sol, read every step of every analyzed run and labeled it for relevance, action type, artifact, interaction modality, and behavior flags, then tied each failed verifier back to the specific steps that sealed it. The explorer below is the raw material. One click selects a run, double-click zooms to it alone, every step dot opens that step's full read, and "Compare judges" stacks both judges' views of the same runs.

Open the interactive full screen.

5.5.1 How the runs behave

Before asking what went wrong, it helps to see what the models actually do with their steps. Every bar below is one run of one model on one harness, built from the same per-step labels you can browse in the explorer: halluminate-harness rows are solid, other harnesses are faded, and hovering any segment shows the exact share. One reading note that matters throughout: the halluminate harness does not allow code execution. When a halluminate run shows code-modality steps, the model was attempting code execution, typing script files and interpreter commands that the harness never runs, and part of what these figures measure is how long each model takes to figure that out and route around it.

5.5.1.1 Step relevance

Green steps advance the task, amber steps are ambiguous, and red steps do not. Halluminate-harness runs compress into a narrow band of steps with mostly green mixes. Native runs spread from under a hundred steps to many hundreds, and the longest lines carry sustained amber and red stretches, compute spent without advancing the deliverable. Because each verifier outcome is linked to the steps that caused it, clicking a step shows which checks that moment earned or lost.

5.5.1.2 Looping

Looping is common, and a significant share of it is detrimental. On their native harnesses, Grok 4.5 and GPT 5.6 Sol loop only to self-verify; Grok 4.5 shows the highest looping share of any configuration, and none of it is judged detrimental. Detrimental looping concentrates in Gemini 3.6 Flash and Opus 5 inside the Halluminate harness. Grok 4.5 inside the Halluminate harness registers no measured looping.

5.5.1.3 Golden trajectory

The blue line in each task is the golden trajectory. For each task we designed an ordered plan of tool calls with complete literal arguments, per-step reasoning, and expected results, constructed from the full ten-run evidence base to satisfy every verifier while remaining realistic and minimal. Each plan was then tested by independent executor agents in the same environment, and a golden was accepted only if it executed with 100 percent accuracy and passed every verifier in the task’s own harness, across two different executor models. The accepted runs land at full credit with a fraction of the tokens the organic runs spend, and each is clickable step by step, including the plan’s reasoning.

What the runs spend their steps on

Share of steps by action type for every run, split by harness.
deepseek-v4-pro / halluminatedeepseek-v4-pro / halluminate · Explore: 10.0%deepseek-v4-pro / halluminate · Build: 79.9%deepseek-v4-pro / halluminate · Verify: 6.5%deepseek-v4-pro / halluminate · Fix: 1.1%deepseek-v4-pro / halluminate · Navigate: 0.9%deepseek-v4-pro / halluminate · Communicate: 1.5%deepseek-v4-pro / halluminate · Narrate: 0.2%deepseek-v4-pro / opencodedeepseek-v4-pro / opencode · Explore: 17.8%deepseek-v4-pro / opencode · Build: 58.0%deepseek-v4-pro / opencode · Verify: 8.2%deepseek-v4-pro / opencode · Fix: 8.5%deepseek-v4-pro / opencode · Navigate: 2.8%deepseek-v4-pro / opencode · Communicate: 1.0%deepseek-v4-pro / opencode · Narrate: 3.6%gemini-3.1-pro / halluminategemini-3.1-pro / halluminate · Explore: 3.7%gemini-3.1-pro / halluminate · Build: 27.1%gemini-3.1-pro / halluminate · Verify: 9.1%gemini-3.1-pro / halluminate · Fix: 3.8%gemini-3.1-pro / halluminate · Navigate: 55.8%gemini-3.1-pro / halluminate · Communicate: 0.3%gemini-3.1-pro / halluminate · Narrate: 0.1%gemini-3.1-pro / nativegemini-3.1-pro / native · Explore: 49.2%gemini-3.1-pro / native · Build: 10.8%gemini-3.1-pro / native · Verify: 16.9%gemini-3.1-pro / native · Fix: 8.5%gemini-3.1-pro / native · Navigate: 2.3%gemini-3.1-pro / native · Communicate: 5.4%gemini-3.1-pro / native · Narrate: 6.9%gemini-3.6-flash / halluminategemini-3.6-flash / halluminate · Explore: 15.9%gemini-3.6-flash / halluminate · Build: 26.1%gemini-3.6-flash / halluminate · Verify: 2.0%gemini-3.6-flash / halluminate · Fix: 5.1%gemini-3.6-flash / halluminate · Navigate: 50.0%gemini-3.6-flash / halluminate · Communicate: 0.5%gemini-3.6-flash / halluminate · Narrate: 0.4%gemini-3.6-flash / nativegemini-3.6-flash / native · Explore: 52.5%gemini-3.6-flash / native · Build: 23.2%gemini-3.6-flash / native · Verify: 10.7%gemini-3.6-flash / native · Fix: 6.6%gemini-3.6-flash / native · Navigate: 3.9%gemini-3.6-flash / native · Communicate: 1.8%gemini-3.6-flash / native · Narrate: 1.4%gpt-5.6-sol / halluminategpt-5.6-sol / halluminate · Explore: 20.0%gpt-5.6-sol / halluminate · Build: 28.8%gpt-5.6-sol / halluminate · Verify: 9.8%gpt-5.6-sol / halluminate · Fix: 3.2%gpt-5.6-sol / halluminate · Navigate: 28.8%gpt-5.6-sol / halluminate · Communicate: 0.9%gpt-5.6-sol / halluminate · Narrate: 8.6%gpt-5.6-sol / nativegpt-5.6-sol / native · Explore: 35.9%gpt-5.6-sol / native · Build: 6.9%gpt-5.6-sol / native · Verify: 14.4%gpt-5.6-sol / native · Fix: 7.8%gpt-5.6-sol / native · Navigate: 32.7%gpt-5.6-sol / native · Communicate: 2.3%grok-4.5 / halluminategrok-4.5 / halluminate · Explore: 9.0%grok-4.5 / halluminate · Build: 10.8%grok-4.5 / halluminate · Verify: 3.9%grok-4.5 / halluminate · Fix: 5.8%grok-4.5 / halluminate · Navigate: 70.1%grok-4.5 / halluminate · Communicate: 0.3%grok-4.5 / nativegrok-4.5 / native · Explore: 37.9%grok-4.5 / native · Build: 3.8%grok-4.5 / native · Verify: 3.5%grok-4.5 / native · Fix: 4.0%grok-4.5 / native · Navigate: 49.1%grok-4.5 / native · Communicate: 1.7%meta-muse-spark-1.1 / halluminatemeta-muse-spark-1.1 / halluminate · Explore: 26.3%meta-muse-spark-1.1 / halluminate · Build: 17.0%meta-muse-spark-1.1 / halluminate · Verify: 9.8%meta-muse-spark-1.1 / halluminate · Fix: 1.0%meta-muse-spark-1.1 / halluminate · Navigate: 44.5%meta-muse-spark-1.1 / halluminate · Communicate: 0.8%meta-muse-spark-1.1 / halluminate · Narrate: 0.6%meta-muse-spark-1.1 / opencodemeta-muse-spark-1.1 / opencode · Explore: 44.6%meta-muse-spark-1.1 / opencode · Build: 34.7%meta-muse-spark-1.1 / opencode · Verify: 6.8%meta-muse-spark-1.1 / opencode · Fix: 6.8%meta-muse-spark-1.1 / opencode · Navigate: 4.5%meta-muse-spark-1.1 / opencode · Communicate: 2.7%opus-5 / halluminateopus-5 / halluminate · Explore: 17.6%opus-5 / halluminate · Build: 34.5%opus-5 / halluminate · Verify: 5.4%opus-5 / halluminate · Fix: 13.3%opus-5 / halluminate · Navigate: 21.6%opus-5 / halluminate · Communicate: 3.8%opus-5 / halluminate · Narrate: 3.7%opus-5 / nativeopus-5 / native · Explore: 50.4%opus-5 / native · Build: 12.2%opus-5 / native · Verify: 13.5%opus-5 / native · Fix: 5.2%opus-5 / native · Navigate: 15.2%opus-5 / native · Communicate: 2.6%opus-5 / native · Narrate: 0.9%ExploreBuildVerifyFixNavigateCommunicateNarrate

Artifact focus by run

Share of steps touching each artifact family, split by harness.
deepseek-v4-pro / halluminatedeepseek-v4-pro / halluminate · excel: 21.9%deepseek-v4-pro / halluminate · docx: 71.4%deepseek-v4-pro / halluminate · email: 1.3%deepseek-v4-pro / halluminate · pdf: 1.5%deepseek-v4-pro / halluminate · filesystem: 3.2%deepseek-v4-pro / halluminate · memory: 0.2%deepseek-v4-pro / halluminate · other: 0.4%deepseek-v4-pro / opencodedeepseek-v4-pro / opencode · excel: 43.6%deepseek-v4-pro / opencode · docx: 40.5%deepseek-v4-pro / opencode · email: 1.0%deepseek-v4-pro / opencode · pdf: 1.0%deepseek-v4-pro / opencode · filesystem: 9.8%deepseek-v4-pro / opencode · memory: 1.5%deepseek-v4-pro / opencode · browser: 0.8%deepseek-v4-pro / opencode · other: 1.8%gemini-3.1-pro / halluminategemini-3.1-pro / halluminate · excel: 12.1%gemini-3.1-pro / halluminate · docx: 12.7%gemini-3.1-pro / halluminate · email: 0.6%gemini-3.1-pro / halluminate · pdf: 0.2%gemini-3.1-pro / halluminate · filesystem: 31.1%gemini-3.1-pro / halluminate · memory: 0.2%gemini-3.1-pro / halluminate · browser: 11.0%gemini-3.1-pro / halluminate · other: 32.2%gemini-3.1-pro / nativegemini-3.1-pro / native · excel: 47.7%gemini-3.1-pro / native · docx: 7.7%gemini-3.1-pro / native · email: 1.5%gemini-3.1-pro / native · pdf: 1.5%gemini-3.1-pro / native · filesystem: 27.7%gemini-3.1-pro / native · memory: 7.7%gemini-3.1-pro / native · browser: 1.5%gemini-3.1-pro / native · other: 4.6%gemini-3.6-flash / halluminategemini-3.6-flash / halluminate · excel: 6.9%gemini-3.6-flash / halluminate · docx: 34.6%gemini-3.6-flash / halluminate · email: 1.1%gemini-3.6-flash / halluminate · pdf: 0.2%gemini-3.6-flash / halluminate · filesystem: 25.4%gemini-3.6-flash / halluminate · memory: 0.3%gemini-3.6-flash / halluminate · browser: 7.6%gemini-3.6-flash / halluminate · other: 23.9%gemini-3.6-flash / nativegemini-3.6-flash / native · excel: 28.6%gemini-3.6-flash / native · docx: 11.6%gemini-3.6-flash / native · email: 6.4%gemini-3.6-flash / native · pdf: 9.5%gemini-3.6-flash / native · filesystem: 27.0%gemini-3.6-flash / native · memory: 1.8%gemini-3.6-flash / native · browser: 3.4%gemini-3.6-flash / native · other: 11.6%gpt-5.6-sol / halluminategpt-5.6-sol / halluminate · excel: 19.4%gpt-5.6-sol / halluminate · docx: 33.2%gpt-5.6-sol / halluminate · email: 0.4%gpt-5.6-sol / halluminate · pdf: 1.4%gpt-5.6-sol / halluminate · filesystem: 10.4%gpt-5.6-sol / halluminate · memory: 12.6%gpt-5.6-sol / halluminate · browser: 21.1%gpt-5.6-sol / halluminate · other: 1.5%gpt-5.6-sol / nativegpt-5.6-sol / native · excel: 10.5%gpt-5.6-sol / native · docx: 19.6%gpt-5.6-sol / native · email: 11.8%gpt-5.6-sol / native · pdf: 8.5%gpt-5.6-sol / native · filesystem: 12.7%gpt-5.6-sol / native · memory: 0.7%gpt-5.6-sol / native · browser: 30.1%gpt-5.6-sol / native · other: 6.2%grok-4.5 / halluminategrok-4.5 / halluminate · excel: 13.8%grok-4.5 / halluminate · pptx: 1.0%grok-4.5 / halluminate · docx: 3.6%grok-4.5 / halluminate · email: 1.2%grok-4.5 / halluminate · pdf: 1.1%grok-4.5 / halluminate · filesystem: 25.4%grok-4.5 / halluminate · memory: 0.3%grok-4.5 / halluminate · browser: 6.2%grok-4.5 / halluminate · other: 47.4%grok-4.5 / nativegrok-4.5 / native · excel: 6.1%grok-4.5 / native · docx: 7.2%grok-4.5 / native · email: 5.8%grok-4.5 / native · pdf: 9.5%grok-4.5 / native · filesystem: 12.7%grok-4.5 / native · browser: 41.0%grok-4.5 / native · other: 17.6%meta-muse-spark-1.1 / halluminatemeta-muse-spark-1.1 / halluminate · excel: 40.6%meta-muse-spark-1.1 / halluminate · docx: 8.0%meta-muse-spark-1.1 / halluminate · email: 3.0%meta-muse-spark-1.1 / halluminate · pdf: 0.5%meta-muse-spark-1.1 / halluminate · filesystem: 17.2%meta-muse-spark-1.1 / halluminate · browser: 11.0%meta-muse-spark-1.1 / halluminate · other: 19.7%meta-muse-spark-1.1 / opencodemeta-muse-spark-1.1 / opencode · excel: 20.3%meta-muse-spark-1.1 / opencode · docx: 43.7%meta-muse-spark-1.1 / opencode · email: 4.5%meta-muse-spark-1.1 / opencode · pdf: 8.6%meta-muse-spark-1.1 / opencode · filesystem: 14.4%meta-muse-spark-1.1 / opencode · browser: 6.3%meta-muse-spark-1.1 / opencode · other: 2.3%opus-5 / halluminateopus-5 / halluminate · excel: 19.6%opus-5 / halluminate · docx: 46.4%opus-5 / halluminate · email: 8.7%opus-5 / halluminate · pdf: 3.2%opus-5 / halluminate · filesystem: 3.4%opus-5 / halluminate · memory: 5.3%opus-5 / halluminate · browser: 8.7%opus-5 / halluminate · other: 4.7%opus-5 / nativeopus-5 / native · excel: 21.7%opus-5 / native · docx: 13.5%opus-5 / native · email: 7.8%opus-5 / native · pdf: 7.8%opus-5 / native · filesystem: 20.9%opus-5 / native · memory: 2.2%opus-5 / native · browser: 20.9%opus-5 / native · other: 5.2%excelpptxdocxemailpdffilesystemmemorybrowserother

How the runs interact with the environment

Share of steps by interaction modality, split by harness: halluminate rows filled, other harnesses plain. Code cells on halluminate rows are attempts, marked with an asterisk: that harness does not execute typed code.
runStructured tool call (API)Code execution (terminal or script)GUI interaction (click, type, scroll)
deepseek-v4-pro / halluminate99.6%0.0%0.4%
deepseek-v4-pro / opencode92.0%4.6%3.4%
gemini-3.1-pro / halluminate17.6%8.6%*73.8%
gemini-3.1-pro / native27.7%69.2%3.1%
gemini-3.6-flash / halluminate23.9%0.0%76.1%
gemini-3.6-flash / native55.2%43.9%0.9%
gpt-5.6-sol / halluminate32.1%0.0%67.9%
gpt-5.6-sol / native3.9%54.6%41.5%
grok-4.5 / halluminate21.2%0.0%78.8%
grok-4.5 / native6.1%32.9%61.0%
meta-muse-spark-1.1 / halluminate54.3%0.0%45.7%
meta-muse-spark-1.1 / opencode67.6%29.7%2.7%
opus-5 / halluminate30.7%0.0%69.3%
opus-5 / native21.3%54.8%23.9%
* Attempted, not executed: the halluminate harness does not allow code execution. On those runs the model types script files and interpreter commands that the harness never runs, and it has to eventually figure that out and route around it. In the taxonomy this shows up as Affordance gap not escalated and Tool format friction when the model keeps trying instead of adapting.

5.5.2 Where the failures concentrate

The failure taxonomy has two axes. Finance reasoning failures are errors in the analysis itself: anchoring on stale sources, copying numbers instead of deriving them, omitting required blocks, work that does not tie out. Long-horizon capability failures are errors in the agent: instructions lost over the horizon, no self verification, premature termination, detrimental looping. Each column below is a distribution and sums to 100 percent, so providers with very different failure volumes stay comparable; the two Geminis pool under Google. Click any failure name for its full description: what the failure is, how it differs from its neighbors, how often and how severely it hit, and real examples from the runs, and use the toggle to switch between the pooled view and each trajectory analyzer alone.

Two more analyzer outputs sit behind these distributions. For every failed run, each trajectory analyzer names a failure point: the single step after which the run could no longer pass, drawn in the explorer as a dashed red ring, with the reasoning in the run panel. And for every task it marks critical regions: the step ranges where the outcome was actually decided, drawn as shaded bands. The tables below count every verifier-linked failure by type; the failure points say which single step sealed each loss, and that sealing step often sits far upstream of where the damage becomes visible.

Finance reasoning failures: where each provider's failure mass goes

For both judges (Opus 4.8 and GPT 5.6 Sol) we show, of all finance-reasoning failure hits for each provider family, the share falling on each type.
OverallAnthropicDeepSeekGoogleMetaOpenAIxAI
Omitted required block (explicit)
Omitted a block the instructions explicitly required.
31.8%33.7%21.8%32.6%30.9%40.7%30.0%
Wrong analytical method
A wrong analytical method substituted for the required one.
23.2%22.4%31.0%19.3%29.4%21.0%21.4%
Copied instead of derived
Copied a number instead of deriving it, with no explicit verification.
20.2%17.3%17.2%19.9%22.1%17.3%30.0%
Omitted expected block (insight)
Omitted analysis a competent analyst would include without being told.
6.5%5.1%9.2%6.1%12.3%5.7%
Reconciliation and tie-out failure
Outputs that should reconcile with each other do not tie out.
4.6%6.1%9.2%3.9%4.4%1.2%2.9%
Stale source anchoring
Anchored on an outdated source figure when a fresher or computable one was available.
4.1%6.1%3.4%2.2%5.9%3.7%5.7%
Wrong presentation of correct work
Correct analysis presented in the wrong form or place.
3.9%4.1%2.3%8.3%2.9%
Missed superseding correction
Missed a correction issued mid task that superseded earlier information.
2.4%2.0%1.1%5.0%1.2%1.4%
Unit and format errors
Unit, currency, magnitude or format errors in figures.
1.7%1.0%3.4%1.7%2.9%1.4%
Formula discipline
Formula discipline errors: wrong ranges, broken references, hardcoded values.
1.5%2.0%1.1%1.1%1.5%2.5%1.4%

Long-horizon capability failures: where each provider's failure mass goes

We show the same construction for the capability axis for both judges (Opus 4.8 and GPT 5.6 Sol).
OverallAnthropicDeepSeekGoogleMetaOpenAIxAI
Instruction loss over horizon
Instructions given early in the task are lost by the time they matter.
42.4%33.7%54.6%28.8%53.2%61.8%43.0%
False verification
Claimed to verify, but the check was shallow, circular or fake.
15.4%13.9%5.9%24.9%8.3%5.9%19.0%
Silent scope drop
Part of the task silently dropped mid run, never delivered or mentioned.
12.9%6.9%21.0%14.9%3.7%15.7%11.0%
No self verification
Finished without checking its own work against the available evidence.
9.5%3.0%2.5%10.0%25.7%2.9%12.0%
Premature discovery closure
Closed discovery too early and missed available material.
5.8%18.8%3.4%6.0%2.9%4.0%
Provided template ignored
A provided template or example was ignored.
3.8%1.0%9.2%2.1%4.6%2.9%5.0%
Premature termination
Stopped early with the final output incomplete.
2.5%19.8%
Delivery mechanics failure
Work was done but delivered wrongly: wrong file, place or naming.
2.1%6.0%
Affordance gap not escalated
Blocked by a missing affordance and did not escalate or rebalance effort to unexplored areas.
2.1%2.0%0.8%3.2%2.0%3.0%
Verbatim fidelity miss
Content that had to be carried verbatim was paraphrased or altered.
1.8%1.7%2.5%0.9%4.9%
Context loss over horizon
Earlier context is forgotten or contradicted later in the run.
0.7%0.8%0.4%1.8%1.0%1.0%
Detrimental looping
Repeated a failing approach without changing strategy.
0.6%1.1%2.0%
Plan drift
Later work drifts away from the stated plan without re-planning.
0.2%1.8%
Fabricated state or success
Fabricated state, data or success that never happened.
0.1%1.0%

5.5.3 Alignment between the two trajectory analyzers

Every label in this analysis was assigned twice, independently, by trajectory analyzers from two different labs: Claude Opus 4.8 and GPT 5.6 Sol. The table below shows, for every label they assign, each analyzer's share, the gap between them, and the misalignment rate: of the steps where at least one analyzer applies the label, the share where the other does not. Self bias gets its own check, comparing how each analyzer reads runs from its own lab against the other analyzer's read of the very same steps.

Label shares and misalignment
labelOpus 4.8GPT 5.6 Solgapmisalignment
Relevant60.9%66.6%-5.7 pp18.0%
Unclear19.3%15.2%+4.1 pp69.1%
Irrelevant19.8%18.2%+1.6 pp40.0%
Looping (any)24.7%26.4%-1.7 pp30.2%
Suboptimal3.0%17.9%-14.8 pp88.1%
Self bias on own-lab runs
trajectory analyzerown-lab stepsown relevant shareother analyzer's readself bias
Opus 4.857493.6%94.6%-1.0 pp
GPT 5.6 Sol84585.6%79.5%+6.0 pp
Shares are the fraction of all double-labeled steps each analyzer gives the label; gap is Opus minus GPT in percentage points.

6. Conclusions and Takeaways

Westworld Finance Diligence Bench asks a model to do a complete deal, not a single task, and the gap that opens up is large: the best configuration clears only 0.51 of the available credit, and most configurations sit well below it. The hardest parts are not exotic. Quality of earnings and modeling are the lowest-scoring categories, and specific chores like reconciling a shifting email thread against a contract are where scores collapse.

The result we did not expect is how little the harness moves the score itself. Model choice explains most of the outcome; the harness explains only about six percent, and for the strongest models it is essentially flat. It bites just two models, Grok 4.5 and Gemini 3.1 Pro, which each lose around 0.15 to 0.16 of mean score moving into the Halluminate harness. Where the harnesses really differ is in wasted motion: how many steps a run takes, how often it loops, and whether that looping is useful self-checking or spinning in place. For teams deploying agents on long-horizon work, that is the actionable takeaway: once you have chosen a model, the scaffolding around it mostly shapes how efficiently it gets there, not whether it succeeds.

7. Next Steps

Given what this benchmark reveals, several directions stand out for future work:

  • 1Post-train open-source models on the benchmark. Because even frontier models clear only about half of the available credit, smaller open models would likely need a curriculum of easier tasks first, whereas a frontier-scale open model could be trained on it directly.
  • 2Turn the error analysis into a training signal. The per-step judgments, the executed golden trajectories, and the identified breakage points are a natural source of process-level reward, crediting how a solution is reached rather than only the final score.
  • 3Measure transfer. Does competence built here carry to adjacent finance work, to other expert domains such as accounting, law, and software engineering, or across modalities, for instance whether training on this process-and-tool-heavy benchmark improves pure coding or web navigation? We leave these questions to future work.

We will release a public subset of the benchmark in a GitHub repository, along with an accompanying paper, soon.

Authors

Alina Hyk, Robert Alward, Victoria Knapp Perez, and Wyatt Marshall

All authors contributed equally.