Skip to content

Paper review 2026-07-30: 17 candidates #347

Description

@github-actions

Reply with commands, one per line:
/approve all · /approve 1,3-5 · /reject 2 · /edit 3 category=world_terminal tags=benchmark (edit implies approve; tags=- clears tags). Valid category keys: see taxonomy.json.

1. Foundational Refinement Proofs for Deployed Bytecode, at the Price of Tokens

Lefteris Lazaropoulos, Zoe Paraskevopoulou · arXiv 2026/07 · paper
proposed: software_testing · tags: none

EquiVM is a Lean-based framework enabling LLM agents to produce foundational, machine-checked refinement proofs between deployed EVM bytecode and high-level specifications, applied end-to-end to 23 real-world smart contracts including MakerDAO. It demonstrates agentic proof development can now certify low-level code correctness at scale.
Formal verification of software correctness performed by an LLM agent falls under software_testing per the boundary rule.

2. How well LLM-based test generation techniques perform with newer LLM versions?

Michael Konstantinou, Renzo Degiovanni, Mike Papadakis · arXiv 2026/07 · paper
proposed: software_testing · tags: empirical

Replicates four state-of-the-art LLM-based unit test generation tools (HITS, SymPrompt, TestSpark, CoverUp) with newer LLMs, finding plain LLM prompting now outperforms these engineered pipelines on coverage and mutation score at comparable cost.
Single-activity empirical study of test generation tools, routed to software_testing plus empirical tag rather than studies (not field-wide).

3. CodeSpec: Dual Executable Specifications for Agentic Long-Horizon Feature Development

Peiding Wang, Li Zhang, Fang Liu, et al. · arXiv 2026/07 · paper
proposed: software_development · tags: none

CodeSpec proposes dual executable architecture/behavior specifications that guide LLM code agents through long-horizon feature development in existing repositories, improving design-implementation consistency and outperforming baselines on FeatureBench.
Targets integrating new functionality into existing codebases, matching the software_development leaf (feature addition, not from-scratch generation).

4. AgenticCANN: Automated Ascend C Operator Generation via Knowledge-Augmented Agentic Evolution

Junhao Qiu, Zidong Wang, Yansong Sun, et al. · arXiv 2026/07 · paper
proposed: systems · tags: none

AgenticCANN is a knowledge-augmented agentic evolution framework that automatically generates and optimizes Ascend C NPU operators, combining structured domain knowledge injection with stage-adaptive LLM-driven evolution. It achieves high feasibility and speedups on Huawei Ascend hardware.
Agent produces low-level hardware/kernel operator code, matching the systems leaf's kernel/runtime code generation.

5. GoGoTB: Agentic RTL Verification with Specification-Grounded Coverage Closure

Xin Xin, Jincheng Lou, Junhui Li, et al. · arXiv 2026/07 · paper
proposed: hardware · tags: none

GoGoTB is an agentic framework for RTL functional verification that generates a full testbench environment with specification-grounded coverage closure, using shared context across an execution-control layer, knowledge system, and coverage framework. It achieves complete verification environment generation and high coverage on 8 RTL designs without human intervention.
The deliverable is hardware verification/testbench code tied to RTL designs, matching the hardware leaf's hardware-verification-via-code scope.

6. ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation

Ekaterina Trofimova, Zosia Shamina, Maria Selifanova, et al. · arXiv 2026/07 · paper
proposed: world_research · tags: benchmark

ML2B is a benchmark evaluating LLMs' ability to generate end-to-end ML pipelines from task descriptions across 14 languages, using 35 Kaggle competitions with isolated evaluation infrastructure.
MLE-bench-style task where an LLM writes and runs ML pipeline code to produce a trained model, matching world_research's machine-learning-engineering agent examples.

7. ClinLens: Towards Long-Horizon Coding Agents for Longitudinal Multimodal Clinical Data Science

Yuan Zhu, Ethan B. Liu, Frank Nie, et al. · arXiv 2026/07 · paper
proposed: world_research · tags: benchmark

CLINLENS is a benchmark of 200 executable clinical data-science tasks over linked MIMIC resources (EHR, notes, ECG, imaging), evaluating coding agents' ability to produce auditable longitudinal analyses. Results show a large gap between runnable code and clinically correct analysis across model-scaffold configurations.
A data-science/data-analysis agent benchmark whose deliverable is clinical insight, fitting world_research's data-analysis and data-science agent scope.

8. Auto Research for Materials: Auditable AI-Scientist Workflows with Held-Out Transfer

Jingjie Ning, Xiaochuan Li, Shanshan Zhong, et al. · arXiv 2026/07 · paper
proposed: world_research · tags: none

Auto Research uses LLM agents to propose, implement, and evaluate ML changes for materials science in a closed loop, validating decisions via held-out transfer across Matbench endpoints. It introduces intervention-centered evaluation distinguishing reusable discoveries from overfit gains.
AI-scientist style ML-engineering agent whose product is validated scientific findings/models, fitting world_research.

9. PUDA: An AI-Native Hardware Harness for Self-Driving Laboratories

Zekun Ren, Hongzhao Tan, Jiaen Yee, et al. · arXiv 2026/07 · paper
proposed: world_research · tags: none

PUDA is an AI-native command-line hardware harness letting agents autonomously observe, decide, and act over physical lab experiments while keeping hardware execution deterministic and auditable. It provides provenance-linked data infrastructure for self-driving laboratories.
Self-driving-lab agentic infrastructure is scientific-discovery agency, placed under world_research per taxonomy examples.

10. EvoPINN: Agentic Discovery of Executable Algorithms for Physics-Informed Neural Networks

Peng Yin, Kai Li, Yifan Zhang, et al. · arXiv 2026/07 · paper
proposed: world_research · tags: none

EvoPINN is an agentic framework where an LLM iteratively proposes and executes code modifications to discover specialized neural PDE-solving algorithms, verifying candidates via structural checks and PDE evaluation. It autonomously discovers a novel architecture (SLRC-PINN) improving accuracy across PDE regimes.
An execution-grounded agent discovering scientific computing methods serves scientific discovery, i.e. world_research.

11. UrbanDS: A Graph-Guided LLM Multi-Agent System for Data-Intensive Urban Tasks

Zhilun Zhou, Jianghao Yu, Yuming Lin, et al. · arXiv 2026/07 · paper
proposed: world_research · tags: benchmark

UrbanDS is a graph-guided LLM multi-agent system that discovers and integrates heterogeneous urban datasets to answer data-intensive urban analysis tasks, evaluated on a new UrbanDS-Bench and deployed in a real city operations platform.
A data-science/data-exploration multi-agent system whose deliverable is analytical insight, matching world_research.

12. SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response

Lehan Wang, Boli Chen, Ruixue Ding, et al. · arXiv 2026/07 · paper
proposed: world_terminal · tags: benchmark

SecRespond is a benchmark evaluating LLM CLI agents on post-compromise incident response, requiring forensic analysis and remediation planning across 10 realistic compromised cloud-host cyber ranges.
Agents act via command-line interfaces to perform security operations/incident response, matching world_terminal's system-administration and security-operations scope.

13. AgentSnare: Learning to Delay, Divert, and Defuse Autonomous Penetration Agents

Ruoyu Wang, Heng Zhao, Renjie Wu, et al. · arXiv 2026/07 · paper
proposed: world_terminal · tags: none

AgentSnare is a trajectory-adaptive deception system that dynamically constructs decoy artifacts to delay, divert, and defuse autonomous LLM penetration-testing agents, tested against 15 CVE-Bench web applications and three attacker models. It substantially reduces real-target exploitation by grounding attacker completion reports in fabricated decoy evidence.
Targets penetration-testing agents that act via tool calls in an offensive-security terminal setting, matching the world_terminal boundary for CTF/pentesting.

14. Embodied Agents Take Control: Minimal-Interface Zero-Shot Agents Rival Industrial-Scale Policies in Vision-and-Language Navigation

Jian Zhou, Xunyi Zhao, Gengze Zhou, et al. · arXiv 2026/07 · paper
proposed: world_physical · tags: empirical

Studies zero-shot vision-and-language navigation using general-purpose software-engineering agent harnesses (e.g. Claude, fable) given only camera input and discrete actions, comparing them to industrial embodied policies. Finds model choice dominates capability, with hybrid tool-augmented agents matching or beating specialized navigation policies at much lower cost.
Agent harnesses act through code/tool-call primitives to control embodied navigation, matching world_physical's code-as-policy control definition.

15. Functional Cache Grafting: Robust and Rapid Code-Policy Synthesis for Embodied Agents

Saehun Chun, Wonje Choi, Sera Choi, et al. · arXiv 2026/07 · paper
proposed: world_physical · tags: none

FCGraft speeds up and robustifies embodied-agent code-policy synthesis by grafting cached key-value states of validated function-level code skeletons instead of fully regenerating control programs. It achieves faster policy synthesis and higher task success than prompt-level caching baselines.
Generated code serves as the executable control policy for embodied agents acting physically, matching world_physical's code-as-policy definition.

16. OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding

Jingbo Zhou, Yusai Zhao, Qi Bao, et al. · arXiv 2026/07 · paper
proposed: world_apps · tags: benchmark

OmegaUse-OfficeVal is a benchmark of 100 long-horizon office-suite tasks paired with human-labor-time and cost signals to evaluate LLM agents against economic baselines. It uses code-based verifiers built from rubrics to assess agent deliverable quality versus human workers.
Agents operate office-suite applications to complete workflows, matching world_apps' office-automation scope.

17. APEX-Accounting

Julien Benchek, Austin Bennett, Jasmin Kern, et al. · arXiv 2026/07 · paper
proposed: world_apps · tags: benchmark

APEX-Accounting is a closed benchmark of 160 expert-authored accounting tasks (reconciling accounts, posting transactions, producing reports) across accounting systems, spreadsheets, and PDFs. It evaluates nine frontier models under varying token budgets, revealing a Simpson's-paradox effect on scores.
Agents operate accounting/office applications to complete professional tasks, matching world_apps' enterprise-systems scope.

machine payload (do not edit)
[{"paper": {"id": "2607.26306", "title": "Foundational Refinement Proofs for Deployed Bytecode, at the Price of Tokens", "authors": ["Lefteris Lazaropoulos", "Zoe Paraskevopoulou"], "venue": "arXiv 2026/07", "category": "", "published": "2026-07-28", "links": {"paper": "https://arxiv.org/abs/2607.26306", "github": "", "website": ""}}, "category": "software_testing", "tags": [], "summary": "EquiVM is a Lean-based framework enabling LLM agents to produce foundational, machine-checked refinement proofs between deployed EVM bytecode and high-level specifications, applied end-to-end to 23 real-world smart contracts including MakerDAO. It demonstrates agentic proof development can now certify low-level code correctness at scale.", "reason": "Formal verification of software correctness performed by an LLM agent falls under software_testing per the boundary rule.", "source": "crawl"}, {"paper": {"id": "2601.09695", "title": "How well LLM-based test generation techniques perform with newer LLM versions?", "authors": ["Michael Konstantinou", "Renzo Degiovanni", "Mike Papadakis"], "venue": "arXiv 2026/07", "category": "", "published": "2026-07-29", "links": {"paper": "https://arxiv.org/abs/2601.09695", "github": "", "website": ""}}, "category": "software_testing", "tags": ["empirical"], "summary": "Replicates four state-of-the-art LLM-based unit test generation tools (HITS, SymPrompt, TestSpark, CoverUp) with newer LLMs, finding plain LLM prompting now outperforms these engineered pipelines on coverage and mutation score at comparable cost.", "reason": "Single-activity empirical study of test generation tools, routed to software_testing plus empirical tag rather than studies (not field-wide).", "source": "crawl"}, {"paper": {"id": "2607.26777", "title": "CodeSpec: Dual Executable Specifications for Agentic Long-Horizon Feature Development", "authors": ["Peiding Wang", "Li Zhang", "Fang Liu", "Taichuan Li", "Yinghao Zhu"], "venue": "arXiv 2026/07", "category": "", "published": "2026-07-29", "links": {"paper": "https://arxiv.org/abs/2607.26777", "github": "", "website": ""}}, "category": "software_development", "tags": [], "summary": "CodeSpec proposes dual executable architecture/behavior specifications that guide LLM code agents through long-horizon feature development in existing repositories, improving design-implementation consistency and outperforming baselines on FeatureBench.", "reason": "Targets integrating new functionality into existing codebases, matching the software_development leaf (feature addition, not from-scratch generation).", "source": "crawl"}, {"paper": {"id": "2607.26661", "title": "AgenticCANN: Automated Ascend C Operator Generation via Knowledge-Augmented Agentic Evolution", "authors": ["Junhao Qiu", "Zidong Wang", "Yansong Sun", "Zhitong Ma", "Ping Guo", "Qingfu Zhang"], "venue": "arXiv 2026/07", "category": "", "published": "2026-07-29", "links": {"paper": "https://arxiv.org/abs/2607.26661", "github": "", "website": ""}}, "category": "systems", "tags": [], "summary": "AgenticCANN is a knowledge-augmented agentic evolution framework that automatically generates and optimizes Ascend C NPU operators, combining structured domain knowledge injection with stage-adaptive LLM-driven evolution. It achieves high feasibility and speedups on Huawei Ascend hardware.", "reason": "Agent produces low-level hardware/kernel operator code, matching the systems leaf's kernel/runtime code generation.", "source": "crawl"}, {"paper": {"id": "2607.26181", "title": "GoGoTB: Agentic RTL Verification with Specification-Grounded Coverage Closure", "authors": ["Xin Xin", "Jincheng Lou", "Junhui Li", "Jinglin Yan", "Panda Xiao", "Di Wu", "Haixiao Li", "Weicong Lu", "Weijian Fan", "Xinyu Qu", "Yuxiang Zhao", "Min Yu", "Zhixiong Di", "Yibo Lin"], "venue": "arXiv 2026/07", "category": "", "published": "2026-07-28", "links": {"paper": "https://arxiv.org/abs/2607.26181", "github": "", "website": ""}}, "category": "hardware", "tags": [], "summary": "GoGoTB is an agentic framework for RTL functional verification that generates a full testbench environment with specification-grounded coverage closure, using shared context across an execution-control layer, knowledge system, and coverage framework. It achieves complete verification environment generation and high coverage on 8 RTL designs without human intervention.", "reason": "The deliverable is hardware verification/testbench code tied to RTL designs, matching the hardware leaf's hardware-verification-via-code scope.", "source": "crawl"}, {"paper": {"id": "2509.22768", "title": "ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation", "authors": ["Ekaterina Trofimova", "Zosia Shamina", "Maria Selifanova", "Artem Zaitsev", "Remi Savchuk", "Maxim Minets", "Daria Ozerova", "Emil Sataev", "Denis Zuenko", "Andrey E. Ustyuzhanin"], "venue": "arXiv 2026/07", "category": "", "published": "2026-07-28", "links": {"paper": "https://arxiv.org/abs/2509.22768", "github": "", "website": ""}}, "category": "world_research", "tags": ["benchmark"], "summary": "ML2B is a benchmark evaluating LLMs' ability to generate end-to-end ML pipelines from task descriptions across 14 languages, using 35 Kaggle competitions with isolated evaluation infrastructure.", "reason": "MLE-bench-style task where an LLM writes and runs ML pipeline code to produce a trained model, matching world_research's machine-learning-engineering agent examples.", "source": "crawl"}, {"paper": {"id": "2607.26155", "title": "ClinLens: Towards Long-Horizon Coding Agents for Longitudinal Multimodal Clinical Data Science", "authors": ["Yuan Zhu", "Ethan B. Liu", "Frank Nie", "Jindong Han"], "venue": "arXiv 2026/07", "category": "", "published": "2026-07-28", "links": {"paper": "https://arxiv.org/abs/2607.26155", "github": "", "website": ""}}, "category": "world_research", "tags": ["benchmark"], "summary": "CLINLENS is a benchmark of 200 executable clinical data-science tasks over linked MIMIC resources (EHR, notes, ECG, imaging), evaluating coding agents' ability to produce auditable longitudinal analyses. Results show a large gap between runnable code and clinically correct analysis across model-scaffold configurations.", "reason": "A data-science/data-analysis agent benchmark whose deliverable is clinical insight, fitting world_research's data-analysis and data-science agent scope.", "source": "crawl"}, {"paper": {"id": "2607.17100", "title": "Auto Research for Materials: Auditable AI-Scientist Workflows with Held-Out Transfer", "authors": ["Jingjie Ning", "Xiaochuan Li", "Shanshan Zhong", "Ji Zeng", "Guolin Ke"], "venue": "arXiv 2026/07", "category": "", "published": "2026-07-29", "links": {"paper": "https://arxiv.org/abs/2607.17100", "github": "", "website": ""}}, "category": "world_research", "tags": [], "summary": "Auto Research uses LLM agents to propose, implement, and evaluate ML changes for materials science in a closed loop, validating decisions via held-out transfer across Matbench endpoints. It introduces intervention-centered evaluation distinguishing reusable discoveries from overfit gains.", "reason": "AI-scientist style ML-engineering agent whose product is validated scientific findings/models, fitting world_research.", "source": "crawl"}, {"paper": {"id": "2607.26464", "title": "PUDA: An AI-Native Hardware Harness for Self-Driving Laboratories", "authors": ["Zekun Ren", "Hongzhao Tan", "Jiaen Yee", "Kedar Hippalgaonkar"], "venue": "arXiv 2026/07", "category": "", "published": "2026-07-29", "links": {"paper": "https://arxiv.org/abs/2607.26464", "github": "", "website": ""}}, "category": "world_research", "tags": [], "summary": "PUDA is an AI-native command-line hardware harness letting agents autonomously observe, decide, and act over physical lab experiments while keeping hardware execution deterministic and auditable. It provides provenance-linked data infrastructure for self-driving laboratories.", "reason": "Self-driving-lab agentic infrastructure is scientific-discovery agency, placed under world_research per taxonomy examples.", "source": "crawl"}, {"paper": {"id": "2607.26490", "title": "EvoPINN: Agentic Discovery of Executable Algorithms for Physics-Informed Neural Networks", "authors": ["Peng Yin", "Kai Li", "Yifan Zhang", "Jian Cheng"], "venue": "arXiv 2026/07", "category": "", "published": "2026-07-29", "links": {"paper": "https://arxiv.org/abs/2607.26490", "github": "", "website": ""}}, "category": "world_research", "tags": [], "summary": "EvoPINN is an agentic framework where an LLM iteratively proposes and executes code modifications to discover specialized neural PDE-solving algorithms, verifying candidates via structural checks and PDE evaluation. It autonomously discovers a novel architecture (SLRC-PINN) improving accuracy across PDE regimes.", "reason": "An execution-grounded agent discovering scientific computing methods serves scientific discovery, i.e. world_research.", "source": "crawl"}, {"paper": {"id": "2607.26724", "title": "UrbanDS: A Graph-Guided LLM Multi-Agent System for Data-Intensive Urban Tasks", "authors": ["Zhilun Zhou", "Jianghao Yu", "Yuming Lin", "yongjun yang", "Sun Yongquan", "Depeng Jin", "Yong Li"], "venue": "arXiv 2026/07", "category": "", "published": "2026-07-29", "links": {"paper": "https://arxiv.org/abs/2607.26724", "github": "", "website": ""}}, "category": "world_research", "tags": ["benchmark"], "summary": "UrbanDS is a graph-guided LLM multi-agent system that discovers and integrates heterogeneous urban datasets to answer data-intensive urban analysis tasks, evaluated on a new UrbanDS-Bench and deployed in a real city operations platform.", "reason": "A data-science/data-exploration multi-agent system whose deliverable is analytical insight, matching world_research.", "source": "crawl"}, {"paper": {"id": "2607.26791", "title": "SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response", "authors": ["Lehan Wang", "Boli Chen", "Ruixue Ding", "Pengjun Xie", "Jinwei Huang", "Zhendong Liu", "Shuo Wang", "Tao Lei", "Xin Ouyang", "Xiaomeng Li"], "venue": "arXiv 2026/07", "category": "", "published": "2026-07-29", "links": {"paper": "https://arxiv.org/abs/2607.26791", "github": "", "website": ""}}, "category": "world_terminal", "tags": ["benchmark"], "summary": "SecRespond is a benchmark evaluating LLM CLI agents on post-compromise incident response, requiring forensic analysis and remediation planning across 10 realistic compromised cloud-host cyber ranges.", "reason": "Agents act via command-line interfaces to perform security operations/incident response, matching world_terminal's system-administration and security-operations scope.", "source": "crawl"}, {"paper": {"id": "2607.26998", "title": "AgentSnare: Learning to Delay, Divert, and Defuse Autonomous Penetration Agents", "authors": ["Ruoyu Wang", "Heng Zhao", "Renjie Wu", "Mengnan Zhao", "Zhixuan Chu", "Wanyu Lin", "Tianhang Zheng"], "venue": "arXiv 2026/07", "category": "", "published": "2026-07-29", "links": {"paper": "https://arxiv.org/abs/2607.26998", "github": "", "website": ""}}, "category": "world_terminal", "tags": [], "summary": "AgentSnare is a trajectory-adaptive deception system that dynamically constructs decoy artifacts to delay, divert, and defuse autonomous LLM penetration-testing agents, tested against 15 CVE-Bench web applications and three attacker models. It substantially reduces real-target exploitation by grounding attacker completion reports in fabricated decoy evidence.", "reason": "Targets penetration-testing agents that act via tool calls in an offensive-security terminal setting, matching the world_terminal boundary for CTF/pentesting.", "source": "crawl"}, {"paper": {"id": "2607.26148", "title": "Embodied Agents Take Control: Minimal-Interface Zero-Shot Agents Rival Industrial-Scale Policies in Vision-and-Language Navigation", "authors": ["Jian Zhou", "Xunyi Zhao", "Gengze Zhou", "Zerui Li", "Sihao Lin", "Jiajun Liu", "Qi Wu"], "venue": "arXiv 2026/07", "category": "", "published": "2026-07-28", "links": {"paper": "https://arxiv.org/abs/2607.26148", "github": "", "website": ""}}, "category": "world_physical", "tags": ["empirical"], "summary": "Studies zero-shot vision-and-language navigation using general-purpose software-engineering agent harnesses (e.g. Claude, fable) given only camera input and discrete actions, comparing them to industrial embodied policies. Finds model choice dominates capability, with hybrid tool-augmented agents matching or beating specialized navigation policies at much lower cost.", "reason": "Agent harnesses act through code/tool-call primitives to control embodied navigation, matching world_physical's code-as-policy control definition.", "source": "crawl"}, {"paper": {"id": "2606.13097", "title": "Functional Cache Grafting: Robust and Rapid Code-Policy Synthesis for Embodied Agents", "authors": ["Saehun Chun", "Wonje Choi", "Sera Choi", "Sanghyun Ahn", "Honguk Woo"], "venue": "arXiv 2026/07", "category": "", "published": "2026-07-29", "links": {"paper": "https://arxiv.org/abs/2606.13097", "github": "", "website": ""}}, "category": "world_physical", "tags": [], "summary": "FCGraft speeds up and robustifies embodied-agent code-policy synthesis by grafting cached key-value states of validated function-level code skeletons instead of fully regenerating control programs. It achieves faster policy synthesis and higher task success than prompt-level caching baselines.", "reason": "Generated code serves as the executable control policy for embodied agents acting physically, matching world_physical's code-as-policy definition.", "source": "crawl"}, {"paper": {"id": "2607.27155", "title": "OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding", "authors": ["Jingbo Zhou", "Yusai Zhao", "Qi Bao", "Jingjia Cao", "Zhenghai Chen", "Chang Gao", "Kaiqi Guo", "Muxin Guo", "Mingxuan Li", "Xinjiang Lu", "Yanru Ma", "Yixiong Xiao", "Zenghui Zhang", "Le Zhang", "Hua Wu"], "venue": "arXiv 2026/07", "category": "", "published": "2026-07-29", "links": {"paper": "https://arxiv.org/abs/2607.27155", "github": "", "website": ""}}, "category": "world_apps", "tags": ["benchmark"], "summary": "OmegaUse-OfficeVal is a benchmark of 100 long-horizon office-suite tasks paired with human-labor-time and cost signals to evaluate LLM agents against economic baselines. It uses code-based verifiers built from rubrics to assess agent deliverable quality versus human workers.", "reason": "Agents operate office-suite applications to complete workflows, matching world_apps' office-automation scope.", "source": "crawl"}, {"paper": {"id": "2607.27189", "title": "APEX-Accounting", "authors": ["Julien Benchek", "Austin Bennett", "Jasmin Kern", "Ryan Stevens", "Rene Sultan", "Charis Ching", "Hayley Popiel", "Vaibhav Mittal", "Felix Mercier", "Brendan Foody", "Bertie Vidgen"], "venue": "arXiv 2026/07", "category": "", "published": "2026-07-29", "links": {"paper": "https://arxiv.org/abs/2607.27189", "github": "", "website": ""}}, "category": "world_apps", "tags": ["benchmark"], "summary": "APEX-Accounting is a closed benchmark of 160 expert-authored accounting tasks (reconciling accounts, posting transactions, producing reports) across accounting systems, spreadsheets, and PDFs. It evaluates nine frontier models under varying token budgets, revealing a Simpson's-paradox effect on scores.", "reason": "Agents operate accounting/office applications to complete professional tasks, matching world_apps' enterprise-systems scope.", "source": "crawl"}]

Metadata

Metadata

Assignees

No one assigned

    Labels

    paper-reviewPapers awaiting the owner's review

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions