MENU

New ‘TAM’ Benchmark Exposes Critical Gaps in GPT-5’s Long-Horizon Procedural Reasoning, Achieving Only 1% Exact Match on ICD-10-CM Clinical Coding

arXiv Unknown
Overview
A new benchmark, Tasks over Application Manuals (TAM), has been introduced on arXiv to accurately assess LLMs’ long-horizon procedural reasoning. Evaluating GPT-5, TAM revealed alarmingly low exact match performance of 1% for ICD-10-CM clinical coding and 15.5% for U.S. federal sentencing tasks. These results suggest current benchmarks may significantly overestimate LLM reasoning capabilities, highlighting a substantial gap for real-world application.
In Depth

Key Findings

A new benchmark, “Tasks over Application Manuals (TAM),” released on arXiv, has exposed significant deficiencies in the long-horizon procedural reasoning capabilities of Large Language Models (LLMs). Evaluations with the advanced GPT-5 model yielded an exceptionally low exact match performance of just 1% on ICD-10-CM clinical coding tasks and only 15.5% on U.S. federal sentencing tasks. This indicates that existing benchmarks may have substantially over-evaluated the true reasoning capacity of LLMs.

Technical Details

The TAM benchmark is meticulously designed to test real-world problem-solving skills, incorporating two complex domains: ICD-10-CM clinical coding from healthcare and U.S. federal sentencing guidelines from the legal sector. These tasks require LLMs to navigate tens of thousands of rules outlined in authoritative manuals, execute a series of interdependent steps, and generate precise final answers. This multi-step reasoning, coupled with the stringent requirement for rule adherence, sets TAM apart from simpler question-answering or text generation benchmarks.

Background & Context

While LLMs have shown remarkable progress in various tasks, their limitations in complex procedural reasoning and strict logical deduction based on extensive knowledge bases have been a recurring concern. The introduction of TAM provides an objective framework to quantify this gap and guide future LLM development. The abysmal performance of current LLMs in high-stakes fields like medical coding and legal analysis underscores that practical deployment in these areas remains a distant prospect, necessitating fundamental research and technological breakthroughs.

Strategic Significance & Outlook

The TAM benchmark will serve as a crucial tool for enhancing LLMs’ ability to reason rigorously in accordance with specialized manuals and regulations. The low scores achieved by GPT-5 clearly define the research and development imperative to improve LLMs’ capacity to navigate complex rule systems and make multi-step decisions akin to human experts. Through this benchmark, the development of more robust and reliable LLMs is anticipated, potentially unlocking broader applications for AI assistants in critical domains such as healthcare, law, and engineering.

Source: https://arxiv.org/abs/2609.13005

Get our weekly technology intelligence — free

Receive an infographic that lets you judge at a glance whether each field’s analysis report is worth reading.

Subscribe Free — Weekly Tech Intelligence

By subscribing, you’ll receive Troy-Technical’s weekly technology intelligence newsletter.

  • Your email and selected fields are used only to deliver the newsletter.
  • We never share your information with third parties.
  • You can unsubscribe anytime via the link in each email.

See our Privacy Policy for details.

Takes about a minute · Unsubscribe anytime

Let's share this post !

Author of this article

Comments

To comment

TOC